ERQA is an evaluation benchmark adapted from embodiedreasoning/ERQA, focusing on real-world scenarios for robotics. The dataset, hosted by FlagEval, was last updated on April 22, 2025, and covers topics related to spatial reasoning and world knowledge. It was originally provided in TFRecord format and has been converted for easier use.
Use Cases
- Benchmarking AI models on spatial reasoning tasks based on real-world scenarios.
- Training robotics systems on world knowledge inference tasks described in the dataset.
- Evaluating multimodal (likely image and text) understanding for embodied agents.
Strengths
- Dataset is explicitly designed as an evaluation benchmark for a focused domain.
- Data was converted from a less accessible TFRecord format to a more user-friendly format.
- Last updated on 2025-04-22, indicating recent maintenance.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and size are unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- FlagEval, adapted from embodiedreasoning/ERQA.
- Collection Method
- Organized and adapted from original TFRecord data.
- Time Range
- null
- Freshness
- Last updated 2025-04-22 03:47:14.
- Geography
- null