Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
FlagEval provides a benchmark for evaluating embodied spatial understanding in large vision-language models. The dataset contains 3,640 question-answer pairs automatically derived from embodied scenes, covering six spatial relationships from an egocentric perspective. It was adapted from an original image-format dataset by Phineas476 and last updated on 2025-04-21.
License is unknown; terms of use must be verified before application.