Human-annotated images paired with referring expressions and specific coordinate points marking referenced locations. The collection includes a diverse range of spatial labels, specifically highlighting expressions with a frequency of 10 or more occurrences.
Use Cases
- Train multimodal models to predict pixel coordinates from referring expressions.
- Benchmark spatial grounding precision using the human-verified point annotations.
- Fine-tune vision models for interactive pointing tasks using the expression and point pairs.
Strengths
- Human-annotated coordinate points marking specific locations for referring expressions.
- Includes high-frequency expressions that appear 10 or more times.
- Serves as a foundational training set for the Molmo family of vision-language models.