1.2 billion object instances across 3.5 million images featuring semantic tags, bounding boxes, and detailed captions. This collection facilitates panoptic visual recognition and general relation comprehension through region-level annotations and image-text pairs.
Use Cases
- Train panoptic segmentation models using the provided object masks and semantic category labels
- Develop visual relationship detection systems using the relation comprehension annotations between specific bounding boxes
- Fine-tune large vision-language models (LVLMs) for region-level captioning using the image-text pairs and localized object descriptions
- Evaluate zero-shot object recognition capabilities across the diverse set of semantic tags
Strengths
- 1.2 billion object instances annotated with semantic tags and bounding boxes
- 3.5 million high-resolution images curated for panoptic understanding
- Includes detailed textual descriptions for object-to-object relations
- Provides region-level question-answering pairs for visual reasoning tasks