D Cube is a vision-language dataset for object detection and segmentation introduced in the NeurIPS 2023 paper 'Described Object Detection: Liberating Object Detection with Flexible Expressions'. Created by the shikras team, it provides labels characterized by intricate and flexible natural language expressions rather than fixed category names. The data supports multi-modal learning tasks where visual grounding is driven by complex descriptive text.
Use Cases
- Referring expression comprehension using the flexible expression labels to ground text in images
- Open-vocabulary object detection using descriptive text queries to identify novel objects
- Instance segmentation based on intricate natural language descriptions for precise mask generation
Strengths
- Peer-reviewed at NeurIPS 2023
- Supports both detection and segmentation tasks
- Features flexible, intricate natural language expressions instead of fixed categories
Limitations
- Unknown total record count and dataset size
- Potential for high label variance due to the flexible nature of natural language expressions
Provenance
- Source
- NeurIPS 2023 publication 'Described Object Detection: Liberating Object Detection with Flexible Expressions'
- Freshness
- Last updated in March 2024 following its NeurIPS 2023 publication.