166,100 human-generated captions describe 15,100 images from the Open Images validation and test sets. The dataset, dubbed NoCaps for novel object captioning at scale, was created by HuggingFaceM4 and last updated in December 2022. Its associated training data consists of COCO image-caption pairs plus Open Images labels and bounding boxes.
Use Cases
- Benchmarking image captioning models on novel object classes based on the 400 classes with few or no training captions.
- Training models for zero-shot captioning based on the combination of COCO pairs and Open Images labels.
- Studying model generalization based on the scale of 166,100 captions across 15,100 images.
- Developing multi-modal AI systems based on the image-text pairs from Open Images and COCO.
Strengths
- Contains 166,100 human-generated captions, providing a substantial corpus for training and evaluation.
- Focuses on 15,100 images from Open Images, introducing nearly 400 novel object classes not well-represented in COCO.
- Leverages associated training data from established sources like COCO and Open Images.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Last updated 2022-12-14 04:08:38; freshness should be verified.
Provenance
- Source
- Open Images validation and test sets, COCO.
- Collection Method
- Human-generated captions combined with existing training pairs, labels, and bounding boxes.
- Time Range
- null
- Freshness
- 2022-12-14 04:08:38
- Geography
- null