Sign in to view source links and access this dataset
Description
A dataset of ImageNet 1K images recaptioned using the Moondream2 vision-language model. The author 'g-ronimo' processed the data on March 9, 2025, generating short, precise captions based on the original ImageNet class names and existing captions. The dataset contains only image IDs and the newly generated text captions.
Use Cases
Fine-tuning image captioning models based on short, descriptive text.
Evaluating the quality of vision-language model outputs against a known benchmark.
Training models for zero-shot image classification using enriched textual descriptions.
Studying the alignment between visual content and machine-generated language.
Strengths
Recaptioned using a specific, named model (Moondream2), providing a consistent generation method.
Focuses on a widely recognized computer vision benchmark (ImageNet 1K).
Captions are generated with a defined prompt structure aimed at brevity and precision.
Limitations
Description metadata is limited; actual data quality requires manual inspection after download.
Column-level documentation is absent; field semantics must be inferred after download.
Row count and dataset size are unknown, which may limit suitability assessment.
Provenance
Source
huggingface
Collection Method
Images from the 'visual-layer/imagenet-1k-vl-enriched' dataset were recaptioned using the 'vikhyatk/moondream2' model.
Freshness
Last updated 2025-03-09 17:38:49.
License is unknown; users must verify permissions before use.