Sign in to view source links and access this dataset
Description
DetailCaps-4870 is an evaluation benchmark for detail image captioning proposed in the paper 'Benchmarking and Improving Detail Image Caption'. It contains 4,870 images curated from various datasets, accompanied by ground truth detail captions generated by GPT-4V, Gemini-1.5-Pro, and GPT-4O. The dataset also includes captions generated by three open-source large vision-language models: LLaVA-1.5, CogVLM, and ShareCaptioner.
Use Cases
Benchmarking detail image captioning models based on the provided ground truth captions.
Comparing the performance of proprietary and open-source vision-language models on a standardized test set.
Training models to generate more detailed captions using the provided reference data.
Strengths
Contains 4,870 images for evaluation.
Provides ground truth captions generated by three advanced proprietary models (GPT-4V, Gemini-1.5-Pro, GPT-4O).
Includes comparative outputs from three open-source models (LLaVA-1.5, CogVLM, ShareCaptioner).
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count for individual caption sets is unknown, which may limit suitability assessment.
Data may reflect bias inherent to the source datasets and model-generated annotations.
Provenance
Source
foundation-multimodal-models
Collection Method
Curated from various datasets; captions generated by proprietary and open-source large vision-language models.
Freshness
Last updated 2025-02-17 04:01:48; freshness should be verified.
License is unknown; terms of use must be verified before application.