ColorSwap is a multimodal dataset of 2,000 unique image-caption pairs, grouped into 1,000 examples. It was created by stanfordnlp and last updated in February 2024. The dataset is designed to assess and improve the proficiency of multimodal models in matching objects with their colors.
Use Cases
- Benchmarking multimodal model performance on color-object binding tasks based on the described color-swapped pairs.
- Training models to improve visual grounding of color words based on the paired image-caption structure.
- Evaluating model robustness to word order changes in captions based on the description of swapped color words.
Strengths
- Contains 2,000 unique image-caption pairs, providing a substantial testbed.
- Structured into 1,000 examples, each with a caption-image pair and a 'color-swapped' pair for controlled comparison.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Last updated 2024-02-06; freshness should be verified.
Provenance
- Source
- stanfordnlp
- Freshness
- Last updated 2024-02-06 22:23:20.