Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Llava recaptioned COCO2014 ValSet. This dataset contains 30,000 images from the COCO 2014 validation set, each paired with its original caption and a new caption generated by the LLaVA vision-language model. It was created by UCSC-VLAA for evaluating text-to-image generation models, as detailed in the associated research paper.
License is unknown and must be verified before use, as it may inherit terms from the original COCO dataset.