100 million metadata records for 99.2 million photos and 0.8 million videos sourced from Flickr under Creative Commons licenses. This subset follows the specific filtering and selection criteria established by OpenAI for the CLIP (Contrastive Language-Image Pre-training) model.
Use Cases
- Train contrastive vision-language models using the specific image-text associations defined by the OpenAI subset
- Conduct large-scale studies on multimedia metadata across 100 million records
- Benchmark data retrieval systems using the 100 million metadata records from the YFCC100M collection
Strengths
- Contains metadata for 99.2 million photos and 0.8 million videos
- Sourced exclusively from Flickr under various Creative Commons licenses
- Filtered specifically to match the subset used in the openai/CLIP research paper