CommonCatalog CC-BY-SA is a collection of high-resolution Creative Commons images sourced from Yahoo Flickr users in 2014. It contains approximately 100 million images with synthetic captions. The dataset was created by the common-canvas organization.
Use Cases
- Generate synthetic captions for high-resolution images using the provided caption data.
- Train image generation models on a large-scale dataset of up to 4k resolution images.
- Analyze the distribution and characteristics of Creative Commons licensed images from Flickr.
Strengths
- Contains approximately 100 million images.
- Images have up to 4k resolution.
Limitations
- Dataset was collected in 2014, making it potentially outdated for current trends.
- Captions are synthetic, not human-written, which may limit quality for certain tasks.
- Specific license composition and image metadata details are unknown.
Provenance
- Source
- Yahoo Flickr users.
- Collection Method
- Collection of Creative Commons images from Flickr.
- Time Range
- 2014.
- Freshness
- null
- Geography
- null