8,563,753 filtered image-text pairs from the LAION-COCO dataset, curated for high visual quality. This subset was created by applying rules for image size, aesthetic score, and watermark probability. The dataset was uploaded by guangyil on Hugging Face in November 2023.
Use Cases
- Training image generation models based on high aesthetic scores.
- Filtering datasets for downstream tasks based on watermark probability.
- Benchmarking aesthetic prediction models using the provided scores.
- Creating clean image-text datasets for research based on size and quality thresholds.
Strengths
- Contains 8,563,753 data instances, providing a substantial sample.
- Each instance includes explicit aesthetic and watermark probability scores.
- Images were filtered by specific rules (size > 384x384, aesthetic > 4.75, watermark probability < 0.5).
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Last updated 2023-11-15 10:34:11; freshness should be verified.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- LAION-COCO dataset
- Collection Method
- Filtered subset of LAION-COCO using text and image rules.
- Freshness
- 2023-11-15 10:34:11