1,440,191 images curated from the YFCC-100M collection that contain no human subjects or identifiable persons. This dataset enables self-supervised pretraining for computer vision models while addressing privacy and ethical concerns related to personal data.
Use Cases
- Train self-supervised models like MoCo or DINO using the 1.4 million images to learn visual representations without human-centric bias
- Evaluate the necessity of human-related features in general-purpose image encoders by comparing downstream task performance against models trained on ImageNet
- Conduct privacy-focused computer vision research using a dataset that bypasses GDPR and facial recognition regulations
Strengths
- 1,440,191 images filtered from the YFCC-100M source
- Complete absence of human subjects, faces, or identifiable body parts
- Sourced from the Yahoo Flickr Creative Commons 100 Million (YFCC-100M) dataset
- Optimized for self-supervised learning (SSL) pretraining tasks