Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
5 million images are each paired with a short caption generated by the Qwen/Qwen2.5-VL-7B-Instruct model. The dataset was created by BLIP3o and last updated on Hugging Face in May 2025. It is intended for pretraining vision-language models.
The dataset uses WebDataset format within .tar archives, requiring specific loading support from the 🤗datasets library.