Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Approximately 1.5 million images from the Conceptual Captions 12M (CC12M) collection, encoded into feature representations using a VQGAN f16 1024 model. The dataset was created by johnowhitaker and last updated in April 2022. A script for further preprocessing is provided in the repository.
The dataset contains only the VQGAN-encoded features, not the original image files. Users must download the VQGAN model checkpoint and configuration files separately if they wish to run the encoding process themselves or decode features.