Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Annotations from the VAST-27M dataset created for the 2024 paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset". The dataset was derived from work by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. It was uploaded to the Hugging Face platform by the user 'it-just-works' on September 10, 2024.
License is unknown; users must verify terms of use before application.