Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
HuggingFaceFV released FineVideo in late 2024, a multimodal collection containing between 10,000 and 100,000 video-text records. The dataset is specifically formatted for English-language tasks such as Visual Question Answering and video-to-text generation.
The dataset is distributed in Parquet format and is compatible with Dask and Polars for distributed processing; users should check the full terms of use on the Hugging Face dataset page.