Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
A partial dataset from the MAmmoTH2 project, containing instruction data primarily sourced from web forums like StackExchange. The data is described as very high-quality and is intended to boost large language model performance through instruction tuning. The dataset was authored by TIGER-Lab and last updated on Hugging Face on October 27, 2024.
Harmonized Landsat and Sentinel-2 multispectral reflectance imagery and MERRA-2 observations centered around eddy covariance flux towers. The dataset includes corresponding Gross Primary Productivity data and is intended to fine-tune geospatial foundation models for GPP regression. It was created by ibm-nasa-geospatial and last updated on October 25, 2024.
lmarena-ai's PPE-GPQA-Best-of-K dataset contains a correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from the GPQA dataset, and the collection is intended for benchmarking and evaluation, not for training. The dataset was last updated on October 22, 2024.
PKU-SafeRLHF is a dataset for AI safety research, particularly for reducing harmful outputs from language models. It was created by the PKU-Alignment Team and was last updated in October 2024. The dataset includes single-dimension preference data, question-answer pairs, and prompts.
A benchmark for evaluating multimodal embedding models, covering 4 meta tasks and 36 datasets. The dataset was created by TIGER-Lab and published in the paper 'VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks'. It was last updated on Hugging Face on October 28, 2024.
Persian VQA is a dataset for Visual Question Answering tasks in the Persian language. It was published by AUT-NLP on the Hugging Face platform and was last updated on December 12, 2024. The dataset's specific content, scale, and structure are not detailed in the available metadata.
Between 1 million and 10 million Japanese-translated vision-language records comprise this collection created by turing-motors in 2024. It adapts the 50-dataset Cauldron collection used for Idefics2 fine-tuning into Japanese using the DeepL API, specifically targeting visual question answering tasks.
Umrb EncyclopediaVQA is a dataset hosted on HuggingFace by author izhx, last updated on December 3, 2024. The title suggests it likely contains visual question-answering tasks based on encyclopedia content. Specifics regarding size, columns, and data format are currently unknown.
OSWorld provides task examples, retrieval documents, and virtual machine snapshots for benchmarking multimodal agents performing open-ended tasks in real computer environments. The dataset was created by xlangai and last updated in October 2024. It supports evaluation on both x86 and arm64 machine architectures using VMware or VirtualBox.
LAV-DF is a dataset for content-driven audio-visual deepfake detection and temporal forgery localization. The dataset was created by ControlNet for research presented at DICTA and submitted to CVIU. It was last updated on Hugging Face on October 16, 2024.
178,510 caption entries and 960,792 open-ended question-answer pairs were compiled by lmms-lab for training the LLaVA-Video model. This multimodal dataset aggregates video-language data from five primary sources. The dataset card was last updated in October 2024.
Visual Haystacks (VHs) is a benchmark dataset designed to evaluate Large Multimodal Models' capability to handle long-context visual information. It is described as the first vision-centric Needle-In-A-Haystack benchmark. The dataset was created by tsunghanwu and was last updated on Hugging Face on October 16, 2024.
25,000 multimodal examples likely containing images paired with text instructions and chain-of-thought reasoning. The dataset was created by author 'tomkld' and last updated on Hugging Face on December 10, 2024. Its columns suggest it contains image and text data for training vision-language models.
VLFeedback contains 80,000 multi-modal instructions and 320,000 model responses annotated by GPT-4V for vision-language preference learning. Developed by MMInstruction in late 2023, the dataset aggregates instructions from diverse sources to evaluate a pool of 12 different Large Vision-Language Models (LVLMs).
3,763 web-collected videos with subtitles and multiple-choice questions comprise this long-context multimodal benchmark. Created for NeurIPS 2024, it evaluates large multimodal models on video-language interleaved inputs with durations reaching up to one hour.
A dataset titled 'Gemex Vqa' was published on the Hugging Face platform by BoKelvin on December 1, 2024. The dataset's title suggests it is related to visual question answering, a multimodal AI task. Specific details on size, format, and content are not provided in the available metadata.
A fine-tuning dataset for the RDT-1B diffusion foundation model, as described in the paper 'RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation'. The dataset was created by robotics-diffusion-transformer and last updated on Hugging Face on 2024-10-13. The associated research paper was published on arXiv in October 2024.
DenseFusion-1M provides 1 million image-text pairs for multi-modal perception, released by the Beijing Academy of Artificial Intelligence (BAAI) in 2024. The dataset uses a Perceptual Fusion approach to combine outputs from specialized vision experts and GPT-4V into detailed descriptions.
S3E provides multimodal sensor data for multi-robot collaborative Simultaneous Localization and Mapping (SLAM), developed by DapengFeng and published in IEEE Robotics and Automation Letters (RA-L). The dataset facilitates research into multi-agent systems by providing synchronized data streams from multiple robotic platforms.
A dataset containing paired images and captions designed for fine-grained multimodal concept understanding. Each data sample contains two images and two corresponding captions that differ only in one object, the color of an object, or the size of an object. The dataset was created by author 'phiyodr' and was last updated on October 2, 2024.