Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
7 visual reasoning tasks comprising geometric primitives designed to test the fundamental perception of Vision-Language Models. The dataset includes categories such as line intersections, circle overlaps, and nested shapes where models frequently fail despite human-level performance.
Mind2Web Live provides approximately 1,000 records for web navigation and interaction tasks, released by iMeanAI in October 2024. The dataset focuses on text-based web environments and is formatted for integration with modern data libraries like Polars and Dask.
CharXiv is a diverse and challenging benchmark for chart understanding, fully curated by human experts. It includes 2,323 high-resolution charts manually sourced from arXiv preprints. The dataset was created by princeton-nlp and released in 2024.
Math-PUMA created this dataset to enhance mathematical reasoning through progressive upward multimodal alignment, as described in a 2024 arXiv preprint. The dataset contains English text focused on mathematics and reasoning tasks. Specific details on size, rows, and columns are not provided in the input.
A dataset associated with the 2024 paper 'Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models' presented at the 8th Annual Conference on Robot Learning. The dataset was uploaded by author 'holgerson' to Hugging Face on October 28, 2024. Its specific content, size, and structure are not detailed in the provided metadata.
Arc2Face contains approximately 21 million facial images representing 1 million unique identities at a resolution of 448x448 pixels. The dataset was created by upsampling half of the WebFace42M database using a blind face restoration network for the Arc2Face foundation model research.
The dataset integrates table images from the AFTdb (Arxiv Figure Table Database) curated by cmarkea. Each image is paired with LaTeX source code and linked to an average of ten questions and answers, half in English and half in French. Questions and answers were generated using Gemini 1.5 Pro and Claude 3.5 Sonnet, and the dataset was last updated on 2024-09-26.
MINT-1T is an open-source multimodal interleaved dataset designed for pretraining research. It contains one trillion text tokens and 3.4 billion images, representing a 10x scale-up from prior open-source collections and includes sources like PDFs and arXiv papers. The dataset was created by a team from the University of Washington and was last updated on the platform in September 2024.
RLAIF-V-Dataset is a large-scale multimodal feedback dataset created by unsloth. It provides 83,132 preference pairs, where instructions are collected from a diverse set of sources. The dataset was last updated on Hugging Face on 2024-09 26.
MINT-1T is an open-source multimodal interleaved dataset containing 1 trillion text tokens and 3.4 billion images, a tenfold increase in scale compared to prior open collections. It was created by a team from the University of Washington and includes data from previously untapped sources like PDFs and arXiv papers. The dataset was uploaded to the platform in September 2024.
A benchmark for synthetic data detection created by bczhou and released on November 5, 2024. The data supports the paper LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models. It is hosted on the Hugging Face platform.
Containing between 100,000 and 1,000,000 Midjourney V6 images re-captioned using the LLaVA-1.6 vision-language model. Released by brivangl in October 2024, the data serves as an augmented version of the CortexLM/midjourney-v6 repository for multimodal research.
MINT-1T contains 1 trillion text tokens and 3.4 billion images, scaling open-source multimodal data by a factor of ten. The dataset was created by a team from the University of Washington and released in 2024, incorporating sources like PDFs and arXiv papers to facilitate research in multimodal pretraining.
10,000+ hyper-detailed image descriptions and object-level annotations derived from the Open Images dataset. The data includes fine-grained attributes, spatial relationships, and dense scene narratives designed to improve vision-language model alignment.
MINT-1T is an open-source multimodal interleaved dataset containing 1 trillion text tokens and 3.4 billion images, representing a 10x scale-up from previous open-source collections. It was created by a team from the University of Washington and includes sources such as PDFs and ArXiv papers to facilitate multimodal pretraining research. The dataset was last updated on the platform in September 2024.
42,678 Vietnamese images paired with detailed text descriptions and visual question-answering pairs generated by GPT-4o. The dataset includes spatial metadata for objects and text, covering specific attributes such as font style, color, and size within a Vietnamese linguistic context.
MINT-1T is an open-source multimodal interleaved dataset containing 1 trillion text tokens and 3.4 billion images, a 10x scale-up from prior open-source collections. It includes previously untapped sources such as PDFs and ArXiv papers and is designed for multimodal pretraining research. The dataset was created by a team from the University of Washington and was last updated on the platform in September 2024.
MINT-1T is an open-source multimodal interleaved dataset containing one trillion text tokens and 3.4 billion images, representing a 10x scale-up from prior open-source collections. It was created by a team from the University of Washington to facilitate research in multimodal pretraining. The dataset was last updated on the platform in September 2024.
Ruozhiba, a popular forum on Baidu Tieba known for short, witty content, provides this raw collection of posts. The dataset was created by user 'kirp' and last updated in October 2024. It contains an unspecified number of posts scraped from the forum up to November 10, 2023.
Psych 101 Test is a text-based evaluation suite containing between 1,000 and 10,000 records for benchmarking human cognition models. Created by Marcel Binz in 2024, it serves as the private test set for the 'Centaur' foundation model research.