Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
AgentsNet is a benchmark for multi-agent reasoning, designed to measure collaborative strategy formation, self-organization, and communication within a network topology. The dataset contains graph instances used in the 'AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs' paper, authored by disco-eth and last updated on 2025-07-17. It draws inspiration from classical problems in distributed systems and graph theory.
Individual-level data from the YOP experiment in Uganda includes geo-referenced coordinates and experimental outcomes. The dataset contains variables for treatment assignment, observed outcomes, and location keys linking to images. It was authored by cjerzak and last updated on July 19, 2025.
DeepShade is a multimodal dataset designed for supervised training of diffusion-based models to simulate sun-shade transitions. It captures realistic outdoor scenes and their corresponding shade conditions over time, based on spatial layout and temporal context. The dataset was introduced by DARL-ASU for an IJCAI 2025 submission and was last updated on July 17, 2025.
A Vietnamese legal text dataset from the VLSP legal evaluation benchmark, designed to assess language models' understanding and reasoning capabilities. The dataset contains three distinct evaluation tasks targeting different aspects of legal reasoning and comprehension. It was created by VLSP2025-LegalSML and last updated on July 15, 2025.
Ukrainian legal data from the States site of Ukraine, listing normative legal acts, acts of individual action, and draft normative legal acts. The dataset was last updated on July 15, 2025, and contains information as of 2023. The data is provided in Excel format.
RoboSense is a large-scale multimodal dataset containing over 133,000 synchronized sensor readings from cameras, LiDAR, and fisheye sensors. It provides 1.4 million annotated 3D bounding boxes and IDs to facilitate robot perception and navigation research in crowded, unstructured environments. The dataset was created by suhaisheng0527 and last updated in July 2025.
DATA.HRSA.GOV provides data on U.S. public health programs, including HRSA-funded health center grants, Health Professional Shortage Areas, and Ryan White HIV/AIDS services. The platform offers searchable data by topic area and geography through APIs and downloadable files. It is maintained by the U.S. Department of Health & Human Services and was last updated in July 2025.
100,000 high-quality samples designed for training multi-layer transparent image generation models. Each sample contains multiple layers including foreground objects and background scenes. The dataset was created by author 'artplus' and was last updated on 2025-06-24.
18 downstream tasks from the Nucleotide Transformer paper, providing a consistent genomics benchmark. The dataset, created by InstaDeepAI, is an updated version following peer review and was last updated on June 30, 2025.
The Carcinogenic Potency Database (CPDB) contains results from 6,153 chronic carcinogenesis bioassays conducted over 45 years. It includes quantitative data on tumor incidence, dose-response, and statistical significance, alongside experimental details like strain, sex, route of administration, and target organ. The database is a standardized resource compiled from the general literature and NCI/NTP technical reports.
Solidity_Code_Graph_Codearena is a dataset of 100 repositories from Codearena parsed into graphs of functions. The dataset includes different types of relationships between functions, such as parent contracts and various call types. It was created by author 'nothingisenough' and last updated on July 10, 2025.
Novovolynsk City Council's organizational structure, as defined by a council decision on 14 February 2025. The dataset is provided by the States site of Ukraine and was last updated on 15 July 2025. It likely contains information on departments, units, and reporting lines within the city administration.
Varash city territorial community in Ukraine provides data on local fairs, including their names, duration, and work schedules. The dataset likely contains information on rental costs, addresses, and the organizers of these events, as well as contracts concluded with them. It was published on the States site of Ukraine and last updated on July 15, 2025.
Diffusion.Studio.Dataset contains 16 high-quality AI-generated images created by author shodiBoy. Each image is accompanied by a detailed text prompt and the name of the diffusion model used for generation. The dataset was last updated on July 12, 2025.
Ukraine's State Forest Resources Agency provides information on its organizational structure. The dataset was last updated on 2025 07 16 and is published on the EU open data platform. The data is available in tabular formats, including Excel and CSV.
Harmonixset Bigvgan is a dataset hosted on Hugging Face by the author m-a-p. The title suggests it is related to the BigVGAN neural vocoder, likely containing audio data for speech or music synthesis. The dataset was last updated on August 21, 2025.
Persian_Arabic_TextLine_Image_Ocr_Small contains between 100,000 and 1,000,000 text-line images for Optical Character Recognition tasks, published by mohajesmaeili in June 2025. The collection covers multiple languages using the Arabic script, including Persian, Arabic, Urdu, and Pashto. It is delivered in Parquet format for compatibility with high-performance data processing libraries.
Information on the organizational structure of the information manager for Zhmerynka City Council and its executive bodies. The data is provided by the States site of Ukraine and was last updated on July 15, 2025. The set contains information defined by the organizational and administrative document 'Structure and staffing'.
Nearly 40,000 books in the Uzbek language form this text corpus, divided into 'original' and 'lat' branches for OCR and Latin versions. Author murodbek released the dataset to support research on low-resource languages. It was last updated on Hugging Face in June 2025.
A processed version of the Aida Calculus Math Handwriting Recognition Dataset, tailored for image-to-LaTeX modeling. The dataset comprises synthetic handwritten calculus expressions, with each image annotated by a ground truth LaTeX sequence. Author 'deepcopy' published this downsampled version on Hugging Face, last updated on June 20, 2025.