Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
Created by blinjrm and last updated in April 2024, this GitHub repository provides a framework for loading and transforming object detection datasets for computer vision. It functions as a utility for deep learning workflows, though specific record counts and file formats are not explicitly detailed in the metadata.
Raw data supporting research on microbial reduction of metal-organic frameworks for synergistic chromium removal. The dataset includes computed averages for reduction rates, biomass conversion, and chromium concentrations, as well as Powder X-ray diffraction spectra. It was authored by Keitz, Benjamin and last updated on March 18, 2024.
Data supporting research on the formation of oxidized organic compounds from chlorine-initiated oxidation of toluene. The dataset was published by author Hildebrandt Ruiz, Lea and harvested by the Texas Data Repository on 2024-03 18. It likely contains the figures and tables referenced in the associated scientific publication.
Data supporting a study on chlorine-initiated oxidation of n-alkanes under high NOx conditions. The dataset provides insights into secondary organic aerosol composition and volatility, as measured using a FIGAERO-CIMS instrument. It was published by author Hildebrandt Ruiz, Lea and last updated on March 18, 2024.
Mason County in Central Texas is the primary location for this collection of unpublished geologic field maps. The collection consists of scanned images (JPG) organized into 13 topographic 7.5-minute quadrangles. It was authored by Emilio Mutis-Duplat and last updated on March 18, -2024.
VegAnn contains 3,775 RGB images for semantic segmentation, each 512x512 pixels with corresponding binary masks. The dataset includes images of 26+ crop species captured under diverse outdoor conditions.
Mixed-methods data collected by Susan Bartels examine local community perceptions of interactions between UN peacekeepers and local women and girls in the Democratic Republic of Congo. The data were collected from six purposively selected UN bases in eastern DRC, representing urban and rural regions. The dataset was last updated on the Borealis Harvested Dataverse platform in February 2024.
French local urban plans (PLUs) digitized according to national CNIG requirements. The dataset informs building rights and can include regulated zoning, surface requirements, and surface information. It was published by the Bureau de Recherches Géologiques et Minières and last updated on February 13, 2024.
1911-1912 licensed grain elevators and warehouses from the Manitoba Grain Inspection Division. The data was generated via XSL transformation of OCR scans of the original list, with the transformation code available on GitHub. The dataset was last updated on the Borealis Harvested Dataverse platform in February 2024.
Over 15 columns detail New York City community-based organizations, including their mission, address, and volunteer programs. The dataset is maintained by data.cityofnewyork.us and was last updated in January 2024. It provides structured information on nonprofit entities and their civic activities across the city's boroughs.
New York City volunteer opportunities and organization listings published by data.cityofnewyork.us. The dataset includes columns for geographic coordinates, addresses, and opportunity details, with a last update timestamp of 2024-01-25 21:39:59. The specific number of records and license information are not provided.
Docreason25K is a dataset published on huggingface by mPLUG. The title suggests it contains data for document-based reasoning tasks. The dataset was last updated on 2024-03-26.
Brady Parrish authored a dataset describing the organization of data within a repository, hosted by the Texas Data Repository Harvested Dataverse. The dataset was last updated on March 18, 2024. Its specific content and structure are described in the provided metadata.
ILSVRC 2012, commonly known as 'ImageNet', organizes images according to the WordNet hierarchy. Each meaningful concept, or 'synset', is illustrated with an average of 1000 quality-controlled, human-annotated images. The dataset was uploaded to Hugging Face by 'timm' and last updated on January 7, 2024.
The Drinking Water Treatability Database (TDB), updated in March 2024, presents referenced information on the control of contaminants in drinking water. It aggregates data from thousands of literature sources for use by water utilities, regulators, and researchers.
Sinhala_synthetic_ocr-large is a dataset for optical character recognition tasks in the Sinhala language, created by Ransaka Ravihara and published on Hugging Face in 2024. The dataset's description suggests it contains synthetic images paired with text, likely generated for training machine learning models. Its specific size, format, and license details are not provided in the available metadata.
Data supporting a 2024 study on atmospheric pollution in New Delhi, India. The dataset, published by author Hildebrandt Ruiz, Lea, likely contains measurements and model outputs used to generate figures in the associated research paper. The specific variables and volume are unknown from the provided metadata.
Data supporting research on the composition of air pollution in Delhi, India. The dataset was published by Hildebrandt Ruiz, Lea and made available on the Texas Data Repository Harvested Dataverse in March 2024. The data likely contains measurements or analyses related to the sources of fine particulate organic matter.
Coco Minitrain provides a reduced subset of the Common Objects in Context (COCO) dataset curated by GitHub user giddyyupp for accelerated model prototyping. Updated in March 2024, it offers a smaller volume of images and annotations to facilitate faster training and debugging cycles than the full COCO release.
Luganda Dataset 2.0 is a language resource for Luganda, a Bantu language spoken primarily in Uganda. It was published on the Hugging Face platform by FarmerlineML and was last updated on March 13, 2024. The specific content, size, and structure of the dataset are not detailed in the available metadata.