Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
A list of childcare organizations and facilities includes company addresses and contact details. The dataset is updated yearly and originates from data.montgomerycountymd.gov. It was last updated on August 12, 2022.
The organizational structure, subdivisions, and service tariffs of the Municipal Enterprise 'Emergency and Dispatcher Service' of Kropyvnytskyi City Council. The dataset is published on the States site of Ukraine and was last updated on September 2, 2022. It is available in CSV, Excel XLSX, and ODS file formats.
Encompassing 2,872 machine-translated news articles and corresponding summaries in Danish. It is derived from the CNN/DailyMail dataset and is intended for Danish text summarization tasks.
Information about fairs, their organizers, and contracts concluded with the organizers of such fairs. The data originates from the States site of Ukraine and was last updated on 2022-08-30 11:13:49.077304. The dataset likely contains details such as fair term, place, number of places, and cost of seats.
Tiny ImageNet contains 100,000 images across 200 classes, with 500 training, 50 validation, and 50 test images per class. The images are downsized to 64x64 colored images, and class labels are in English. It was created by zh-plus and last updated in July 2022.
70 health tracking indicators and sub-indicators measure progress across five priority areas of the Prevention Agenda 2019-2024. The trend dataset provides historical and recent county-level data for New York State, compiled by health.data.ny.gov. Data was last updated in June 2022.
70 health tracking indicators provide county-level metrics for New York State's 2019-2024 health improvement plan. Data includes the most recent values, historical trends where available, and 2024 state objectives for each indicator. The dataset was published by health.data.ny.gov and last updated in June 2022.
The LLD dataset contains over 600,000 logos crawled from the web for exploring machine learning in creative design tasks. It was created by author diwank and used to train Generative Adversarial Networks for logo synthesis and manipulation. The dataset was last updated in August 2022.
This repository contains small synthetic data for the MNIST, SVHN, and CIFAR-10 image datasets. The data consists of images and corresponding labels, with sizes ranging from 1, 10, to 50 images per class (IPC). It was created by ICML2022 for research on dataset condensation.
Aggregating 4,988,099 full-text pages from 28,909 OCR-processed historical works. The texts originate from the digital collections of the Berlin State Library, covering a time period from 1470 to 1945. The language of each page has been automatically identified.
Version 8.0 of this NIST database contains 225 additional ligands and features revisions to 30% of previous data entries. It provides critically selected stability constants for interactions in aqueous systems involving organic and inorganic ligands, protons, and metal ions. The database is maintained by the National Institute of Standards and Technology and was last updated in July 2022.
Encompassing 17,868 books from the original BookCorpus, with text examples rendered as images at a resolution of 16x8464 pixels. It was created by Team-PIXEL and last updated in August 2022.
Triplets of Quora questions for training semantic equivalence models. It is structured with anchor, positive, and negative question examples. The dataset size is categorized as between 100,000 and 1,000,000 rows.
The Oxford-IIIT Pet Dataset provides images of cats and dogs for computer vision tasks. It was uploaded to Hugging Face by user 'enterprise-explorers' in August 2022, containing only images and labels from the original dataset. Segmentation annotations from the original source were not included in this upload.
Featuring five textual captions for each image from the COCO dataset, designed for sentence similarity tasks. It was released by the embedding-data organization on Hugging Face in August 2022.
Recommended values for chemical thermodynamic properties of inorganic and C1-C2 organic substances, compiled by the National Bureau of Standards. The dataset provides enthalpy of formation, Gibbs energy of formation, entropy, and heat capacity at 298.15 K, with values for gaseous, liquid, crystalline, and aqueous states. It represents a collective edition of the serially issued 'Selected Values of Chemical Thermodynamic Properties' from 1965 to 1981.
A collection of the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution. Each rendered example contains a subset of one full Wikipedia article.
New York City's OneNYC 2050 strategic plan indicators track progress toward eight city-wide goals, including Vibrant Democracy and Livable Climate. The Mayorβs Office of Operations collects data from relevant agencies, with the dataset last updated in May 2022. It includes baseline values, targets, and annual measurements for each indicator.
A National Institute of Standards and Technology tool estimates costs for U.S. manufacturers, categorized by standardized systems like NAICS and SOC. It supports gauging potential returns on research projects aimed at cost reductions in areas like labor, materials, and energy. The dataset is designed to answer specific cost questions for manufacturing industry research.
A collection of images of WeChat QR codes, likely for optical character recognition (OCR) tasks. It was created by breezedeus and last updated in September 2022. The specific number of images, features, and data structure are unknown.