Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
Giving access to line-level handwritten English text images and transcriptions categorized for text recognition and writer identification. All image instances are resized to an uniform height of 128 pixels to support standardized input for neural network architectures.
A directory listing enterprises, institutions, and organizations managed by the State Geology and Mineral Resources authority of Ukraine. The data was last updated on March 21, 2024, and is published via the eu_open_data platform. The specific scope and scale of the entities listed are not detailed in the available metadata.
A dataset hosted on HuggingFace by amaye15, last updated on April 22, 2024. The title suggests it contains invoice images and the text extracted from them using Google's Optical Character Recognition technology. The specific scale, file formats, and column structure are not detailed in the available metadata.
ManiSkill is a unified benchmark for learning generalizable robotic manipulation skills powered by the SAPIEN simulator. It features 20 out-of-box task families with over 2000 diverse object models and more than 4 million demonstration frames. The dataset, created by haosulab and last updated in March 2024, is designed to enable fast visual input learning, supporting CNN-based policies that can collect samples at about 2000 FPS.
Span Propaganda Techniques is a dataset published on HuggingFace by anismahmahi on April 24, 2024. The title suggests it contains examples of propaganda techniques, likely in Spanish text. The dataset's specific content, size, and structure are not detailed in the provided metadata.
69,247 images categorized into classes including normal, sexual, violence, and weapon, collected for the Mod-RWKV Bluesky Hackathon project. The dataset totals 10.57 GB and is intended for image classification tasks in the NSFW domain.
A deduplicated collection of AI-generated portrait images from CivitAI, each described as a character by the Llava1.6-34b model. Each image is taller than it is wide and contains exactly one face, with bounding box coordinates provided. The dataset was created by kubernetes-bad and last updated on February 27, 2024.
The Julia Ward Howe collection includes manuscripts and letters by the American poet, abolitionist, and women's rights advocate. It was digitized as part of Project REVEAL (Read and View English & American Literature) and is hosted by the Texas Data Repository. The collection contains letters to individuals such as Elizabeth Fairchild and E. S. Goodwin.
Project REVEAL digitized a collection of materials from nineteenth-century author William Makepeace Thackeray. The collection includes manuscripts of writings, drawings, and letters, with items like a pocket sketchbook from 1840 and character illustrations for Vanity Fair. The dataset was contributed by Ballou, Jullianne and last updated on 2024 03 18.
The Unsloth-DPO dataset contains question and answer pairs intended for Direct Preference Optimization (DPO) training of large language models. It was created by NeuralNovel and is inspired by the orca_dpo_pairs dataset, with a specific focus on content related to Unsloth.ai. The dataset was last updated on March 4, 2024.
Blockface-level data from the 2015 TreesCount! street tree census organized by NYC Parks & Recreation and volunteers. It provides tree counts and collection status for city blocks, linking to detailed tree-level records. The dataset was last updated in February 2024.
Tabular data from the manuscript "Secondary organic aerosol from chlorine-initiated oxidation of isoprene" is contained in .txt files. The dataset supports research into atmospheric chemical mechanisms, specifically the formation of secondary organic aerosols from isoprene reacting with chlorine radicals. It is openly accessible via the Texas Data Repository and the PapersWithCode platform.
Published on March 8, 2024, this dataset contains information about the organizational structure of the CEC "CPMSD Ü2" in the Darnytskyi district of Kyiv. It likely includes a list of structural units and staff position names, sourced from the States site of Ukraine. The data is available in Excel formats.
The Jack London collection includes manuscripts and letters by the American author and journalist Jack London. Letters are directed to Edward Martin Moore, H. Ray Peck, and Vincent Starrett, and there are also three letters by his wife, Charmian Kittredge London, to poet John Myers O'Hara. This collection was digitized as part of Project REVEAL (Read and View English & American Literature) and is hosted by the Texas Data Repository.
A 1928 map of Rio de Janeiro, Brazil, georectified and paired with contemporary OpenStreetMap data. It contains latitude and longitude coordinates for locations like churches, museums, and government buildings. The dataset was created by Joshua Ortiz Baco and is hosted by the Texas Data Repository.
Scans of political posters produced during the Salvadoran Civil War in support of anti-government and international solidarity activities. The collection is housed at the Museo de la Palabra y la Imagen in San Salvador and includes metadata created for visualization. This dataset was last updated on March 18, 2024.
Water samples from sea, streams, ponds, and snow in Horseshoe Island, Marguerite Bay, Antarctica, contain microorganisms. The dataset aims to identify the origin of airborne microorganisms by linking them to these water sources. It was collected by organization SCIOPS and published in 2024.
2024 samples of microorganisms collected from soils, lake sediments, microbial mats, lichens, moss, and algae in Horseshoe Island, Margarite Bay, Antarctica. The dataset is intended to identify the origin of airborne microorganisms found in air samples. It was contributed by the organization SCIOPS and published via NASA EarthData in February 2024.
MICROAIRPOLAR 2 contains data on microorganisms identified in soils 400-600 meters east of the Juan Carlos I Spanish Base on Livingston Island. The dataset was created by SCIOPS to trace the origin of airborne microorganisms collected in air samples. It was last updated in February 2024.
How2Sign provides video data for American Sign Language (ASL) recognition. The dataset contains frontal and side-view RGB videos, with English translation text available for the frontal view. It was uploaded by aipieces and last updated on February 7, 2024.