Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,701 datasets
A dataset for Word Sense Linking (WSL), requiring systems to disambiguate text spans to senses from a reference inventory. Annotations are provided as sense keys from WordNet, a large lexical database of English, and all terms have been fully manually annotated. The dataset sentences are taken from the ALL split of the Raganato framework and was authored by Babelscape.
65,000 images across 100 classes from the ImageNet-1k dataset, curated by the timm library author in 2024. It provides a high-fidelity subset for classification tasks by maintaining original image dimensions instead of downsampling to low resolutions.
90,000 multi-view RGB-D frames capturing 3D human-object interactions across 8 subjects and 20 object categories. The dataset includes temporal sequences of SMPL human body meshes and 3D object poses with precise contact annotations.
MedMNIST consists of 18 standardized subsets for 2D and 3D biomedical image classification, maintained by the MedMNIST organization and updated as of January 2025. It provides a large-scale benchmark for medical image analysis across multiple modalities including 2D and 3D data.
A curated collection of song lyrics from modern pop artists, sourced from Genius. The dataset contains lyrics in multiple languages, including English and Spanish, and was curated by Scott Griffin. It was last updated in December 2024.
Information about the organizational structure of the Main Directorate of the State Tax Service in the Rivne region of Ukraine. The dataset was published by the States site of Ukraine and was last updated on December 9, 2024. Available file formats include CSV.
Shot2Story provides a benchmark for multi-shot video understanding featuring video-level summaries and shot-level captions. The dataset facilitates the development of models that can synthesize information across temporal segments to form a cohesive narrative.
A geospatial dataset from the Bundesamt für Kartographie und Geodäsie, last updated on 2024-12-11. It appears to contain addendums related to the georeferencing of municipal servers and likely addresses the association of communities. The data is served via the WMS format.
8 images from the COCO train 2017 set, split into 4 for training and 4 for validation. Ultralytics created this dataset, which was last updated on the platform in November 2024. It is designed for testing and debugging object detection models.
The Directory of enterprises, institutions and organisations of Cultural Institution "Vinnytsia city centralised library system" contains information about libraries within the CBS VM. It was published on the States site of Ukraine and last updated on November 27, 2024. The dataset is available in CSV formats.
mOSCAR is a multilingual web-crawled text corpus developed by the oscar-corpus organization. The dataset includes additional filtering steps to remove toxic content, a complete Spanish split, and face detection in images to blur them. The dataset page was last updated on 2024-11-23.
Ultralytics Tiger-pose is a specialized dataset for pose estimation of tigers, containing 210 training and 53 validation images. Each image is annotated with 12 keypoints per tiger. The dataset was created by Ultralytics and was last updated on November 13, 2024.
Godseye Violence Detection Dataset is a composite resource assembled by valiantlynxz, merging three existing video datasets: RWF-2000, Hockey Violence Dataset, and Airtlab Violence Dataset. It contains videos depicting real-world fights, non-fight scenarios, violent hockey interactions, and violent incidents in various settings. The dataset was last updated on HuggingFace in November 2024.
A training set distilled from GPT-4o for the paper 'Automatic Evaluation for Text-to-Image Generation: Fine-grained Framework, Distilled Evaluation Model and Meta-Evaluation Benchmark'. The dataset was created by DataHammer and last updated on December 16, 2024.
A Web Map Service (WMS) layer provides geospatial data for the development plan 163 Ganderkesee (origin). The data conforms to the INSPIRE PLU Version 4.0.1 standard for land use planning. The dataset was published by Bundesamt für Kartographie und Geodäsie and was last updated on 2024-11-27.
A Web Map Service (WMS) providing geospatial data for the development plan 057 Ganderkesee (origin). The data is formatted according to the INSPIRE PLU (Planning Land Use) version 4.0.1 specification. It was last updated on 2024-11-27 and is provided by the Bundesamt für Kartographie und Geodäsie.
A Web Feature Service (WFS) providing the municipal land use plan (FNP) for Ganderkesee, Germany, in the XPlanGML 5.4 data format. The data is provided by the Bundesamt für Kartographie und Geodäsie and was last updated on 2024-11-27. It likely contains geospatial features and attributes related to zoning and land use regulations.
Ganderkesee municipality in Germany provides land use planning data via an ATOM feed. The dataset is published by the Bundesamt für Kartographie und Geodäsie and was last updated on November 27, 2024. It conforms to the INSPIRE PLU (Land Use) data specification for standardized geospatial information.
A Web Feature Service (WFS) provides geospatial data for the development plan 084 in Ganderkesee, Germany. The data conforms to the INSPIRE PLU (Planned Land Use) version 4.0.1 specification. It was published by the Bundesamt für Kartographie und Geodäsie and last updated on November 27, 2024.
A Web Map Service (WMS) layer provides geospatial data for development plan 103 in Ganderkesee, Germany. The data conforms to the INSPIRE PLU (Land Use) data specification version 4.0.1, a European standard for spatial data interoperability. It was published by the Bundesamt für Kartographie und Geodäsie and last updated on November 27, 2024.