Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
Organizational structure data for the Ukrainian institution KZ 'DNZ No 3 VMR'. It likely contains identifiers, names, and staffing numbers for positions within structural subdivisions. The data was published on the States site of Ukraine and last updated on November 27, 2024.
OCRBench is a multimodal dataset containing images and text, accepted for publication in Science China Information Sciences. It was created by author echo840 and last updated in December 2024. Specific details on row count, column structure, and file formats are not provided in the input.
A dataset for post-OCR error correction in the Sanskrit language, generated using the RoundTripOCR technique. It was created by cfilt and uploaded to HuggingFace on December 8, 2024. The dataset includes train, test, and validation splits.
Pairs of original news articles and their AI-generated completions. The dataset was created by ilyasoulk and last updated on December 6, 2024. It likely contains articles from the CNN Daily Mail corpus, with AI completions generated using GPT-3.5 Turbo.
Between 10,000 and 100,000 image-text pairs comprise this 1% sample of the LaTeX_OCR collection for mathematical formula recognition. Released by unsloth in late 2024, the data provides paired visual representations and LaTeX source code formatted for Polars and Pandas.
A Web Feature Service (WFS) providing the urban development plan 'Ilben (Ganterhof)' for the city of Furtwangen in the Black Forest, Germany. The data is transformed according to the INSPIRE directive and based on an XPlanung dataset in version 5.0. The service was last updated on 2024-11-19.
A cross-modal dataset for Non-Fungible Token retrieval, developed by author shuxunoo. The dataset, referred to as NFT1000, was collected and organized by November 2023, and a related paper was accepted by ACM Multimedia in July 2024. The dataset page was last updated on November 14, 2024.
Roundtripocr Nepali is a text dataset for post-OCR error correction in the Nepali language, generated using the RoundTripOCR technique. The dataset, created by cfilt, includes training, validation, and test splits for machine learning tasks.
An ATOM feed provides the articles of association Sentence 15 Ganderkesee (1st amendment) in the INSPIRE PLU Version 4.0.1 data format. The Bundesamt für Kartographie und Geodäsie published this feed, which was last updated on November 15, 2024. The feed likely contains structured geospatial data related to land use planning regulations.
54,000 professional resumes aggregated from public sources such as Kaggle and GitHub. The resumes were originally in PDF format, converted to text via OCR, and structured using LLMs. The dataset was uploaded by Suriyaganesh on November 14, 2024.
This curated repository by coderonion provides a collection of public object detection and recognition datasets as of January 2025. It aggregates resources across specialized domains including thermal imaging, aerial imagery, and synthetic aperture radar (SAR) for computer vision tasks.
Ganderkesee municipality in Germany contains the Falkenburg compensation area, designated as area 169. The dataset is provided by the German Federal Agency for Cartography and Geodesy (BKG) and was last updated in November 2024. It is served via a Web Feature Service (WFS), indicating it is a vector-based geospatial dataset.
A Web Feature Service (WFS) dataset from the German Federal Agency for Cartography and Geodesy, last updated on November 14, 2024. It contains geospatial data for the origin point of Grüppenbührener Straße (Sentence 26) within the municipality of Ganderkesee. The specific features and attributes are not detailed in the available metadata.
ATOM-Feed provides geospatial data for the district of Bookholzberg, which originates from the municipality of Ganderkesee. The dataset is published by the German Federal Agency for Cartography and Geodesy (BKG). It was last updated on November 14, 2024.
ATOM-Feed provides structured data for the municipality of Ganderkesee, Germany. The data originates from the Bundesamt für Kartographie und Geodäsie and was last updated on November 14, 2024. The specific content and volume of records are not detailed in the available metadata.
Cadastral data for the municipality of Ganderkesee, Germany, served via a Web Feature Service (WFS). The dataset is provided by the Bundesamt für Kartographie und Geodäsie and was last updated on November 14, 2024. The specific number of features and attributes is unknown.
WMS) - 131 Schierbrok (origin) of the municipality of Ganderkesee. This geospatial dataset is provided by the Bundesamt für Kartographie und Geodäsie (BKG), Germany's federal agency for cartography and geodesy. It was last updated on 2024-11-14.
Ganderkesee municipality in Germany provides geospatial data for the Falkenburg area's first amendment. The dataset originates from the Bundesamt für Kartographie und Geodäsie and was last updated on November 14, 2024. It is served via a Web Feature Service (WFS) protocol.
5,179,510 aligned face images of 93,431 individuals form this dataset introduced for the Lightweight Face Recognition Challenge at ICCV 2019. All images were processed using RetinaFace for landmark detection and resized to 112x112 pixels. The dataset is hosted by gaunernst on Hugging Face and was last updated in October 2024.
Monthly aggregated water vapor column data derived from MODIS MCD19A2 v061 satellite imagery, covering the entire globe from 2000 to 2022. The dataset provides mean and standard deviation values at a 1 km spatial resolution, with gaps filled using the TMWM algorithm and smoothed using the Whittaker method. It is provided by CERN and the European Commission and was last updated in August 2024.