Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
An image dataset of Javanese script characters, created by akbarrizqi167 and last updated on 2025-05-26. The description suggests it is intended for training an Optical Character Recognition model for Javanese script.
The dataset is a directory of enterprises, institutions, and organizations managed by the Black Sea City Council. It contains three resources: a main organization directory, a directory of structural subdivisions, and a directory of officials and employees. The data was published on the States site of Ukraine and last updated on 2025-04-25.
OctoNet is a multi-modal dataset of human activity recordings from multiple sensor modalities, including inertial measurement units (IMU), motion capture, and mmWave/Radar data. The dataset is hosted by the author 'hku-aiot' on Hugging Face and was last updated on 2025-05-16. It is intended for research in activity recognition, pose estimation, and multi-modal data fusion.
Three MicroED datasets containing 106 crystals total: 61 for Lysozyme, 20 for MTH1-8DG collected with high fluence, and 25 for MTH1-8DG collected with low fluence. Data was collected using a Titan Krios G2/G3i microscope with a CetaD CMOS detector at 77 Kelvin, processed via AutoLEI with XDS. The dataset is hosted on the eu_open_data platform under a CC-BY-4.0 license and was last updated on April 11, 2025.
ALLVB is a benchmark dataset for evaluating multimodal large language models on long video understanding tasks. The dataset was created by the ALLVB organization and was last updated on the Hugging Face platform in May 2025. It is licensed for research purposes only under CC-BY-NC-SA-4.0.
The State site of Ukraine provides this dataset on the organizational structure of the Main Directorate of the State Tax Service in the Sumy region. It was last updated on 2025 -05-14. The data is available in CSV format.
Comprising 102,476 face images from 1,507 Chinese individuals (762 males, 745 females). Each subject contributed 62 multi-pose and 6 multi-expression images, captured under varying angles, poses, and lighting conditions.
A directory of the information manager and subordinate organizations within the Main Department of the Pension Fund of Ukraine in the Volyn region. The dataset includes identification codes from the Unified State Register, official websites, email addresses, telephone numbers, and location details. It was published by the States site of Ukraine and last updated on May 15, 2025.
Aggregating a sample of 1,930 people, with 200 images collected per subject. The images feature 4 lighting conditions, 10 occlusion cases, and 5 face poses. It is designed for computer vision tasks like occluded face detection.
105,941 images of text in natural scenes. The images cover 12 languages, including 6 Asian and 6 European languages, with line-level quadrilateral bounding box and transcription annotations.
Featuring 5,030 face images from 7 images per person, featuring individuals wearing gauze masks. It includes multiple mask types, age groups, lighting conditions, and scenes. The data is intended for occluded face detection and recognition tasks.
From 1778 to 1894, this dataset contains 752 digitized polygon features representing land cessions and reservations from 67 historical maps compiled by Charles C. Royce. It was created by the U.S. Forest Service from the 1896-1897 Bureau of American Ethnology report, georeferenced and digitized with attributes linking to treaty dates and tribal names.
Lutsk City Council provides a directory of enterprises, institutions, and organizations within its jurisdiction. The data likely includes identification codes from the Ukrainian Unified State Register of Legal Entities, Individual Entrepreneurs and Public Organizations. It was published on the eu_open_data platform and last updated on 2025-05-21.
The dataset describes the organizational structure of the Main Directorate of the State Tax Service in the Khmelnytsky region of Ukraine. It was last updated on May 13, 2025, and is provided by the States site of Ukraine. The data is available in document formats such as .DOC and .XLS.
IndoorOutdoorNet-20K is a labeled image dataset containing 19,998 images for binary classification between indoor and outdoor scenes. It was created by prithivMLmods and uploaded to Hugging Face Datasets on 2025-04-22. The dataset is intended for tasks like scene understanding, transfer learning, and model benchmarking.
A dataset of startup pitch decks, organized into folders and files. The dataset includes evaluation files in JSON format. It was created by skyforclouds and last updated on June 6, 2025.
An annual pay gap report from the London Borough of Camden, published by the Government Digital Service. The analysis examines pay disparities by gender, ethnic origin, and disability status, following UK legislation mandating gender pay reporting from April 2017. The dataset was last updated on May 6, 2025.
FEMA Organization Codes is a reference list based on the USDA TMGT Table 005, detailing the agency's organizational structure. The data is maintained by the Department of Homeland Security and was last updated in June 2025. FEMA uses up to eight hierarchical layers to segment its programs and personnel.
Over 1,800 federal, state, local, tribal, and territorial agencies are designated to alert the public of disasters and threats using the Integrated Public Alert & Warning System (IPAWS). The dataset lists these Alerting Authorities, which are managed by the Department of Homeland Security and were last updated in June 2025.
14,492 annotated underwater images form a large-scale dataset for training AI models in marine biodiversity research. The dataset is designed for multi-label image classification tasks and was created by lombardata. It was last updated on the platform in April 2025.