Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,690 datasets
A directory of enterprises, institutions, and organizations within the Shostka City Territorial Community in Ukraine, last updated on 2025-01-09. The dataset likely contains identification codes, official websites, email addresses, telephone numbers, and physical addresses. It is sourced from the Handbook of the State Registration Department of Shostka City Council and published on the eu_open_data platform.
A large-scale dataset curated for training unified image generation models. The dataset is associated with the OmniGen project and was last updated in December 2024. It was created by author yzwang to address the lack of readily available data for multi-task image processing.
CC-OCR is a benchmark dataset for evaluating large multimodal models in optical character recognition and literacy tasks. The dataset is hosted in a TSV format for use with the VLMEvalKit evaluation framework. It was created by author wulipc and last updated in December 2024.
TIGER-Lab's VISTA-400K is a synthetic dataset designed to enhance video understanding in large multimodal models. The data is generated by the VISTA pipeline, which creates long-duration, high-resolution video instruction-following examples. The dataset was last updated on the Hugging Face platform in December 2024.
VLGuard consists of image-text pairs categorized into training and testing sets for safety fine-tuning of Vision Large Language Models. The dataset includes metadata files in JSON format and corresponding image archives to facilitate the development of safety baselines for multimodal AI systems.
48,384 high-resolution aerial images containing 1,041,714 annotated instances across 18 object categories. The dataset focuses on oriented object detection (OBB) and was introduced in the IEEE TPAMI journal for remote sensing research.
Image segmentation, object detection, and keypoint labels are generated through this web-based framework created by jsbroks and last updated in January 2025. It facilitates the creation of structured computer vision datasets in the COCO format from user-provided image files.
A directory listing enterprises, institutions, and organizations managed by a specific information manager in Ukraine, likely including identification codes. The dataset is published on the eu_open_data platform by the States site of Ukraine and was last updated in January 2025. Its exact scope and size are not detailed in the provided metadata.
Information on the organizational and personnel work of the Executive Committee of Starokostiantyniv City Council, socio-economic development and socio-political situation in Starokostiantyniv. The dataset is published on the eu_open_data platform by the States site of Ukraine and was last updated on 2025-01-14. The raw description indicates it is a certificate, and the primary file format is WORD DOC.
BioTrove comprises a large collection of images for biodiversity research, curated by BGLab. The dataset includes well-processed metadata with full taxa information and URLs pointing to image files. The dataset page was last updated on December 13, 2024.
TurkishLLaVA OCR Enhancement Dataset is a specialized collection of 100,000 books sourced entirely from Turkish materials. Created by ytu-ce-cosmos, it was last updated on December 17, 2024. The dataset is designed to improve the Turkish Optical Character Recognition capabilities of the Turkish-LLaVA-v0.1 model.
92,421 lines of Chinese game dialogue and narration extracted from Honkai Impact 3rd gameplay videos, covering story chapters from 'Main Story 1: Dusk, Maiden, Warship' to 'Main Story Part 2 Chapter 03 Interlude: A Sleepwalker's Pain'. The dataset was created by author mrzjy using an AI pipeline involving video frame extraction, OCR, and visual language model processing, and was last updated on December 17, 2024.
600 PubChem organic molecules structured into 5 multimodal question-and-answering tasks for chemistry reasoning. The content focuses on atom counting for carbon and hydrogen and molecular weight calculations, divided into validation and evaluation subsets.
Infinity-MM provides between 10 million and 100 million multimodal instruction samples, developed by the Beijing Academy of Artificial Intelligence (BAAI) in 2024. It utilizes a synthetic generation pipeline to create detailed image annotations and diverse question-answer pairs for bilingual model training in English and Chinese.
Lukas Blecher's pix2tex project provides data for converting images of mathematical equations into LaTeX code using a Vision Transformer (ViT) architecture. Updated in January 2025, the repository facilitates the im2markup task specifically for scientific and mathematical notation. The data pairs visual representations of formulas with their corresponding LaTeX markup strings.
Human-annotated images paired with referring expressions and specific coordinate points marking referenced locations. The collection includes a diverse range of spatial labels, specifically highlighting expressions with a frequency of 10 or more occurrences.
MuLAn provides over 44,000 multi-layer RGBA decompositions and 100,000 instance images derived from COCO and LAION sources. Released by the mulan-dataset team in late 2024, it serves as a photorealistic resource for instance-wise image manipulation and generation.
The dataset contains 8,500 rows, each representing the full text of a single Arabic book, extracted using the arabic-large-nougat model. It spans approximately 1.1 billion tokens, as calculated by the GPT-4 tokenizer. The dataset was authored by MohamedRashad and last updated on 2024-11-28.
A corpus of line-level text images paired with word-segmented Burmese text labels for optical character recognition. The dataset format pairs image file paths with text labels using a tab delimiter. The dataset was created by LULab and last updated on December 20, 2024.
The dataset lists recipients of the Staffing for Adequate Fire and Emergency Response (SAFER) grants, which provide funding to fire departments and volunteer organizations. It includes information such as organization name, state, program, award amount, and award date to enhance compliance with NFPA staffing standards.