Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
A dataset describing livestock, offspring, wool, milk, and cheese production in Herrera de Alcántara, likely focusing on traditional agricultural practices. The dataset was coordinated by Álvarez Pérez, Xosé Afonso and harvested by e-cienciaDatos from a Dataverse source, with metadata last updated on May 5, -2024. The description suggests a focus on the decline of cheese-making and mentions slaughtering practices.
A multilingual description of historical livestock uses in La Alamedilla, authored by Álvarez Pérez, Xosé Afonso. The record was last updated on May 5, 2024, and is hosted by the e-cienciaDatos Harvested Dataverse platform. It details cattle for plowing and carting, sheep and goats for milk and cheese, and pigs for slaughter.
A dataset from the e-cienciaDatos Harvested Dataverse, coordinated by Álvarez Pérez, Xosé Afonso, and last updated on May 5, 2024. The description discusses cattle and wool, noting distinctions in terminology between Spanish and Portuguese languages, such as the words 'toro' and 'boi'. The data appears to focus on linguistic and cultural aspects of animal husbandry.
An interview with an informant in Laranjeiras, covering cattle farming practices. The dataset likely contains discussions on shearing, milking, cheese preparation, pig slaughter, and typical meals made from the products. It was coordinated by Álvarez Pérez, Xosé Afonso and harvested by e-cienciaDatos, with a metadata update on 2024-05-05.
A dataset coordinated by Álvarez Pérez, Xosé Afonso, last updated on May 5, 2024. It describes types of livestock (cows, sheep, goats, pigs), fairs, milk and derived products, and slaughtering practices. The data appears to be harvested from a Dataverse repository by e-cienciaDatos.
Informante 1 is a dataset coordinated by Álvarez Pérez, Xosé Afonso, harvested from the e-cienciaDatos Dataverse and last updated on May 5, 2024. The description indicates it covers cattle, sheep, goats, and pigs, along with related economic activities and products. The data likely contains information on livestock fairs, wool, dairy products, and slaughtering practices.
Charmve curated this repository of scene text detection and recognition resources, last updated in May 2024. It organizes research papers, source code, and links to various benchmark datasets specifically for OCR tasks.
2 categories of objects, mirrors and glass, are featured for detection tasks in real-world scenes. The content focuses on the visual identification of reflective and transparent surfaces in unconstrained environments.
Sentinel-1 satellite data reveals surface deformation across an 80,000 square kilometer oil-producing region between 2014 and 2019. The dataset provides cumulative line-of-sight deformation maps at 120-meter pixel spacing for three time intervals, with an uncertainty of approximately 1 centimeter. It includes GPS station locations and digital elevation models for validation and further geospatial analysis.
A dataset harvested on 2024-05-05, coordinated by Álvarez Pérez, Xosé Afonso, and hosted by e-cienciaDatos. It likely contains information on livestock, including cows, sheep, goats, and pigs, and related traditional practices such as wool production and slaughtering. The data appears to be focused on the region of Hermisende in Zamora, Spain.
Muñoz-Organero's dataset contains simulated physiological data for Type 1 Diabetes Mellitus (T1DM) patients. The data was generated using the AIDA diabetes simulator to produce 10 days of records for validating a deep physiological model for blood glucose prediction. The dataset was harvested by e-cienciaDatos and last updated on May 5, 2024.
Amparo López (La Alamedilla) dataset, part 2, focuses on the slaughter process and derived products. The dataset was coordinated by Álvarez Pérez, Xosé Afonso and harvested by e-cienciaDatos from the Dataverse platform, with metadata last updated on 2024-05-05. It likely contains information on traditional butchery and sausage-making practices.
The Federico-Tena World Population Historical Database provides historical demographic data for Afghanistan. This dataset was developed by Giovanni Federico of New York University Abu Dhabi and Antonio Tena Junguito of Universidad Carlos III de Madrid. The record was last updated on May 5, 2024.
The Industry Documents Library (IDL) is a multimodal dataset containing 19 million pages of documents filtered from the UCSF library. Each document is available as a PDF, a TIFF image, a JSON file with Textract OCR annotations, and an older .ocr file. The dataset was created by pixparse and last updated on March 29, 2024.
8,000 images annotated with 157,000 bounding boxes across five object classes, collected from 18 wide-angle fisheye cameras. The dataset was created by a team of researchers including Munkhjargal Gochoo, Munkh-Erdene Otgonbold, and others. It was last updated on the Hugging Face platform in April 2024.
2024 samples of airborne microorganisms collected from multiple locations across the Antarctic Peninsula and Drake Passage. The MICROAIRPOLAR research group gathered the data using specialized air collectors designed for polar conditions. The dataset was published by SCIOPS in March 2024.
Approximately 45,500 images of Russian car license plates are provided for training neural networks to recognize plate text. The dataset is derived from the Nomeroff Net project and contains a single plate type. It was created by AY000554 and last updated in April 2024.
Information about the organizational structure of the Department of Capital Construction of Zhytomyr City Council. The dataset was published on the States site of Ukraine and last updated on 2024-04-12 14:09:35.996383. The data is provided in an EXCEL XLSX file format.
PDFA Eng Wds contains between 1,000 and 10,000 English document pages filtered from the SafeDocs CC-MAIN-2021-31-PDF-UNTRUNCATED corpus by pixparse. Released in 2024, it provides a curated subset specifically formatted for vision-language models with pre-processed bounding box annotations. The data serves as a machine-learning-ready version of the original SafeDocs document analysis collection.
A text dataset for analyzing hallucinations in generated language. Each sample contains a ground truth reference document followed by a generated text segment for comparison. The dataset was created by clam004 and last updated on 2024-04-10.