Loading...
Loading...
Predictive maintenance, quality control, supply chain, industrial sensors, factory automation
1,040 datasets
20,000 preference pairs for Direct Preference Optimization (DPO) training, sourced from four established Hugging Face datasets. The collection includes 10,000 Chinese and 10,000 English examples, each filtered by quality scores. Author llamafactory uploaded this multilingual mix on June 7, 2024.
Alpaca GPT4 Zh provides Chinese-language instruction-response pairs derived from GPT-4 outputs for model fine-tuning. The dataset was curated by llamafactory, who removed over 6,100 mistruncated examples to improve quality. It was last updated in June 2024.
A Chinese-language dataset derived from GlaiveAI's function calling v2 dataset, translated by GPT-3.5. It is hosted by llamafactory and was last updated on June 7, 2024. The dataset is intended for use in LLaMA Factory by specifying the dataset name.
A collection of 4,000 high-quality Chinese instruction-response pairs for training large language models. The dataset was created by author 'rqq' by translating questions from the 'Sao10K/Claude-3-Opus-Instruct-5K' dataset into Chinese and generating responses using the GLM-4 model. It was last updated on May 6, 2024.
A dataset of long conversation data, with portions translated using Google Translate. The data format is compatible with the LLaMA-Factory framework. The dataset was published by user 'zgce' on the Hugging Face platform and was last updated on June 14, 2024.
A 2024 dataset from dataverse describes a chemical synthesis approach for chiral phthalocyanines. The work by De la Torre, Gema details a method to enhance the reactivity of binaphthol-bridged phthalonitriles, leading to a specific ethynyl-containing Zn(II)Pc compound in remarkable yields. This amphiphilic compound was further modified and shown to form nanostructures in aqueous solutions.
Updated in May 2024 by Charmve, this repository aggregates what is described as the largest collection of open-source industrial surface defect datasets and their corresponding research papers. It provides a centralized index for computer vision tasks, specifically targeting image segmentation and anomaly detection within manufacturing domains such as PCB inspection.
A filtered subset of the bigcode/the-stack-dedup dataset containing source code files in 41 programming languages, including Ada, Assembly, C, C++, Python, Java, and Rust. It was created by MaLA-LM and last updated on Hugging Face in April 2024. The dataset's exact size, row count, and column structure are not specified in the provided metadata.
Davide Gotti published this dataset on 2024-05 05 as complementary material for a topology identification method. It includes input measurements, binary classification outputs, D+ elements, and PowerFactory source files for the modified New England and IEEE 118-bus test systems.
CNC reports on the implementation of the financial plan are provided by the States site of Ukraine. The dataset was last updated on 2024-04-12. The data is available for download in a ZIP file format.
London Assembly Constituencies in England are listed with their official names and codes as of 31st December 2023. The dataset is provided by the Government Digital Service and was last updated on the platform on 21st November 2023. It contains two text fields for each constituency.
Yhyu13 uploaded this dataset to Hugging Face on 2024-01 03. It contains processed JSON files for the ToolBench project, intended for training and evaluating language models on tool-use tasks. The data is derived from the original ToolBench repository and has been tailored for use with frameworks like LLaMA-Factory.
Freight Facts and Figures is a statistical collection from the Bureau of Transportation Statistics, last updated in December 2023. It contains charts and tables on U.S. freight movement, system performance, and economic contributions. The data covers safety, energy, and environmental impacts of the freight transportation industry.
26,707 shape programs derived from parametric CAD models were used to train the PlankAssembly model. The dataset was created by manycore-research and was last updated on September 21, 2023. It is structured as a directory containing individual JSON files for each model.
AUSRIVAS (Australian River Assessment System) data from 13 monitoring locations in the Australian Capital Territory, maintained by the ACT Government. The dataset contains O/E taxa scores, which compare observed to expected macroinvertebrate fauna to assess biological condition, recorded biannually in spring and autumn. The data was last updated on the platform on 2023-07-28.
Digital vector boundaries for the National Assembly for Wales Electoral Regions, representing the administrative geography as of December 2017. The Government Digital Service provides this data, which is generalised to a 20-meter resolution and clipped to the coastline. Last updated on the platform in June 2023, the dataset contains both Ordnance Survey and ONS intellectual property rights.
December 2018 digital vector boundaries for London Assembly Constituencies in England. The data contains both Ordnance Survey and Office for National Statistics intellectual property rights and is available in multiple formats including GeoPackage, GeoJSON, and ESRI File Geodatabase. The dataset was last updated on the platform in March 2023.
2.6GB of text comprises a 0.1% subset of the-stack dataset, containing 10,000 random samples for each of 30 programming languages. The dataset was created by bigcode and released in May 2023.
This large-scale image dataset features wood surface defects captured at a consistent resolution of 2800 x 1024 pixels for industrial quality control. It provides dual-layer annotations including bounding boxes in YOLO format and semantic maps for precise defect localization.
1,016 U.S. commodities defined by the 2017 NAICS system have greenhouse gas emission factors based on 2019 data. The dataset provides three factor types per commodity: Supply Chain Emissions without Margins (SEF), Margins (MEF), and their combined total (SEF+MEF). Factors are expressed as kg CO2e per 2021 USD or as kg of individual GHGs per dollar.