Sign in to view source links and access this dataset
Description
Magpie-Align released the Magpie Qwen2 Pro 200K Chinese dataset to advance the democratization of AI alignment data. The dataset contains 200,000 high-quality instruction-response pairs in Chinese, intended for aligning large language models. It was published on Hugging Face on August 22, 2024.
Use Cases
Instruction-tuning of Chinese LLMs based on the described high-quality instruction-response pairs.
Research into LLM alignment methodologies using open-source data.
Benchmarking model performance on Chinese instruction-following tasks.
Strengths
Contains 200,000 instruction-response pairs, indicating a substantial scale.
Described as 'high-quality' instruction data in the abstract.
Explicitly targets Chinese language alignment, filling a specific niche.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is known, but other specifics like file formats and license are unknown.
Data may reflect biases inherent to its unspecified collection methodology.
Provenance
Source
Magpie-Align
Collection Method
Methodology described in the associated technical report (arXiv:2406.08464).
Freshness
Last updated 2024-08-22 21:12:11.
License is unknown; users must verify permissions before commercial use.