A collection of Chinese extractive Machine Reading Comprehension (MRC) datasets aggregated into a unified format. It provides context, question, and answer triplets designed for training models to identify specific text spans within Chinese documents.
Use Cases
- Train neural networks to predict answer spans within a given Chinese context passage.
- Benchmark Chinese language understanding (NLU) performance across various extractive MRC sources.
- Fine-tune pre-trained Chinese language models for precise information extraction from document contexts.
Strengths
- Consolidates multiple Chinese-language extractive reading comprehension datasets into one repository.
- Features context, question, and answer triplets required for span-based supervised learning.
- Focuses exclusively on the Chinese language domain for natural language processing tasks.