Sign in to view source links and access this dataset
Description
A 2026 dataset contains 669 Chinese novels with raw text, annotations, and training data for narrative agent systems. It was created by mikuhhn1239 and includes 2.2 GB of data across raw books, processed chapters, and training files. The training data includes supervised fine-tuning examples for novel continuation.
Use Cases
Training supervised fine-tuning models for Chinese novel continuation based on the provided SFT data.
Analyzing narrative structures and patterns in Chinese fiction based on the annotated and processed text.
Developing agents capable of understanding and generating long-form Chinese narrative text.
Strengths
Includes 669 distinct Chinese novels as raw source material.
Provides 2.2 GB of total data across raw, processed, and training splits.
Offers structured training data for supervised fine-tuning, including 72K novel continuation examples.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Last updated 2026-06-17 10:53:26; freshness should be verified.
The source and collection methodology for the 669 novels are not specified, which may affect reproducibility.
Provenance
Source
huggingface
Collection Method
Collection and processing of Chinese novels for narrative understanding projects.
Freshness
Last updated 2026-06-17 10:53:26.
Geography
China (implied by Chinese-language content)
License is unknown; users should verify usage rights before download.