Sign in to view source links and access this dataset
Description
An industry-specific instruction dataset created by BAAI to address gaps in domain knowledge for AI models. It contains Chinese-English paired data across 12 sectors including automobiles, aerospace, finance, and healthcare. The dataset was last updated on November 12, 2024.
Use Cases
Fine-tuning large language models for industry-specific question-answering based on the described sector knowledge.
Training multilingual AI assistants for professional domains like law, finance, or healthcare using the bilingual instruction pairs.
Benchmarking model performance on domain-specific reasoning tasks across different industries.
Augmenting pre-training corpora with high-value industry knowledge extracted from the BAAI/IndustryCorpus2 source.
Strengths
Covers 12 distinct high-value industry sectors with bilingual (Chinese-English) data.
Derived from the BAAI/IndustryCorpus2 pre-training corpus, suggesting a foundation of curated text.
Provides data quality distribution curves and word cloud visualizations per industry for initial assessment.
Limitations
Description metadata is limited; actual data quality requires manual inspection after download.
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Provenance
Source
BAAI (Beijing Academy of Artificial Intelligence)
Collection Method
Likely extracted and curated from the BAAI/IndustryCorpus2 pre-training dataset.
Freshness
Last updated 2024-11-12 08:25:33; freshness should be verified.
License is unknown; terms of use must be verified before application.