A multimodal dataset containing authentic human speech and AI-generated speech from multiple synthesis and voice cloning systems. The dataset is intended for research on deepfake voice detection and scam call detection in Vietnamese. It was created by vietkemmai and last updated on July 10, 2026.
Use Cases
- Train deep learning models for deepfake voice detection based on the described authentic and AI-generated speech samples.
- Develop scam call detection systems based on the dataset's focus on Vietnamese scam call audio.
- Benchmark speech synthesis and voice cloning systems using the provided AI-generated speech data.
- Research the characteristics of Vietnamese speech in both human and synthetic contexts.
Strengths
- Contains both authentic human speech and AI-generated speech, providing a basis for comparative analysis.
- Designed specifically for the Vietnamese language, addressing a specific linguistic domain.
- Intended for developing machine learning and deep learning models, suggesting a research-oriented structure.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- huggingface user vietkemmai
- Collection Method
- Collected from multiple speech synthesis and voice cloning systems.
- Freshness
- Last updated 2026-07-10 16:23:36; freshness should be verified.
- Geography
- Vietnam (inferred from language focus)