Sign in to view source links and access this dataset
Description
Per-channel human-transcript exports for audio from eight YouTube channels, including FAPTV and AnhThamTuTV. The dataset, created by Meddies, contains sequential WAV audio files, VTT transcripts, and manifests within a 'processed/' folder, along with audit and summary files in an 'analysis/' folder. It was last updated on July 1, 2026.
Use Cases
Training automatic speech recognition models based on human-labeled VTT transcripts.
Evaluating ASR model performance across different content channels mentioned in the description.
Analyzing speech patterns or transcription quality using the provided dataset audit files.
Building audio-text alignment tools based on the sequential WAV and transcript file structure.
Strengths
Human-transcribed labels, which likely provide high-quality ground truth for ASR tasks.
Includes dataset audit and summary files for quality analysis.
Contains data from eight distinct YouTube channels, suggesting some diversity in content.
Limitations
Description metadata is limited; actual data quality requires manual inspection after download.
Column-level documentation is absent; field semantics must be inferred after download.
Row count and total audio duration are unknown, which may limit suitability assessment.
Provenance
Source
huggingface
Collection Method
Human-transcript exports from audio sourced from eight listed YouTube channels.
Freshness
Last updated 2026-07-01 12:52:25; freshness should be verified.
License is unknown; terms of use must be verified before application.