1,581 podcast episodes and 2,444 transcript chunks processed from raw ASR output to cleaned text. The dataset includes 76,096 language model traces documenting the cleaning and evaluation process. Author hudsongouge updated the collection in July 2026.
Use Cases
- Train or evaluate automatic speech recognition (ASR) correction models based on the raw-to-cleaned transcript pairs.
- Analyze language model reasoning patterns based on the documented prompts, chains-of-thought, and tool outputs.
- Benchmark the performance of text cleaning algorithms using the provided pass/fail evaluations.
- Study podcast content and metadata across different shows using the episode_id and show_id fields.
Strengths
- Contains 76,096 detailed language model traces, providing transparency into the cleaning process.
- Includes 1,581 full-episode transcripts and 2,444 chunk-level pairs, offering data at multiple granularities.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count for the primary 'episodes' and 'chunks' configs is known, but overall dataset size and file formats are unknown.
Provenance
- Source
- hudsongouge on Hugging Face
- Collection Method
- Raw ASR output processed into cleaned transcripts, with LM traces documenting the steps.
- Freshness
- Last updated 2026-07-14 13:57:49; freshness should be verified.