Sign in to view source links and access this dataset
Description
90,334 hours of processed English (UK) dual-channel call center audio recordings, part of a larger multilingual collection totaling 2,065,026 hours. The dataset consists of real-world customer and agent speech from call center environments, created by InfoBayAI and last updated on Hugging Face in June 2026. It is designed to support the development of advanced speech and conversational AI systems.
Use Cases
Train automatic speech recognition (ASR) models based on real-world call center conversations.
Develop speaker diarization systems based on the dual-channel format separating agent and customer audio.
Build and fine-tune conversational AI agents based on natural customer service dialogues.
Conduct acoustic analysis or accent modeling based on UK English speech samples from call centers.
Strengths
Contains 90,334 hours of processed UK English call center audio.
Features a dual-channel format, which likely separates agent and customer speech.
Part of a larger multilingual collection of over 2 million hours of call center audio.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count and file formats are unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
InfoBayAI via Hugging Face.
Collection Method
Real-world customer and agent speech recordings collected from call center environments.
Freshness
Last updated 2026-06-03 05:49:59; freshness should be verified.
Geography
United Kingdom (based on language and title).
License is unknown; terms of use must be verified before application.