Sign in to view source links and access this dataset
Description
A Russian speech dataset containing 600 audio chunks with transcriptions, intended for fine-tuning neural models on oil and gas industry terminology. The dataset was created by author N07P and last updated on Hugging Face in July 2026. Audio samples are classified into four context types: industry terms, connectors, questions, and noise.
Use Cases
Fine-tuning speech recognition models based on domain-specific oil and gas vocabulary mentioned in the description.
Training audio classifiers to distinguish between speech, questions, and noise based on the provided context labels.
Building robust Russian ASR pipelines for technical environments using the pre-split train/validation/test subsets.
Studying the performance of speech models on specialized connector phrases and question intonations in Russian.
Strengths
Contains 600 audio chunks with manual transcriptions.
Data is pre-split into train, validation, and test sets in an 80/10/10 proportion.
Each audio sample is classified into one of four specific context categories (term, connector, question, noise).
Limitations
Description metadata is limited; actual data quality requires manual inspection after download.
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Provenance
Source
huggingface
Freshness
Last updated 2026-07-15 09:40:20; freshness should be verified.
License is unknown; users should verify permissions before use.