A lightweight, text-only derivative of the original tiny-aya-translate/hinglish-casual dataset. The original dataset is designed for simultaneous translation and contains many columns including audio references, speaker metadata, and duration. It was created by author bingbangboom and last updated on HuggingFace on 2026-07-01.
Use Cases
Train text-only machine translation models based on the described translation pairs.
Benchmark translation quality for casual Hinglish-English language pairs.
Process and clean text data containing embedded paralinguistic tags like <sigh> or <laugh>.
Develop lightweight NLP applications where audio and metadata columns are not required.
Strengths
Derived from a dataset designed for simultaneous translation, suggesting aligned text pairs.
Stripped of non-text columns like audio references and speaker metadata, focusing the data on core translation tasks.
Last updated on 2026-07-01, indicating recent maintenance.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
The dataset is a derivative; the original data's quality and potential biases are inherited.
Provenance
Source
huggingface
Collection Method
Derived from the tiny-aya-translate/hinglish-casual dataset.
Freshness
Last updated 2026-07-01.
License is unknown; terms of use must be verified before application.