Sign in to view source links and access this dataset
Description
LIEPA-3 is a large, open corpus of Lithuanian speech built for automatic speech recognition, text-to-speech, and linguistic research. The corpus contains approximately 10,000 hours of audio across about 7.5 million files, spanning read, spontaneous, phonetically-annotated, and dialectal speech. It was created by meldynamics and last updated on the platform in July 2026.
Use Cases
Train automatic speech recognition models based on the large volume of Lithuanian audio files.
Develop text-to-speech synthesis systems based on the corpus's diverse speech types.
Conduct linguistic research on Lithuanian phonetics and dialects based on the annotated and dialectal speech data.
Study speech variation across recording conditions based on the studio, dictaphone, radio, TV, telephone, and audiobook sources mentioned.
Strengths
Approximately 10,000 hours of audio data provides a substantial resource for model training.
The corpus includes about 7.5 million individual audio files, offering fine-grained data points.
It covers diverse speech types including read, spontaneous, phonetically-annotated, and dialectal speech.
Audio was recorded under a wide range of conditions such as studio, dictaphone, radio, TV, telephone, and audiobooks.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
License information is unknown, which may restrict usage.
The description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
meldynamics
Collection Method
Likely compiled from various recorded sources including studio, dictaphone, radio, TV, telephone, and audiobooks.
Freshness
Last updated 2026-07-03 20:15:27; freshness should be verified.
Geography
Lithuania (implied by focus on Lithuanian language)
License restrictions are unknown and should be verified before use.