Sign in to view source links and access this dataset
Description
Project Euphonia, a public initiative led by Google, aims to improve Automatic Speech Recognition for individuals with atypical speech. The Vaani corpus expands this work beyond English to include languages such as French, Spanish, Japanese, and Hindi. This dataset is hosted by ARTPARK-IISc and was last updated on March 18, 2026.
Use Cases
Training personalized speech recognition models based on atypical speech patterns mentioned in the description
Improving ASR system robustness for underrepresented speech types based on the project's focus
Benchmarking multilingual speech recognition performance across languages like French, Spanish, Japanese, and Hindi
Developing assistive technology applications, similar to the Android app Project Relate referenced in the description
Strengths
Part of Google's public Project Euphonia initiative, suggesting a structured collection effort
Includes speech data from multiple languages: French, Spanish, Japanese, and Hindi
Limitations
Column-level documentation is absent; field semantics must be inferred after download
Row count is unknown, which may limit suitability assessment
Freshness should be verified; the last update timestamp is March 18, 2026
Provenance
Source
ARTPARK-IISc, associated with Google's Project Euphonia
Collection Method
Public data collection initiative for atypical speech, as described
Freshness
Last updated 2026-03-18 13:52:52
License is unknown; users should verify terms of use before downloading.