Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A multimodal dataset of 70,000 samples constructed by pairing handwritten digit images with spoken digit audio clips. The handwritten data is sourced from the MNIST database, and the spoken data is extracted from the Google Speech Commands dataset, with audio pre-processed into Mel Frequency Cepstral Coefficients. The dataset was created by Lyes Khacef and colleagues for research in multimodal fusion.
If a shuffle is performed on the training or test subsets, it must be performed in unison with the same order for the written digits, spoken digits, and labels to maintain alignment.