MOCHA-TIMIT Corpus
The MOCHA-TIMIT corpus is a multi-modal English speech dataset recorded at the University of Edinburgh. It contains simultaneous acoustic, articulatory (EMA), electropalatographic (EPG), and laryngographic (LAR) signals from 8 speakers (~460 utterances each). While the original corpus contains 10 speakers, only 2 had been previously cleaned; we curated the remaining speakers, with 8 featured in this demo. This interactive explorer presents the time-aligned multimodal data with 8-tier TextGrid annotations.
{Project page} · {GitHub repo} · {Dataset download} · {How to cite}
Tap legend to toggle sensors · Curves show tongue trajectory (EMA y-axis) through the utterance
8×8 electrode grid · Orange = tongue–palate contact · Row order: alveolar (top) → velum (bottom)
Laryngograph waveform showing vocal fold vibration · Higher oscillation = voiced segments · 1.6 kHz display rate