0.000 / 2.218s

MOCHA-TIMIT Corpus

The MOCHA-TIMIT corpus is a multi-modal English speech dataset recorded at the University of Edinburgh. It contains simultaneous acoustic, articulatory (EMA), electropalatographic (EPG), and laryngographic (LAR) signals from 8 speakers (~460 utterances each). While the original corpus contains 10 speakers, only 2 had been previously cleaned; we curated the remaining speakers, with 8 featured in this demo. This interactive explorer presents the time-aligned multimodal data with 8-tier TextGrid annotations.

Audio waveform
EMA — electromagnetic articulography

Tap legend to toggle sensors · Curves show tongue trajectory (EMA y-axis) through the utterance

EPG — electropalatography

8×8 electrode grid · Orange = tongue–palate contact · Row order: alveolar (top) → velum (bottom)

LAR — laryngograph

Laryngograph waveform showing vocal fold vibration · Higher oscillation = voiced segments · 1.6 kHz display rate

Praat-style TextGrid — 8 tiers
← swipe to explore →
1.0×
Tier
Label
Time range
Duration
© Zheng Yuan, Leonardo Lancia, Allan Wrench — MOCHA-TIMIT Corpus