Interspeech 2026 · Interactive Results

Do Machines Listen Like Humans?

A Temporal Benchmark for Phonological Competition in End-to-End ASR

Human listeners recognize speech incrementally, activating and suppressing competing word candidates as sound unfolds. This page lets you explore how the time course of lexical activation in ASR models compares to human eyetracking data from the Visual World Paradigm — and shows that high transcription accuracy does not imply human-like temporal dynamics.

Key finding

Causal architectures (LSTM, causal CNN, causal RCNN) replicate the hallmark human pattern: early cohort competition followed by later rhyme activation.

Non-causal models with look-ahead (BiLSTM, Transformer, ConvTransformer) and large pretrained ASR models (wav2vec 2.0, HuBERT, Whisper) fail to capture these dynamics despite higher transcription accuracy.

These results raise a caution against simply claiming that a high-accuracy model is brain-like without evaluating its temporal dynamics.

* NonCausal-RCNN is a hybrid: a unidirectional (causal) LSTM on top of a non-causal, centered CNN front-end. The CNN's receptive field is 25 frames (±12 frames ≈ ±120 ms of look-ahead at 10 ms/frame), which is what makes the overall model non-causal.

How human-like is each model?

Overall RMSE between each model's activation trajectories and human VWP fixation proportions. Click any bar to open it in the explorer.

Causal Non-causal Foundation (pretrained)

When more than one competitor type is shown, line color = model and dash style = competitor type (Target solid, Cohort dashed, Rhyme dotted, Unrelated dash-dot, Cross long-dash).

Click a column header to sort; click a row to open that model in the explorer.