Key finding
Causal architectures (LSTM, causal CNN, causal RCNN) replicate the hallmark human pattern: early cohort competition followed by later rhyme activation.
Non-causal models with look-ahead (BiLSTM, Transformer, ConvTransformer) and large pretrained ASR models (wav2vec 2.0, HuBERT, Whisper) fail to capture these dynamics despite higher transcription accuracy.
These results raise a caution against simply claiming that a high-accuracy model is brain-like without evaluating its temporal dynamics.
* NonCausal-RCNN is a hybrid: a unidirectional (causal) LSTM on top of a non-causal, centered CNN front-end. The CNN's receptive field is 25 frames (±12 frames ≈ ±120 ms of look-ahead at 10 ms/frame), which is what makes the overall model non-causal.
How human-like is each model?
Overall RMSE between each model's activation trajectories and human VWP fixation proportions. Click any bar to open it in the explorer.
When more than one competitor type is shown, line color = model and dash style = competitor type (Target solid, Cohort dashed, Rhyme dotted, Unrelated dash-dot, Cross long-dash).
Click a column header to sort; click a row to open that model in the explorer.