Zhang, S., Li, S., Yang, R., Chen, G., Tian, X., Wang, Q., Fang, F.
Applying the human brain's local sequential phonology-to-semantics transformation mechanism to AI models improves speech comprehension performance.
AI speech comprehension models have reached human-level performance, but whether they rely on human-like mechanisms remains unknown.
Researchers used a phonology-semantics confusion paradigm comparing 12 speech language models with human brain data (SEEG). They identified two mechanisms in the human brain (local sequential transformation, global cross-regional hierarchy) and performed targeted lesioning and activation steering of model units.
Local sequential transformation predicted model performance, and manipulating those units improved performance. This mechanism was found across languages, suggesting a fundamental computational principle shared between biological and artificial intelligence.