bioRxiv · 10.64898/2026.03.12.711349
From sound to source: Human and model recognition of environmental sounds
Abstract
Our ability to recognize sound sources in the world is critical to daily life, but is not well documented or understood in computational terms. We developed a large-scale behavioral benchmark of human environmental sound recognition, built models of sound recognition that operate on the acoustic signal, and used the benchmark to compare models to humans. The behavioral benchmark measured how sound recognition varied across source categories, audio distortions, and concurrent sound sources, all of which influenced recognition performance in humans. Artificial neural network models trained to recognize sounds in multi-source scenes reached near-human accuracy and qualitatively matched human patterns of performance in many conditions. By contrast, classifiers trained to recognize sounds from traditional models of the cochlea and auditory cortex produced worse matches to human performance. Models trained on larger datasets exhibited stronger alignment with both human behavior and brain responses. The results suggest that many aspects of human sound recognition emerge in systems optimized for the problem of real-world recognition.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alavilli, S., McDermott, J. H.. 2026-03-14. From sound to source: Human and model recognition of environmental sounds. https://doi.org/10.64898/2026.03.12.711349
Cite the original work for its findings. Save a collection to share your selection of sources.