Search bioRxiv⌕ Search

Biology subjects

Dudipala, K. R.

Publications and source records attributed to Dudipala, K. R..

2 recordsLinked to original sources

Standardized human tests show that most vision models are face-blind

Face recognition ability varies enormously across humans. Individuals with face-blindness (prosopagnosia) struggle to recognize even close family members while super-recognizers can identify strangers with exceptional accuracy. Where do ANNs fall within the human face recognition spectrum? Here, we administered the standardized tests used to characterize human face recognition ability to a diverse set of models, allowing us to contextualize a model's performance within the distribution of human behavior. We found that the majority 55% of models to be classified as face-blind and that even the best face-trained models do not cross the human threshold to be considered a super-recognizer. Interrogating the internal representations of these models showed that models that performed well on standardized face recognition tasks were more identity-selective and viewpoint-invariant. We also found that low-performing models did contain some identity information in independent representational subspaces. Removing viewpoint-dependent subspace improved face recognition abilities in 49 of the 53 models tested. Targeted unit ablations further identified opposing contributions, with viewpoint-dependent units disrupting identity coding and viewpoint-tolerant units supporting it. Together, our results show that most AI models are face-blind with worse face recognition ability than humans, and demonstrate how differences across AI models can be used generate testable hypotheses about the computational basis of human face recognition in humans which can then be probed in future studies.

neuroscience↗

Zero-shot evaluations expose structured generalization limits in predictive models of the brain

Artificial neural networks (ANNs) can predict neural responses to natural images with remarkable accuracy, making them promising models of human vision. However, prediction scores alone reveal little about what computations these models have captured or how well they generalize beyond the images on which they are trained. Here, we introduce a principled zero-shot model evaluation framework in which ANN-based encoding models are frozen before being tested on not just held out images but entirely new datasets. Our strongest tests repurpose decades of cognitive neuroscience experiments as diagnostic benchmarks for identifying which computations current models capture and where they fail. To implement this framework we built encoding models of human category-selective regions (FFA, PPA, and EBA) using several pretrained ANNs and evaluated these fixed models across 7 independent fMRI datasets and 31 cognitive experiments from 10 published studies. This framework revealed three broad patterns. First, zero-shot evaluations exposed systematic limits to model predictivity that were largely hidden by standard within-dataset cross-validation. These limits were structured varying across brain regions and model classes and depending strongly on the similarity between the images used to build and test the models. Second, current models reproduced some, but not all, classical findings from cognitive neuroscience, providing a diagnosis of the computations that current models have yet to capture. Finally, models that performed well on prediction tests also tended to succeed at reproducing cognitive neuroscience findings, suggesting that prediction and explanation are closely linked. Unlike prediction scores however, cognitive tests help diagnose where and why models fail. Together, the zero-shot evaluation framework provides a scalable and unified approach for evaluating brain models across prediction and cognitive neuroscience tests, which can help us better understand what current models capture and what still remains to be explained.

neuroscience↗