bioRxiv · 10.1101/2025.08.29.673066
Shared texture-like representations, not global form, underlie deep neural network alignment with human visual processing
Abstract
Deep neural networks (DNNs) excel at predicting neural responses across the visual hierarchy 1-5, a success widely interpreted as evidence of shared object recognition computations 6,7. Yet improving DNN object recognition accuracy does not reliably increase neural predictivity 8,9, and even untrained networks predict brain responses above chance 9-11. This disconnect suggests that object recognition may not drive DNN-brain alignment. Texture-like statistics are represented in both DNNs and mid-level visual cortex 12-17. In natural images, these statistics are carried by objects and backgrounds, shaping representations and recognition in both systems 18-21. Does DNN-brain alignment reflect shared sensitivity to object-related information or texture-like statistics? To dissociate these factors, we recorded EEG from 57 participants viewing natural scenes, texture-synthesized images preserving local statistics while disrupting global form and object-only images with backgrounds removed. If alignment reflects texture-like statistics, it should peak for texture-synthesized images. If it reflects object-related processing, alignment should be strongest for natural and object-only conditions, which preserve object information. We compared EEG responses with DNN activations via weighted representational similarity analysis 22,23. Texture-synthesized images yielded the strongest DNN-EEG alignment, peaking in early responses (<200 ms) and explaining up to [~]85% of noise-ceiling-normalized explainable variance versus [~]44% for natural and [~]55% for isolated objects. Crucially, object categories were more decodable for natural and object-only images than texture-synthesized images, yet these object-rich conditions showed weaker alignment. This dissociation reveals that DNNs capture the texture-statistical component of early visual responses while failing to explain later, object-related variance. HighlightsO_LITexture-synthesized images yield the strongest DNN-brain alignment across DNNs. C_LIO_LIObject categories decode better from object-rich than texture-synthesized EEG. C_LIO_LIDNN-brain alignment reflects shared texture statistics, not object-related processing. C_LI
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Loke, J. L., Soerensen, L. K. A., Groen, I. I. A., Cappaert, N., Scholte, H. S.. 2025-09-04. Shared texture-like representations, not global form, underlie deep neural network alignment with human visual processing. https://doi.org/10.1101/2025.08.29.673066
Cite the original work for its findings. Save a collection to share your selection of sources.