Search bioRxiv⌕ Search

Biology subjects

Bartnik, C. G.

Publications and source records attributed to Bartnik, C. G..

2 recordsLinked to original sources

Temporal misalignment in scene perception: Divergent representations of locomotive action affordances in human brain responses and DNNs

The human visual system processes scenes with remarkable speed, enabling the extraction of essential information to navigate our surroundings in a single glance. To elucidate how the brain transforms visual inputs into neural representations of navigationally relevant information, we collected electroencephalography (EEG) responses to diverse indoor and outdoor scenes along with behavioral annotations of locomotive action affordances (e.g., walking, cycling), object annotations, and low-level image features to model distinct types of scene information. Using representational similarity analysis, we examined the neural representation of locomotive action affordances over time, their co-localization within scene-selective cortex, and their computational alignment with deep neural networks (DNNs). Our results show that locomotive action affordance representations emerge within 200 ms of visual processing, showing unique contributions to EEG responses at temporally distinct time-points from objects and low-level properties. Spatiotemporal fusion with functional magnetic resonance imaging (fMRI) recordings in scene-selective brain regions reveals that both the parahippocampal (PPA) and occipital place region (OPA), but not the medial place region (MPA), contribute to locomotive action affordance representations, with a distinct temporal hierarchy between them. While DNNs exhibit good predictivity of early EEG responses, they primarily capture low-level features and show limited alignment with affordance processing. These findings reveal a temporally distinct neural representation of action affordances and highlight a limitation of current DNNs in modeling affordance perception.

neuroscience↗

Distinct representation of navigational action affordances in human behavior, brains and deep neural networks

To decide how to move around the world, we must determine which locomotive actions (e.g., walking, swimming, or climbing) are afforded by the immediate visual environment. The neural basis of our ability to recognize locomotive affordances is unknown. Here, we compare human behavioral annotations, functional magnetic resonance imaging (fMRI) measurements, and deep neural network (DNN) activations to both indoor and outdoor real-world images to demonstrate that human visual cortex represents locomotive action affordances in complex visual scenes. Hierarchical clustering of behavioral annotations of six possible locomotive actions show that humans group environments into distinct affordance clusters using at least three separate dimensions. Representational similarity analysis of multi-voxel fMRI responses in scene-selective visual cortex shows that perceived locomotive affordances are represented independently from other scene properties such as objects, surface materials, scene category or global properties, and independent of the task performed in the scanner. Visual feature activations from DNNs trained on object or scene classification as well as a range of other visual understanding tasks correlate comparatively lower with behavioral and neural representations of locomotive affordances than with object representations. Training DNNs directly on affordance labels or using affordance-centered language embeddings increases alignment with human behavior, but none of the tested models fully captures locomotive action affordance perception. These results uncover a new type of representation in the human brain that reflects locomotive action affordances. SignificanceTo navigate the world around us, we can use different actions, such as walking, swimming or climbing. How does our brain compute and represent such locomotive action affordances? Here, we show that activation patterns in high-level visual regions in the human brain represent information about affordances independent of other visual elements such as surface materials and objects, and do so in an automatic manner. We also demonstrate that commonly used models of visual processing in human brains, namely object- and scene- classification trained deep neural networks, do not strongly represent this information. Our results suggest that locomotive action affordance perception in scenes relies on specialized neural representations different from those used for other visual understanding tasks.

neuroscience↗