Search bioRxivSearch

Biology subjects

Groen, I. I. A.

Publications and source records attributed to Groen, I. I. A..

4 recordsLinked to original sources

Similarity judgments and cortical visual responses reflect different properties of object and scene categories in naturalistic images

Numerous factors have been reported to underlie the representation of complex images in high-level human visual cortex, including categories (e.g. faces, objects, scenes), animacy, and real-world size, but the extent to which this organization is reflected in behavioral judgments of real-world stimuli is unclear. Here, we compared representations derived from explicit similarity judgments and ultra-high field (7T) fMRI of human visual cortex for multiple exemplars of a diverse set of naturalistic images from 48 object and scene categories. Behavioral judgements revealed a coarse division between man-made (including humans) and natural (including animals) images, with clear groupings of conceptually-related categories (e.g. transportation, animals), while these conceptual groupings were largely absent in the fMRI representations. Instead, fMRI responses tended to reflect a separation of both human and non-human faces/bodies from all other categories. This pattern yielded a statistically significant, but surprisingly limited correlation between the two representational spaces. Further, comparison of the behavioral and fMRI representational spaces with those derived from the layers of a deep neural network (DNN) showed a strong correspondence with behavior in the top-most layer and with fMRI in the mid-level layers. These results suggest that there is no simple mapping between responses in high-level visual cortex and behavior - each domain reflects different visual properties of the images and responses in high-level visual cortex may correspond to intermediate stages of processing between basic visual features and the conceptual categories that dominate the behavioral response.\n\nSignificance StatementIt is commonly assumed there is a correspondence between behavioral judgments of complex visual stimuli and the response of high-level visual cortex. We directly compared these representations across a diverse set of naturalistic object and scene categories and found a surprisingly and strikingly different representational structure. Further, both types of representation showed good correspondence with a deep neural network, but each correlated most strongly with different layers. These results show that behavioral judgments reflect more conceptual properties and visual cortical fMRI responses capture more general visual features. Collectively, our findings highlight that great care must be taken in mapping the response of visual cortex onto behavior, which clearly reflect different information.

neuroscience

Scene complexity modulates degree of feedback activity during object recognition in natural scenes

Object recognition is thought to be mediated by rapid feed-forward activation of object-selective cortex, with limited contribution of feedback. However, disruption of visual evoked activity beyond feed-forward processing stages has been demonstrated to affect object recognition performance. Here, we unite these findings by reporting that the detection of target objects in natural scenes is selectively characterized by enhanced feedback when these objects are embedded in high complexity scenes. Human participants performed an animal target detection task on scenes with low, medium or high complexity as determined by a biologically plausible computational model of low-level contrast statistics. Three converging lines of evidence indicate that feedback was enhanced during categorization of scenes with high, but not low or medium complexity. First, functional magnetic resonance imaging (fMRI) activity in early visual cortex (V1) was selectively enhanced for target objects in scenes with high complexity. Second, event-related potentials (ERPs) evoked by high complexity scenes were selectively enhanced from 220 ms after stimulus-onset. Third, behavioral performance deteriorated for highly complex scenes when participants were pressed for time, but not when they could process the scenes fully and thereby benefit from the enhanced feedback. Formal modeling of the reaction time distributions revealed that object information accumulated more slowly for high complexity scenes (resulting in more errors especially for fast decisions), and directly related to the build-up of the feedback activity that was observed exclusively for high complexity scenes. Together, these results suggest that while feed-forward activity may suffice for simple scenes, the brain employs recurrent processing more adaptively in naturalistic settings, using minimal feedback for sparse, coherent scenes and increasing feedback for complex, fragmented scenes.\n\nAuthor summaryHow much neural processing is required to detect objects of interest in natural scenes? The astonishing speed of object recognition suggests that fast feed-forward buildup of perceptual activity is sufficient. However, this view is contradicted by findings that show that disruption of slower neural feedback leads to decreased detection performance. Our study unites these discrepancies by identifying scene complexity as a critical driver of neural feedback. We show how feedback is enhanced for complex, cluttered scenes compared to simple, well-organized scenes. Moreover, for complex scenes, more feedback is associated with better performances. These findings relate the flexibility of neural processes to perceptual decision-making by demonstrating that the brain dynamically directs neural resources based on the complexity of real-world visual inputs.

neuroscience

The temporal evolution of conceptual object representations revealed through models of behavior, semantics and deep neural networks

Visual object representations are commonly thought to emerge rapidly, yet it has remained unclear to what extent early brain responses reflect purely low-level visual features of these objects and how strongly those features contribute to later categorical or conceptual representations. Here, we aimed to estimate a lower temporal bound for the emergence of conceptual representations by defining two criteria that characterize such representations: 1) conceptual object representations should generalize across different exemplars of the same object, and 2) these representations should reflect high-level behavioral judgments. To test these criteria, we compared magnetoencephalography (MEG) recordings between two groups of participants (n = 16 per group) exposed to different exemplar images of the same object concepts. Further, we disentangled low-level from high-level MEG responses by estimating the unique and shared contribution of models of behavioral judgments, semantics, and different layers of deep neural networks of visual object processing. We find that 1) both generalization across exemplars as well as generalization of object-related signals across time increase after 150 ms, peaking around 230 ms; 2) behavioral judgments explain the most unique variance in the response after 150 ms. Collectively, these results suggest a lower bound for the emergence of conceptual object representations around 150 ms following stimulus onset.

neuroscience

Characterizing the temporal dynamics of object recognition by deep neural networks: role of depth

Convolutional neural networks (CNNs) have recently emerged as promising models of human vision based on their ability to predict hemodynamic brain responses to visual stimuli measured with functional magnetic resonance imaging (fMRI). However, the degree to which CNNs can predict temporal dynamics of visual object recognition reflected in neural measures with millisecond precision is less understood. Additionally, while deeper CNNs with higher numbers of layers perform better on automated object recognition, it is unclear if this also results into better correlation to brain responses. Here, we examined 1) to what extent CNN layers predict visual evoked responses in the human brain over time and 2) whether deeper CNNs better model brain responses. Specifically, we tested how well CNN architectures with 7 (CNN-7) and 15 (CNN-15) layers predicted electro-encephalography (EEG) responses to several thousands of natural images. Our results show that both CNN architectures correspond to EEG responses in a hierarchical spatio-temporal manner, with lower layers explaining responses early in time at electrodes overlying early visual cortex, and higher layers explaining responses later in time at electrodes overlying lateral-occipital cortex. While the explained variance of neural responses by individual layers did not differ between CNN-7 and CNN-15, combining the representations across layers resulted in improved performance of CNN-15 compared to CNN-7, but only after 150 ms after stimulus-onset. This suggests that CNN representations reflect both early (feed-forward) and late (feedback) stages of visual processing. Overall, our results show that depth of CNNs indeed plays a role in explaining time-resolved EEG responses.

neuroscience