Search bioRxivSearch

Biology subjects

Scholte, H. S.

Publications and source records attributed to Scholte, H. S..

4 recordsLinked to original sources

Scene complexity modulates degree of feedback activity during object recognition in natural scenes

Object recognition is thought to be mediated by rapid feed-forward activation of object-selective cortex, with limited contribution of feedback. However, disruption of visual evoked activity beyond feed-forward processing stages has been demonstrated to affect object recognition performance. Here, we unite these findings by reporting that the detection of target objects in natural scenes is selectively characterized by enhanced feedback when these objects are embedded in high complexity scenes. Human participants performed an animal target detection task on scenes with low, medium or high complexity as determined by a biologically plausible computational model of low-level contrast statistics. Three converging lines of evidence indicate that feedback was enhanced during categorization of scenes with high, but not low or medium complexity. First, functional magnetic resonance imaging (fMRI) activity in early visual cortex (V1) was selectively enhanced for target objects in scenes with high complexity. Second, event-related potentials (ERPs) evoked by high complexity scenes were selectively enhanced from 220 ms after stimulus-onset. Third, behavioral performance deteriorated for highly complex scenes when participants were pressed for time, but not when they could process the scenes fully and thereby benefit from the enhanced feedback. Formal modeling of the reaction time distributions revealed that object information accumulated more slowly for high complexity scenes (resulting in more errors especially for fast decisions), and directly related to the build-up of the feedback activity that was observed exclusively for high complexity scenes. Together, these results suggest that while feed-forward activity may suffice for simple scenes, the brain employs recurrent processing more adaptively in naturalistic settings, using minimal feedback for sparse, coherent scenes and increasing feedback for complex, fragmented scenes.\n\nAuthor summaryHow much neural processing is required to detect objects of interest in natural scenes? The astonishing speed of object recognition suggests that fast feed-forward buildup of perceptual activity is sufficient. However, this view is contradicted by findings that show that disruption of slower neural feedback leads to decreased detection performance. Our study unites these discrepancies by identifying scene complexity as a critical driver of neural feedback. We show how feedback is enhanced for complex, cluttered scenes compared to simple, well-organized scenes. Moreover, for complex scenes, more feedback is associated with better performances. These findings relate the flexibility of neural processes to perceptual decision-making by demonstrating that the brain dynamically directs neural resources based on the complexity of real-world visual inputs.

neuroscience

How to control for confounds in decoding analyses of neuroimaging data

Over the past decade, multivariate pattern analyses and especially decoding analyses have become a popular alternative to traditional mass-univariate analyses in neuroimaging research. However, a fundamental limitation of decoding analyses is that the source of information driving the decoder is ambiguous, which becomes problematic when the to-be-decoded variable is confounded by variables that are not of primary interest. In this study, we use a comprehensive set of simulations and analyses of empirical data to evaluate two techniques that were previously proposed and used to control for confounding variables in decoding analyses: counterbalancing and confound regression. For our empirical analyses, we attempt to decode gender from structural MRI data when controlling for the confound brain size. We show that both methods introduce strong biases in decoding performance: counterbalancing leads to better performance than expected (i.e., positive bias), which we show in our simulations is due to the subsampling process that tends to remove samples that are hard to classify; confound regression, on the other hand, leads to worse performance than expected (i.e., negative bias), even resulting in significant below-chance performance in some scenarios. In our simulations, we show that below-chance accuracy can be predicted by the variance of the distribution of correlations between the features and the target. Importantly, we show that this negative bias disappears in both the empirical analyses and simulations when the confound regression procedure performed in every fold of the cross-validation routine, yielding plausible model performance. From these results, we conclude that foldwise confound regression is the only method that appropriately controls for confounds, which thus can be used to gain more insight into the exact source(s) of information driving ones decoding analysis.\n\nHIGHLIGHTSO_LIThe interpretation of decoding models is ambiguous when dealing with confounds;\nC_LIO_LIWe evaluate two methods, counterbalancing and confound regression, in their ability to control for confounds;\nC_LIO_LIWe find that counterbalancing leads to positive bias because it removes hard-to-classify samples;\nC_LIO_LIWe find that confound regression leads to negative bias, because it yields data with less signal than expected by chance;\nC_LIO_LIOur simulations demonstrate a tight relationship between model performance in decoding analyses and the sample distribution of the correlation coefficient;\nC_LIO_LIWe show that the negative bias observed in confound regression can be remedied by cross-validating the confound regression procedure;\nC_LI

neuroscience

Characterizing the temporal dynamics of object recognition by deep neural networks: role of depth

Convolutional neural networks (CNNs) have recently emerged as promising models of human vision based on their ability to predict hemodynamic brain responses to visual stimuli measured with functional magnetic resonance imaging (fMRI). However, the degree to which CNNs can predict temporal dynamics of visual object recognition reflected in neural measures with millisecond precision is less understood. Additionally, while deeper CNNs with higher numbers of layers perform better on automated object recognition, it is unclear if this also results into better correlation to brain responses. Here, we examined 1) to what extent CNN layers predict visual evoked responses in the human brain over time and 2) whether deeper CNNs better model brain responses. Specifically, we tested how well CNN architectures with 7 (CNN-7) and 15 (CNN-15) layers predicted electro-encephalography (EEG) responses to several thousands of natural images. Our results show that both CNN architectures correspond to EEG responses in a hierarchical spatio-temporal manner, with lower layers explaining responses early in time at electrodes overlying early visual cortex, and higher layers explaining responses later in time at electrodes overlying lateral-occipital cortex. While the explained variance of neural responses by individual layers did not differ between CNN-7 and CNN-15, combining the representations across layers resulted in improved performance of CNN-15 compared to CNN-7, but only after 150 ms after stimulus-onset. This suggests that CNN representations reflect both early (feed-forward) and late (feedback) stages of visual processing. Overall, our results show that depth of CNNs indeed plays a role in explaining time-resolved EEG responses.

neuroscience

Visual pathways from the perspective of cost functions and multi-task deep neural networks

Vision research has been shaped by the seminal insight that we can understand the higher-tier visual cortex from the perspective of multiple functional pathways with different goals. In this paper, we try to give a computational account of the functional organization of this system by reasoning from the perspective of multi-task deep neural networks. Machine learning has shown that tasks become easier to solve when they are decomposed into subtasks with their own cost function. We hypothesize that the visual system optimizes multiple cost functions of unrelated tasks and this causes the emergence of a ventral pathway dedicated to vision for perception, and a dorsal pathway dedicated to vision for action. To evaluate the functional organization in multi-task deep neural networks, we propose a method that measures the contribution of a unit towards each task, applying it to two networks that have been trained on either two related or two unrelated tasks, using an identical stimulus set. Results show that the network trained on the unrelated tasks shows a decreasing degree of feature representation sharing towards higher-tier layers while the network trained on related tasks uniformly shows high degree of sharing. We conjecture that the method we propose can be used to analyze the anatomical and functional organization of the visual system and beyond. We predict that the degree to which tasks are related is a good descriptor of the degree to which they share downstream cortical-units.

neuroscience