bioRxiv · 10.64898/2026.04.24.720581
Image-grounded encoding models reveal distinct temporal profiles of naturalistic object and scene processing in the human brain
Abstract
In real-world vision, the human brain needs to process large amounts of information to effectively interact with its environment. It is well established that our visual system has specialized regions to process information efficiently, such as scene-, face-, and object-selective areas, which can be uncovered using functional magnetic resonance imaging (fMRI). However, functional localizers typically use artificially separated stimuli (e.g. scene backgrounds vs. object cutouts), and fMRI signals are insensitive to potential temporal processing differences. Here, we identify temporal signatures of object and scene processing in natural visual environments by building image-grounded brain-predictive encoding models of human electroencephalography (EEG) responses. In a large set of high-resolution natural images, we first separate object from scene information on a per-image basis, and then feed this information to train separate deep neural network-based encoding models to predict EEG responses to the full, intact images. We find that encoding models that receive only object information consistently exhibit a delayed temporal encoding profile compared to models that only receive scene information. Control analyses confirm the robustness of this delayed object encoding, showing that consistent selection of object or scene information is needed to achieve high encoding performance. Using these distinct encoding profiles as processing templates, we reveal which image parts are seen by the brain as an object, versus a background scene element. These findings demonstrate that object and scene processing in the human brain unfold differentially over time and indicate that image-grounded encoding models are powerful tools for isolating components of naturalistic perception. SignificanceThe neural mechanisms underlying visual perception are typically studied by experimentally separating different kinds of information, for example presenting either isolated objects or scene backgrounds. Since information naturally co-occurs in the world -- objects are typically embedded within scenes -- this artificial separation reduces ecological validity. We leverage image-computable neural encoding models to identify distinct visual processing profiles within ecologically valid, natural images. Models that receive stimuli in which only objects are visible and scene information is masked predict EEG responses at later time points than models receiving the inverse, "scene-only" stimuli. By quantifying how items such as buildings, trees or the ground are encoded in EEG signals, we reveal which elements of visual environments drive neural responses to intact, real-world images.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mueller, N., Scholte, H. S., Groen, I. I. A.. 2026-04-24. Image-grounded encoding models reveal distinct temporal profiles of naturalistic object and scene processing in the human brain. https://doi.org/10.64898/2026.04.24.720581
Cite the original work for its findings. Save a collection to share your selection of sources.