Search bioRxiv⌕ Search

Biology subjects

Adeli, H.

Publications and source records attributed to Adeli, H..

2 recordsLinked to original sources

A brain-inspired object-based attention network for multi-object recognition and visual reasoning

The visual system uses sequences of selective glimpses to objects to support goal-directed behavior, but how is this attention control learned? Here we present an encoder-decoder model inspired by the interacting bottom-up and top-down visual pathways making up the recognitionattention system in the brain. At every iteration, a new glimpse is taken from the image and is processed through the "what" encoder, a hierarchy of feedforward, recurrent, and capsule layers, to obtain an object-centric (object-file) representation. This representation feeds to the "where" decoder, where the evolving recurrent representation provides top-down attentional modulation to plan subsequent glimpses and impact routing in the encoder. We demonstrate how the attention mechanism significantly improves the accuracy of classifying highly overlapping digits. In a visual reasoning task requiring comparison of two objects, our model achieves near-perfect accuracy and significantly outperforms larger models in generalizing to unseen stimuli. Our work demonstrates the benefits of object-based attention mechanisms taking sequential glimpses of objects.

animal behavior and cognition↗

Readers move their eyes mindlessly using midbrain visuo-motor principles

Saccadic eye movements rapidly shift our gaze over 100,000 times daily, enabling countless tasks ranging from driving to reading. Long regarded as a window to the mind1 and human information processing2, they are thought to be cortically/cognitively controlled movements aimed at objects/words of interest3-10. Saccades however involve a complex cerebral network11-13 wherein the contribution of phylogenetically older sensory-motor pathways14-15 remains unclear. Here we show using a neuro-computational approach16 that mindless visuo-motor computations, akin to reflexive orienting responses17 in neonates18-19 and vertebrates with little neocortex15,20, guide humans eye movements in a quintessentially cognitive task, reading. These computations occur in the superior colliculus, an ancestral midbrain structure15, that integrates retinal and (sub)cortical afferent signals13 over retinotopically organized, and size-invariant, neuronal populations21. Simply considering retinal and primary-visual-cortex afferents, which convey the distribution of luminance contrast over sentences (visual-saliency map22), we find that collicular population-averaging principles capture readers prototypical word-based oculomotor behavior2, leaving essentially rereading behavior unexplained. These principles reveal that inter-word spacing is unnecessary23-24, explaining metadata across languages and writing systems using only print size as a predictor25-26. Our findings demonstrate that saccades, rather than being a window into cognitive/linguistic processes, primarily reflect rudimentary visuo-motor mechanisms in the midbrain that survived brain-evolution pressure27.

neuroscience↗