Correlation-based binocular disparity computations induce representational bottlenecks at the population level
Binocular stereopsis depends on comparing the images seen by the two eyes. Although correlation-based models explain responses of individual binocular neurons in primary visual cortex (V1), it remains elusive whether population representations formed by local correlation activity patterns can support depth perception under ambiguous inputs. Using psychophysics, fMRI, and neural network modeling, we tested human stereopsis with dynamic anticorrelated stimuli that were dominated by binocular mismatches. Humans reliably perceived reversed depth as predicted by correlation-based computations, yet population representations consistent with this percept were detected in V3A, not V1. Shallow and deep neural networks constrained by correlation-like binocular interactions did not capture the full pattern of human depth judgments. Analyses of their internal representations showed greater representational overlap, whereas deep architectures not constrained to explicit correlation interactions exhibited less entangled representations and better aligned with human behavior. These findings suggest that biological stereopsis may rely on population coding beyond correlation-like computations. Significance StatementThe brain must infer depth from binocular inputs that are inherently ambiguous. Although correlation-based models explain disparity tuning of individual neurons in primary visual cortex (V1), whether these local mechanisms support perceptual inference at the population level remains unclear. Using psychophysics, fMRI, and neural network modeling, we show that population representations consistent with perceived depth under ambiguity were detected in mid-dorsal area V3A, not V1. Analyses of how neural networks encode multiple features showed that correlation-based computations represent features into overlapping activity patterns that may constrain downstream readout and degrade depth estimates. In contrast, models not constrained by explicit interocular correlation maintained more distinct population codes and closely matched human perception. These findings suggest that architectures constrained to explicit correlation-like processing can form population representations that are suboptimal to explain human stereopsis, motivating hybrid mechanisms.