bioRxiv · 10.64898/2026.05.13.724755
Predictive coding video models capture dorsal parietal representations and human judgments for surfaces defined by motion
Abstract
Interacting with the world relies on knowing its three-dimensional structure, yet vision receives a stream of two-dimensional images. Recovering 3D geometry from dynamic input is therefore foundational to vision, but lacks a computational account. Here we identify one, using artificial neural networks to predict human behavior and neuronal responses from the two major visual pathways in macaques. Using videos of camouflaged objects whose shape is revealed through motion, we found both pathways carry information for motion-defined surfaces, while neurons in parietal cortex most closely track what humans perceive. Only predictive coding video models--trained to fill in missing content in natural videos--reproduced both neuronal responses and human behavior. Predictive coding thus emerges as a principle that yields a brain-like geometric world model of the physical world.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bai, Y. H., O'Connell, T. P., Friedman, Y., Ayvazian-Hancock, A., Maver, H., Tenenbaum, J. B., DiCarlo, J.. 2026-05-18. Predictive coding video models capture dorsal parietal representations and human judgments for surfaces defined by motion. https://doi.org/10.64898/2026.05.13.724755
Cite the original work for its findings. Save a collection to share your selection of sources.