Search bioRxiv⌕ Search

bioRxiv · 10.64898/2025.12.10.693611

Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space

Abstract

The representational structure of large-scale human color perception remains incompletely understood. While classical studies measured numerous color pairs, these measurements compared only similar colors, and exploring the global relationships among thousands of colors has been infeasible due to the time costs of psychophysical experiments. Given these constraints, deep neural networks (DNNs) have attracted attention as a promising tool for providing proxies or predictions of human perception beyond the scope of psychophysical experiments. However, it remains unclear which DNNs possess embeddings that geometrically align with human color perception. Furthermore, it is unclear which learning paradigm enables DNNs to acquire a color representation that aligns with that of humans. Here, we systematically investigate which learning paradigm enables DNNs to produce a color representation that is structurally congruent with that of humans, with a focus on three types: self-supervised learning (SSL) that trains on images alone, supervised learning (SL) that trains on images with category labels, and contrastive language-image pre-training (CLIP) that trains on image-text pairs. We compared the embeddings of DNNs with the human similarity judgments of 93 colors using a rigorous unsupervised method termed Gromov-Wasserstein Optimal Transport (GWOT). Our results show that, while each learning paradigm acquires color representations that strongly align with human data at the fine-item level in early layers, only CLIP sustains such a representation at the output. Furthermore, when we leveraged a key advantage of DNNs and investigated the representational structure of 4096 colors, the early layers of each learning paradigm and the output of CLIP consistently converged on their own characteristic structures. These structures present plausible predictions for the large-scale human color representation. Our work demonstrates an approach for exploring unknown territories of human perception through the use of computational models validated in a limited empirical space, and provides predictions for future large-scale psychophysical experiments. Author summaryHow do we perceive the vast world of color? Despite extensive research into human color perception, studies evaluating many colors have mostly captured differences between similar colors, while those mapping global relationships are restricted to a few dozen. Consequently, we still do not know the global structure of the massive "color map" that might underlie our perception of thousands of colors, as testing this directly is practically impossible. To explore this space, we turned to deep neural networks, a form of AI. Our first step was to identify models that "see" color in a way that matches humans. We compared models against human data capturing the global relationships among all possible pairs of 93 colors. Using a powerful geometric comparison method, we found the models that matched the human color map. This allowed us to use these models as reliable computational proxies. We then used them to do what human experiments currently cannot: chart a vast global map of 4,096 colors. The human-aligned models consistently converged on two distinct structures. Our work provides the first plausible, testable predictions for the large-scale structure of human color perception and demonstrates a new way to explore otherwise unreachable territories of our perceptual world.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wickramanayaka, N. R., Oizumi, M.. 2025-12-13. Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space. https://doi.org/10.64898/2025.12.10.693611

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Different hippocampal subfield volumes predict source memory performance and general cognitive ability in an adult lifespan sample

Modest positive associations between episodic memory performance and whole hippocampal and hippocampal subfield volumes have been reported in numerous prior studies. A smaller number of studies have reported associations between hippocampal volume and performance on tests of non-mnemonic cognition. The present study examined whether these associations were evident in a lifespan sample of cognitively healthy adults. Of particular interest was whether any identified associations were sensitive to age, and whether associations between subfield volumes and mnemonic and non-mnemonic performance were subfield dependent. We acquired high-resolution T1- and T2-weighted structural images from 163 adults (18-87 years of age). Participants also undertook a comprehensive neuropsychological test battery and an in-scanner test of source memory. Principal components analysis was employed to reduce the neuropsychological test scores to 5 cognitive components. Two components reflected memory performance while the other three reflected different aspects of non-mnemonic cognition. Hippocampal subfields (Cornu Ammonis (CA)1, CA2-3, dentate gyrus (DG) and subiculum) were segmented and measured with the Automated Segmentation of Hippocampus Subfields (ASHS) package. Source memory performance was selectively associated across participants with CA2-3 volume. By contrast, both mnemonic and non-mnemonic component scores derived from the test battery were associated exclusively with the volume of the DG. All associations were age-invariant. The findings indicate that different cognitive domains can be dissociated by virtue of their associations with different hippocampal subfields. Of importance, these associations appear to be life-long and hence are unlikely to reflect individual differences in age-related decline in structural integrity.

neuroscience↗

Cell type specific astrocytic feedback regulates excitation inhibition balance and cortical network dynamics

Astrocytes actively regulate synaptic transmission and neuronal excitability, yet their role in orchestrating macroscopic cortical network regimes and slow-wave oscillations remains an active area of reasearch. This study investigates how bidirectional neuron astrocyte interactions shape emergent population dynamics using a computational network model of excitatory and inhibitory neurons coupled to an astrocyte. The results identify astrocytic feedback topology, rather than astrocytic coupling strength alone, as a key determinant of emergent cortical network dynamics. By systematically dissecting pathway-specific connectivity, it has been shown that the neuronal population driving astrocytic activation and the neuronal population receiving gliotransmission jointly determine whether the network occupies asynchronous irregular (AI), synchronous irregular (SI), synchronous regular(SR), asynchronous regular(AR) or quiescent regimes.Directing gliotransmission selectively onto excitatory neurons consistently promotes population synchrony regardless of the population influencing astrocytic dynamics, whereas selective modulation of inhibitory interneurons induces network quiescence via strong suppression. Under dual-target gliotransmission, network synchrony is dictated by the population driving astrocytic dynamics: excitatory-only drive promotes synchrony, while combined or inhibitory-specific drive preserves asynchronous states. Furthermore, the model reveals that astrocytic signaling kinetics provide an additional temporal control mechanism that regulates the frequency and persistence of self sustained up states.

neuroscience↗

VCP inhibition prevents cone photoreceptor degeneration in the cpfl1 mouse model of achromatopsia

Achromatopsia (ACHM) is a rare autosomal recessive retinal disorder characterized by absent cone photoreceptor function from early life, leading to severe visual impairment. Mutations in genes involved in the cone phototransduction cascade frequently result in elevated cyclic guanosine monophosphate (cGMP) levels and activation of stress pathways, including endoplasmic reticulum (ER) stress and the unfolded protein response. Targeting common downstream mechanisms rather than individual mutations may provide a broadly applicable therapeutic strategy. Here, we investigated whether pharmacological inhibition of valosin-containing protein (VCP), a key regulator of ER and protein homeostasis, can prevent cone degeneration in the spontaneous cone photoreceptor function loss 1 (cpfl1) mouse model of ACHM. Organotypic culture of retinal explants from cpfl1 mice were treated with the selective VCP inhibitor ML240. Cone survival, cell death, opsin expression and localization were assessed by TUNEL assay, immunohistochemistry, and quantitative image analysis. ML240 treatment significantly increased cone density and improved cone opsin expression and trafficking to the outer segments (OSs) in cpfl1 explants compared to controls. Importantly, rhodopsin trafficking in rod photoreceptors was unaffected, indicating that VCP inhibition did not impair normal rod phototransduction. These findings demonstrate that VCP inhibition by ML240 effectively preserves cone photoreceptors and improves cone-specific functional markers in the cpfl1 model. Targeting VCP may represent a mutation-independent therapeutic strategy for preventing cone death in ACHM.

neuroscience↗