Search bioRxiv⌕ Search

Biology subjects

Kuchibhotla, K.

Publications and source records attributed to Kuchibhotla, K..

2 recordsLinked to original sources

Rapid emergence of latent knowledge in the sensory cortex drives learning

Rapid learning confers significant advantages to animals in ecological environments. Despite the need for speed, animals appear to only slowly learn to associate rewarded actions with predictive cues1-4. This slow learning is thought to be supported by a gradual expansion of predictive cue representation in the sensory cortex2,5. However, evidence is growing that animals learn more rapidly than classical performance measures suggest6-8, challenging the prevailing model of sensory cortical plasticity. Here, we investigated the relationship between learning and sensory cortical representations. We trained mice on an auditory go/no-go task that dissociated the rapid acquisition of task contingencies (learning) from its slower expression (performance) 7. Optogenetic silencing demon-strated that the auditory cortex (AC) drives both rapid learning and slower performance gains but becomes dispensable at expert. Rather than enhancement or expansion of cue representations9, two-photon calcium imaging of AC excitatory neurons throughout learning revealed two higher-order signals that were causal to learning and performance. First, a reward prediction (RP) signal emerged rapidly within tens of trials, was present after action-related errors only early in training, and faded at expert levels. Strikingly, silencing at the time of the RP signal impaired rapid learning, suggesting it serves an associative and teaching role. Second, a distinct cell ensemble encoded and controlled licking suppression that drove the slower performance improvements. These two ensembles were spatially clustered but uncoupled from underlying sensory representations, indicating a higher-order functional segregation within AC. Our results reveal that the sensory cortex manifests higher-order computations that separably drive rapid learning and slower performance improvements, reshaping our understanding of the fundamental role of the sensory cortex. Despite the value of rapid learning in ecological environments, most laboratory models of rodent learning show that linking sensory cues with reinforced actions is a slow, gradual process1-4,10. An alternative view suggests that animals, including humans, rapidly infer relationships between cues, actions, and reinforcement (i.e. learning)6 even if they continue to make ongoing performance errors 7,8,11. Recent behavioral studies in rodents have begun to reconcile these views, arguing that latent task knowledge (i.e. discriminative contingencies) can emerge rapidly even though behavioral performance appears to improve only gradually7. How are these two dissociable behavioral processes--rapid acquisition of contingencies versus slower performance improvements--implemented in the brain? An attractive brain region to consider is the sensory cortex as it is thought to subserve instrumental learning by enhancing or attenuating the representation of sensory cues that drive behavior. Plasticity of cue-related responses in the sensory cortex is thought to subserve learning as it mirrors the slow and gradual improvements in behavioral performance 1,2,5,10. This raises a fundamental challenge: if animals learn discriminative contingencies rapidly but cue representations in the sensory cortex change slowly1,2,9, the causal model linking cue-related plasticity to learning becomes problematic. One possible solution is that the sensory cortex plays a role beyond cue-related representational plasticity and directly represents high-order signals that associate reinforced actions with predictive cues. Here we focus on the auditory cortex (AC) and asked whether and how it plays a higher-order role in cue-guided learning. We trained head-fixed, water-restricted mice to lick to a target tone (S+) for water reward and to withhold licking to a foil tone (S-) to avoid a timeout (auditory go/no-go task, Fig. 1a). We used simple pure tones to prevent the AC from being recruited for complex sensory processing. To confirm this, two-photon imaging of AC excitatory neurons showed that stimulus identity could accurately be decoded from AC activity from the first training day with no subsequent improvement throughout training (Supplementary Figure 1), suggesting that the AC was indeed not needed for perceptual sharpening in the task and thereby allowing us to identify possible associative functions. Performance was evaluated in each session in reinforced and non-reinforced ( probe) trials (Fig. 1b). Performance in probe trials revealed a rapid acquisition of task contingency knowledge which was only expressed much later in reinforced trials (Fig. 1c)7. Reinforcement feedback, although critical for learning, paradoxically masked the underlying task knowledge. By combining this behavioral procedure with optogenetics and longitudinal two-photon imaging, we aimed to determine how quickly animals learn stimulus-action contingencies and to define the fundamental role of the auditory cortex in sound-guided learning. O_FIG O_LINKSMALLFIG WIDTH=123 HEIGHT=200 SRC="FIGDIR/small/597946v1_fig1.gif" ALT="Figure 1"> View larger version (44K): org.highwire.dtl.DTLVardef@951aadorg.highwire.dtl.DTLVardef@10a6bbforg.highwire.dtl.DTLVardef@127e8eeorg.highwire.dtl.DTLVardef@12d84ba_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 1.C_FLOATNO Auditory cortex silencing impairs sound-guided learning and performance during learning.a, Head-fixed mice were trained on an auditory go/no-go task with 3 -spaced pure tones. H: hit, M: miss, FA: false alarm, CR: correct reject. b, Every day during training, task knowledge is probed by omitting reinforcement for 20 trials. c, Two distinct learning trajectories are revealed: a fast acquisition of task contingencies (measured in probe trials; green) and a slower knowledge expression (measured in reinforced trials; black). d, Probabilistic optogenetic silencing of the auditory cortex over learning. e, Testing conditions. f, Accuracy in reinforced light-on trials (two-way ANOVA, p < 10-8). g, Action rate in reinforced light-on trials (HIT, p = 0.07; FA, p < 10-33). See also Supplementary Figure 4. h, Accuracy in probe light-off trials (two-way ANOVA, p < 10-4). i, Tone response index in S+ trials (see Methods; two-way ANOVA, p < 10-101). Black and gray lines are individual mice and dots indicate change points (see Methods). j, Maximal difference between hit and FA rates in probe light-off trials over the first 6 days (t-test, p < 10-3). k, Hit lick latency in probe light-off trials (median {+/-} s.e.median; Wilcoxon test, p = 0.007). l, Accuracy in reinforced light-off trials (two-way ANOVA, p < 10-8). m, Action rate in reinforced light-off trials (two-way ANOVA, HIT: p = 0.57, FA: p < 10-8). n, Accuracy in reinforced light-off trials with inter-subject alignment to the day where probe accuracy[&ge;] 0.65 (green triangle) (two-way ANOVA, p < 10-5). Supplementary Figure 3a-c. o, Comparison of light-off versus light-on trials to measure auditory cortex silencing effect on on-line performance. p, Session density plot of accuracy in reinforced light-on against light-off. Top, control; bottom, PV-ChR2. See also Supplementary Figure 3d-g. q, Within subject accuracy difference in reinforced light-on and light-off trials, aligned to the day where FA rate < 0.3 in reinforced light-off (two-way ANOVA, p < 10-15). r, Within subject accuracy difference in reinforced light-on and light-off when silencing started at expert level (n = 4; t-test, p = 0.58). See also Supplementary Figure 6. mean {+/-} s.e.m.; *p < 0.05; **p < 0.01; ***p < 0.001, n.s.: not significant. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=38 SRC="FIGDIR/small/597946v1_figs1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@164efbdorg.highwire.dtl.DTLVardef@1b784e2org.highwire.dtl.DTLVardef@1754599org.highwire.dtl.DTLVardef@2c7b07_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 1.C_FLOATNO Stimulus decoding in the auditory cortex is at ceiling from Day 1 of learning. a, Stimulus decoding is at ceiling on Day 1 and remains high throughout learning (example mouse) Only the cells tracked across all days were used to decode tone identity. b, Stimulus decoding is at ceiling on day 1 and remains high throughout passive exposure over 15 days (example mouse). c, Average decoding accuracy for all Learning mice (n = 5). d, Average decoding accuracy for all Passive mice (n = 3). e, Evolution of tone decoding accuracy in the tone-evoked window across days for Learning and Passive mice compared to chance level (trial shuffle, see Methods). C_FIG O_FIG O_LINKSMALLFIG WIDTH=125 HEIGHT=200 SRC="FIGDIR/small/597946v1_figs4.gif" ALT="Figure 4"> View larger version (42K): org.highwire.dtl.DTLVardef@4114f7org.highwire.dtl.DTLVardef@c76c77org.highwire.dtl.DTLVardef@a20971org.highwire.dtl.DTLVardef@19dc17_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 4.C_FLOATNO Effect of AC full trial silencing on lick patterns a, Example control (top) and ChR2 (bottom) mice accuracy in probe light-off, reinforced light-off and reinforced light-on trials across day. Dashed rectangle indicates day where licks in b are extracted from. b, Lick raster plots from day 4 from the example mouse from A in probe light-off (left), reinforced light-off (middle) and reinforced light-on (right) trials, split into target (black, left) and foil (red, right) trials. Green and red dots indicates correct and incorrect trials, respectively. Note the difference in discrimination in all contexts between control and PV-ChR2 mice. c, Average lick probability across training days for control (n = 8) and ChR2 (n = 8) mice in response to target (vertical green line) and foil (vertical red line) tones, in reinforced light-off (black) and light-on (blue) trials. d, Insets showing faster lick latencies (red arrow heads) in response to both tones and higher lick probability in response to the foil (incorrect licking) in reinforced light-on compared to light-off in ChR2 mice (right). Light has no effect on lick structure in control mice (left). e, Lick latencies (top) and lick rate (bottom) in response to target (HIT trials; left) and foil (false alarm (FA) trials; right) tones in reinforced light-off trials (HIT lick latencies, Days: F (20, 256) = 8.2738, p < 10-17, Groups: F (1, 256) = 8.1568, p = 0.0046, Days*Groups: F (20, 256) = 0.9176, p = 0.56; FA Lick latencies, Days: F (20, 190) = 2.2393, p = 0.0027, Groups: F (1, 190) = 1.8422, p = 0.18, Days*Group: F (20, 190) = 1.5563, p = 0.067; HIT lick rate, Days: F (20, 256) = 4.3619, p < 10-8, Groups: F (1, 256) = 2.9549, p = 0.087, Days*Groups: F (20, 256) = 0.2927, p = 0.99; FA lick rate, Days: F (20, 190) = 4.04477, p < 10-6, Groups: F (1, 190) = 7.4070, p = 0.0071, Days*Groups: F (20, 190) = 1.1944, p = 0.26). f, Lick latencies (top) and lick rate (bottom) in response to target (HIT trials; left) and foil (false alarm (FA) trials; right) tones in reinforced light-on trials (HIT lick latencies, Days: F (20, 256) = 10.5303, p < 10-22, Groups: F (1, 256) = 11.2328, p < 10-3, Days*Groups: F (20, 256) = 0.6211, p = 0.90; FA Lick latencies, Days: F (20, 254) = 3.9111, p < 10-6, Groups: F (1, 254) = 450.4358, p < 10-57, Days*Group: F (20, 254) = 2.1947, p = 0.0029; HIT lick rate, Days: F (20, 256) = 2.6372, p < 10-3, Groups: F (1, 256) = 3.7748, p = 0.0531, Days*Groups: F (20, 256) = 0.4520, p = 0.98; FA lick rate, Days: F (20, 254) = 6.4469, p < 10-13, Groups: F (1, 254) = 301.2679, p < 10-44, Days*Groups: F (20, 254) = 0.6326, p = 0.89). g, Lick latencies (left) and lick rate (right) in response to target (HIT) and foil (FA) tones in probe light-off trials (HIT lick latencies, Days: F (5, 83) = 6.4522, p < 10-4, Groups: F (1, 83) = 11.7734, p < 10-3, Days*Groups: F (5, 83) = 0.2878, p = 0.92; FA Lick latencies, Days: F (5, 58) = 2.9217, p = 0.020, Groups: F (1, 58) = 0.9337, p = 0.338, Days*Group: F (5, 58) = 2.1909, p = 0.068; HIT lick rate, Days: F (5, 83) = 2.0103, p = 0.086, Groups: F (1, 83) = 5.9422, p = 0.017, Days*Groups: F (5, 83) = 0.5721, p = 0.72; FA lick rate, Days: F (5, 58) = 5.6386, p < 10-3, Groups: F (1, 58) = 0.0192, p = 0.89, Days*Groups: F (5, 58) = 1.6182, p = 0.17). C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=74 SRC="FIGDIR/small/597946v1_figs3.gif" ALT="Figure 3"> View larger version (23K): org.highwire.dtl.DTLVardef@1c07f99org.highwire.dtl.DTLVardef@f919c8org.highwire.dtl.DTLVardef@bd263org.highwire.dtl.DTLVardef@217412_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 3.C_FLOATNO AC full trial silencing impairs expression and on-line performance a, Assessment of the impact of AC full trial silencing over learning on Expression by controlling for the delay in Acquisition. b, Cumulative distribution function (CDF) of mice as function of the day to reach an accuracy[&ge;] 0.65 in probe trials. c, Cumulative distribution function (CDF) of mice as function of the relative number of days to reach accuracy (acc.) criteria of >0.7 (left), >0.8 (middle), and >0.9 (right) in reinforced light-off trials after reaching an accuracy[&ge;] 0.65 in probe trials. Black and dark gray vertical lines correspond to when CDF was reach for acc.>0.7 and >0.8, respectively. d, Comparing action rate and accuracy between reinforced light-off versus reinforced light-on trials to assess the impact of AC silencing on on-line performance. e, Hit (solid line) and FA (dashed line) of an example control mouse (top) and an example PV-ChR2 mouse (bottom) in reinforced light-off (black) and reinforced light-on (blue) trials across learning. f, Averaged action rate in reinforced light-off (black) and reinforced light-on (blue) trials per day for control (top) and PV-ChR2 (bottom) groups. g, Accuracy in light-on reinforced trials from the day when FA<0.3 in light-off reinforced trials. Note how PV-ChR2 mice (gray lines) increase accuracy (positive slopes) with light-on, showing that performance impairment fades away. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=43 SRC="FIGDIR/small/597946v1_figs6.gif" ALT="Figure 6"> View larger version (14K): org.highwire.dtl.DTLVardef@4b4bdcorg.highwire.dtl.DTLVardef@1617149org.highwire.dtl.DTLVardef@547fb7org.highwire.dtl.DTLVardef@18ce33e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 6.C_FLOATNO AC full trial silencing at expert level a, Probabilistic optogenetic silencing of the auditory cortex at expert level. Silencing starts once stable performance is reached. b, Accuracy in probe light-off (green), reinforced light-off (black) and reinforced light-on (blue) trials. Silencing is performed from day 19 to 23. c, Accuracy in reinforced light-off and light-on trials (paired t-test, p = 0.602). C_FIG The auditory cortex is the default pathway for sound-guided learningLesion studies have suggested that the AC may not be essential to learn or execute cue-guided tasks with simple sensory stimuli12-15. However, permanent lesions cannot determine whether the AC is normally used for, or causally produces16, learning in an intact brain. To address this, we exploited a transient silencing approach to prevent the recruitment of alternative pathways15,17-20 while also using a probabilistic design to allow assessment of learning as distinct from performance by measuring behavior on non-silenced trials, thereby avoiding direct effects of silencing on performance. We examined the impact of bilateral cortical silencing of the AC throughout learning (Fig. 1a). We probabilistically silenced the AC on 90% of reinforced trials throughout learning ( light-on reinforced, Fig. 1d), leaving 10% of reinforced ( light-off reinforced) and 100% of probe trials ( light-off probe) with intact AC activity. Silenced trials were pseudo-randomly sequenced and equally split between S+ and S-. Silencing was achieved by shining blue light bilaterally through cranial windows implanted above the AC of double transgenic mice (n=8) expressing channel rhodopsin (ChR2) in parvalbumin (PV) interneurons14,21 (Fig. 1d). We confirmed that the excitatory network was effectively silenced using this approach by combining two-photon calcium imaging of excitatory neurons and full-field optogenetic stimulation in PV-ChR2 mice (Supplementary Figure 2). Control mice (n=8) received the same light stimulation but did not express ChR2. This experimental design allowed us to assay the impact of cortical silencing on performance (control vs PV-ChR2 performance on light-on reinforced trials) versus acquisition learning (control vs PV-ChR2 performance on light-off probe trials) and expression learning (control vs PV-ChR2 performance on light-off reinforced trials) (Fig. 1e). O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=190 SRC="FIGDIR/small/597946v1_figs2.gif" ALT="Figure 2"> View larger version (61K): org.highwire.dtl.DTLVardef@9b2f88org.highwire.dtl.DTLVardef@4d8c14org.highwire.dtl.DTLVardef@127aa8borg.highwire.dtl.DTLVardef@12daa04_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 2.C_FLOATNO Activating PV+ neurons in the auditory cortex robustly suppresses stimulus-evoked activity of excitatory neurons. a, PV-ChR2 mice (n = 2) were injected with AAV-CaMKII-GCaMP6f to allow simultaneous one-photon excitation of PV cells and two-photon recordings of pyramidal cell population. b, Schematic of simultaneous widefield optogenetics and two-photon imaging. c, Optogenetic activation was locked to frame acquisition. d, Trial-averaged{Delta} F/F aligned to tone onset (black vertical line) of an example neuron at different intensity of LED power (blue scale). Yellow rectangle indicates period of light delivery. mean {+/-} s.e.m. e, Effect of optogenetic silencing as a function of LED power (n = 454 neurons; Friedman test, p [~] 0).{Delta} F/F at powers 0-0.26 mW/mm2 are all significantly different from{Delta} F/F at powers 0.84-3.15 mW/mm2 (post hoc comparisons with Tukey-Kramer test, ***p < 0.001). Black line is the logistic fit. median {+/-} s.e.median. f, Immunostaining of PV-ChR2 mice auditory cortex showing ChR2 expression in PV cells (PV+ and ChR2+ colocalization). g, Post-task imaging of a representative control (top) and a representative test (PV-ChR2, bottom) mouse used in AC silencing experiments. Note that no fluorescence below the dura is detected in control mice. C_FIG We first compared performance in light-on reinforced trials between PV-ChR2 and control mice (Fig. 1e) and observed a large performance impairment in PV-ChR2 mice (Fig. 1f,g). To address whether this performance reduction was accompanied by an impairment in rapid learning, we compared performance in PV-ChR2 and control animals in light-off probe trials (Fig. 1e,h-k) when the AC was not silenced and knowledge acquisition can be accurately measured7. Accuracy was lower during probe trials in PV-ChR2 mice (Fig. 1h), with delayed S+-response learning (Fig. 1i), lower discrimination (Fig. 1j), and longer lick latency on hit trials (Fig. 1k). Rapid acquisition of task knowledge was therefore impaired in PV-ChR2 mice. Accuracy was also lower in reinforced light-off trials in PV-ChR2 mice (Fig. 1l,m). This remained true even after controlling for their slower task acquisition (Figs.1n, Supplementary Figure 3a-c). These impairments were also apparent in response latency and response vigor (Supplementary Figure 4). Together, these results suggest that the AC is the default pathway for sound-guided reward learning, even when not needed for perceptual sharpening. The auditory cortex is used during learning but becomes dispensable at expert levelsWe next sought to understand the contribution of AC activity for the expression of the learned behavior as animals transitioned to expert performance. Transient inactivation of auditory cortex in expert animals has led to conflicting results, with some reports showing degradation of sound-guided behavior14,17,22,23 and others not14,24,25. We exploited our probabilistic silencing strategy and compared performance in light-on (AC silenced) versus light-off (AC functional) reinforced trials within subjects (Fig. 1o). Performance on these two trial types was similar at early periods of training, as performance was poor overall (Fig. 1p). As training progressed, performance remained poor on light-on trials but improved on light-off trials (Fig. 1p), demonstrating that the AC is used for task performance at early and intermediate time-point during learning. Surprisingly, this deficit in performance on light-on trials gradually waned (Fig. 1p,q), suggesting that while the AC was used during learning, it became dispensable once the mice had mastered the task. These results could be explained by three alternative explanations. First, the optogenetic manipulation per se may not be interfering with a task-relevant process but instead could be distracting the animal, necessitating more time to increase performance in light-on trials. We reasoned that bilateral silencing of another cortical region that is nominally unrelated to the task would serve as an important control. We bilaterally silenced the visual cortex throughout learning in PV-ChR2 mice and found no evidence of performance impairment in light-on trials (Supplementary Figure 5), demonstrating that the performance impairment was specific to AC silencing. Second, it is possible that AC silencing altered tone perception, increasing task difficulty at the perceptual level in light-on trials. Third, the reduction of impairment during light-on trials could be driven by a reduction of the silencing effect with time due, for example, to brain damage induced by repeated silencing. To address the second and third possibilities, we trained a separate cohort of PV-ChR2 mice without daily inactivation and, instead, inactivated the AC only after they reached expert performance (see Methods). We observed no impact from AC silencing (Figs.1r, Supplementary Figure 6)14. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=113 SRC="FIGDIR/small/597946v1_figs5.gif" ALT="Figure 5"> View larger version (28K): org.highwire.dtl.DTLVardef@f501baorg.highwire.dtl.DTLVardef@1449653org.highwire.dtl.DTLVardef@1e923cborg.highwire.dtl.DTLVardef@12d042e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 5.C_FLOATNO Silencing of the visual cortex does not impair performance throughout learning a, Silencing of the visual cortex in 90% of the reinforced trials throughout learning (n = 8 PV-ChR2 mice). b, Comparison of reinforced light-off versus light-on trials shows no deficit when silencing the VC demonstrating the specificity of the effects of AC silencing. c, Accuracy in reinforced light-off and light-on trials across days (two-way repeated measures ANOVA, Group: F (1, 140) = 0.5093, p = 0.50). d, Accuracy in reinforced light-off and light-on trials (n = 168 sessions; Wilcoxon signed rank, p = 0.41). e, Difference in accuracy in reinforced light-on versus light-off trials per session. f, Difference in accuracy in reinforced light-on versus light-off trials across days in visual cortex PV-ChR2 mice (dashed line) versus auditory cortex control mice (solid line) (two-way ANOVA, Days: F (20, 271) = 1.5547, p = 0.06, Groups: F (1, 271) = 2.3072, p = 0.13, Days*Groups: F (20, 271) = 1.1540, p = 0.2950). C_FIG Altogether, these results show that the AC is engaged during learning but is dispensable at expert levels, potentially tutoring subcortical structures that take over once the associations are learned. Unsupervised discovery of learning-related dynamics by low-rank tensor decompositionWe next sought to understand the nature and dynamics of auditory cortical activity underlying learning and performance. To do so, we performed longitudinal, two-photon calcium imaging of thousands of excitatory neurons in mice learning the auditory go/no-go task (n = 5). A separate group of water-restricted mice was passively exposed to two pure tones over the same duration but with no association with reinforcement (n = 3, see Methods; Supplementary Figure 7). This design allowed us to use the passive network as a base-case model to isolate learning-related neural dynamics. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=172 SRC="FIGDIR/small/597946v1_figs7.gif" ALT="Figure 7"> View larger version (17K): org.highwire.dtl.DTLVardef@6d72deorg.highwire.dtl.DTLVardef@1906cd7org.highwire.dtl.DTLVardef@d99914org.highwire.dtl.DTLVardef@1d1311b_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 7.C_FLOATNO Experimental design and timeline of imaging experiments. a, After surgery, animals underwent a 10-day recovery period after which water restriction started. Tonotopic mapping (tuning curve session) of the auditory cortex took place 5 days later under the two-photon microscope, followed by two days of lick training under the two-photon microscope. These two sessions also allowed for habituation to head fixation and context. Behavior sessions started the following day for 15 or 16 days, after which tonotopic mapping sessions took place at day +1, +7 and +15 post learning. b, One behavioral session consisted of three blocks of 80 or 100 trials, and a baseline session (no tone presented). Two groups of mice were imaged under the two-photon microscope: the Passive group (top; n = 3) was presented with two pure tones but was never rewarded (lick tube out), and the Learning group (n = 5) was rewarded (3{micro}l water drop) if licking in the response window after the S+ tone. Two probe blocks of 10 trials each were introduced in two of the three reinforced blocks. c, Trial structure. After a no-lick period of 1s, a 100-ms tone was played, followed by a 200-ms dead period and a[&le;] 2.5s response period. The length of the delay period was of 2s after a miss (M, no lick after S+) or a correct reject (CR, no lick after S-), 4s after a hit (H, lick after S+) and 7s after a false alarm (FA, lick after S-). C_FIG We expressed the genetically encoded calcium indicator GCaMP6f under the CaMKII pro-moter, targeting AC layer 2/3 pyramidal neurons. We imaged two planes[~] 50{micro}m apart (Fig. 2a), allowing us to record simultaneously hundreds of neurons per animal (n=7,137 neurons in 8 mice). All mice were passively presented with a series of pure tones (4 to 64kHz, quarter-octave spaced) to characterize auditory tuning properties within the local area of expression. We computed single-neuron tuning curves and then constructed a best frequency map confirming the location in the AC (Fig. 2b). For each mouse, we chose two stimuli that were similarly represented in the recorded population and were 3/4 octaves apart (Fig. 2c). We used a custom head-fixation system that allowed for kinematic registration and tracked the activity of the same neurons across weeks, including pre- and post-learning tuning curve sessions (n = 4, 643 neurons in 8 mice, see Methods; Fig. 2d-g). O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=191 SRC="FIGDIR/small/597946v1_fig2.gif" ALT="Figure 2"> View larger version (85K): org.highwire.dtl.DTLVardef@ef252borg.highwire.dtl.DTLVardef@716b28org.highwire.dtl.DTLVardef@321266org.highwire.dtl.DTLVardef@155e6a3_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 2.C_FLOATNO Low-rank tensor decomposition reveals learning-related network dynamics. a, Multi-plane, longitudinal two-photon calcium imaging of layer 2/3 excitatory network in the auditory cortex during learning (n = 5 mice) or passive exposure (n = 3 mice; see Methods). b, Tonotopic organization of the field of view of one example mouse before learning (left). Cells are colored according to their best frequency and tone-evoked responses of example cells circled in black to 17 pure tones ranging from 4 to 64 kHz are displayed on the right. c, Tone-evoked activity (top) and proportion of responsive cells (bottom) to pure tones. S+ and S- (filled and unfilled triangles, respectively) are chosen for training in the task based on their equal representation in the field of view in b. d, Six example cells tracked everyday across weeks. e, Two planes recorded in one example mouse. Cells are colored according to the number of days tracked among the 19 recording sessions in this mouse. f, Distribution of number of tracked days per cells in e. g, Cumulative distribution of tracked cells according to the percentage of recording sessions. Data for mouse in e is the light blue line. h, Calcium data is arranged by neurons x time within trial (-1 to +4s relative to tone onset, vertical line) x trials over time x trial outcomes. i, Activity from all Learning and Passive cells are concatenated together to create a fourth-order tensor (megamouse; left). In the 3rd, across trials dimension, data is aligned across mice according to learning phases: Acquisition (performance increases in probe trials), Expression (performance increases in reinforced trials), and Expert (high, stable performance in reinforced trials; see Methods and Supplementary Figure 8). j, Megamouse tensor decomposition identifies six neuronal dynamics (numbered; see Methods) that are characterized by a set of four factors: Neuron, Within trial, Across trial, and Outcome (see also Supplementary Figure 10). k, Projection of the tensor decomposition output onto principal subspace. WNr, WW r and WAr indicate neuronal, within trial and across trial weights for a component r, respectively. l, t-distributed stochastic neighbor embedding (t-SNE) projections of neuronal weights. Each dot represents a cell, colored according to the neuronal dynamic it contributed in the most. Bars (right) display the proportion of learning and passive cells among the highest contributors for each dynamic. Dynamics 1 and 2 are driven by the passive network (burgundy), while Dynamics 3 to 6 are driven by the learning network (blue). m, In the passive network, the highest contributing cells in Dynamic 1 define cell ensemble 1, and highest contributing cells in Dynamic 2 define cell ensemble 2. Similarly, in the learning network, cell ensembles 3 to 6 are constituted of the highest contributing cells to Dynamics 3 to 6, respectively. n, Absolute weights of cell ensembles across the six identified dynamics. Neurons can participate in more than one dynamic. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=145 SRC="FIGDIR/small/597946v1_figs8.gif" ALT="Figure 8"> View larger version (42K): org.highwire.dtl.DTLVardef@990420org.highwire.dtl.DTLVardef@1ddf61corg.highwire.dtl.DTLVardef@148c3e4org.highwire.dtl.DTLVardef@34ef05_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 8.C_FLOATNO Inter-subject performance alignment for megamouse tensor. a, Accuracy in probe and reinforced contexts across days of all Learning mice. b, Action rate in reinforced context across days of all Learning mice. c, Action rate in probe context across days of all Learning mice. Please note that we fixed the probe performance at the maximum discrimination that was followed by a decrease in hit rate do to extinction. d, After the alignment procedure, action rate from the megamouse (all learning mice pooled) in reinforced context across learning phases. e, Megamouse accuracy in reinforced context across learning phases. f, Accuracy difference between the start and the end of the three learning phases in probe (green) and reinforced (black) contexts. Acquisition is characterized by an increase of accuracy in probe trials (paired t-test, p = 5.47.10-4) but not in reinforced trials (paired t-test, p = 0.07), Expression corresponds to an increase of accuracy in reinforced trials (paired t-test, p = 0.008) and Expert is when accuracy in reinforced trials is high and stable (paired t-test, p = 0.27). C_FIG O_FIG O_LINKSMALLFIG WIDTH=147 HEIGHT=200 SRC="FIGDIR/small/597946v1_figs10.gif" ALT="Figure 10"> View larger version (36K): org.highwire.dtl.DTLVardef@1383264org.highwire.dtl.DTLVardef@747519org.highwire.dtl.DTLVardef@1b3d529org.highwire.dtl.DTLVardef@16f8e2f_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 10.C_FLOATNO Low-rank tensor decomposition. a, Similarity score as a function of model components. Each dot shows the similarity of a single optimization run compared to the best-fit model within each category. b, Model reconstruction error as a function of the number of components, where each dot corresponds to a different optimization run. c, Neuronal contribution (Learning vs Passive cells) per components (binomial proportion tests, all p < 0.001). d, Positive and negative neuronal weights across components in cell population recorded in learning mice (Learning) or in passive mice (Passive) (Wilcoxon tests). e, Positive and negative neuronal weights across components and individual mice. f, t-SNE of neuronal weights. Note how Learning and Passive cell populations are largely non-overlapping. g, Projection of neuronal x within trial weights of Learning and Passive network activity into principal component space. h, Projection of neuronal x within trial x trial outcome weights of Learning and Passive network activity into principal component space. i, Projection of neuronal x within trial x across trials x trial outcome (H/M and CR only) weights of Learning and Passive network activity into principal component space. C_FIG From this high-dimensional dataset, we sought to identify single neurons and neuronal ensembles carrying learning-related information, resolve stimulus and non-stimulus related activity within a given trial, identify changes in representation across trials, and determine outcome-specific dynamics. To do so, we organized our data into a 4-dimensional array containing neurons x time in trial x trials across learning x trial outcome (Fig. 2h). To identify shared and distinct variability in neuronal populations recorded in passive mice (n = 2, 339, passive network) and in learning mice (n = 2, 304, learning network), we created a megamouse by combining data from all mice and aligning neural activity to learning phase (n=4,643 neurons, see Methods; Fig. 2i; Supplementary Figure 8). We then used low-rank tensor decomposition to allow unsupervised identification of demixed, low-dimensional neural dynamics across multiple (> 2) dimensions26,27 (Supplementary Figure 9 and Supplementary Figure 10a,b; see Methods). The tensor decomposition revealed six neuronal dynamics, each characterized by the four factors of the original tensor (see Methods; Figs.2j, Supplementary Figure 10c,d, Supplementary Figure 11d). These six dynamics represented independent computations performed by the auditory cortical networks. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/597946v1_figs9.gif" ALT="Figure 9"> View larger version (18K): org.highwire.dtl.DTLVardef@dfedf2org.highwire.dtl.DTLVardef@17eb071org.highwire.dtl.DTLVardef@7205f4org.highwire.dtl.DTLVardef@1e4d76d_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 9.C_FLOATNO Tensor representation of neural data. a, Data are organized into a fourth-order tensor with dimensions NxWxAxO. Tensor decom-position approximates the data as a sum of outer products of four vectors. Each outer product contains a neuron factor (green rectangles), within trial factor (pink rectangles), across trial factor (blue rectangles) and outcome factor (purple rectangles). Each set of low-dimensional factors (i.e. component) describes the activity of group of neurons within and across trials according to trial outcomes. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=120 SRC="FIGDIR/small/597946v1_figs11.gif" ALT="Figure 11"> View larger version (22K): org.highwire.dtl.DTLVardef@fb3b34org.highwire.dtl.DTLVardef@1ec029corg.highwire.dtl.DTLVardef@19f8a32org.highwire.dtl.DTLVardef@129f839_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 11.C_FLOATNO Defining unique cell ensembles based on neuronal weights. a, Neuronal weights in the four components. Each neuron is attributed to a given dynamic according to its highest absolute weights, i.e. highest contribution. As a result, each dynamic is attributed to a unique cell ensemble (gray rectangles). b, Neuronal weights distribution before (raw, black) and after unique contribution attribution (gray). c, Learning and Passive cell proportion among components after unique attribution (binomial proportion tests). d, Learning and Passive cell proportion among components and given neuronal weight sign after unique attribution. In other words, proportion of cells from Learning and Passive networks describing the tensor-revealed neuronal dynamics (binomial proportion tests). ***p < 0.001, n.s.: not significant. C_FIG Projecting the product of the decomposition into principal component subspace showed that learning and passive networks exhibit almost orthogonal dynamics (Fig. 2k; Supplementary Figure 10f,g) and that the neural dynamics of different trial types evolved further apart in the learning network than in the passive network (Supplementary Figure 10h,i). Importantly, we ensured that the identified dynamics were not driven by isolated mice (Supplementary Figure 10e). Therefore, decomposition of the megamouse tensor discovered distinct dynamics exhibited by passive versus learning networks. For further analyses, we attributed each dynamic to individual neurons based on the neurons maximum weight ( unique participation; Fig. 2l; see Methods and Supplementary Figure 11). This allowed us to map the six dynamics onto six distinct cell ensembles, i.e. groups of neurons maximally encoding a particular network-specific dynamic (Fig. 2m and Supplementary Figure 11d). It is important to note that individual neurons (and corresponding ensembles) could exhibit mixed selectivity for the six dynamics, which allows an individual neurons to contribute to multiple, independent computations (Fig. 2n). Learning counteracts tone-evoked habituation by maintaining stimulus selectivity in distinct cell populationsA prevailing view in sensory systems holds that sensory cortices subserve associative learning through plasticity of the cue representation5,28-36. This model posits that individual neurons (via changes in sensory tuning) and neural populations (via cortical map expansion) enhance the representation of behaviorally relevant cues for use by downstream regions37-39. These studies, however, measure neural tuning and map expansion outside of the task context in a pre and post learning design and infer that plasticity of cue representations reflects the mechanistic role of the sensory cortex. To assess this model, we initially focused on the cell ensembles that exhibited classical stimulus-evoked activity (Fig. 2j), namely cell ensembles 1-4. We observed a prominent signature of stimulus-evoked habituation over hundreds to thousands of trials. This habituation dominated activity in passive networks, as seen in cell ensembles 1 and 2 which represented[~] 77% (1, 803/2, 339) of all passive cells (Fig. 3a,d). These neurons exhibited stimulus-evoked activation (cell ensemble 1) or suppression (cell ensemble 2), both of which decreased in amplitude over time (Fig. 3b-c,e-f). These cell ensembles were not stimulus selective and displayed the same dynamic in both stimulus 1 (S1) and stimulus 2 (S2) trials (Fig. 3b,e). These ensembles thus reflected the broad-based suppression of non-selective neurons after long-term repeated presentation of the same sounds. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=166 SRC="FIGDIR/small/597946v1_fig3.gif" ALT="Figure 3"> View larger version (53K): org.highwire.dtl.DTLVardef@678173org.highwire.dtl.DTLVardef@163d156org.highwire.dtl.DTLVardef@447847org.highwire.dtl.DTLVardef@1349c3a_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 3.C_FLOATNO Learning counteracts tone-evoked habituation by maintaining stimulus selectivity in distinct populations. a, Representation of cell ensemble 1 in the Passive network. b, Average activity of cell ensemble 1 in S1 (black) and S2 (gray) trials across time in 80-trial blocks. Black triangles indicate tone onset, gray lines delimit averaged trial blocks. Black dashed lines separate time phases indicated by light to dark gray rectangles at the top: early, middle and late (see Methods). c, Cell ensemble 1 tone-evoked calcium responses across time phases for S1 and S2 trials combined (Friedman test, p = 1.26.10-291). d, Representation of cell ensemble 2 in the Passive network. e, Average activity of cell ensemble 2 in S1 and S2 trials across time. f, Cell ensemble 2 tone-evoked calcium responses across time phases for S1 and S2 trials combined (Friedman test, p = 7.32.10-121). g, Representation of cell ensemble 3 in the Learning network. h, Average activity of cell ensemble 3 in hit (green) and CR (yellow) trials across learning in 80-trial blocks. Black triangles indicate tone onset, gray lines delimit averaged trial blocks. Black dashed lines separate learning phases indicated by colored rectangles at the top: Acquisition, Expression and Expert (see Methods). i, Representation of cell ensemble 4 in the Learning network. j, Average activity of cell ensemble 4 in hit and CR trials across learning. k, Response index (response probability over learning; see Methods) of cell ensembles 1 and 2 (red) vs cell ensembles 3 and 4 (blue) (Wilcoxon test, p = 1.23.10-30). l, Selectivity index (see Methods) of cell ensembles 1 and 2 (red) vs cell ensembles 3 and 4 (blue) (Wilcoxon test, p = 1.37.10-94). m, Pre (top raw) and post (bottom raw) learning tonotopic maps (left), after spatial binning (middle) and restricted to surface with S+ (filled triangle) and S- (open triangle) best frequency (right) of one example mouse. n, Change in surface representation of S+ and S- pre- vs post-task learning (Learning) or pre- vs post-passive exposure (Passive) (binomial proportion tests). o, Pre vs post-learning change in percentage of neurons responsive to S+ and S-(binomial proportion tests). p, Pre vs post-learning change in tone-evoked responses of pre-task S+ and S- responsive neurons (KW test, p = 2.77.10-5). q, Pre- vs post-learning comparison of local best frequency differences in tonotopic maps. r, Distribution of local differences (from difference maps in q) in Learning versus Passive. median {+/-} s.e.median; *p < 0.05; **p < 0.01; ***p < 0.001, n.s.: not significant. C_FIG Stimulus-evoked responses in learning networks were observed in cell ensembles 3 and 4 (Fig. 3g-j). This includes a high selectivity for the S- (cell ensemble 3) or S+ (cell ensemble 4) cues (Fig. 3g-j). Cell ensemble 3 consisted of 19% of the Learning cell population (Fig. 3g), and displayed a slight habituation but mainly a strong preference for the S- throughout learning (Fig. 3h), while cell ensemble 4 (12% of total learning cells; Fig. 3j) exhibited S+ selectivity throughout learning (Fig. 3j). Cell ensembles 3 and 4 were more tone responsive and tone selective than cell ensembles 1 and 2 (Fig. 3k,l). Stimulus-evoked activity analyses across days of all recorded neurons (n = 7, 137) also support these results (Supplementary Figure 12, Supplementary Figure 13). Therefore, learning counteracted tone-evoked habituation by maintaining distinct ensembles that encoded either the S+ or S- selectively. O_FIG O_LINKSMALLFIG WIDTH=154 HEIGHT=200 SRC="FIGDIR/small/597946v1_figs12.gif" ALT="Figure 12"> View larger version (69K): org.highwire.dtl.DTLVardef@fcbe19org.highwire.dtl.DTLVardef@1247daaorg.highwire.dtl.DTLVardef@b654b2org.highwire.dtl.DTLVardef@729d91_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 12.C_FLOATNO Evolution of tone-evoked responses across days. a, Tone-evoked responses to S+ and S- in Learning mice across days for all cells recorded. b, Tone-evoked responses to S1 and S2 in Passive mice across days for all cells recorded. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=33 SRC="FIGDIR/small/597946v1_figs13.gif" ALT="Figure 13"> View larger version (10K): org.highwire.dtl.DTLVardef@c6a200org.highwire.dtl.DTLVardef@b5b179org.highwire.dtl.DTLVardef@96ad75org.highwire.dtl.DTLVardef@5604a2_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 13.C_FLOATNO Learning counteracts tone-evoked habituation. a, Proportion of tone-responsive cells across days among Passive and Learning cells. b, Averaged proportion of tone-responsive cells in Passive and Learning networks (mean {+/-} s.e.m.; t-test, p = 3.89.10-5). c, Proportion of tone-responsive cells in days 1-5 versus days 11-15 in Learning and Passive networks (mean {+/-} s.e.m.; two-way ANOVA, Time x Group, p = 1.73.10-7). d, Proportion of cells responsive to S+ and S- in Learning network and S1, S2 or S1 or S2 (S) in Passive network. e, Averaged proportion of cells responsive to S+, S- or S (mean {+/-} s.e.m.; ANOVA, p = 1.93.10-6). C_FIG Learning was not associated with cortical map expansionTo directly test representational expansion and tuning shifts, we conducted a series of analyses focusing on stimulus-evoked responses before (pre-task) and after (post-task) learning, akin to classical measures of tuning and tonotopy. We computed the change in surface area occupied by S+ and S- preferring cells in tuning curve sessions, outside the task (Fig. 3m). Surprisingly, we observed no increase in the map-level representation of the S+ or S- after learning, and instead, observed a modest decrease (Fig. 3m-n). In addition to the best frequency representation, the fraction of neurons responding to the S+ and S- decreased (Fig. 3o) and the response amplitude of neurons that were initially tuned to the S+ and S- was lower after learning (Fig. 3p). Interestingly, while we observed no increase in representation to the S+ and S-, learning networks favored the representation of frequencies in between S+ and S-, but not higher or lower as seen in passive networks (Fig. 3n). Finally, using our passive networks as a base-case comparison, we calculated the local changes in the tonotopic map structure (Fig. 3q). Learning networks were surprisingly stable and exhibited less local changes than passive networks (Fig. 3r). These pre- vs post-learning changes in responsiveness and tonotopy thus mirrored the responsiveness observed online during learning (in dynamics 1 and 2) in a stable, tracked network (n=4,643 neurons, Fig. 3a-l), as well as when we include all neurons from each session (n=7,137 neurons) (Supplementary Figure 13). Altogether, our results suggest that cortical map expansion and changes in single-neuron tuning are unlikely to be the substrate for associative learning40,41. Tone-restricted silencing only partially impairs learning and performanceWe next sought to understand the extent to which the maintenance of stimulus-selectivity by learning networks was important to learning and performing the task. We performed daily bilateral silencing of AC during stimulus presentation throughout learning (Supplementary Figure 14a). Tone-restricted AC silencing impaired task performance throughout learning (Supplementary Figure 14b-e), task acquisition (Supplementary Figure 14f-i), and online performance during learning, with gradual fading of the effect at expert performance (Supplementary Figure 14n-q). Accuracy and action rate were not affected in reinforced light-off trials (Supplementary Figure 14j-k), but PV-ChR2 mice lick more and faster to the S- (Supplementary Figure 14l-m), suggesting that tone-restricted AC silencing also impaired expression, but to a lesser extent than full-trial silencing. Altogether, these results showed that information carried by the AC network in the tone-evoked window is used during learning. Interestingly, tone-restricted silencing impacted learning less than full trial silencing across nearly all measures (Fig. 1, Supplementary Figure 14), suggesting that activity after the tone-evoked window was critical for rapid contingency acquisition and performance during learning. O_FIG O_LINKSMALLFIG WIDTH=186 HEIGHT=200 SRC="FIGDIR/small/597946v1_figs14.gif" ALT="Figure 14"> View larger version (51K): org.highwire.dtl.DTLVardef@14a2aedorg.highwire.dtl.DTLVardef@486bd0org.highwire.dtl.DTLVardef@9e6054org.highwire.dtl.DTLVardef@1c64aea_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 14.C_FLOATNO AC silencing restricted to sound presentation impairs audiomotor learning and on-line performance during learning. a, Probabilistic optogenetic silencing of the auditory cortex during learning. Light-on periods were restricted to sound presentation only (see Methods). b, Accuracy in reinforced light-on trials (two-way ANOVA, Days: F (17, 86) = 5.4950, p < 10-7; Groups: F (1, 86) = 50.5343, p < 10-9; Days*Groups: F (17, 86) = 0.70700, p = 0.79). c, Action rate in reinforced light-on trials (HIT, two-way ANOVAs, HIT, Days: F (17, 86)10.68010, p < 10-14; Groups: F (1, 86) = 0.0200, p = 0.89; Days*Groups: F (17, 86) = 1.0647, p = 0.40; FA, Days: F (17, 86) = 2.7330, p = 0.0012; Groups: F (1, 86) = 41.5010, p < 10-8; Days*Groups: F (17, 86) = 0.7255, p = 0.77). d, False alarm lick rate in reinforced light-on trials (two-way ANOVA, Days: F (17, 86) = 0.8663, p = 0.6140; Groups: F (1, 86) = 89.3004, p < 10-14; Days*Groups: F (17, 86) = 3.2285, p < 10-3). e, False alarm lick latency in reinforced light-on trials (two-way ANOVA, Days: F (17, 86) = 2.0216, p = 0.018; Groups: F (1, 86) = 251.7387, p < 10-26; Days*Groups: F (17, 86) = 4.8600, p < 10-6). f, Accuracy in probe light-off trials (two-way ANOVA, Days: F (5, 30) = 8.3041, p < 10-4; Groups: F (1, 30) = 4.7288, p = 0.038; Days*Groups: F (5, 30) = 0.7288, p = 0.619). g, Action rate in probe light-off trials (two-way ANOVAs, HIT, Days: F (5, 30) = 5.4632, p = 0.0011; Groups: F (1, 30) = 6.3510, p = 0.017; Days*Groups: F (5, 30) = 1.2158, p = 0.33; FA, Days: F (5, 30) = 5.5019, p = 0.0010; Groups: F (1, 30) = 0, p = 1; Days*Groups: F (5, 30) = 1.1320, p = 0.37). h, HIT lick latency in probe light-off trials (two-way ANOVA, Days: F (5, 29) = 6.0308, p < 10-3; Groups: F (1, 29) = 10.3058, p = 0.0032; Days*Groups: F (5, 29) = 0.1542, p = 0.98). i, Maximal difference between hit and false alarm rates in probe light-off trials over the first 6 days (t-test, p = 0.40). j,Accuracy in reinforced light-off trials (two-way ANOVA, Days: F (17, 86) = 8.3579, p < 10-11; Groups: F (1, 86) = 1.6832, p = 0.20; Days*Groups: F (17, 86) = 0.2356, p = 1). k, Action rate in reinforced light-off trials (two-way ANOVAs, HIT, Days: F (17, 86) = 11.1314, p < 10-14; Groups: F (1, 86) = 2.1423, p = 0.15; Days*Groups: F (17, 86) = 0.9107, p = 0.56; FA, Days: F (17, 86) = 4.2760, p < 10-5; Groups: F (1, 86) = 0.5043, p = 0.48; Days*Groups: F (17, 86) = 0.3026, p = 1). l, FA lick latency in reinforced light-off trials (two-way ANOVA, Days: F (17, 78) = 1.7364, p = 0.053; Groups: F (1, 78) = 9.0848, p = 0.0035; Days*Groups: F (17, 78) = 1.3749, p = 0.17). m, FA lick rate in reinforced light-off trials (two-way ANOVA, Days: F (17, 78) = 0.7983, p = 0.69; Groups: F (1, 78) = 13.4564, p < 10-3; Days*Groups: F (17, 78) = 1.4494, p = 0.14). n, Comparison of light-off versus light-on trials to measure auditory cortex silencing effect on on-line performance. o, Session density plot of accuracy in reinforced light-on against light-off. Top, control; bottom, PV-ChR2. p, Accuracy in light-on reinforced trials from day where FA< 0.3 in light-off reinforced trials. Note the general trend for ChR2 mice (gray lines) to increase accuracy (positive slopes), i.e. performance impairment fades away. q, Within subject difference between accuracy in reinforced light-on and light-off aligned to the day where false alarm rate < 0.3 in reinforced light-off. C_FIG Rapid emergence of reward prediction activity in the auditory cortexThe sensory cortex is widely considered to be specialized for perception by interpreting complex sensory objects42,43 or adjusting representations of behaviorally-relevant stimuli2,33,37,44,45. Recent evidence, however, suggests that sensory cortical neurons directly encode non-sensory variables such as movement46-49, reward timing50-53, expectation54,55, and context23,45,56-63. Conjoint representations of sensory and non-sensory variables in the same network could further hone perception or, alternatively, subserve more integrative associative processes. Inspection of the within-trial dynamics of learning-driven cell ensembles 5 and 6 suggested that these neurons exhibited non-canonical activity in the form of a signal that occurred late in the trial, delayed from the tone-evoked response (Fig. 2j). This late-in-trial signal increased over learning and was trial type selective (Fig. 2j). We next sought to further explore the encoding properties of these two cell ensembles. Cell ensemble 5 (n = 155 cells from the learning network), exhibited late-in-trial activity on hit trials (licking to the S+) that increased with learning (Fig. 4a). This delayed activity was not apparent on correct S-trials (correct reject, CR), where neurons exhibited classical stimulus-evoked response that habituated over learning (Fig. 4b). O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=179 SRC="FIGDIR/small/597946v1_fig4.gif" ALT="Figure 4"> View larger version (76K): org.highwire.dtl.DTLVardef@16050bdorg.highwire.dtl.DTLVardef@54ba70org.highwire.dtl.DTLVardef@9c1355org.highwire.dtl.DTLVardef@b96c3b_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 4.C_FLOATNO Rapid emergence of reward prediction encoding drives learning. a, Heat map of cell ensemble 5 activity (n = 155 cells) across learning phases (delimited by horizontal white dashed lines) in hit trials (20-trial blocks). White trace represents the average trial trace. Inserts (right) show average activity at time indicated by black triangles. Colored rectangles indicate learning phases: Acquisition (green), Expression (black) and Expert (blue). b, Heat map of cell ensemble 5 activity across learning phases (delimited by horizontal white dashed lines) in CR trials (20-trial blocks). c, Heat map of the activity of a fraction of cells from cell ensemble 5 (n = 20 cells) from one example mouse across consecutive S+ trials. Black dots indicate licks. Trial outcome is represented on the right (green circle: hit; blue stars: miss). d, Cell ensemble 5 activity in hit vs miss trials (time and number matched, see Methods and Supplementary Figure 15a). e, Area under the curve (AUC) quantification of data in gray rectangle in d (Wilcoxon signed rank test, p = 6.78.10-21). f, Procedure of reinforced and probe hit trial (H) matching. g, Average cell ensemble 5 activity in reinforced hit trials immediately before (black) or after (gray) probe hit trials (green). h, AUC quantification of data in h (Friedman test, p = 0.3071). i, Lick PSTHs in reinforced hit trials immediately before (black) or after (gray) probe hit trials (green). j, Quantification of number of licks in 1-s window post-tone (KW test, p = 3.18.10-56). k, Average activity of cell ensemble 5 over the first five blocks of 40-reinforced hit trials in learning. l, Late peak activity in HIT trials across learning phases of cell ensemble 5 (green) and low weighted cells (null, black). m, Procedure of reinforced and probe FA trial (fa) matching (top) and corresponding local accuracy quantification (bottom; see Methods; repeated measures ANOVA, p = 3.16.10-4). n, Average cell ensemble 5 activity in FA trials in the probe, non-reinforced context (orange). AUC late-in-trial (gray rectangle) compared to zero (Wilcoxon signed rank test, p = 1.46.10-8). o, Average activity of cell ensemble 5 (n = 51 cells) from one example mouse in FA trials in the reinforced context (n = 423) after classification based on the detection of a reward prediction signal. Bottom, average activity of FA trials with (RP+, n = 101) or without (RP-, n = 322) reward prediction signal, and activity during FA trials in the probe context (n = 19 trials, orange) reflecting knowledge errors (see also Supplementary Figure 16). p, Heat map of the activity of a fraction of cells from cell ensemble 5 (n = 51 cells) from the same example mouse in o across consecutive FA trials in the reinforced context. Identification of a RP signal is represented by a black dot (right). q, Distribution of RP+ and RP- FA trials over learning in learning mice (binomial proportion tests, Acquisition, p = 1.65.10-7, Expression, p = 3.32.10-10, Expert, p = 0.22). r, Trial-specific closed-loop optogenetic AC inactivation over learning. s, Performance index (left, see Methods; two-way ANOVA, p = 2.11.10-21) and hit lick latency (right; two-way ANOVA, p = 0.013) in probe context in post-hit silencing experiments. t, Performance index (left, see Methods; two-way ANOVA, p = 6.36.10-5) and hit lick latency (right; two-way ANOVA, p = 0.008) in probe context in post-FA silencing experiments. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC="FIGDIR/small/597946v1_figs15.gif" ALT="Figure 15"> View larger version (26K): org.highwire.dtl.DTLVardef@170c926org.highwire.dtl.DTLVardef@1a60375org.highwire.dtl.DTLVardef@2cfc4corg.highwire.dtl.DTLVardef@168027e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 15.C_FLOATNO Emergence of reward prediction signal. a, Procedure of hit and miss trial matching. b, Heat map of members of cell ensemble 5 (n = 105) activity aligned to lick bout onset outside task events in day 1 of training. Lick PSTH is represented above. c, Quantification of z-scored calcium activity 1s pre- vs 1s post-lick bout onset (Wilcoxon test, p = 0.11). d, Average cell ensemble 5 activity in reinforced hit (green) and FA (orange) trials over Expression phase. e, Lick PSTHs aligned to tone onset of FA trials in Expression and hit trials in probe context. f, Cell ensemble 5 activity over the first 300 hit trials (20-trial blocks). Only significant activity (and higher than null population, see Methods) is represented. Note the emergence of a stable late-on-trial signal after 40 hit trials onwards. g, Quantification of Fig. 4l, i.e. evolution of late-in-trial signal of cell ensemble 5 across learning, taking first and last two 40-hit trial blocks (KW test, p = 1.05.10-23). *p < 0.05, **p < 0.01, ***p < 0.001, n.s.: not significant. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=92 SRC="FIGDIR/small/597946v1_figs16.gif" ALT="Figure 16"> View larger version (26K): org.highwire.dtl.DTLVardef@17bc3d8org.highwire.dtl.DTLVardef@76c09forg.highwire.dtl.DTLVardef@601c80org.highwire.dtl.DTLVardef@1ef45da_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 16.C_FLOATNO Reward prediction signal on error trials. a, Classification of hit versus CR trials in the reinforced context from the AUC post-tone of a fraction of cell ensemble 5 (n = 51) recorded in the example mouse showed in Fig. 4p,q. Right: posterior probability of being part of CR class. b, Proportion of RP+ and RP- FA trials from the example mouse showed in Fig. 4o,p. c, No difference in lick latency was observed between RP+ and RP- FA trials (Wilcoxon test, p = 0.83). d, AUC quantification of RP+, RP- and probe FA trials (KW, p = 9.76.10-28). e, Proportion of RP+ among all FA trials and misclassification rate in each learning mice. *p < 0.05, **p < 0.01, ***p < 0.001, n.s.: not significant. C_FIG To understand the nature of the late-in-trial activity, we exploited our multiple trial types to disambiguate the contribution of sensory, motor, and reward signals. To assess whether the late-in-trial signal was a delayed form of sensory activity, we compared activity in hit trials to activity in trials where the same stimulus was presented but the mice did not lick and did not get rewarded (miss trials, Figs.1a and 4c-e). To ensure an appropriate comparison between hit and miss trials, we generated a balanced set of trials that were matched in number (given that miss trials were less frequent) and occurred within the same time period (given that the signal amplitude evolved with learning) (Supplementary Figure 15a). Cell ensemble 5 did not exhibit late-in-trial activity on miss trials (Fig. 4c-e), discarding the possibility that it reflected a delayed sensory response. We then asked whether this activity reflected reward consumption. We compared cell ensemble activity during hit trials in the reinforced context to the activity during hit trials in the probe context (Fig. 4f), where the mice expected reward and thus correctly licked to the S+ but the reward was omitted (Fig. 1b). We matched the number of trials between reinforced and probe contexts and controlled for within-session and across-session changes by comparing probe hit trials to reinforced hit trials immediately before and after the probe block (Fig. 4f). Strikingly, late-in-trial activity was preserved in probe trials (Fig. 4g,h), indicating that it did not reflect reward consumption. Finally, although movement has been reported to decrease auditory cortical activity46,64-66, we sought to understand the degree to which this late-in-trial signal could be driven by licking itself. To do this, we first exploited probe hit trials where the lick rate was strongly reduced compared to reinforced hit trials (Fig. 4i,j). We observed no difference in the late-in-trial neural signal and could thus conclude that the signal was not due to ongoing licking (Fig. 4i,j). Second, we tested the possibility that this late-in-trial signal was driven by the initiation of a lick bout as compared to the ongoing licking activity. We isolated spontaneous lick bouts in between training blocks and observed that the cell ensemble was not lick-responsive (Supplementary Figure 15b,c). In addition, if lick initiation drove this activity, we would also expect to see it on false alarm trials (incorrect licking to the S-). For this analysis, we focused on false alarms that occurred after task acquisition, as these errors are unlikely to be errors due to imperfect task knowledge. We observed no systematic late-in-trial activity on these trials (Supplementary Figure 15d) even though the licking pattern in false alarm trials was similar to that during probe hit trials (Supplementary Figure 15e). Taken together, the late-in-trial activity did not reflect stimulus, reward consumption, licking, nor lick initiation. Instead, these results showed that cell ensemble 5 encoded the higher-order process of reward prediction (RP). We next sought to identify the precise moment when a contingency is formed by identifying the trials when this reward prediction signal emerged. Initially, these neurons exhibited classical tone-evoked responses but then abruptly and within only 40 hit trials, developed a robust reward prediction activity (Fig. 4k, Supplementary Figure 15f). This reward prediction signal continued to develop over Acquisition, strengthened during Expression, and then surprisingly receded at Expert level when learning is nominally complete (Fig. 4a,l, Supplementary Figure 15g). This longitudinal temporal dynamic mirrored our optogenetic results which demonstrates that the AC is the default pathway for learning but then becomes dispensable at expert levels. Altogether, these results show that a reward prediction signal rapidly emerges at the timescale of Acquisition in auditory cortical networks. Revealing the underlying cognitive drivers of errorsIdentifying the cognitive drivers of errors is particularly challenging during learning 4. Errors during learning are typically considered mistakes while discriminative contingencies (task knowledge) are still forming. However, errors arise not only from knowledge-related mistakes (for which animals incorrectly expect reward), but also from factors such as impulsivity, disengagement, and exploration (for which animals do not expect reward). While detailed behavioral inspection has been a promising route to uncover the nature of errors11, an alternative approach is to use neural activity itself. Given our findings of reward prediction encoding on correct trials, we hypothesized that the same signal would be present when animals make knowledge-related errors, when animals incorrectly expected rewards on S- trials. To address this, we first focused on the occasional false alarms (FA) that occurred during probe trials, as they reflected errors of task knowledge (Fig. 4m)7. Strikingly, we observed a robust reward prediction activity in these trials (Fig. 4n), strongly suggesting that animals were indeed expecting reward. We next reasoned that such knowledge errors should be present not only on probe trials, but also in a subset of reinforced trials, interspersed with non-knowledge errors. We classified individual FA trials in the reinforced context based on the presence of a reward prediction signal (see Methods; Supplementary Figure 16a). We identified a significant proportion of trials that exhibited robust reward prediction activity, but also many that did not (Fig. 4o, Supplementary Figure 16b). The reward prediction signal was identical to that observed in probe trials (Fig. 4o, Supplementary Figure 16d), providing further confidence that these were indeed knowledge errors. These data suggest that we could isolate knowledge errors using neural data, which was not possible from behavioral inspection alone (Supplementary Figure 16c). Interestingly, we found that knowledge errors were interspersed with errors that did not elicit reward prediction activity (Fig. 4p). Finally, we hypothesized that knowledge errors should predominantly occur during the Acquisition phase of behavior, when animals are still learning the discriminative contingencies. We computed the fraction of RP+ (knowledge-related errors) and RP-(non-knowledge errors) over time and found that RP+ errors peaked during the Acquisition phase of learning, and rar

neuroscience↗

Contributions and synaptic basis of diverse cortical neuron responses to task performance

Neuronal responses during behavior are diverse, ranging from highly reliable classical responses to irregular or seemingly-random non-classically responsive firing. While a continuum of response properties is frequently observed across neural systems, little is known about the synaptic origins and contributions of diverse response profiles to network function, perception, and behavior. Here we use a task-performing, spiking recurrent neural network model incorporating spike-timing-dependent plasticity that captures heterogeneous responses measured from auditory cortex of behaving rodents. Classically responsive and non-classically responsive model units contributed to task performance via output and recurrent connections, respectively. Excitatory and inhibitory plasticity independently shaped spiking responses and task performance. Local patterns of synaptic inputs predicted spiking response properties of network units as well as the responses of auditory cortical neurons from in vivo whole-cell recordings during behavior. Thus a diversity of neural response profiles emerges from synaptic plasticity rules with distinctly important functions for network performance.

neuroscience↗