Search bioRxiv⌕ Search

Biology subjects

Winchester, M. M.

Publications and source records attributed to Winchester, M. M..

4 recordsLinked to original sources

Exploring the impact of social relevance on the cortical tracking of speech: viability and temporal response characterisation

Human speech is inherently social. Yet our understanding of the neural substrates underlying continuous speech perception relies largely on neural responses to monologues, leaving substantial uncertainty about how social interactions shape the neural encoding of speech. Here, we bridge this gap by studying how EEG responses to speech change when the input includes a social element. In Experiment 1, we compared the neural encoding of synthesised undirected monologues, directed monologues, and dialogues. In Experiment 2, we extended this by using podcasts, addressing the additional challenges of real speech dialogue, such as dysfluency. Using temporal response function analyses, we show that the presence of a social component strengthens the cortical tracking of the speech envelope, despite identical acoustic properties. Neural responses to synthesised speech showed a strong correlation with those for real speech podcasts, with a stronger alignment emerging for more socially-relevant speech material. In addition, we demonstrate that robust neural indices of sound and lexical-level processing can be derived using real podcast recordings despite the presence of dysfluencies. Finally, we present a simulation to put to the test the robustness of temporal response function analyses under increasing levels of dysfluency. Together, these findings highlighting the impact of social elements in shaping auditory neural processing, providing a framework for future investigation and analysis of social speech listening and speech interaction. Significance StatementHuman speech is rarely produced or processed in a social vacuum. Yet, our understanding of continuous speech neurophysiology mostly comes from experiments involving speech monologues. This study reveals how social context modulates the neural encoding of speech. We directly contrast neural signals recorded when participants listened to monologues and dialogues, using controlled material from speech synthesis and real podcast recordings. We found that the social element amplifies the neural encoding of speech features, reflecting greater engagement. We also show strong correlation between synthetic and real podcast neural responses, scaling with social relevance. Finally, we demonstrate that lexical processing can be measured robustly even amid natural dysfluencies. These insights advance our understanding of speech neurophysiology, informing future research on social speech.

neuroscience↗

Where is the melody? Spontaneous attention orchestrates melody formation during polyphonic music listening

Humans seamlessly process multi-voice music into a coherent perceptual whole. Yet the neural strategies supporting this experience remain unclear. One fundamental component of this process is the formation of melody, a core structural element of music. Previous work on monophonic listening has provided strong evidence for the neurophysiological basis of melody processing, for example indicating predictive processing as a foundational mechanism underlying melody encoding. However, considerable uncertainty remains about how melodies are formed during polyphonic music listening, as existing theories (e.g., divided attention, figure-ground model, stream integration) fail to unify the full range of empirical findings. Here, we combined behavioural measures with non-invasive electroencephalography (EEG) to probe spontaneous attentional bias and melodic expectation while participants listened to two-voice classical excerpts. Our uninstructed listening paradigm eliminated a major experimental constraint, creating a more ecologically valid setting. We found that attention bias was significantly influenced by both the high-voice superiority effect and intrinsic melodic statistics. We then employed transformer-based models to generate next-note expectation profiles and test competing theories of polyphonic perception. Drawing on our findings, we propose a weighted-integration framework in which attentional bias calibrates the overall degree of integration of the competing streams. In doing so, the proposed framework reconciles previous divergent accounts by showing that, even under free-listening conditions, melodies emerge through an attention-guided statistical integration mechanism. HighlightsO_LIEEG can be used to decode spontaneous attention during the uninstructed listening of polyphonic music. C_LIO_LIBehavioural and neural data indicate that spontaneous attention is influenced by both high-voice superiority and melodic contour. C_LIO_LIAttention bias impacts the neural encoding of the polyphonic streams, with strongest effects within 200 ms after note onset. C_LIO_LIStimuli that produced a stronger attention bias aligned with monophonic-model expectations, whereas stimuli with a weaker bias aligned with the Stream-Integration model. C_LIO_LIWe propose a bi-directional influence between attention and prediction mechanisms, with horizontal statistics impacting attention (i.e., salience), and attention impacting melody extraction. C_LI

neuroscience↗

No Risky Bets: The Brain Avoids All-In Predictions During Naturalistic Multitalker Listening

Listeners daily adapt to new talkers, yet how acoustic variability challenges the brains predictive mechanisms during speech processing remains unclear. Here, we used EEG and Temporal Response Function to examine neural responses to continuous speech narrated by a single talker (Single), enabling stable acoustic model formation, or multiple talkers (Multi), introducing acoustic uncertainty. We assessed whether talker variability influences phoneme recognition and predictive processing, indexed by neural responses to phonemes, phonemic and semantic surprisal. In the Multi condition, responses to phonemes increased but responses to phonemic surprisal decreased, indicating greater speech perception demands and weaker phonemic predictions. Conversely, semantic surprisal responses were stronger, suggesting increased reliance on lexical-semantic predictions. These findings reveal a trade-off in the brains predictive mechanisms, where acoustic uncertainty reduces lower-level phonemic anticipation but promotes higher-level semantic prediction. Adaptive processing underscored the brains ability to dynamically adjust predictions across linguistic levels, promoting speech comprehension in variable environments. SignificanceHumans attend to many different talkers daily, switching between them apparently without any effort. This stems from the brains ability to construct adaptive, probabilistic models of speech. In this study, we provide novel evidence that predictions in language comprehension are sensitive to acoustic uncertainty. Specifically, under conditions of increased acoustic uncertainty, listeners rely more on bottom-up information in phoneme recognition and reduce the anticipation of phonemic information. This is accompanied by enhanced semantic prediction, suggesting that listeners compensate for acoustic uncertainty by increasing reliance on higher-level contextual representations. This flexible approach allows for robust comprehension, highlighting the brains capacity to dynamically adjust its predictive processing to accommodate varying acoustic environments. TeaserTalker-driven acoustic uncertainty reduces phoneme predictions but boosts semantic inference.

neuroscience↗

Robust assessment of the cortical encoding of word-level expectations using the temporal response function

Speech comprehension involves detecting words and interpreting their meaning according to the preceding semantic context. This process is thought to be underpinned by a predictive neural system that uses that context to anticipate upcoming words. Recent work demonstrated that such a predictive process can be probed from neural signals recorded during ecologically-valid speech listening tasks by using linear lagged models, such as the temporal response function. This is typically done by extracting stimulus features, such as the estimated word-level surprise, and relate such features to the neural signal. While modern large language models (LLM) have led to a substantial leap forward on how word-level features and predictions are modelled, there has been little progress made towards the metrics used for evaluating how well a model is relating stimulus features and neural signals. In fact, previous studies relied on evaluation metrics that were designed for studying continuous univariate sound features, such as the sound envelope, without considering the different requirements of word-level features, which are discrete and sparse in nature. As a result, studies probing lexical prediction mechanisms in ecologically-valid experiments typically exhibit small effect-sizes, severely limiting the type of observations that can be drawn and leaving considerable uncertainty on how exactly our brains build lexical predictions. First, the present study discusses and quantifies these limitations on both simulated and actual electroencephalography signals capturing responses to a speech comprehension task. Second, we tackle the issue by introducing two assessment metrics for the neural encoding of lexical surprise that substantially improve the state-of-the-art. The new metrics were tested on both the simulated and actual electroencephalography datasets, demonstrating effect-sizes over 140% larger than those for the vanilla temporal response function evaluation.

neuroscience↗