Representation of emotional expressions across the face and voice brain networks
Recognizing emotion expressions from facial and vocal signals is crucial for optimal social interactions. We comprehensively characterized regions within the face and voice processing networks to probe their contribution to unimodal and crossmodal emotion representation. Using multivariate pattern classification analyses, we found that emotional expressions could be reliably decoded from the dominant sensory modality in all individually defined face- and voice-selective regions. Emotion expressions from the non-dominant modality (vocal expressions in the face networks, and vice versa) could also be decoded in all areas except the occipital face areas. A shared neural code for facial and vocal emotions is implemented in the temporal voice areas (TVA), the posterior superior temporal sulcus (pSTS) and the precentral gyrus (PCG). The simultaneous presentation of congruent facial and vocal expressions elicited distinct activity patterns across most regions, highlighting that multisensory inputs reshape brain responses relative to unisensory stimulation across the entire face and voice brain network. These findings suggest that face and voice-selective regions broadly encode emotion expressions within and even across the senses, relying on a multisensory gradient that converges in temporo-frontal regions where their distributed multisensory responses align to create a supramodal representation of specific emotion expressions.