Search bioRxiv⌕ Search

Biology subjects

Andreas, J.

Publications and source records attributed to Andreas, J..

3 recordsLinked to original sources

WhaleLM: Finding Structure and Information in Sperm Whale Vocalizations and Behavior with Machine Learning

Language models (LMs), which are neural sequence predictors trained to model distributions over natural language texts, have come to play a central role in human language technologies like machine translation and information retrieval. They have also contributed to the scientific study of human language itself, enabling progress on long-standing questions about the learnability, optimality, and universality of key features of human languages. Many analogous questions exist in the study of communication between non-human animals--for which, in many cases, we have only a preliminary understanding of signals structure and use. Can neural sequence models help us understand these animal communication systems as well? We use these models to characterize the structure and information content of sperm whale vocalizations. Sperm whales (Physeter macrocephalus) engage in complex, coordinated behaviours like foraging and navigation in the darkness of the ocean while exchanging sequences of rhythmic clicks known as codas. However, little is known about whether there are any systematic patterns governing coda production, or how codas influence group decision-making and behaviour. To begin to answer these questions, we first train a neural sequence model (a sperm whale language model) to predict whales future vocalizations from their conversational history. By systematically manipulating the information available to this model, and measuring the change in predictive accuracy, we show that sperm whale vocalizations exhibit order dependence, long-range dependencies on up to the past eight codas in an exchange, and predictable turn-taking. Second, we train the sequence model to predict whales behaviour from their vocal exchanges, and find that both current behavioural context and future actions are predictable, with accuracies of 72% and 86% respectively, from coda sequences. Our study provides the first evidence that sperm whale vocalizations contain information that could be used to coordinate behaviour. More generally, it offers a framework for using modern machine learning tools for hypothesis generation and to assist in investigating the structure and function of unknown communication systems.

animal behavior and cognition↗

Contextual and Combinatorial Structure in Sperm Whale Vocalisations

Sperm whales (Physeter macrocephalus) are long-lived and highly social mammals that engage in complex group behaviours, including navigation, foraging, and child-rearing. During these behaviours, sperm whales communicate primarily using sequences of short bursts of clicks with varying inter-click intervals, known as codas. Past research has identified around 150 discrete coda types globally, with 21 in the Caribbean. A subset of these have been shown to encode information about caller and clan identity. However, almost everything else about the sperm whale communication system, including basic questions about its structure and information-carrying capacity, remains unknown. In this study, we show that codas exhibit contextual and combinatorial structure with key similarities to aspects of human language and other primate communication systems. First, we report previously undescribed variations in coda structure that are sensitive to the conversational context in which they occur. We call these rubato and ornamentation, by analogy to musical terminology. These variations are systematically controlled and imitated across individual whales. Second, we show that coda types are not defined by arbitrary sequences of inter-click intervals, but instead form a combinatorial coding system in which rubato and ornamentation combine with two categorical, context-independent features that we call rhythm and tempo to give rise to a large inventory of distinguishable codas. In a dataset of 8,719 codas from the sperm whales of the Eastern Caribbean clan, this sperm whale phonetic alphabet makes it possible to systematically explain observed variability in coda structure. Sperm whale vocalisations are more expressive and structured than previously believed, and are built from a repertoire comprising nearly an order of magnitude more distinguishable codas. These results show contextsensitive and combinatorial vocalisation systems extend beyond humans, and can appear in an organism with a divergent evolutionary lineage and vocal apparatus.

animal behavior and cognition↗

Lexical semantic content, not syntactic structure, is the main contributor to ANN-brain similarity of fMRI responses in the language network

Representations from artificial neural network (ANN) language models have been shown to predict human brain activity in the language network. To understand what aspects of linguistic stimuli contribute to ANN-to-brain similarity, we used an fMRI dataset of responses to n=627 naturalistic English sentences (Pereira et al., 2018) and systematically manipulated the stimuli for which ANN representations were extracted. In particular, we i) perturbed sentences word order, ii) removed different subsets of words, or iii) replaced sentences with other sentences of varying semantic similarity. We found that the lexical semantic content of the sentence (largely carried by content words) rather than the sentences syntactic form (conveyed via word order or function words) is primarily responsible for the ANN-to-brain similarity. In follow-up analyses, we found that perturbation manipulations that adversely affect brain predictivity also lead to more divergent representations in the ANNs embedding space and decrease the ANNs ability to predict upcoming tokens in those stimuli. Further, results are robust to whether the mapping model is trained on intact or perturbed stimuli, and whether the ANN sentence representations are conditioned on the same linguistic context that humans saw. The critical result--that lexical- semantic content is the main contributor to the similarity between ANN representations and neural ones--aligns with the idea that the goal of the human language system is to extract meaning from linguistic strings. Finally, this work highlights the strength of systematic experimental manipulations for evaluating how close we are to accurate and generalizable models of the human language network.

neuroscience↗