Search bioRxiv⌕ Search

Biology subjects

Serikov, O.

Publications and source records attributed to Serikov, O..

2 recordsLinked to original sources

SIGNAL: Dataset for Semantic and Inferred Grammar Neurological Analysis of Language

Recently, the idea of comparison of models representations and human brain signals has been a topic of several works. Consequently, several datasets with text data and EEG representations have been published. However, most of the datasets are based on normal reading task with grammatical sentences. At the same time, in the interpretability studies of LLMs, more and more attention is paid to thoroughly designed linguistic tasks based on acceptability measures. In this paper, we present SIGNAL, a dataset for Semantic and Inferred Grammar Neurological Analysis of Language. Our dataset contains a group of sentences with a combination of a fully acceptable sentence and a grammatically or/and semantically incongruent sentences. The dataset has been approved by native speakers and later used for an EEG experiment. In total, our dataset contains recordings of 21 participants, each of whom read 600 sentences. In addition, we present a pilot study where we compare EEG analysis with simple probing experiments.

neuroscience↗

Representational dissimilarity component analysis (ReDisCA)

The principle of Representational Similarity Analysis (RSA) posits that neural representations reflect the structure of encoded information, allowing exploration of spatial and temporal organization of brain information processing. Traditional RSA when applied to EEG or MEG data faces challenges in accessing activation time series at the brain source level due to modeling complexities and insufficient geometric/anatomical data. To address this, we introduce Representational Dissimilarity Component Analysis (ReDisCA), a method for estimating spatial-temporal components in EEG or MEG responses aligned with a target representational dissimilarity matrix (RDM). ReDisCA yields informative spatial filters and associated topographies, offering insights into the location of "representationally relevant" sources. Applied to evoked response time series, ReDisCA produces temporal source activation profiles with the desired RDM. Importantly, while ReDisCA does not require inverse modeling its output is consistent with EEG and MEG observation equation and can be used as an input to rigorous source localization procedures. Demonstrating ReDisCAs efficacy through simulations and comparison with conventional methods, we show superior source localization accuracy and apply the method to real EEG and MEG datasets, revealing physiologically plausible representational structures without inverse modeling. ReDisCA adds to the family of inverse modeling free methods such as independent component analysis [34], Spatial spectral decomposition [41], and Source power comodulation [9] designed for extraction sources with desired properties from EEG or MEG data. Extending its utility beyond EEG and MEG analysis, ReDisCA is likely to find application in fMRI data analysis and exploration of representational structures emerging in multilayered artificial neural networks.

neuroscience↗