Search bioRxivSearch

Biology subjects

Maeda, N.

Publications and source records attributed to Maeda, N..

2 recordsLinked to original sources

Behavioral correlates of cortical semantic representations modeled by word vectors

The quantitative modeling of semantic representations in the brain plays a key role in understanding the neural basis of semantic processing. Previous studies have demonstrated that word vectors, which were originally developed for use in the field of natural language processing, provide a powerful tool for such quantitative modeling. However, whether semantic representations revealed by the word vector-based models actually capture our perception of semantic information remains unclear, as there has been no study explicitly examining the behavioral correlates of the modeled semantic representations. To address this issue, we compared the semantic structure of nouns and adjectives estimated from word vector-based brain models with that evaluated from human behavior. The brain models were constructed using voxelwise modeling to predict the functional magnetic resonance imaging (fMRI) response to natural movies from semantic contents in each movie scene through a word vector space. The semantic dissimilarity of word representations was then evaluated using the brain models. Meanwhile, data on human behavior reflecting the perception of semantic dissimilarity between words were collected in psychological experiments. We found a significant correlation between brain model- and behavior-derived semantic dissimilarities of words. This finding suggests that semantic representations in the brain modeled via word vectors appropriately capture our perception of word meanings. Author summeryWord vectors, which have been originally developed in the field of engineering (natural language processing), have been extensively leveraged in neuroscience studies to model semantic representations in the human brain. These studies have attempted to model brain semantic representations by associating them with the meanings of thousands of words via a word vector space. However, there has been no study explicitly examining whether the modeled semantic representations actually capture our perception of semantic information. To address this issue, we compared the semantic representational structure of words estimated from word vector-based brain models with that evaluated from behavioral data in psychological experiments. The results revealed a significant correlation between these model- and behavior-derived semantic representational structures of words. This indicates that the brain semantic representations modeled using word vectors actually reflect the human perception of word meanings. Our findings contribute to the establishment of word vector-based brain modeling as a useful tool in studying human semantic processing.

neuroscience

Multi-sample Full-length Transcriptome Analysis of 22 Breast Cancer Clinical Specimens with Long-Read Sequencing

Although transcriptome alteration is considered as one of the essential drivers of carcinogenesis, conventional short-read RNAseq technology has limited researchers from directly exploring full-length transcripts, only focusing on individual splice sites. We developed a pipeline for Multi-Sample long-read Transcriptome Assembly, MuSTA, and showed through simulations that it enables construction of transcriptome from the transcripts expressed in target samples and more accurate evaluation of transcript usage. We applied it to 22 breast cancer clinical specimens to successfully acquire cohort-wide full-length transcriptome from long-read RNAseq data. By comparing isoform existence and expression between estrogen receptor positive and triple-negative subtypes, we obtained a comprehensive set of subtype-specific isoforms and differentially used isoforms which consisted of both known and unannotated isoforms. We have also found that exon-intron structure of fusion transcripts tends to depend on their genomic regions, and have found three-piece fusion transcripts that were transcribed from complex structural rearrangements. For example, a three-piece fusion transcript resulted in aberrant expression of an endogenous retroviral gene, ERVFRD-1, which is normally expressed exclusively in placenta and supposed to protect fetus from maternal rejection, and expression of which were increased in several TCGA samples with ERVFRD-1 fusions. Our analyses of real clinical specimens and simulated data provide direct evidence that full-length transcript sequencing in multiple samples can add to our understanding of cancer biology and genomics in general.

genomics