Search bioRxiv⌕ Search

Biology subjects

Makinde, E.

Publications and source records attributed to Makinde, E..

2 recordsLinked to original sources

A Multi-Agent Approach to Generating Context-Rich Gene Sets

Gene sets are collections of genes that share a common biological function, process, or component that can be used to get insight into the biological relevance of genomic data. Databases containing these gene sets aids in a wide array of analytical methods. The results of these methods, such as gene set analysis or phenotype-based gene prioritization, depend on the quality of the gene sets. Despite the extensive literature and genetic data available for constructing these databases, they often lack sufficient biological context. Current curation methods rely on labour-intensive expert manual curation from literature and datasets, as well as automated methods that are not context-aware. Therefore, there is a significant opportunity to utilize publicly available literature to bridge this gap and create more precise gene sets. With the advancement of natural language processing technologies, particularly large language models, this task can be performed more efficiently. In this work, we present a multi-agent system that utilizes the Llama 3, DeepSeek, and Qwen open-source large language models to analyze PubMed abstracts, allowing us to reconstruct gene sets in existing databases that better reflect specific biological contexts. Our approach consists of two pipelines. One verifies the inclusion of genes in a gene set by proof of evidence in the abstracts showing the association between the gene and the gene set. The second pipeline parses through the abstracts to identify genes not already included in the gene set for potential inclusion. To evaluate the proposed approach, we reconstructed a random selection of gene sets within the Human Ontology Phenotype (HPO). Our analysis shows that 149 of these gene sets have a similarity of 65.18% when compared to the original HPO gene sets, aligning well with the current HPO database. Additionally, we found an average of 3.15 new genes not included in the HPO gene sets, each supported by verified literature linking them to their respective gene sets. This highlights that our updated gene set database better reflects the current state of biological findings.

bioinformatics↗

Neuron type-specific mRNA translation programs provide a gateway for memory consolidation.

Long-term memory consolidation is a dynamic process that requires a heterogeneous ensemble of neurons, each with a highly specialized molecular signature. Considerable effort has been devoted to identifying molecular changes that accompany the process of consolidation, but mostly hours or days after training, when memory consolidation is already complete. Studies have shown that protein synthesis is elevated during the early stages of consolidation, but how this increase impacts neuronal function remains unclear. We hypothesize that mRNAs translated during the early stages of consolidation could provide information on how diverse neurons involved in memory formation restructure their molecular signatures to support memory. Here, we generate a landscape of the translatome of three neuron types in the dorsal hippocampus during the first hour of contextual memory consolidation. Our results reveal that translation programs associated with consolidation are different among neurons, fueling the reconfiguration of specific biological processes. We further demonstrate the patterned translation of mRNAs in different neuron types during consolidation is explained by features hard-coded in the mRNA sequence, suggesting ubiquitous mechanisms controlling activity-induced neuronal translation. Altogether, our work uncovers previously unknown mechanisms controlling activity-induced translation in neurons and provides a large, readily available resource for scientists interested in the role of memory formation in health and disease.

neuroscience↗