Search bioRxiv⌕ Search

Biology subjects

Galindez, G.

Publications and source records attributed to Galindez, G..

2 recordsLinked to original sources

Nucleotide dependency analysis of DNA language models reveals genomic functional elements

Deciphering how nucleotides in genomes encode regulatory instructions and molecular machines is a long-standing goal in biology. DNA language models (LMs) implicitly capture functional elements and their organization from genomic sequences alone by modeling probabilities of each nucleotide given its sequence context. However, using DNA LMs for discovering functional genomic elements has been challenging due to the lack of interpretable methods. Here, we introduce nucleotide dependencies which quantify how nucleotide substitutions at one genomic position affect the probabilities of nucleotides at other positions. We generated genome-wide maps of pairwise nucleotide dependencies within kilobase ranges for animal, fungal, and bacterial species. We show that nucleotide dependencies indicate deleteriousness of human genetic variants more effectively than sequence alignment and DNA LM reconstruction. Regulatory elements appear as dense blocks in dependency maps, enabling the systematic identification of transcription factor binding sites as accurately as models trained on experimental binding data. Nucleotide dependencies also highlight bases in contact within RNA structures, including pseudoknots and tertiary structure contacts, with remarkable accuracy. This led to the discovery of four novel, experimentally validated RNA structures in Escherichia coli. Finally, using dependency maps, we reveal critical limitations of several DNA LM architectures and training sequence selection strategies by benchmarking and visual diagnosis. Altogether, nucleotide dependency analysis opens a new avenue for discovering and studying functional elements and their interactions in genomes.

genomics↗

Identification of transcriptional regulators using a combined disease module identification and prize-collecting Steiner tree approach

Transcription factors play important roles in maintaining normal biological function, and their dys-regulation can lead to the development of diseases. Identifying candidate transcription factors involved in disease pathogenesis is thus an important task for deriving mechanistic insights from gene expression data. We developed Transcriptional Regulator Identification using Prize-collecting Steiner trees (TRIPS), a workflow for identifying candidate transcriptional regulators from case-control expression data. In the first step, TRIPS combines the results of differential expression analysis with a disease module identification step to retrieve perturbed subnetworks comprising an expanded gene list. TRIPS then solves a prize-collecting Steiner tree problem on a gene regulatory network, thereby identifying candidate transcriptional modules and transcription factors. We compare TRIPS to relevant methods using publicly available disease datasets and show that the proposed workflow can recover known disease-associated transcription factors with high precision. Network perturbation analyses demonstrate the reliability of TRIPS results. We further evaluate TRIPS on Alzheimers disease, diabetic kidney disease, and prostate cancer single-cell omics datasets. Overall, TRIPS is a useful approach for prioritizing transcriptional mechanisms for further downstream analyses.

bioinformatics↗