Search bioRxivSearch

Biology subjects

Mejia Guerra, M. K.

Publications and source records attributed to Mejia Guerra, M. K..

2 recordsLinked to original sources

Evolutionarily informed deep learning methods: Predicting transcript abundance from DNA sequence

Deep learning methodologies have revolutionized prediction in many fields, and show potential to do the same in molecular biology and genetics. However, applying these methods in their current forms ignores evolutionary dependencies within biological systems and can result in false positives and spurious conclusions. We developed two novel approaches that account for evolutionary relatedness in machine learning models: 1) gene-family guided splitting, and 2) ortholog contrasts. The first approach accounts for evolution by constraining the models training and testing sets to include different gene families. The second, uses evolutionarily informed comparisons between orthologous genes to both control for and leverage evolutionary divergence during the training process. The two approaches were explored and validated within the context of mRNA expression level prediction, and have prediction auROC values ranging from 0.72 to 0.94. Model weight inspections showed biologically interpretable patterns, resulting in the novel hypothesis that the 3 UTR is more important for fine tuning mRNA abundance levels while the 5 UTR is more important for large scale changes.

molecular biology

k-mer grammar uncovers maize regulatory architecture

Only a small percentage of the genome sequence is involved in regulation of gene expression, but to biochemically identify this portion is expensive and laborious. In species like maize, with diverse intergenic regions and lots of repetitive elements, this is an especially challenging problem. While regulatory regions are rare, they do have characteristic chromatin contexts and sequence organization (the grammar) with which they can be identified. We developed a computational framework to exploit this sequence arrangement. The models learn to classify regulatory regions based on sequence features - k-mers. To do this, we borrowed two approaches from the field of natural language processing: (1) \"bag-of-words\" which is commonly used for differentially weighting key words in tasks like sentiment analyses, and (2) a vector-space model using word2vec (vector-k-mers), that captures semantic and linguistic relationships between words. We built \"bag-of-k-mers\" and \"vector-k-mers\" models that distinguish between regulatory and non-regulatory regions with an accuracy above 90%. Our \"bag-of-k-mers\" achieved higher overall accuracy, while the \"vector-k-mers\" models were more useful in highlighting key groups of sequences within the regulatory regions. These models now provide powerful tools to annotate regulatory regions in other maize lines beyond the reference, at low cost and with high accuracy.

plant biology