Search bioRxiv⌕ Search

Biology subjects

Townsend, H. A.

Publications and source records attributed to Townsend, H. A..

4 recordsLinked to original sources

Improving confidence of differential transcription calls in enhancers

MotivationMost disease-associated genetic variants reside within transcribed regulatory elements (tREs). Patterns of differential transcription at tREs can be leveraged to identify upstream regulators and link enhancers to their target genes. But the low transcription levels and high variability in tREs makes identifying high confidence differentially transcribed elements challenging. ResultsWe present Mu Counts and TFEA-LE, two algorithms for robust identification of differentially transcribed tREs. The first step in accurate identification of differentially transcribed tREs is to obtain accurate RNA lengths and therefore counts over these regions. To this end we developed a method of accurate length inference (LIET-EMG) as wll as a rapid method for counting reads over tREs (Mu Counts). Armed with newly identified and quantified tREs, TFEA-LE then integrates motif information to simultaneously identify responsive tREs and their likely upstream regulators. We show improved precision and recall over general-purpose tools (e.g. DESeq2) in detecting p53-responsive tREs. We then clarify TF-specific responses within multi-TF perturbations in lung cells. Finally we show that the TFEA-LE approach improves TF activity inference, including in complex perturbations where many TFs respond. TFEALE is especially effective in technically challenging datasets, whether due to highly specific or broad responses, outliers, or high GC content. Ultimately, these 1 methods advance the systematic characterization of individual tREs, enabling their integration with regulators, target genes, and disease-associated variants for translational research. Availability and ImplementationTFEA-LE: https://github.com/Dowell-Lab/TFEA/tree/Lead_edge. Nextflow pipeline to run Mu Counts: https://github.com/Dowell-Lab/Bidir_Counting_Analysis. LIET (including modifications for tREs): https://github.com/Dowell-Lab/LIET/tree/LIET_EMGtoo. Source code for this work: https://github.com/Dowell-Lab/Improving_tRE_Analysis_Paper Contactrobin.dowell@colorado.edu

bioinformatics↗

LIET Model: Capturing the kinetics of RNA polymerase from loading to termination

Transcription by RNA polymerases is an exquisitely regulated step of the central dogma. Transcription is the primary determinant of cell-state, and most cellular perturbations impact transcription by altering polymerase activity. Thus, detecting changes in polymerase activity yields insight into most cellular processes. Nascent run-on sequencing provides a direct readout of polymerase activity, but no tools exist to model this activity at genes. We focus on RNA polymerase II--responsible for transcribing protein-coding genes. We present the first model to capture the complete process of gene transcription. For individual genes, this model parameterizes each distinct stage of transcription--Loading, Initiation, Elongation, and Termination, hence LIET--in a biologically interpretable Bayesian mixture, which is applied to nascent run-on data. Our improved modeling of Loading /Initiation demonstrates these are characteristically different between sense and antisense strands. Applying LIET to 24 human cell-types, our analysis indicates the position of dissociation (the last step of Termination) appears to be highly consistent, indicative of a highly regulated process. Furthermore, applying LIET to perturbation experiments, we demonstrate its ability to detect specific changes in pausing (5'end), strand-bias, and dissociation location (3'end)--opening the door to differential assessment of transcription at individual stages of individual genes.

bioinformatics↗

Single-cell based integrative analysis of transcriptomics and genetics reveals robust associations and complexities for inflammatory diseases

BackgroundUnderstanding genetic underpinnings of immune-mediated inflammatory diseases is crucial to improve treatments. Single-cell RNA sequencing (scRNA-seq) identifies cell states expanded in disease, but often overlooks genetic causality due to cost and small genotyping cohorts. Conversely, large genome-wide association studies (GWAS) are commonly accessible. MethodsWe present a 3-step robust benchmarking analysis of integrating GWAS and scRNA-seq to identify genetically relevant cell states and genes in inflammatory diseases. First, we applied and compared the results of two recent algorithms, based on networks (scGWAS) or single-cell disease scores (scDRS), according to accuracy/sensitivity and interpretability (M. J. Zhang et al. 2022; Jia et al. 2022). While previous studies focused on coarse cell types, we used disease-specific, fine-grained single-cell atlases (183,742 and 228,211 cells) and GWAS data (Ns of 97,1,73 and 45,975) for rheumatoid arthritis (RA) and ulcerative colitis (UC) (F. Zhang et al. 2023; Ishigaki et al. 2022; Smillie et al. 2019; de Lange et al. 2017). Second, given the lack of scRNA-seq for many diseases with GWAS, we further tested the tools resolution limits by differentiating between similar diseases with only one fine-grained scRNA-seq atlas. Lastly, we provide a novel evaluation of noncoding SNP incorporation methods by testing which enabled the highest sensitivity/accuracy of known cell-state calls. ResultsWe first found that single-cell based tool scDRS called superior numbers of supported cell states, like MERTK+ myeloid cells in RA, which were overlooked by network-based scGWAS. While scGWAS was advantageous for gene exploration, scDRS captured cellular heterogeneity of disease-relevance without single-cell genotyping. For noncoding SNP integration, we found a key trade-off between statistical power and confidence with positional (e.g. MAGMA) and non-positional approaches (e.g. chromatin-interaction, eQTL). Even when directly incorporating noncoding SNPs through 5 scRNA-seq measures of regulatory elements, non disease-specific atlases gave misleading results by not containing disease-tissue specific transcriptomic patterns. Despite this criticality of tissue-specific scRNA-seq, we showed that scDRS enabled deconvolution of two similar diseases with a single fine-grained scRNA-seq atlas and separate GWAS. Indeed, we identified supported and novel genetic-phenotype linkages separating RA and ankylosing spondylitis, and UC and crohns disease. Overall, while noting evolving single-cell technologies, our study provides key findings for integrating expanding fine-grained scRNA-seq, GWAS, and noncoding SNP resources to unravel the complexities of inflammatory diseases.

bioinformatics↗

Atlas of nascent RNA transcripts reveals enhancer to gene linkages

Gene transcription is controlled and modulated by regulatory regions, including enhancers and promoters. These regions are abundant in unstable, non-coding bidirectional transcription. Using nascent RNA transcription data across hundreds of human samples, we identified over 800,000 regions containing bidirectional transcription. We then identify highly correlated transcription between bidirectional and gene regions. The identified correlated pairs, a bidirectional region and a gene, are enriched for disease associated SNPs and often supported by independent 3D data. We present these resources as an SQL database which serves as a resource for future studies into gene regulation, enhancer associated RNAs, and transcription factors.

genomics↗