Search bioRxivSearch

Biology subjects

Chung, M.

Publications and source records attributed to Chung, M..

7 recordsLinked to original sources

FADU: A Feature Counting Tool for Prokaryotic RNA-Seq Analysis

MotivationThe major algorithms for quantifying transcriptomics data for differential gene expression analysis were designed for analyzing data from human or human-like genomes, specifically those with single gene transcripts and distinct transcriptional boundaries that extend beyond the coding sequence (CDS) as identified through expressed sequence tags (ESTs) or EST-like sequence data. Some eukaryotic genomes and all, or nearly all, bacterial genomes require alternate methods of quantification since they lack annotation of transcriptional boundaries with EST or EST-like data, have overlapping transcriptional boundaries, and/or have polycistronic transcripts.\n\nResultsAn algorithm was developed and tested that better quantifies transcriptomics data for differential gene expression analysis in organisms with overlapping transcriptional units and polycistronic transcripts. Using data from standard libraries originating from Escherichia coli and Ehrlichia chaffeensis, and strand-specific libraries from the Wolbachia endosymbiont wBm, FADU can derive counts for genes that are missed by HTSeq and featurecounts. Using the default parameters with the E. coli data, FADU can detect transcription of 51 more genes than HTSeq in union mode and 21 genes more than featurecounts, with 42 and 18 of these features being <300 bp, respectively. Due to its ability to derive counts for otherwise unrepresented genes without overstating their abundance, we believe FADU to be an improved tool for quantifying transcripts in prokaryotic systems for RNA-Seq analyses.\n\nAvailability and implementationFADU is available at https://github.com/adkinsrs/FADU. FADU was implemented using Python3 and requires the PySAM module (version 0.12.0.1 or later).\n\nContactjdhotopp@som.umaryland.edu

genomics

Using Core Genome Alignments to Assign Bacterial Species

With the exponential increase in the number of bacterial taxa with genome sequence data, a new standardized method is needed to assign bacterial species designations using genomic data that is consistent with the classically-obtained taxonomy. This is particularly acute for unculturable obligate intracellular bacteria like those in the Rickettsiales, where classical methods like DNA-DNA hybridization cannot be used to define species. Within the Rickettsiales, species designations have been applied inconsistently, often obfuscating the relationship between organisms and the context for experimental results. In this study, we generated core genome alignments for a wide range of genera with classically defined species, including Arcobacter, Caulobacter, Erwinia, Neisseria, Polaribacter, Ralstonia, Thermus, as well as genera within the Rickettsiales including Rickettsia, Orientia, Ehrlichia, Neoehrlichia, Anaplasma, eorickettsia, and Wolbachia. A core genome alignment sequence identity (CGASI) threshold of 96.8% was found to maximize the prediction of classically-defined species. Using the CGASI cutoff, the Wolbachia genus can be delineated into species that differ from the currently used supergroup designations, while the Rickettsia genus is delineated into nine species, as opposed to the current 27 species. Additionally, we find that core genome alignments cannot be constructed between genomes belonging to different genera, establishing a bacterial genus cutoff that suggests the need to create new genera from the Anaplasma and Neorickettsia. By using core genome alignments to assign taxonomic designations, we aim to provide a high-resolution, robust method for bacterial nomenclature that is aligned with classically-obtained results.

microbiology

Targeted enrichment outperforms other enrichment techniques and enables more multi-species RNA-Seq analyses

Enrichment methodologies enable analysis of minor members in multi-species transcriptomic analyses. We compared standard enrichment of bacterial and eukaryotic mRNA to targeted enrichment with Agilent SureSelect (AgSS) capture for Brugia malayi, Aspergillus fumigatus, and the Wolbachia endosymbiont of B. malayi (wBm). Without introducing significant systematic bias, the AgSS quantitatively enriched samples, resulting in more reads mapping to the target organism. The AgSS-enriched libraries consistently had a positive linear correlation with its unenriched counterpart (r2=0.559-0.867). Up to a 2,242-fold enrichment of RNA from the target organism was obtained following a power law (r2=0.90), with the greatest fold enrichment achieved in samples with the largest ratio difference between the major and minor members. While using a single total library for prokaryote and eukaryote in a single sample could be beneficial for samples where RNA is limiting, we observed a decrease in reads mapping to protein coding genes and an increase of multi-mapping reads to rRNAs in AgSS enrichments from eukaryotic total RNA libraries as opposed to eukaryotic poly(A)-enriched libraries. Our results support a recommendation of using Agilent SureSelect targeted enrichment on poly(A)-enriched libraries for eukaryotic captures and total RNA libraries for prokaryotic captures to increase the robustness of multi-species transcriptomic studies.

genomics

Transcription factor dynamics reveals a circadian code for fat cell differentiation

Glucocorticoid and other adipogenic hormones are secreted in mammals in circadian oscillations. Loss of this circadian oscillation pattern during stress and disease correlates with increased fat mass and obesity in humans, raising the intriguing question of how hormone secretion dynamics affect the process of adipocyte differentiation. By using live, single-cell imaging of the key adipogenic transcription factors CEBPB and PPARG, endogenously tagged with fluorescent proteins, we show that pulsatile circadian hormone stimuli are rejected by the adipocyte differentiation control system, leading to very low adipocyte differentiation rates. In striking contrast, equally strong persistent signals trigger maximal differentiation. We identify the mechanism of how hormone oscillations are filtered as a combination of slow and fast positive feedback centered on PPARG. Furthermore, we confirm in mice that flattening of daily glucocorticoid oscillations significantly increases the mass of subcutaneous and visceral fat pads. Together, our study provides a molecular mechanism for why stress, Cushings disease, and other conditions for which glucocorticoid secretion loses its pulsatility can lead to obesity. Given the ubiquitous nature of oscillating hormone secretion in mammals, the filtering mechanism we uncovered may represent a general temporal control principle for differentiation.\n\nHIGHLIGHTO_LIWe found that the fraction of differentiated cells is controlled by rhythmic and pulsatile hormone stimulus patterns.\nC_LIO_LITwelve hours is the cutoff point for daily hormone pulse durations below which cells fail to differentiate, arguing for a circadian code for hormone-induced cell differentiation.\nC_LIO_LIIn addition to fast positive feedback such as between PPARG and CEBPA, the adipogenic transcriptional architecture requires added parallel slow positive feedback to mediate temporal filtering of circadian oscillatory inputs\nC_LI

cell biology

Therapeutically advantageous secondary targets of abemaciclib identified by multi-omics profiling of CDK4/6 inhibitors

FDA approval of multiple drugs differing in chemical structures but targeting the same protein raises the question whether such drugs have sufficiently similar mechanisms of action to be considered functionally equivalent. In this paper we compare three recently approved inhibitors of the cyclin-dependent kinases CDK4/6 - palbociclib, ribociclib, and abemaciclib - that are becoming important therapies for the treatment of hormone-receptor positive breast and potentially other cancers. We find that transcriptional and proteomic changes induced by the three drugs differ significantly and that abemaciclib has unique cellular activities including induction of cell death (even in pRb-deficient cells), arrest in the G2 phase of the cell cycle, and reduced drug adaptation. These activities appear to arise from inhibition of kinases other than CDK4/6 including CDK2/Cyclin A/E and CDK1/Cyclin B.\n\nSIGNIFICANCEThe target profiles of most drugs are established relatively early in their development and are not systematically revisited at the time of approval. Scattered reports suggest that palbociclib, ribociclib, and abemaciclib differ in pharmacokinetics, dosing, and adverse effects but the three drugs are generally regarded as similar. Our finding that the drugs differ substantially in mechanism of action - abemaciclib retains activities of the earlier-generation drug alvocidib - suggests the potential for different uses in the clinic: in particular, abemaciclib may show activity in patients progressing on palbociclib or ribociclib. More generally, our approach relying on data from five distinct phenotypic and biochemical assays strongly suggests that a multi-faceted approach is necessary to get a reliable picture the target spectrum of kinase inhibitors.

cancer biology

A multi-center study on factors influencing the reproducibility of in vitro drug-response studies

Evidence that some influential biomedical results cannot be repeated has increased interest in practices that generate data meeting findable, accessible, interoperable and reproducible (FAIR) standards. Multiple papers have identified examples of irreproducibility, but practical steps for increasing reproducibility have not been widely studied. Here, seven research centers in the NIH LINCS Program Consortium investigate the reproducibility of a prototypical perturbational assay: quantifying the responsiveness of cultured cells to anti-cancer drugs. Such assays are important for drug development, studying cell biology, and patient stratification. While many experimental and computational factors have an impact on intra- and inter-center reproducibility, the factors most difficult to identify and correct are those with a strong dependency on biological context. These factors often vary in magnitude with the drug being analyzed and with growth conditions. We provide ways of identifying such context-sensitive factors, thereby advancing the conceptual and practical basis for greater experimental reproducibility.

cancer biology

Heritability of hierarchical structural brain network

We present a new structural brain network parcellation scheme that can subdivide existing parcellations into smaller subregions in a hierarchically nested fashion. The hierarchical parcellation was used to build multilayer convolutional structural brain networks that preserve topology across different network scales. As an application, we applied the method to diffusion weighted imaging study of 111 twin pairs. The genetic contribution of the whole brain structural connectivity was determined. We showed that the overall heritability is consistent across different network scales.

neuroscience