Search bioRxivSearch

Biology subjects

Huang, Y.-F.

Publications and source records attributed to Huang, Y.-F..

4 recordsLinked to original sources

Estimation of allele-specific fitness effects across human protein-coding sequences and implications for disease

A central challenge in human genomics is to understand the cellular, evolutionary, and clinical significance of genetic variants. Here we introduce a unified population-genetic and machine-learning model, called Linear Allele-Specific Selection InferencE (LASSIE), for estimating the fitness effects of all potential single-nucleotide variants, based on polymorphism data and predictive genomic features. We applied LASSIE to 51 high-coverage genome sequences annotated with 33 genomic features, and constructed a map of allele-specific selection coefficients across all protein-coding sequences in the human genome. We show that this map is informative about both human evolution and disease.

genomics

The ligand-mediated affinity of escort proteins determines the directionality of lipophilic cargo transport

Intracellular cargo transport is a highly dynamic process. In eukaryotic cells, the uptake and release of lipophilic ligands are executed by escort proteins. However, how these carriers control the directionality of cargo trafficking remains unclear. Here, we have elucidated the unliganded structure of an archetypal fatty acid-binding protein (FABP) and found that it possesses stronger binding affinity than its liganded counterpart towards empty nanodiscs. Titrating unliganded FABP and nanodiscs with long-chain fatty acids (LCFAs) rescued the broadening of FABP cross-peak intensities in HSQC spectra due to decreased protein-membrane interaction. Crystallographic studies revealed that the tails of bound LCFAs obstructed the charged interfaces of the FABP-nanodisc complexes. We conclude that the lipophilic ligands, by taking advantage of escort proteins with high conformational homogeneity and nanodiscs as the third interaction partner involved in this transport study, participate directly in the control of their own transportation in an irreversible, unidirectional fashion.

biophysics

Scikit-ribo: Accurate estimation and robust modeling of translation dynamics at codon resolution

Ribosome profiling (Riboseq) is a powerful technique for measuring protein translation, however, sampling errors and biological biases are prevalent and poorly understand. Addressing these issues, we present Scikit-ribo (https://github.com/hanfang/scikit-ribo), the first open-source software for accurate genome-wide A-site prediction and translation efficiency (TE) estimation from Riboseq and RNAseq data. Scikit-ribo accurately identifies A-site locations and reproduces codon elongation rates using several digestion protocols (r = 0.99). Next we show commonly used RPKM-derived TE estimation is prone to biases, especially for low-abundance genes. Scikit-ribo introduces a codon-level generalized linear model with ridge penalty that correctly estimates TE while accommodating variable codon elongation rates and mRNA secondary structure. This corrects the TE errors for over 2000 genes in S. cerevisiae, which we validate using mass spectrometry of protein abundances (r = 0.81) and allows us to determine the Kozak-like sequence directly from Riboseq. We conclude with an analysis of coverage requirements needed for robust codon-level analysis, and quantify the artifacts that can occur from cycloheximide treatment.

bioinformatics

Nascent RNA sequencing reveals a dynamic global transcriptional response at genes and enhancers to the natural medicinal compound celastrol

Most studies of responses to transcriptional stimuli measure changes in cellular mRNA concentrations. By sequencing nascent RNA instead, it is possible to detect changes in transcription in minutes rather than hours, and thereby distinguish primary from secondary responses to regulatory signals. Here, we describe the use of PRO-seq to characterize the immediate transcriptional response in human cells to celastrol, a compound derived from traditional Chinese medicine that has potent anti-inflammatory, tumor-inhibitory and obesity-controlling effects. Our analysis of PRO-seq data for K562 cells reveals dramatic transcriptional effects soon after celastrol treatment at a broad collection of both coding and noncoding transcription units. This transcriptional response occurred in two major waves, one within 10 minutes, and a second 40-60 minutes after treatment. Transcriptional activity was generally repressed by celastrol, but one distinct group of genes, enriched for roles in the heat shock response, displayed strong activation. Using a regression approach, we identified key transcription factors that appear to drive these transcriptional responses, including members of the E2F and RFX families. We also found sequence-based evidence that particular TFs drive the activation of enhancers. We observed increased polymerase pausing at both genes and enhancers, suggesting that pause release may be widely inhibited during the celastrol response. Our study demonstrates that a careful analysis of PRO-seq time course data can disentangle key aspects of a complex transcriptional response, and it provides new insights into the activity of a powerful pharmacological agent.

genomics