Search bioRxivSearch

Biology subjects

Siepel, A.

Publications and source records attributed to Siepel, A..

7 recordsLinked to original sources

Estimation of allele-specific fitness effects across human protein-coding sequences and implications for disease

A central challenge in human genomics is to understand the cellular, evolutionary, and clinical significance of genetic variants. Here we introduce a unified population-genetic and machine-learning model, called Linear Allele-Specific Selection InferencE (LASSIE), for estimating the fitness effects of all potential single-nucleotide variants, based on polymorphism data and predictive genomic features. We applied LASSIE to 51 high-coverage genome sequences annotated with 33 genomic features, and constructed a map of allele-specific selection coefficients across all protein-coding sequences in the human genome. We show that this map is informative about both human evolution and disease.

genomics

phastWeb: a web interface for evolutionary conservation scoring of multiple sequence alignments using phastCons and phyloP

The Phylogenetic Analysis with Space/Time models (PHAST) package is a widely used software package for comparative genomics that has been freely available for download since 2002. Here we introduce a web interface (phastWeb) that makes it possible to use two of the most popular programs in PHAST, phastCons and phyloP, without downloading and installing the PHAST software. This interface allows users to upload a sequence alignment and either upload a corresponding phylogeny or have one estimated from the alignment. After processing, users can visualize alignments and conservation scores as genome browser tracks, and download estimated tree models and raw scores for further analysis. Altogether, this resource makes key features of the PHAST package conveniently available to a broad audience.\n\nAVAILABILITYphastWeb is freely available on the web at http://compgen.cshl.edu/phastweb/. The website provides instructions as well as examples.\n\nCONTACTphasthelp@cshl.edu

evolutionary biology

How Much Information is Provided by Human Epigenomic Data? An Evolutionary View

Here, we ask the question, \"How much information do available epigenomic data sets provide about human genomic function, individually or in combination?\" We consider nine epigenomic and annotation features across 115 cell types and measure genomic function by using signatures of natural selection as a proxy. We measure information as the reduction in entropy under a probabilistic evolutionary model that describes genetic variation across [~]50 diverse humans and several nonhuman primates. We find that several genomic features yield more information in combination than they do individually, with DNase-seq displaying particularly strong synergy. Most of the entropy in human genetic variation, by far, reflects mutation and neutral drift; the genome-wide reduction in entropy due to selection is equivalent to only a small fraction of the storage requirements of a single human genome. Based on this framework, we produce cell-type-specific maps of the probability that a mutation at each nucleotide will have fitness consequences (FitCons scores). These scores are predictive of known functional elements and disease-associated variants, they reveal relationships among cell types, and they suggest that [~]8% of nucleotide sites are constrained by natural selection.

genomics

Scikit-ribo: Accurate estimation and robust modeling of translation dynamics at codon resolution

Ribosome profiling (Riboseq) is a powerful technique for measuring protein translation, however, sampling errors and biological biases are prevalent and poorly understand. Addressing these issues, we present Scikit-ribo (https://github.com/hanfang/scikit-ribo), the first open-source software for accurate genome-wide A-site prediction and translation efficiency (TE) estimation from Riboseq and RNAseq data. Scikit-ribo accurately identifies A-site locations and reproduces codon elongation rates using several digestion protocols (r = 0.99). Next we show commonly used RPKM-derived TE estimation is prone to biases, especially for low-abundance genes. Scikit-ribo introduces a codon-level generalized linear model with ridge penalty that correctly estimates TE while accommodating variable codon elongation rates and mRNA secondary structure. This corrects the TE errors for over 2000 genes in S. cerevisiae, which we validate using mass spectrometry of protein abundances (r = 0.81) and allows us to determine the Kozak-like sequence directly from Riboseq. We conclude with an analysis of coverage requirements needed for robust codon-level analysis, and quantify the artifacts that can occur from cycloheximide treatment.

bioinformatics

Deep experimental profiling of microRNA diversity, deployment, and evolution across the Drosophila genus

Comparative genomic analyses of microRNAs (miRNAs) have yielded myriad insights into their biogenesis and regulatory activity. While miRNAs have been deeply annotated in a small cohort of model organisms, evolutionary assessments of miRNA flux are clouded by the functional uncertainty of orthologs in related species, and insufficient data regarding the extent of species-specific miRNAs. We address this by generating a comparative small RNA (sRNA) catalog of unprecedented breadth and depth across the Drosophila genus, extending our extant deep analyses of D. melanogaster with sRNA data from multiple tissues of 11 other fly species. Aggregate analysis of several billion sRNA reads permits curation of accurate and holistic compendia of miRNAs across this genus, providing abundant opportunities to identify species- and clade-specific variation in miRNA identity, abundance, and processing. Amongst well-conserved miRNAs, we observe unexpected cases of clade-specific variation in 5' end precision, occasional antisense loci, and some putatively non-canonical loci. We also employ strict criteria to identify a massive set (649) of novel, evolutionarily-restricted miRNAs. Amongst the bulk collection of species-restricted miRNAs, two notable subpopulations of rapidly-evolving miRNAs are splicing-derived mirtrons and testis-restricted, clustered (TRC) canonical miRNAs. We quantify rates of miRNA birth and death using our annotation and a phylogenetic model for estimating rates of miRNA turnover in the presence of annotation uncertainty. We show striking differences in birth and death rates across miRNA classes defined by biogenesis pathway, genomic clustering, and tissue restriction, and even identify variation heterogeneity amongst Drosophila clades. In particular, distinct molecular rationales underlie the distinct evolutionary behavior of different miRNA classes. We broaden observations made from D. melanogaster as Drosophilid-wide principles for opposing evolutionary viewpoints for miRNA maintenance. Mirtrons are associated with a high rate of 3' untemplated addition, a mechanism that impedes their biogenesis, whereas TRC miRNAs appear to evolve under positive selection. Altogether, these data reveal miRNA diversity amongst Drosophila species and permit future discoveries in understanding their emergence and evolution.

genomics

Nascent RNA sequencing reveals a dynamic global transcriptional response at genes and enhancers to the natural medicinal compound celastrol

Most studies of responses to transcriptional stimuli measure changes in cellular mRNA concentrations. By sequencing nascent RNA instead, it is possible to detect changes in transcription in minutes rather than hours, and thereby distinguish primary from secondary responses to regulatory signals. Here, we describe the use of PRO-seq to characterize the immediate transcriptional response in human cells to celastrol, a compound derived from traditional Chinese medicine that has potent anti-inflammatory, tumor-inhibitory and obesity-controlling effects. Our analysis of PRO-seq data for K562 cells reveals dramatic transcriptional effects soon after celastrol treatment at a broad collection of both coding and noncoding transcription units. This transcriptional response occurred in two major waves, one within 10 minutes, and a second 40-60 minutes after treatment. Transcriptional activity was generally repressed by celastrol, but one distinct group of genes, enriched for roles in the heat shock response, displayed strong activation. Using a regression approach, we identified key transcription factors that appear to drive these transcriptional responses, including members of the E2F and RFX families. We also found sequence-based evidence that particular TFs drive the activation of enhancers. We observed increased polymerase pausing at both genes and enhancers, suggesting that pause release may be widely inhibited during the celastrol response. Our study demonstrates that a careful analysis of PRO-seq time course data can disentangle key aspects of a complex transcriptional response, and it provides new insights into the activity of a powerful pharmacological agent.

genomics

Natural Selection has Shaped Coding and Non-coding Transcription in Primate CD4+ T-cells

Transcriptional regulatory changes have been shown to contribute to phenotypic differences between species, but many questions remain about how gene expression evolves. Here we report the first comparative study of nascent transcription in primates. We used PRO-seq to map actively transcribing RNA polymerases in resting and activated CD4+ T-cells in multiple human, chimpanzee, and rhesus macaque individuals, with rodents as outgroups. This approach allowed us to measure transcription separately from post-transcriptional processes. We observed general conservation in coding and non-coding transcription, punctuated by numerous differences between species, particularly at distal enhancers and non-coding RNAs. We found evidence that transcription factor binding sites are a primary determinant of transcriptional differences between species, that stabilizing selection maintains gene expression levels despite frequent changes at distal enhancers, and that adaptive substitutions have driven lineage-specific transcription. Finally, we found strong correlations between evolutionary rates and long-range chromatin interactions. These observations clarify the role of primary transcription in regulatory evolution.

genomics