Search bioRxivSearch

Biology subjects

Song, Y. S.

Publications and source records attributed to Song, Y. S..

13 recordsLinked to original sources

The key parameters that govern translation efficiency

Translation of mRNA into protein is a fundamental yet complex biological process with multiple factors that can potentially affect its efficiency. In particular, different genes can have quite different initiation rates, while site-specific elongation rates can vary substantially along a given transcript. Here, we analyze a stochastic model of translation dynamics to identify the key parameters that govern the overall rate of protein synthesis and the efficiency of ribosome usage. The mathematical model we study is an interacting particle system that generalizes the Totally Asymmetric Simple Exclusion Process (TASEP), where particles correspond to ribosomes. While the TASEP and its variants have been studied for the past several decades through simulations and mean field approximations, a general analytic solution has remained challenging to obtain. By analyzing the so-called hydrodynamic limit, we here obtain exact closed-form expressions for stationary currents and particle densities that agree well with Monte Carlo simulations. In addition, we provide a complete characterization of phase transitions in the system. Surprisingly, phase boundaries depend on only four parameters: the particle size, and the first, last and minimum particle jump rates. Relating these theoretical results to translation, we formulate four design principles that detail how to tune these parameters to optimize translation efficiency in terms of protein production rate and resource usage. We then analyze ribosome profiling data of S. cerevisiae and demonstrate that its translation system is generally efficient, consistent with the design principles we found. We discuss implications of our findings on evolutionary constraints and codon usage bias.

genomics

Differences in the path to exit the ribosome across the three domains of life

Recent advances in biological imaging have led to a surge of fine-resolution structures of the ribosome from diverse organisms. Comparing these structures, especially the exit tunnel, to characterize the key similarities and differences across species is essential for various important applications, such as designing antibiotic drugs and understanding the intricate details of translation dynamics. Here, we compile and compare 20 fine-resolution cryo-EM and X-ray crystallography structures of the ribosome recently obtained from all three domains of life (bacteria, archaea and eukarya). We first show that a hierarchical clustering of tunnel shapes closely reflects the species phylogeny. Then, by analyzing the ribosomal RNAs and proteins localized near the tunnel, we explain the observed geometric variations and show direct association between the conservations of the geometry, structure, and sequence. We find that the tunnel is more conserved in its upper part, from the polypeptide transferase center to the constriction site. In the lower part, tunnels are significantly narrower in eukaryotes than in bacteria, and we provide evidence for the existence of a second constriction site in eukaryotic tunnels. We also show that ribosomal RNA and protein sequences are more likely to be conserved closer to the tunnel, as is the presence of positively charged amino acids. Overall, our comparative analysis shows how the geometric and biophysical properties of the exit tunnel play an important role in ensuring proper transit of the nascent polypeptide chain, and may explain the differences observed in several co-translational processes across species.

biophysics

Efficiently inferring the demographic history of many populations with allele count data

The sample frequency spectrum (SFS), or histogram of allele counts, is an important summary statistic in evolutionary biology, and is often used to infer the history of population size changes, migrations, and other demographic events affecting a set of populations. The expected multipopulation SFS under a given demographic model can be efficiently computed when the populations in the model are related by a tree, scaling to hundreds of populations. Admixture, back-migration, and introgression are common natural processes that violate the assumption of a tree-like population history, however, and until now the expected SFS could be computed for only a handful of populations when the demographic history is not a tree. In this article, we present a new method for efficiently computing the expected SFS and linear functionals of it, for demographies described by general directed acyclic graphs. This method can scale to more populations than previously possible for complex demographic histories including admixture. We apply our method to an 8-population SFS to estimate the timing and strength of a proposed \"basal Eurasian\" admixture event in human history. We implement and release our method in a new open-source software package momi2.

evolutionary biology

Evaluating Tumor Evolution via Genomic Profiling of Individual Tumor Spheroids in a Malignant Ascites from a Patient with Ovarian Cancer Using a Laser-aided Cell Isolation Technique

BackgroundEpithelial ovarian cancer (EOC) is a silent but mostly lethal gynecologic malignancy. Most patients present with malignant ascites and peritoneal seeding at diagnosis. In the present study, we used a laser-aided isolation technique to investigate the clonal relationship between the primary tumor and tumor spheroids found in the malignant ascites of an EOC patient. Somatic alteration profiles of ovarian cancer-related genes were determined for eight spatially separated samples from primary ovarian tumor tissues and ten tumor spheroids from the malignant ascites using next-generation sequencing.\n\nResultsWe observed high levels of intra-tumor heterogeneity (ITH) in copy number alterations (CNAs) and single-nucleotide variants (SNVs) in the primary tumor and the tumor spheroids. As a result, we discovered that tumor cells in the primary tissues and the ascites were genetically different lineages. We categorized the CNAs and SNVs into clonal and subclonal alterations according to their distribution among the samples. Also, we identified focal amplifications and deletions in the analyzed samples. For SNVs, a total of 171 somatic mutations were observed, among which 66 were clonal mutations present in both the primary tumor and the ascites, and 61 and 44 of the SNVs were subclonal mutations present in only the primary tumor or the ascites, respectively.\n\nConclusionsBased on the somatic alteration profiles, we constructed phylogenetic trees and inferred the evolutionary history of tumor cells in the patient. The phylogenetic trees constructed using the CNAs and SNVs showed that two branches of the tumor cells diverged early from an ancestral tumor clone during an early metastasis step in the peritoneal cavity. Our data support the monophyletic spread of tumor spheroids in malignant ascites.

cancer biology

High-throughput inference of pairwise coalescence times identifies signals of selection and enriched disease heritability

Interest in reconstructing demographic histories has motivated the development of methods to estimate locus-specific pairwise coalescence times from whole-genome sequence data. We developed a new method, ASMC, that can estimate coalescence times using only SNP array data, and is 2-4 orders of magnitude faster than previous methods when sequencing data are available. We were thus able to apply ASMC to 113,851 phased British samples from the UK Biobank, aiming to detect recent positive selection by identifying loci with unusually high density of very recent coalescence times. We detected 12 genome-wide significant signals, including 6 loci with previous evidence of positive selection and 6 novel loci, consistent with coalescent simulations showing that our approach is well-powered to detect recent positive selection. We also applied ASMC to sequencing data from 498 Dutch individuals (Genome of the Netherlands data set) to detect background selection at deeper time scales. We observed highly significant correlations between average coalescence time inferred by ASMC and other measures of background selection. We investigated whether this signal translated into an enrichment in disease and complex trait heritability by analyzing summary association statistics from 20 independent diseases and complex traits (average N=86k) using stratified LD score regression. Our background selection annotation based on average coalescence time was strongly enriched for heritability (p = 7x10-153) in a joint analysis conditioned on a broad set of functional annotations (including other background selection annotations), meta-analyzed across traits; SNPs in the top 20% of our annotation were 3.8x enriched for heritability compared to the bottom 20%. These results underscore the widespread effects of background selection on disease and complex trait heritability.

genetics

A Likelihood-Free Inference Framework for Population Genetic Data using Exchangeable Neural Networks

An explosion of high-throughput DNA sequencing in the past decade has led to a surge of interest in population-scale inference with whole-genome data. Recent work in population genetics has centered on designing inference methods for relatively simple model classes, and few scalable general-purpose inference techniques exist for more realistic, complex models. To achieve this, two inferential challenges need to be addressed: (1) population data are exchangeable, calling for methods that efficiently exploit the symmetries of the data, and (2) computing likelihoods is intractable as it requires integrating over a set of correlated, extremely high-dimensional latent variables. These challenges are traditionally tackled by likelihood-free methods that use scientific simulators to generate datasets and reduce them to hand-designed, permutation-invariant summary statistics, often leading to inaccurate inference. In this work, we develop an exchangeable neural network that performs summary statistic-free, likelihood-free inference. Our frame-work can be applied in a black-box fashion across a variety of simulation-based tasks, both within and outside biology. We demonstrate the power of our approach on the recombination hotspot testing problem, outperforming the state-of-the-art.

evolutionary biology

Geometry of the sample frequency spectrum and the perils of demographic inference

The sample frequency spectrum (SFS), which describes the distribution of mutant alleles in a sample of DNA sequences, is a widely used summary statistic in population genetics. The expected SFS has a strong dependence on the historical population demography and this property is exploited by popular statistical methods to infer complex demographic histories from DNA sequence data. Most, if not all, of these inference methods exhibit pathological behavior, however. Specifically, they often display runaway behavior in optimization, where the inferred population sizes and epoch durations can degenerate to 0 or diverge to infinity, and show undesirable sensitivity of the inferred demography to perturbations in the data. The goal of this paper is to provide theoretical insights into why such problems arise. To this end, we characterize the geometry of the expected SFS for piecewise-constant demographic histories and use our results to show that the aforementioned pathological behavior of popular inference methods is intrinsic to the geometry of the expected SFS. We provide explicit descriptions and visualizations for a toy model with sample size 4, and generalize our intuition to arbitrary sample sizes n using tools from convex and algebraic geometry. We also develop a universal characterization result which shows that the expected SFS of a sample of size n under an arbitrary population history can be recapitulated by a piecewise-constant demography with only{kappa} n epochs, where{kappa} n is between n/2 and 2n - 1. The set of expected SFS for piecewise-constant demographies with fewer than{kappa} n epochs is open and non-convex, which causes the above phenomena for inference from data.

evolutionary biology

Three-way clustering of multi-tissue multi-individual gene expression data using constrained tensor decomposition

The advent of next generation sequencing methods has led to an increasing availability of large, multi-tissue datasets which contain gene expression measurements across different tissues and individuals. In this setting, variation in expression levels arises due to contributions specific to genes, tissues, individuals, and interactions thereof. Classical clustering methods are illsuited to explore these three-way interactions, and struggle to fully extract the insights into transcriptome complexity and regulation contained in the data. Thus, to exploit the multi-mode structure of the data, new methods are required. To this end, we propose a new method, called MultiCluster, based on constrained tensor decomposition which permits the investigation of transcriptome variation across individuals and tissues simultaneously. Through simulation and application to the GTEx RNA-seq data, we show that our tensor decomposition identifies three-way clusters with higher accuracy, while being 11x faster, than the competing Bayesian method. For several age-, race-, or gender-related genes, the tensor projection approach achieves increased significance over single-tissue analysis by two orders of magnitude. Our analysis finds gene modules consistent with existing knowledge while further detecting novel candidate genes exhibiting either tissue-, individual-, or tissue-by-individual specificity. These identified genes and gene modules offer bases for future study, and the uncovered multi-way specificities provide a finer, more nuanced snapshot of transcriptome variation than previously possible.

genomics

Model-based detection and analysis of introgressed Neanderthal ancestry in modern humans

Genetic evidence has revealed that the ancestors of modern human populations outside of Africa and their hominin sister groups, notably the Neanderthals, exchanged genetic material in the past. The distribution of these introgressed sequence-tracts along modern-day human genomes provides insight into the ancient structure and migration patterns of these archaic populations. Furthermore, it facilitates studying the selective processes that lead to the accumulation or depletion of introgressed genetic variation. Recent studies have developed methods to localize these introgressed regions, reporting long regions that are depleted of Neanderthal introgression and enriched in genes, suggesting negative selection against the Neanderthal variants. On the other hand, enriched Neanderthal ancestry in hair- and skin-related genes suggests that some introgressed variants facilitated adaptation to new environments. Here, we present a model-based method called diCal-admix and apply it to detect tracts of Neanderthal introgression in modern humans. We demonstrate its efficiency and accuracy through extensive simulations. We use our method to detect introgressed regions in modern human individuals from the 1000 Genomes Project, using a high coverage genome from a Neanderthal individual from the Altai mountains as reference. Our introgression detection results and findings concerning their functional implications are largely concordant with previous studies, and are consistent with weak selection against Neanderthal ancestry. We find some evidence that selection against Neanderthal ancestry was due to higher genetic load in Neanderthals, resulting from small effective population size, rather than Dobzhansky-Muller incompatibilities. Finally, we investigate the role of the X-chromosome in the divergence between Neanderthals and modern humans.

evolutionary biology

Worldwide genetic variation of the IGHV and TRBV immune receptor gene families in humans

The immunoglobulin heavy variable (IGHV) and T cell beta variable (TRBV) loci are among the most complex and variable regions in the human genome. Generated through a process of gene duplication/deletion and diversification, these loci can vary extensively between individuals in copy number and contain genes that are highly similar, making their analysis technically challenging. Here, we present a comprehensive study of the functional gene segments in the IGHV and TRBV loci, quantifying their copy number and single nucleotide variation in a globally diverse sample of 109 (IGHV) and 286 (TRBV) humans from over a hundred populations. We find that the IGHV and TRBV gene families exhibit starkly different patterns of variation. In particular, with hundreds of copy number haplotypes (instances that have differences in the number of functional gene segments), the IGHV locus has undergone more frequent gene duplication/deletion compared to the TRBV locus, which has only a few copy number haplotypes. In contrast, the TRBV locus has a greater or at least equal propensity to mutate, as evidenced by greater single nucleotide variation, compared to the IGHV locus. Thus, despite common molecular and functional characteristics, the genes that comprise the IGHV and TRBV loci have evolved in strikingly different ways. As well as providing insight into the different evolutionary paths the IGHV and TRBV loci have taken, our results are also important to the adaptive immune repertoire sequencing community, where the lack of frequencies of common alleles and copy number variants is hampering existing analytical pipelines.

genomics

Theoretical quantification of interference in the TASEP: Application to mRNA translation shows near-optimality of termination rates

The Totally Asymmetric Exclusion Process (TASEP) is a classical stochastic model for describing the transport of interacting particles, such as ribosomes moving along the mRNA during translation. Although this model has been widely studied in the past, the extent of collision between particles and the average distance between a particle to its nearest neighbor have not been quantified explicitly. We provide here a theoretical analysis of such quantities via the distribution of isolated particles. In the classical form of the model in which each particle occupies only a single site, we obtain an exact analytic solution using the Matrix Ansatz. We then employ a refined mean field approach to extend the analysis to a generalized TASEP with particles of an arbitrary size. Our theoretical study has direct applications in mRNA translation and the interpretation of experimental ribosome profiling data. In particular, our analysis of data from S. cerevisiae suggests a potential bias against the detection of nearby ribosomes with gap distance less than ~ 3 codons, which leads to some ambiguity in estimating the initiation rate and protein production flux for a substantial fraction of genes. Despite such ambiguity, however, we demonstrate theoretically that the interference rate associated with collisions can be robustly estimated, and show that approximately 1% of the translating ribosomes get obstructed.

systems biology

Identification and quantitative analysis of the major determinants of translation elongation rate variation

Previous studies have shown that translation elongation is regulated by multiple factors, but the observed heterogeneity remains only partially explained. To dissect quantitatively the different determinants of elongation speed, we use probabilistic modeling to estimate initiation and local elongation rates from ribosome profiling data. This model-based approach allows us to quantify the extent of interference between ribosomes on the same transcript. We show that neither interference nor the distribution of slow codons is sufficient to explain the observed heterogeneity. Instead, we find that electrostatic interactions between the ribosomal exit tunnel and specific parts of the nascent polypeptide govern the elongation rate variation as the polypeptide makes its initial pass through the tunnel. Once the N-terminus has escaped the tunnel, the hydropathy of the nascent polypeptide within the ribosome plays a major role in modulating the speed. We show that our results are consistent with the biophysical properties of the tunnel.

genomics

Time-resolved proteomics vs. ribosome profiling reveals translation dynamics under stress

Many small molecule chemotherapeutics induce stresses that globally inhibit mRNA translation, remodeling the cancer proteome and governing response to treatment. Here we measured protein synthesis in multiple myeloma cells treated with low-dose bortezomib by coupling pulsed-SILAC (pSILAC) with high-accuracy targeted quantitative proteomics. We found that direct measurement of protein synthesis by pSILAC correlated well with the indirect measurement of protein synthesis by ribosome profiling under conditions of robust translation. By developing a statistical model integrating longitudinal proteomic and mRNA-seq measurements, we found that proteomics could directly detect global alterations in translational rate as a function of therapy-induced stress after prolonged bortezomib exposure. Finally, the model we develop here, in combination with our experimental data including both protein synthesis and degradation, predicts changes in proteome remodeling under a variety of cellular perturbations. pSILAC therefore provides an important complement to ribosome profiling in directly measuring proteome dynamics under conditions of cellular stress.

cancer biology