Search bioRxivSearch

Biology subjects

Lee, S. R.

Publications and source records attributed to Lee, S. R..

3 recordsLinked to original sources

Resource: Scalable whole genome sequencing of 40,000 single cells identifies stochastic aneuploidies, genome replication states and clonal repertoires

Essential features of cancer tissue cellular heterogeneity such as negatively selected genome topologies, sub-clonal mutation patterns and genome replication states can only effectively be studied by sequencing single-cell genomes at scale and high fidelity. Using an amplification-free single-cell genome sequencing approach implemented on commodity hardware (DLP+) coupled with a cloud-based computational platform, we define a resource of 40,000 single-cell genomes characterized by their genome states, across a wide range of tissue types and conditions. We show that shallow sequencing across thousands of genomes permits reconstruction of clonal genomes to single nucleotide resolution through aggregation analysis of cells sharing higher order genome structure. From large-scale population analysis over thousands of cells, we identify rare cells exhibiting mitotic mis-segregation of whole chromosomes. We observe that tissue derived scWGS libraries exhibit lower rates of whole chromosome anueploidy than cell lines, and loss of p53 results in a shift in event type, but not overall prevalence in breast epithelium. Finally, we demonstrate that the replication states of genomes can be identified, allowing the number and proportion of replicating cells, as well as the chromosomal pattern of replication to be unambiguously identified in single-cell genome sequencing experiments. The combined annotated resource and approach provide a re-implementable large scale platform for studying lineages and tissue heterogeneity.

genomics

Genetic Diversity Patterns and Domestication Origin of Soybean

Understanding diversity and evolution of a crop is an essential step to implement a strategy to expand its germplasm base for crop improvement research. Samples intensively collected from Korea, which is a small but central region in the distribution geography of soybean, were genotyped to provide sufficient data to underpin genome-wide population genetic questions. After removing natural hybrids and duplicated or redundant accessions, we obtained a non-redundant set comprising 1,957 domesticated and 1,079 wild accessions to perform population structure analyses. Our analysis demonstrates that while wild soybean germplasm will require additional sampling from diverse indigenous areas to expand the germplasm base, the current domesticated soybean germplasm is saturated in terms of genetic diversity. We then showed that our genome-wide polymorphism map enabled us to detect genetic loci underling flower color, seed-coat color, and domestication syndrome. A representative soybean set consisting of 194 accessions were divided into one domesticated subpopulation and four wild subpopulations that could be traced back to their geographic collection areas. Population genomics analyses suggested that the monophyletic group of domesticated soybeans was originated in eastern Japan. The results were further substantiated by a phylogenetic tree constructed from domestication-associated single nucleotide polymorphisms identified in this study.

plant biology

Natural Genetic Variation Modifies Gene Expression Dynamics at the Protein Level During Pheromone Response in Saccharomyces cerevisiae

Heritable variation in gene expression patterns plays a fundamental role in trait variation and evolution, making understanding the mechanisms by which genetic variation acts on gene expression patterns a major goal for biology. Both theoretical and empirical work have largely focused on variation in steady-state mRNA levels and mRNA synthesis rates, particularly of protein-coding genes. Yet in order for this variation to affect higher order traits it must lead to differences at the protein level. Variation in protein-specific processes including protein synthesis rates and protein decay rates could amplify, mask, or even reverse effects transmitted from the transcript level, but the extent to which this happens is unclear. Moreover, mechanisms that underlie protein expression variation under dynamic conditions have not been examined. To address this challenge, we analyzed how mRNA and protein expression dynamics covary between two strains of Saccharomyces cerevisiae during mating pheromone response. Although divergent steady-state mRNA expression levels explained divergent steady-state protein levels for four out of five genes in our study, the same was true for only one out of five genes for expression dynamics. By integrating decay rate and allele-specific protein expression analyses, we resolved that expression divergence for Fig1p was caused by genetic variation acting in trans on protein synthesis rate, expression divergence for Ina1p was caused by cis-by-trans epistatic effects on transcript level and protein synthesis rate, and expression divergence for Fus3p and Tos6p were caused by divergence in protein synthesis rates. Our study demonstrates that steady-state analysis of gene expression is insufficient to understand the impact of genetic variation on gene expression variation. An integrated and dynamic approach to gene expression analysis - comparing mRNA levels, protein levels, protein decay rates, and allele-specific protein expression - allows for a detailed analysis of the genetic mechanisms underlying protein expression divergences.

genetics