Search bioRxivSearch

Biology subjects

Church, G. M.

Publications and source records attributed to Church, G. M..

At least 19 recordsLinked to original sources

The whale shark genome reveals how genomic and physiological properties scale with body size

The endangered whale shark (Rhincodon typus) is the largest fish on Earth and is a long-lived member of the ancient Elasmobranchii clade. To characterize the relationship between genome features and biological traits, we sequenced and assembled the genome of the whale shark and compared its genomic and physiological features to those of 81 animals and yeast. We examined scaling relationships between body size, temperature, metabolic rates, and genomic features and found both general correlations across the animal kingdom and features specific to the whale shark genome. Among animals, increased lifespan is positively correlated to body size and metabolic rate. Several genomic features also significantly correlated with body size, including intron and gene length. Our large-scale comparative genomic analysis uncovered general features of metazoan genome architecture: GC content and codon adaptation index are negatively correlated, and neural connectivity genes are longer than average genes in most genomes. Focusing on the whale shark genome, we identified multiple features that significantly correlate with lifespan. Among these were very long gene length, due to large introns highly enriched in repetitive elements such as CR1-like LINEs, and considerably longer neural genes of several types, including connectivity, activity, and neurodegeneration genes. The whale sharks genome had an expansion of gene families related to fatty acid metabolism and neurogenesis, with the slowest evolutionary rate observed in vertebrates to date. Our comparative genomics approach uncovered multiple genetic features associated with body size, metabolic rate, and lifespan, and showed that the whale shark is a promising model for studies of neural architecture and lifespan.

genomics

Enzymatic DNA synthesis for digital information storage

DNA is an emerging storage medium for digital data but its adoption is hampered by limitations of phosphoramidite chemistry, which was developed for single-base accuracy required for biological functionality. Here, we establish a de novo enzymatic DNA synthesis strategy designed from the bottom-up for information storage. We harness a template-independent DNA polymerase for controlled synthesis of sequences with user-defined information content. We demonstrate retrieval of 144-bits, including addressing, from perfectly synthesized DNA strands using batch-processed Illumina and real-time Oxford Nanopore sequencing. We then develop a codec for data retrieval from populations of diverse but imperfectly synthesized DNA strands, each with a ~30% error tolerance. With this codec, we experimentally validate a kilobyte-scale design which stores 1 bit per nucleotide. Simulations of the codec support reliable and robust storage of information for large-scale systems. This work paves the way for alternative synthesis and sequencing strategies to advance information storage in DNA.

synthetic biology

Toward machine-guided design of proteins

Proteins--molecular machines that underpin all biological life--are of significant therapeutic and industrial value. Directed evolution is a high-throughput experimental approach for improving protein function, but has difficulty escaping local maxima in the fitness landscape. Here, we investigate how supervised learning in a closed loop with DNA synthesis and high-throughput screening can be used to improve protein design. Using the green fluorescent protein (GFP) as an illustrative example, we demonstrate the opportunities and challenges of generating training datasets conducive to selecting strongly generalizing models. With prospectively designed wet lab experiments, we then validate that these models can generalize to unseen regions of the fitness landscape, even when constrained to explore combinations of non-trivial mutations. Taken together, this suggests a hybrid optimization strategy for protein design in which a predictive model is used to explore difficult-to-access but promising regions of the fitness landscape that directed evolution can then exploit at scale.

synthetic biology

Establishing a cell-free Vibrio natriegens expression system

The fast growing bacterium Vibrio natriegens is an emerging microbial host for biotechnology. Harnessing its productive cellular components may offer a compelling platform for rapid protein production and prototyping of metabolic pathways or genetic circuits. Here, we report the development of a V. natriegens cell-free expression system. We devised a simplified crude extract preparation protocol and achieved >260 g/mL of super-folder GFP in a small-scale batch reaction after three hours. Culturing conditions, including growth media and cell density, significantly affect translation kinetics and protein yield of extracts. We observed maximal protein yield at incubation temperatures of 26{degrees}C or 30{degrees}C, and show improved yield by tuning ions crucial for ribosomal stability. This work establishes an initial V. natriegens cell-free expression system, enables probing of V. natriegens biology, and will serve as a platform to accelerate metabolic engineering and synthetic biology applications.

synthetic biology

Circularization of genes and chromosome by CRISPR in human cells

Extrachromosomal circular DNA (eccDNA) and ring chromosomes are genetic alterations found in humans with genetic disorders and diseases such as cancer. However, there is a lack of genetic engineering tool to recapitulate these features. Here, we report the discovery that delivery of pairs of CRISPR/Cas9 guide RNAs into human cells generate functional eccDNAs and ring chromosomes. We generated a dual-fluorescence eccDNA biosensor system, which allows us to study CRISPR deletion, inversion, and circularization of genes inside cells. Analysis after CRISPR editing at intergenic and genic loci in human embryonic kidney 293T cells and human mammary fibroblasts reveal that CRISPR deleted DNA readily form eccDNA in human cells. DNA in sizes from a few hundred base pairs up to a 47.4 megabase-sized ring chromosome (chr18) can be circularized. Our discoveries advance and expand CRISPR-Cas9 technology applications for genetic engineering, modeling of human diseases, and chromosome engineering.\n\nOne Sentence Summary: CRISPR circularization of DNA offers new tools for studying eccDNA biogenesis, function, chromosome engineering, and synthetic biology.

genomics

Enhanced bacterial immunity and mammalian genome editing via RNA polymerase-mediated dislodging of Cas9 from double strand DNA breaks.

The ability to target the Cas9 nuclease to DNA sequences via Watson-Crick base pairing with a single guide RNA (sgRNA) has provided a dynamic tool for genome editing and an essential component of adaptive immune systems in bacteria. After generating a double strand break (DSB), Cas9 remains stably bound to it. Here we show persistent Cas9 binding blocks access to DSB by repair enzymes, reducing genome editing efficiency. Cas9 can be dislodged by translocating RNA polymerases, but only if the polymerase approaches one direction towards the Cas9-DSB complex. By exploiting these RNA polymerase-Cas9 interactions, Cas9 can be conditionally converted into a multi-turnover nuclease, mediating increased mutagenesis frequencies in mammalian cells and enhancing bacterial immunity to bacteriophages. These consequences of a stable Cas9-DSB complex provide insights into the evolution of PAM sequences and a simple method of improving selection of highly active sgRNA for genome editing.

molecular biology

Barcoded oligonucleotides ligated on RNA amplified for multiplex and parallel in-situ analyses

We present Barcoded Oligonucleotides Ligated On RNA Amplified for Multiplexed and parallel In-Situ analysis (BOLORAMIS), a reverse-transcription (RT)-free method for spatially-resolved, targeted, in-situ RNA identification of single or multiple targets. For this proof of concept, we have profiled 154 distinct coding and small non-coding transcripts ranging in sizes 18 nucleotides in length and upwards, from over 200, 000 individual human induced pluripotent stem cells (iPSC) and demonstrated compatibility with multiplexed detection, enabled by fluorescent in-situ sequencing. We use BOLORAMIS data to identify differences in spatial localization and cell-to-cell expression heterogeneity. Our results demonstrate BOLORAMIS to be a generalizable toolset for targeted, in-situ detection of coding and small non-coding RNA for single or multiplexed applications.

cell biology

A homing CRISPR mouse resource for barcoding and lineage tracing

Cellular barcoding using nuclease-induced genetic mutations is an effective approach that is emerging for recording biological information, including developmental lineages. We have previously introduced the homing CRISPR system as a promising methodology for generating such barcodes with scalable diversity and without crosstalk. Here, we present a mouse line (MARC1) with multiple genomically-integrated and heritable homing guide RNAs (hgRNAs). We determine the genomic locations of these hgRNAs, their activity profiles during gestation, and the diversity of their mutants. We apply the line for unique barcoding of mouse embryos and differential barcoding of embryonic tissues. We conclude that this mouse line can address the unique challenges associated with in vivo barcoding in mammalian model organisms and is thus an enabling platform for recording and lineage tracing applications in a mammalian model system.

synthetic biology

Multiplexed imaging using same species primary antibodies with signal amplification

Immunofluorescence (IF) imaging using antibodies to visualize specific biomolecules is a widely used technique in both biological and clinical laboratories. Standard IF imaging methods using primary antibodies followed by secondary antibodies have low multiplexing capability due to limited availability of primary antibodies raised in different animal species. Here, we used a DNA-based signal amplification method, Hybridization Chain Reaction (HCR), to replace secondary antibodies to achieve multiplexed imaging using primary antibodies of the same species with superior signal intensity. To enable imaging with DNA-conjugated antibodies, we developed a new antibody staining protocol to minimize nonspecific binding of antibodies caused by conjugated DNA oligonucleotides. We also expanded the HCR hairpin pool from previously published 5 to 13 for highly multiplexed in situ imaging. We finally demonstrated multiplexed in situ protein imaging using the technique in both cultured cells and mouse retina sections.

cell biology

Significant abundance of cis configurations of mutations in diploid human genomes

To fully understand human genetic variation, one must assess the specific distribution of variants between the two chromosomal homologues of genes, and any functional units of interest, as the phase of variants can significantly impact gene function and phenotype. To this end, we have systematically analyzed 18,121 autosomal protein-coding genes in 1,092 statistically phased genomes from the 1000 Genomes Project, and an unprecedented number of 184 experimentally phased genomes from the Personal Genome Project. Here we show that mutations predicted to functionally alter the protein, and coding variants as a whole, are not randomly distributed between the two homologues of a gene, but do occur significantly more frequently in cis-than trans-configurations, with cis/trans ratios of [~]60:40. Significant cis-abundance was observed in virtually all individual genomes in all populations. Nearly all variable genes exhibited either cis, or trans configurations of protein-altering mutations in significant excess, allowing distinction of cis- and trans-abundant genes. These common patterns of phase were largely constituted by a shared, global set of phase-sensitive genes. We show significant enrichment of this global set with gene sets indicating its involvement in adaptation and evolution. Moreover, cis- and trans-abundant genes were found functionally distinguishable, and exhibited strikingly different distributional patterns of protein-altering mutations. This work establishes common patterns of phase as key characteristics of diploid human exomes and provides evidence for their potential functional significance. Thus, it highlights the importance of phase for the interpretation of protein-coding genetic variation, challenging the current conceptual and functional interpretation of autosomal genes.

genomics

Current CRISPR gene drive systems are likely to be highly invasive in wild populations

Recent reports have suggested that CRISPR-based gene drives are unlikely to invade wild populations due to drive-resistant alleles that prevent cutting. Here we develop mathematical models based on existing empirical data to explicitly test this assumption. We show that although resistance prevents drive systems from spreading to fixation in large populations, even the least effective systems reported to date are highly invasive. Releasing a small number of organisms often causes invasion of the local population, followed by invasion of additional populations connected by very low gene flow rates. Examining the effects of mitigating factors including standing variation, inbreeding, and family size revealed that none of these prevent invasion in realistic scenarios. Highly effective drive systems are predicted to be even more invasive. Contrary to the National Academies report on gene drive, our results suggest that standard drive systems should not be developed nor field-tested in regions harboring the host organism.

synthetic biology

Single-cell sequencing reveals αβ chain pairing shapes the T cell repertoire

A diverse T cell repertoire is a critical component of the adaptive immune system, providing protection against invading pathogens and neoplastic changes, relying on the recognition of foreign antigens and neoantigen peptides by T cell receptors (TCRs). However, the statistical properties and function of the T cell pool in an individual, under normal physiological conditions, are poorly understood. In this study, we report a comprehensive, quantitative characterization of the T cell repertoire from over 1.9 million cells, yielding over 200,000 high quality paired {beta} sequences in 5 healthy human subjects. The dataset was obtained by leveraging recent biotechnology developments in deep RNA sequencing of lymphocytes via single-cell barcoding in emulsion. We report non-random associations and non-monogamous pairing between the and {beta} chains, lowering the theoretical diversity of the T cell repertoire, and increasing the frequency of public clones shared among individuals. T cell clone size distributions closely followed a power law, with markedly longer tails for CD8+ cytotoxic T cells than CD4+ helper T cells. Furthermore, clonality estimates based on paired chains from single T cells were lower than that from single chain data. Taken together, these results highlight the importance of sequencing {beta} pairs to accurately quantify lymphocyte receptor diversity.

immunology

Biosensor libraries harness large classes of binding domains for allosteric transcription regulators

Bacterias ability to specifically sense small molecules in their environment and trigger metabolic responses in accordance is an invaluable biotechnological resource. While many transcription factors (TFs) mediating these processes have been studied, only a handful has been leveraged for molecular biology applications. To expand this panel of biotechnologically important sensors here we present a strategy for the construction and testing of chimeric TF libraries, based on the fusion of highly soluble periplasmic binding proteins (PBPs) with DNA-binding domains (DBDs). We validated this strategy by constructing and functionally testing two unique sense-and-response regulators for benzoate, an environmentally and industrially relevant metabolite. This work will enable the development of tailored biosensors for synthetic regulatory circuits.

synthetic biology

Efficient in situ barcode sequencing using padlock probe-based BaristaSeq

Cellular DNA/RNA tags (barcodes) allow for multiplexed cell lineage tracing and neuronal projection mapping with cellular resolution. Conventional approaches to reading out cellular barcodes trade off spatial resolution with throughput. Bulk sequencing achieves high throughput but sacrifices spatial resolution, whereas manual cell picking has low throughput. In situ sequencing could potentially achieve both high spatial resolution and high throughput, but current in situ sequencing techniques are inefficient at reading out cellular barcodes. Here we describe BaristaSeq, an optimization of a targeted, padlock probe-based technique for in situ barcode sequencing compatible with Illumina sequencing chemistry. BaristaSeq results in a five-fold increase in amplification efficiency, with a sequencing accuracy of at least 97%. BaristaSeq could be used for barcode-assisted lineage tracing, and to map long-range neuronal projections.\n\nKey PointsO_LIIn situ sequencing by gap-filling padlock probes is limited by the strand displacement of DNA polymerases\nC_LIO_LIIllumina sequencing chemistry offers superior signal-to-noise ratio in situ compared to sequencing by ligation\nC_LIO_LIBaristaSeq as an accurate method for barcode sequencing in situ with improved gap-filling efficiency\nC_LI

molecular biology

Long-term adaptive evolution of genomically recoded Escherichia coli

Efforts are underway to construct several recoded genomes anticipated to exhibit multi-virus resistance, enhanced non-standard amino acid (NSAA) incorporation, and capability for synthetic biocontainment. Though we succeeded in pioneering the first genomically recoded organism (Escherichia coli strain C321.{Delta}A), its fitness is far lower than that of its non-recoded ancestor, particularly in defined media. This fitness deficit severely limits its utility for NSAA-linked applications requiring defined media such as live cell imaging, metabolic engineering, and industrial-scale protein production. Here, we report adaptive evolution of C321.{Delta}A for more than 1,000 generations in independent replicate populations grown in glucose minimal media. Evolved recoded populations significantly exceed the growth rates of both the ancestral C321.{Delta}A and non-recoded strains, permitting use of the recoded chassis in several new contexts. We use next-generation sequencing to identify genes mutated in multiple independent populations, and we reconstruct individual alleles in ancestral strains via multiplex automatable genome engineering (MAGE) to quantify their effects on fitness. Several selective mutations occur only in recoded evolved populations, some of which are associated with altering the translation apparatus in response to recoding, whereas others are not apparently associated with recoding, but instead correct for off-target mutations that occurred during initial genome engineering. This report demonstrates that laboratory evolution can be applied after engineering of recoded genomes to streamline fitness recovery compared to application of additional targeted engineering strategies that may introduce further unintended mutations. In doing so, we provide the most comprehensive insight to date into the physiology of the commonly used C321.{Delta}A strain.\n\nSignificance StatementAfter demonstrating construction of an organism with an altered genetic code, we sought to evolve this organism for many generations to improve its fitness and learn what unique changes natural selection would bestow upon it. Although this organism initially had impaired fitness, we observed that adaptive laboratory evolution resulted in several selective mutations that corrected for insufficient translation termination and for unintended mutations that occurred when originally altering the genetic code. This work further bolsters our understanding of the pliability of the genetic code, it will help guide ongoing and future efforts seeking to recode genomes, and it results in a useful strain for non-standard amino acid incorporation in numerous contexts relevant for research and industry.

evolutionary biology

Engineering post-translational proofreading to discriminate non-standard amino acids

Progress in genetic code expansion requires accurate, selective, and high-throughput detection of non-standard amino acid (NSAA) incorporation into proteins. Here, we discover how the N-end rule pathway of protein degradation applies to commonly used NSAAs. We show that several NSAAs are N-end stabilizing and demonstrate that other NSAAs can be made stabilizing by rationally engineering the N-end rule adaptor protein ClpS. We use these insights to engineer a synthetic quality control method, termed \"Post-Translational Proofreading\" (PTP). By implementing PTP, false positive proteins resulting from misincorporation of structurally similar standard amino acids or undesired NSAAs rapidly degrade, enabling high-accuracy discrimination of desired NSAA incorporation. We illustrate the utility of PTP during evolution of the biphenylalanine orthogonal translation system used for synthetic biocontainment. Our new OTS is more selective and confers lower escape frequencies and greater fitness in all tested biocontained strains. Our approach presents a new paradigm for molecular recognition of amino acids in target proteins.

synthetic biology

The experimental design and data interpretation in “Unexpected mutations after CRISPR–Cas9 editing in vivo” by Schaefer et al. are insufficient to support the conclusions drawn by the authors

To the Editor To the Editor Conflict of Interest Statement References References The recent correspondence to the Editor of Nature Methods by Schaefer et al.1 has garnered significant attention since its publication as a result of its strong conclusions contradicting numerous publications in the field using similar analytical approaches and methods2-4. The authors suggest that the CRISPR-Cas9 system is highly mutagenic in genomic regions not expected to be targeted by the gRNA. We believe that the conclusions drawn from this study are unsubstantiated by the disclosed experiments as they were designed and carried out. Further, it is impossible to ascribe the observed differences in the subject mice to the effects of CRISPR per se. The genetic differences seen in ...

genetics

Physiological Assembly Of Functionally Active 30S Ribosomal Subunits From In Vitro Synthesized Parts

Synthetic ribosomes in vitro can facilitate engineering translation of novel polymers, identifying ribosome biogenesis central components, and paving the road to constructing replicating systems from defined biochemical components. Here, we report functional synthetic Escherichia coli 30S ribosomal subunits constructed using a defined, purified cell free system under physiological conditions. We test hypotheses about key components of natural ribosome biogenesis pathway as required for efficient function - including integration of 16S rRNA modification, cofactors facilitated ribosome assembly and protein synthesis in the same compartment in vitro. We observe ~17% efficiency for fully synthetic 30S and ~70% efficiency from in vitro transcribed 16S rRNA assembled with natural proteins. We observe up to 5 fold improvement over previous crude extracts. We suggest extending the minimal list of components required for central-dogma replication from the 151 gene products previously reported to at least 180 to allow the speed and accuracy of macromolecular synthesis to approach native E. coli values.

biochemistry