Search bioRxivSearch

Biology subjects

Taylor, J.

Publications and source records attributed to Taylor, J..

At least 19 recordsLinked to original sources

High throughput droplet single-cell Genotyping of Transcriptomes (GoT) reveals the cell identity dependency of the impact of somatic mutations

Defining the transcriptomic identity of clonally related malignant cells is challenging in the absence of cell surface markers that distinguish cancer clones from one another or from admixed non-neoplastic cells. While single-cell methods have been devised to capture both the transcriptome and genotype, these methods are not compatible with droplet-based single-cell transcriptomics, limiting their throughput. To overcome this limitation, we present single-cell Genotyping of Transcriptomes (GoT), which integrates cDNA genotyping with high-throughput droplet-based single-cell RNA-seq. We further demonstrate that multiplexed GoT can interrogate multiple genotypes for distinguishing subclonal transcriptomic identity. We apply GoT to 26,039 CD34+ cells across six patients with myeloid neoplasms, in which the complex process of hematopoiesis is corrupted by CALR-mutated stem and progenitor cells. We define high-resolution maps of malignant versus normal hematopoietic progenitors, and show that while mutant cells are comingled with wildtype cells throughout the hematopoietic progenitor landscape, their frequency increases with differentiation. We identify the unfolded protein response as a predominant outcome of CALR mutations, with significant cell identity dependency. Furthermore, we identify that CALR mutations lead to NF-{kappa}B pathway upregulation specifically in uncommitted early stem cells. Collectively, GoT provides high-throughput linkage of single-cell genotypes with transcriptomes and reveals that the transcriptional output of somatic mutations is heavily dependent on the native cell identity.

cancer biology

TADs pair homologous chromosomes to promote interchromosomal gene regulation

Homologous chromosomes colocalize to regulate gene expression in processes including genomic imprinting and X-inactivation, but the mechanisms driving these interactions are poorly understood. In Drosophila, homologous chromosomes pair throughout development, promoting an interchromosomal gene regulatory mechanism called transvection. Despite over a century of study, the molecular features that facilitate chromosome-wide pairing are unknown. The \"button\" model of pairing proposes that specific regions along chromosomes pair with a higher affinity than their surrounding regions, but only a handful of DNA elements that drive homologous pairing between chromosomes have been described. Here, we identify button loci interspersed across the fly genome that have the ability to pair with their homologous sequences. Buttons are characterized by topologically associated domains (TADs), which drive pairing with their endogenous loci from multiple locations in the genome. Fragments of TADs do not pair, suggesting a model in which combinations of elements interspersed along the full length of a TAD are required for pairing. Though DNA-binding insulator proteins are not associated with pairing, buttons are enriched for insulator cofactors, suggesting that these proteins may mediate higher order interactions between homologous TADs. Using a TAD spanning the spinelessd gene as a paradigm, we find that pairing is necessary but not sufficient for transvection. spineless pairing and transvection are cell-type-specific, suggesting that local buttoning and unbuttoning regulates transvection efficiency between cell types. Together, our data support a model in which specialized TADs button homologous chromosomes together to facilitate cell-type-specific interchromosomal gene regulation.

molecular biology

Response of extremophile microbiome to a rare rainfall reveals a two-step adaptation mechanism

Understanding the mechanisms underlying microbial resistance and resilience to perturbations is essential to predict the impact of climate change on Earths ecosystems. However, the resilience and adaptation mechanisms of microbial communities to natural perturbations remain relatively unexplored, particularly in extreme environments. The response of an extremophile community inhabiting halite (salt rocks) in the Atacama Desert to a catastrophic rainfall provided the opportunity to characterize and de-convolute the temporal response of a highly specialized community to a major disturbance. With shotgun metagenomic sequencing, we investigated the halite microbiome taxonomic composition and functional potential over a 4-year longitudinal study, uncovering the dynamics of the initial response and of the recovery of the community after a rainfall event. The observed changes can be recapitulated by two general modes of community shifts - a rapid Type 1 shift and a more gradual Type 2 adjustment. In the initial response, the community entered an unstable intermediate state after stochastic niche re-colonization, resulting in broad predicted protein adaptations to increased water availability. In contrast, during recovery, the community returned to its former functional potential by a gradual shift in abundances of the newly acquired taxa. The general characterization and proposed quantitation of these two modes of community response could potentially be applied to other ecosystems, providing a theoretical framework for prediction of taxonomic and functional flux following environmental changes.

microbiology

Global climate changes will lead to regionally divergent trajectories for ectomycorrhizal communities in North American Pinaceae forests

Ectomycorrhizal fungi (ECMF) are partners in a globally distributed tree symbiosis that enhanced ecosystem carbon (C)-sequestration and storage. However, resilience of ECMF to future climates is uncertain. We sampled ECMF across a broad climatic gradient in North America, modeled climatic drivers of diversity and community composition, and then forecast ECMF response to climate changes over the next 50 years. We predict ECMF richness will decline over nearly half of North American Pinaceae forests, with median species losses as high as 21%. Mitigation of greenhouse gas emissions can reduce these declines, but not prevent them. Warming of forests along the boreal-temperate ecotone results in projected ECMF species loss and declines in the relative abundance of C demanding, long-distance foraging ECMF species, but warming of eastern temperate forests has the opposite effect. Sites with more ECMF species had higher activities of nitrogen-mineralizing enzymes, suggesting that ECMF species-losses will compromise their associated ecosystem functions.

ecology

Point of care influenza testing using Alere-i Influenza A & B assay: a practical assessment

We assessed the utility of the Alere i Influenza A & B point of care influenza test (Ai-POCIT) with laboratory testing using RT-PCR. 270 adult hospital patients had both Ai-POCIT and laboratory influenza tests conducted on the same sample. Overall, 30% and 32% influenza tests were positive by Ai-POCIT and RT-PCR, respectively. The sensitivity of the Ai-POCIT for influenza A, influenza B and any influenza were 93%, 100%, and 95%, respectively. Specificity was 100% for both viruses, but an 11% test failure rate indicates the need for better training of users. We believe that the use of nasopharyngeal (NP) swabs resulted in the observed high performance of the Ai-POCIT in comparison to other published studies. Ai-POCIT was regarded as very useful by front line clinical staff for clinical decision making and acute bed management.

microbiology

Optical and physical mapping with local finishing enables megabase-scale resolution of agronomically important regions in the wheat genome

BackgroundNumerous scaffold-level sequences for wheat are now being released and, in this context, we report on a strategy for improving the overall assembly to a level comparable to that of the human genome.\n\nResultsUsing chromosome 7A of wheat as a model, sequence-finished megabase scale sections of this chromosome were established by combining a new independent assembly based on a BAC-based physical map, BAC pool paired end sequencing, chromosome arm specific mate-pair sequencing and Bionano optical mapping with the IWGSC RefSeq v1.0 sequence and its underlying raw data. The combined assembly results in 18 super-scaffolds across the chromosome. The value of finished genome regions is demonstrated for two approximately 2.5 Mb regions associated with yield and the grain quality phenotype of fructan carbohydrate grain levels. In addition, the 50 Mb centromere region analysis incorporates cytological data highlighting the importance of non-sequence data in the assembly of this complex genome region.\n\nConclusionsSufficient genome sequence information is shown to be now available for the wheat community to produce sequence-finished releases of each chromosome of the reference genome. The high-level completion identified that an array of seven fructosyl transferase genes underpins grain quality and yield attributes are affected by five f-box-only-protein-ubiquitin ligase domain and four root-specific lipid transfer domain genes. The completed sequence also includes the centromere.

genomics

Thyroid hormone signaling specifies cone subtypes in human retinal organoids

The mechanisms underlying the specification of diverse neuronal subtypes within the human nervous system are largely unknown. The blue (shortwavelength/S), green (medium-wavelength/M) and red (long-wavelength/L) cone photoreceptors of the human retina enable high-acuity daytime vision and trichromatic color perception. Cone subtypes are specified in a poorly understood two-step process, with a first decision between S and L/M fates, followed by a decision between L and M fates. To determine the mechanism controlling S vs. L/M fates, we studied the differentiation of human retinal organoids. We found that human organoids and retinas have similar distributions, gene expression profiles, and morphologies of cone subtypes. We found that S cones are specified first, followed by L/M cones, and that thyroid hormone signaling is necessary and sufficient for this temporal switch. Temporally dynamic expression of thyroid hormone degrading and activating proteins supports a model in which the retina itself controls thyroid hormone levels, ensuring low signaling early to specify S cones and high signaling late to produce L/M cones. This work establishes organoids as a model for determining the mechanisms of cell fate specification during human development.\n\nOne sentence summaryCone specification in human organoids

developmental biology

Kluyveromyces marxianus as a robust synthetic biology platform host

Throughout history, the yeast Saccharomyces cerevisiae has played a central role in human society due to its use in food production and more recently as a major industrial and model microorganism, because of the many genetic and genomic tools available to probe its biology. However S. cerevisiae has proven difficult to engineer to expand the carbon sources it can utilize, the products it can make, and the harsh conditions it can tolerate in industrial applications. Other yeasts that could solve many of these problems remain difficult to manipulate genetically. Here, we engineer the thermotolerant yeast Kluyveromyces marxianus to create a new synthetic biology platform. Using CRISPR-Cas9 mediated genome editing, we show that wild isolates of K. marxianus can be made heterothallic for sexual crossing. By breeding two of these mating-type engineered K. marxianus strains, we combined three complex traits- thermotolerance, lipid production, and facile transformation with exogenous DNA-into a single host. The ability to cross K. marxianus strains with relative ease, together with CRISPR-Cas9 genome editing, should enable engineering of K. marxianus isolates with promising lipid production at temperatures far exceeding those of other fungi under development for industrial applications. These results establish K. marxianus as a synthetic biology platform comparable to S. cerevisiae, with naturally more robust traits that hold potential for the industrial production of renewable chemicals.

synthetic biology

Understanding trivial challenges of microbial genomics: An assembly example

The perceived \"simplicity\" of bacterial genomics (these genomes are small and easy to assemble) feeds the decentralized state of the field where computational analysis standards have been slow to evolve. This situation has a historical explanation. In cases of human, mouse, fly, worm and other model organisms there have been large sustained multinational genome sequencing efforts and analysis consortia such as the 1,000 genomes, ENCODE, modENCODE, GTEx and others. These resulted in development and proliferation of common tools, workflows, and data standards. This is not the case in microbiology. After the development of highly parallel sequencing methodologies in mid-2000s bacterial genomes no longer required initiatives of such scale. The flipside of this is the extreme heterogeneity of approaches to many well established microbial genomic analysis problems such as genome assembly. While competition amongst different methods is good, we argue that the quality of data analyses will improve if cutting edge tools are more accessible and microbiologists become more computationally savvy. Here we use genome assembly as an example to highlight current challenges and to provide a possible solution.

microbiology

Correct CYFIP1 dosage is essential for synaptic inhibition and the excitatory/inhibitory balance.

Altered excitatory/inhibitory balance is implicated in neuropsychiatric disorders but the genetic aetiology of this is still poorly understood. Copy number variations in CYFIP1 are associated with autism, schizophrenia and intellectual disability but the role of CYFIP1 in regulating synaptic inhibition or excitatory/inhibitory balance remains unclear. We show, CYFIP1, and its paralogue CYFIP2, are enriched at inhibitory postsynaptic sites. While upregulation of CYFIP1 or CYFIP2 increased excitatory synapse number and the frequency of miniature excitatory postsynaptic currents (mEPSCs), it had the opposite effect at inhibitory synapses, decreasing their size and the amplitude of miniature inhibitory postsynaptic currents (mIPSCs). Contrary to CYFIP1 upregulation, its loss in vivo, upon conditional knockout in neocortical principal cells, increased expression of postsynaptic GABAA receptor {beta}2/3-subunits and neuroligin 3 and enhanced synaptic inhibition. Thus, CYFIP1 dosage can bi-directionally impact inhibitory synaptic structure and function, potentially leading to altered excitatory/inhibitory balance and circuit dysfunction in CYFIP1-associated neurodevelopmental disorders.

neuroscience

MetaWRAP - a flexible pipeline for genome-resolved metagenomic data analysis

BackgroundThe study of microbiomes using whole-metagenome shotgun sequencing enables the analysis of uncultivated microbial populations that may have important roles in their environments. Extracting individual draft genomes (bins) facilitates metagenomic analysis at the single genome level. Software and pipelines for such analysis have become diverse and sophisticated, resulting in a significant burden for biologists to access and use them. Furthermore, while bin extraction algorithms are rapidly improving, there is still a lack of tools for their evaluation and visualization.\n\nResultsTo address these challenges, we present metaWRAP, a modular pipeline software for shotgun metagenomic data analysis. MetaWRAP deploys state-of-the-art software to handle metagenomic data processing starting from raw sequencing reads and ending in metagenomic bins and their analysis. MetaWRAP is flexible enough to give investigators control over the analysis, while still being easy-to-install and easy-to-use. It includes hybrid algorithms that leverage the strengths of a variety of software to extract and refine high-quality bins from metagenomic data through bin consolidation and reassembly. MetaWRAPs hybrid bin extraction algorithm outperforms individual binning approaches and other bin consolidation programs in both synthetic and real datasets. Finally, metaWRAP comes with numerous modules for the analysis of metagenomic bins, including taxonomy assignment, abundance estimation, functional annotation, and visualization.\n\nConclusionsMetaWRAP is an easy-to-use modular pipeline that automates the core tasks in metagenomic analysis, while contributing significant improvements to the extraction and interpretation of high-quality metagenomic bins. The bin refinement and reassembly modules of metaWRAP consistently outperform other binning approaches. Each module of metaWRAP is also a standalone component, making it a flexible and versatile tool for tackling metagenomic shotgun sequencing data. MetaWRAP is open-source software available at https://github.com/bxlab/metaWRAP.

microbiology

Lamins organize the global three-dimensional genome from the nuclear periphery

Lamins are structural components of the nuclear lamina (NL) that regulate genome organization and gene expression, but the mechanism remains unclear. Using Hi-C, we show that lamins maintain proper interactions among the topologically associated chromatin domains (TADs) but not their overall architecture. Combining Hi-C with fluorescence in situ hybridization (FISH) and analyses of lamina-associated domains (LADs), we reveal that lamin loss causes expansion or detachment of specific LADs in mouse ES cells. The detached LADs disrupt 3D interactions of both LADs and interior chromatin. 4C and epigenome analyses further demonstrate that lamins maintain the active and repressive chromatin domains among different TADs. By combining these studies with transcriptome analyses, we found a significant correlation between transcription changes and the changes of active and inactive chromatin domain interactions. These findings provide a foundation to further study how the nuclear periphery impacts genome organization and transcription in development and NL-associated diseases.\n\nHighlightsO_LILamin loss does not affect the overall TAD structure but alters TAD-TAD interactions\nC_LIO_LILamin null ES cells exhibit decondensation or detachment of specific LAD regions\nC_LIO_LIExpansion and detachment of LADs can alter genome-wide 3D chromatin interactions\nC_LIO_LIAltered chromatin domain interactions are correlated with altered transcription\nC_LI

genomics

QuASAR: Quality Assessment of Spatial Arrangement Reproducibility in Hi-C Data

Hi-C has revolutionized global interrogation of chromosome conformation, however there are few tools to assess the reliability of individual experiments. Here we present a new approach, QuASAR, for measuring quality within and between Hi-C samples. We show that QuASAR can detect even tiny fractions of noise and estimate both return on additional sequencing and quality upper bounds. We also demonstrate QuASAR's utility in measuring replicate agreement across feature resolutions. Finally, QuASAR can estimate resolution limits based on both internal and replicate quality scores. QuASAR provides an objective means of Hi-C sample comparison while providing context and limits to these measures.

bioinformatics

Practical computational reproducibility in the life sciences

Many areas of research suffer from poor reproducibility. This problem is particularly acute in computationally intensive domains where results rely on a series of complex methodological decisions that are not well captured by traditional publication approaches. Various guidelines have emerged for achieving reproducibility, but practical implementation of these practices remains difficult. This is because reproducing published computational analyses requires installing many software tools plus associated libraries, connecting tools together into the complete pipeline, and specifying parameters. Here we present a suite of recently emerged technologies which make computational reproducibility not just possible, but, finally, practical in both time and effort. By combining a system for building highly portable packages of bioinformatics software, containerization and virtualization technologies for isolating reusable execution environments for these packages, and an integrated workflow system that automatically orchestrates the composition of these packages for entire pipelines, an unprecedented level of computational reproducibility can be achieved.

bioinformatics

A near complete haplotype-phased genome of the dikaryotic wheat stripe rust fungus Puccinia striiformis f. sp. tritici reveals high inter-haplome diversity

A long-standing biological question is how evolution has shaped the genomic architecture of dikaryotic fungi. To answer this, high quality genomic resources that enable haplotype comparisons are essential. Short-read genome assemblies for dikaryotic fungi are highly fragmented and lack haplotype-specific information due to the high heterozygosity and repeat content of these genomes. Here we present a diploidaware assembly of the wheat stripe rust fungus Puccinia striiformis f. sp. tritici based on long-reads using the FALCON-Unzip assembler. RNA-seq datasets were used to infer high quality gene models and identify virulence genes involved in plant infection referred to as effectors. This represents the most complete Puccinia striiformis f. sp. tritici genome assembly to date (83 Mb, 156 contigs, N50 1.5 Mb) and provides phased haplotype information for over 92% of the genome. Comparisons of the phase blocks revealed high inter-haplotype diversity of over 6%. More than 25% of all genes lack a clear allelic counterpart. When investigating genome features that potentially promote the rapid evolution of virulence, we found that candidate effector genes are spatially associated with conserved genes commonly found in basidiomycetes. Yet candidate effectors that lack an allelic counterpart are more distant from conserved genes than allelic candidate effectors, and are less likely to be evolutionarily conserved within the P. striiformis species complex and Pucciniales. In summary, this haplotype-phased assembly enabled us to discover novel genome features of a dikaryotic plant pathogenic fungus previously hidden in collapsed and fragmented genome assemblies.\n\nImportanceCurrent representations of eukaryotic microbial genomes are haploid, hiding the genomic diversity intrinsic to diploid and polyploid life forms. This hidden diversity contributes to the organisms evolutionary potential and ability to adapt to stress conditions. Yet it is challenging to provide haplotype-specific information at a whole-genome level. Here, we take advantage of long-read DNA sequencing technology and a tailored-assembly algorithm to disentangle the two haploid genomes of a dikaryotic pathogenic wheat rust fungus. The two genomes display high levels of nucleotide and structural variations, which leads to allelic variation and the presence of genes lacking allelic counterparts. Non-allelic candidate effector genes, which likely encode important pathogenicity factors, display distinct genome localization patterns and are less likely to be evolutionary conserved than those which are present as allelic pairs. This genomic diversity may promote rapid host adaptation and/or be related to the age of the sequenced isolate since last meiosis.

microbiology

Paternal-age-related de novo mutations and risk for five disorders

BackgroundThere are well-established epidemiologic associations between advanced paternal age and increased offspring risk for several psychiatric and developmental disorders. These associations are commonly attributed to age-related de novo mutations. However, the actual magnitude of risk conferred by age-related de novo mutations in the male germline is unknown. Quantifying this risk would clarify the clinical and public health significance of delayed paternity.\n\nMethodsUsing results from large, parent-child trio whole-exome-sequencing studies, we estimated the relationship between paternal-age-related de novo single nucleotide variants (dnSNVs) and offspring risk for five disorders: autism spectrum disorders (ASD), congenital heart disease (CHD), neurodevelopmental disorders with epilepsy (EPI), intellectual disability (ID), and schizophrenia (SCZ). Using Danish national registry data, we then investigated the degree to which the epidemiologic association between each disorder and advanced paternal age was consistent with the estimated role of de novo mutations.\n\nResultsIncidence rate ratios comparing dnSNV-based risk to offspring of 45 versus 25-year-old fathers ranged from 1.05 (95% confidence interval 1.01-1.13) for SCZ to 1.29 (95% CI 1.13-1.68) for ID. Epidemiologic estimates of paternal age risk for CHD, ID and EPI were consistent with the dnSNV effect. However, epidemiologic effects for ASDs and SCZ significantly exceeded the risk that could be explained by dnSNVs alone (p<2e-4 for both comparisons).\n\nConclusionIncreasing dnSNVs due to advanced paternal age confer a small amount of offspring risk for psychiatric and developmental disorders. For ASD and SCZ, epidemiologic associations with delayed paternity largely reflect factors that cannot be assumed to increase with age.

genetics

Measuring the reproducibility and quality of Hi-C data

Hi-C is currently the most widely used assay to investigate the 3D organization of the genome and to study its role in gene regulation, DNA replication, and disease. However, Hi-C experiments are costly to perform and involve multiple complex experimental steps; thus, accurate methods for measuring the quality and reproducibility of Hi-C data are essential to determine whether the output should be used further in a study. Using real and simulated data, we profile the performance of several recently proposed methods for assessing reproducibility of population Hi-C data, including HiCRep, GenomeDISCO, HiC-Spector and QuASAR-Rep. By explicitly controlling noise and sparsity through simulations, we demonstrate the deficiencies of performing simple correlation analysis on pairs of matrices, and we show that methods developed specifically for Hi-C data produce better measures of reproducibility. We also show how to use established (e.g., ratio of intra to interchromosomal interactions) and novel (e.g., QuASAR-QC) measures to identify low quality experiments. In this work, we assess reproducibility and quality measures by varying sequencing depth, resolution and noise levels in Hi-C data from 13 cell lines, with two biological replicates each, as well as 176 simulated matrices. Through this extensive validation and benchmarking of Hi-C data, we describe best practices for reproducibility and quality assessment of Hi-C experiments. We make all software publicly available at http://github.com/kundajelab/3DChromatin_ReplicateQC to facilitate adoption in the community.

genomics

Natural variation in stochastic photoreceptor specification and color preference in Drosophila

Each individual perceives the world in a unique way, but little is known about the genetic basis of variation in sensory perception. Here we investigated natural variation in the development and function of the color vision system of Drosophila. In the fly eye, the random mosaic of color-detecting R7 photoreceptor subtypes is determined by stochastic expression of the transcription factor Spineless (Ss). Individual R7s randomly choose between SsON or SsOFF fates at a ratio of 65:35, resulting in unique patterns but consistent proportions of cell types across genetically identical retinas. In a genome wide association study, we identified a naturally occurring insertion in a regulatory DNA element in the ss gene that lowers the ratio of SsON to SsOFF cells. This change in photoreceptor fates shifts the innate color preference of flies from green to blue. The genetic variant increases the binding affinity for Klumpfuss (Klu), a zinc finger transcriptional repressor that regulates ss expression. Klu is expressed at intermediate levels to determine the normal ratio of SsON to SsOFF cells. Thus, binding site affinity and transcription factor levels are finely tuned to regulate stochastic on/off gene expression, setting the ratio of alternative cell fates and ultimately determining color preference.

developmental biology