Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

GenomeScope: Fast reference-free genome profiling from short reads

SummaryGenomeScope is an open-source web tool to rapidly estimate the overall characteristics of a genome, including genome size, heterozygosity rate, and repeat content from unprocessed short reads. These features are essential for studying genome evolution, and help to choose parameters for downstream analysis. We demonstrate its accuracy on 324 simulated and 16 real datasets with a wide range in genome sizes, heterozygosity levels, and error rates.\n\nAvailability and Implementationhttp://genomescope.org, https://github.com/schatzlab/genomescope.git\n\nContactmschatz@jhu.edu.\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Comparative analysis of 2D and 3D distance measurements to study spatial genome organization

The spatial organization of eukaryotic genomes is non-random, cell-type specific, and has been linked to cellular function. The investigation of spatial organization has traditionally relied extensively on fluorescence microscopy. The validity of the imaging methods used to probe spatial genome organization often depends on the accuracy and precision of distance measurements. Imaging-based measurements may either use 2 dimensional datasets or 3D datasets including the z-axis information in image stacks. Here we compare the suitability of 2D versus 3D distance measurements in the analysis of various features of spatial genome organization. We find in general good agreement between 2D and 3D analysis with higher convergence of measurements as the interrogated distance increases, especially in flat cells. Overall, 3D distance measurements are more accurate than 2D distances, but are also more prone to noise. In particular, z-stacks are prone to error due to imaging properties such as limited resolution along the z-axis and optical aberrations, and we also find significant deviations from unimodal distance distributions caused by low sampling frequency in z. These deviations can be ameliorated by sampling at much higher frequency in the z-direction. We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.

Cell Biology

PEMapper / PECaller: A simplified approach to whole-genome sequencing

The analysis of human whole-genome sequencing data presents significant computational challenges. The sheer size of datasets places an enormous burden on computational, disk array, and network resources. Here we present an integrated computational package, PEMapper/PECaller, that was designed specifically to minimize the burden on networks and disk arrays, create output files that are minimal in size, and run in a highly computationally efficient way, with the single goal of enabling whole-genome sequencing at scale. In addition to improved computational efficiency, we implement a novel statistical framework that allows for a base-by-base error model, allowing this package to perform as well or better than the widely used Genome Analysis Toolkit (GATK) in all key measures of performance on human whole-genome sequences.

Genetics

Comparative analysis of variation and selection in the HCV genome

Genotype 1 of the hepatitis C virus (HCV) is the most prevalent of the variants of this virus. Its two main subtypes, HCV-1a and HCV-1b, are associated to differences in epidemic features and risk groups, despite sharing similar features in most biological properties. We have analyzed the impact of positive selection on the evolution of these variants using complete genome coding regions, and compared the levels of genetic variability and the distribution of positively selected sites. We have also compared the distributions of positively selected and conserved sites considering different factors such as RNA secondary structure, the presence of different epitopes (antibody, CD4 and CD8), and secondary protein structure. Less than 10% of the genome was found to be under positive selection, and purifying selection was the main evolutionary force in both subtypes. We found differences in the number of positively selected sites between subtypes in several genes (Core, HVR2 in E2, P7, helicase in NS3 and NS4a). Heterozygosity values in positively selected sites and the rate of non-synonymous substitutions were significantly higher in subtype HCV-1b. Logistic regression analyses revealed that similar selective forces act at the genome level in both subtypes: RNA secondary structure and CD4 T-cell epitopes are associated with conservation, while CD8 T-cell epitopes are associated with positive selection in both subtypes. These results indicate that similar selective constraints are acting along HCV-1a and HCV-1b genomes, despite some differences in the distribution of positively selected sites at independent genes.

Evolutionary Biology

Contrasting patterns of genome-level diversity across distinct co-occurring bacterial populations

To understand the forces driving differentiation and diversification in wild bacterial populations, we must be able to delineate and track ecologically relevant units through space and time. Mapping metagenomic sequences to reference genomes derived from the same environment can reveal genetic heterogeneity within populations, and in some cases, be used to identify boundaries between genetically similar, but ecologically distinct, populations. Here we examine population-level heterogeneity within abundant and ubiquitous freshwater bacterial groups such as the acI Actinobacteria and LD12 Alphaproteobacteria (the freshwater sister clade to the marine SAR11) using 33 single cell genomes and a 5-year metagenomic time series. The single cell genomes grouped into 15 monophyletic clusters (termed \"tribes\") that share at least 97.9% 16S rRNA identity. Distinct populations were identified within most tribes based on the patterns of metagenomic read recruitments to single-cell genomes representing these tribes. Genetically distinct populations within tribes of the acI actinobacterial lineage living in the same lake had different seasonal abundance patterns, suggesting these populations were also ecologically distinct. In contrast, sympatric LD12 populations were less genetically differentiated. This suggests that within one lake, some freshwater lineages harbor genetically discrete (but still closely related) and ecologically distinct populations, while other lineages are composed of less differentiated populations with overlapping niches. Our results point at an interplay of evolutionary and ecological forces acting on these communities that can be observed in real time.

Evolutionary Biology

FIDDLE: An integrative deep learning framework for functional genomic data inference

Numerous advances in sequencing technologies have revolutionized genomics through generating many types of genomic functional data. Statistical tools have been developed to analyze individual data types, but there lack strategies to integrate disparate datasets under a unified framework. Moreover, most analysis techniques heavily rely on feature selection and data preprocessing which increase the difficulty of addressing biological questions through the integration of multiple datasets. Here, we introduce FIDDLE (Flexible Integration of Data with Deep LEarning) an open source data-agnostic flexible integrative framework that learns a unified representation from multiple data types to infer another data type. As a case study, we use multiple Saccharomyces cerevisiae genomic datasets to predict global transcription start sites (TSS) through the simulation of TSS-seq data. We demonstrate that a type of data can be inferred from other sources of data types without manually specifying the relevant features and preprocessing. We show that models built from multiple genome-wide datasets perform profoundly better than models built from individual datasets. Thus FIDDLE learns the complex synergistic relationship within individual datasets and, importantly, across datasets.

bioinformatics

Folding Principle of Chromosome Emerges from Mapping of Genome Features onto its 3D Structure

How chromosomes fold into 3D structures and how genome functions are affected or even controlled by their spatial organization remain challenging questions. Hi-C experiment has provided important structural insights for chromosome, and Hi-C data are used here to construct the 3D chromatin structure which are characterized by two spatially segregated chromatin compartments A and B. By mapping a plethora of genome features onto the constructed 3D chromatin model, we show vividly the close connection between genome properties and the spatial organization of chromatin. We are able to dissect the whole chromatin into two types of chromatin domains which have clearly different Hi-C contact patterns as well as different sizes of chromatin loops. The two chromatin types can be respectively regarded as the basic units of chromatin compartments A and B, and also spatially segregate from each other as the two chromatin compartments. Therefore, the chromatin loops segregate in the space according to their sizes, suggesting the excluded volume or entropic effect in chromatin compartmentalization as well as chromosome positioning. Taken together, these results provide clues to the folding principles of chromosomes, their spatial organization, and the resulted clustering of many genome features in the 3D space.

biophysics

Evolutionary forces affecting synonymous variations in plant genomes

Base composition is highly variable among and within plant genomes, especially at third codon positions, ranging from GC-poor and homogeneous species to GC-rich and highly heterogeneous ones (particularly Monocots). Consequently, synonymous codon usage is biased in most species, even when base composition is relatively homogeneous. The causes of these variations are still under debate, with three main forces being possibly involved: mutational bias, selection and GC-biased gene conversion (gBGC). So far, both selection and gBGC have been detected in some species but how their relative strength varies among and within species remains unclear. Population genetics approaches allow to jointly estimating the intensity of selection, gBGC and mutational bias. We extended a recently developed method and applied it to a large population genomic datasets based on transcriptome sequencing of 11 angiosperm species spread across the phylogeny. We found that base composition is far from mutation-drift equilibrium in most genomes and that gBGC is a widespread and stronger process than selection. gBGC could strongly contribute to base composition variation among plant species, implying that it should be taken into account in plant genome analyses, especially for GC-rich ones.

evolutionary biology

Optimizing complex phenotypes through model-guided multiplex genome engineering

Optimization of complex phenotypes in engineered microbial strains has traditionally been accomplished by laboratory evolution. However, only a subset of the resulting mutations may affect the phenotype of interest and many others may have unintended effects. Multiplexed genome editing can complement evolutionary approaches by creating diverse combinations of targeted changes, but in both cases it remains challenging to identify which alleles influence the desired phenotype. We present a method for identifying a minimal set of genomic modifications that optimizes a complex phenotype by combining iterative cycles of multiplex genome engineering and predictive modeling. We applied our method to the 63-codon E. coli strain C321.{Delta}A, which has 676 mutations relative to its wild-type ancestor, and identified six single nucleotide mutations that together recover 59% of the fitness defect exhibited by the strain. The resulting optimized strain, C321.DA.opt, is an improved chassis for production of proteins containing non-standard amino acids. Our data reveal how multiple cycles of multiplex automated genome engineering (MAGE) and inexpensive sequencing can generate rich genotypic and phenotypic diversity that can be combined with linear regression techniques to quantify individual allelic effects. While laboratory evolution relies on enrichment as a proxy for allelic effect, our model-guided approach is less susceptible than enrichment to bias from population dynamics and recombination efficiency. We also show that the method can identify beneficial de novo mutations that arise adventitiously. Beyond improving the fitness of C321, {Delta}A, our work provides a proof-of-principle for high-throughput quantification of individual allelic effects which can be used with any method for generating targeted genotypic diversity.

synthetic biology

DNA-polymerase guided elimination of paternal mitochondrial genomes:An escape-proof obstacle to their transmission

Mitochondrial DNA is predominantly inherited from only one parent. In animals this is usually the mother. This program is not in the interest of the paternal mitochondrial genome whose potential to contribute to future generations is restricted. However, in a dramatic example of genetic conflict, nuclear programs ensure the outcome. Two large mitochondria extend the length of Drosophila sperm tails. The hundreds of nucleoids in these mitochondria vanish during spermatogenesis eliminating their potential for transmission. Our previous work showed that mutational inactivation of EndoG, a nuclear encoded mitochondrial endonuclease, slows elimination of mitochondrial genomes. Here, we show that knockdown of the nuclearly encoded mitochondrial DNA polymerase, Tamas, produces a much more complete block of mtDNA loss. Recruitment of Tamas to the nucleoid at the time of its disappearance suggests a direct contribution to the elimination, but the 3'-exonuclease function of the polymerase is not needed. While DNA elimination is a surprising function for DNA polymerase, its use to restrict paternal genomes provides a strategy that cannot easily be evaded by the mitochondrial genome without compromising its replication.

developmental biology

Genome-wide analysis of phased nucleosomal arrays reveals the functional characteristic of the nucleosome remodeler ACF

Regular successions of positioned nucleosomes - phased nucleosome arrays (PNAs) - are predominantly known from transcriptional start sites (TSS). It is unclear whether PNAs occur elsewhere in the genome. To generate a comprehensive inventory of PNAs for Drosophila, we applied spectral analysis to nucleosome maps and identified thousands of PNAs throughout the genome. About half of them are not near TSS and strongly enriched for a novel sequence motif. Through genome-wide reconstitution of physiological chromatin in Drosophila embryo extracts we uncovered the molecular basis of PNA formation. We identified Phaser, an unstudied zinc finger protein that positions nucleosomes flanking the new motif. It also revealed how the global activity of the chromatin remodeler CHRAC/ACF, together with local barrier elements, generates islands of regular phasing throughout the genome. Our work demonstrates the potential of chromatin assembly by embryo extracts as a powerful tool to reconstitute chromatin features on a global scale in vitro.

molecular biology

Deconstructing isolation-by-distance: the genomic consequences of limited dispersal

Geographically limited dispersal can shape genetic population structure and result in a correlation between genetic and geographic distance, commonly called isolation-bydistance. Despite the prevalence of isolation-by-distance in nature, to date few studies have empirically demonstrated the processes that generate this pattern, largely because few populations have direct measures of individual dispersal and pedigree information. Intensive, long-term demographic studies and exhaustive genomic surveys in the Florida Scrub-Jay (Aphelocoma coerulescens) provide an excellent opportunity to investigate the influence of dispersal on genetic structure. Here, we used a panel of genome-wide SNPs and extensive pedigree information to explore the role of limited dispersal in shaping patterns of isolation-by-distance in both sexes, and at an exceedingly fine spatial scale (within ~10 km). Isolation-by-distance patterns were stronger in male-male and male-female comparisons than in female-female comparisons, consistent with observed differences in dispersal propensity between the sexes. Using the pedigree, we demonstrated how various genealogical relationships contribute to fine-scale isolation-by-distance. Simulations using field-observed distributions of male and female natal dispersal distances showed good agreement with the distribution of geographic distances between breeding individuals of different pedigree relationship classes. Furthermore, we extended Malecots theory of isolation-by-distance by building coalescent simulations parameterized by the observed dispersal curve, population density, and immigration rate, and showed how incorporating these extensions allows us to accurately reconstruct observed sex-specific isolation-by-distance patterns in autosomal and Z-linked SNPs. Therefore, patterns of fine-scale isolation-by-distance in the Florida Scrub-Jay can be well understood as a result of limited dispersal over contemporary timescales.\n\nAuthor SummaryDispersal is a fundamental component of the life history of most organisms and therefore influences many biological processes. Dispersal is particularly important in creating genetic structure on the landscape. We often observe a pattern of decreased genetic relatedness between individuals as geographic distances increases, or isolation-by-distance. This pattern is particularly pronounced in organisms with extremely short dispersal distances. Despite the ubiquity of isolation-by-distance patterns in nature, there are few examples that explicitly demonstrate how limited dispersal influences spatial genetic structure. Here we investigate the processes that result in spatial genetic structure using the Florida Scrub-Jay, a bird with extremely limited dispersal behavior and extensive genome-wide data. We take advantage of the long-term monitoring of a contiguous population of Florida Scrub-Jays, which has resulted in a detailed pedigree and measurements of dispersal for hundreds of individuals. We show how limited dispersal results in close genealogical relatives living closer together geographically, which generates a strong pattern of isolation-by-distance at an extremely small spatial scale (<10 km) in just a few generations. Given the detailed dispersal, pedigree, and genomic data, we can achieve a fairly complete understanding of how dispersal shapes patterns of genetic diversity over short spatial scales.

evolutionary biology

Genome-wide association study in Collaborative Cross mice revealed a skeletal role for Rhbdf2

Osteoporosis, the most common bone disease, is characterized by a low bone mass and increased risk of fractures. Importantly, individuals with the same bone mineral density (BMD), as measured on two dimensional (2D) radiographs, have different risks for fracture, suggesting that microstructural architecture is an important determinant of skeletal strength. Here we took advantage of the rich phenotypic and genetic diversity of the Collaborative Cross (CC) mice. Using microcomputed tomography, we examined key structural parameters in the femoral cortical and trabecular compartments of male and female mice from 34 CC lines. These traits included the trabecular bone volume fraction, number, thickness, connectivity, and spacing, as well as structural morphometric index. In the mid-diaphyseal cortex, we recorded cortical thickness and volumetric BMD.\n\nThe broad-sense heritability of these traits ranged between 50 to 60%. We conducted a genome-wide association study to unravel 5 quantitative trait loci (QTL) significantly associated with 6 of the traits. We refined each locus by combining information obtained from the known ancestry of the mice and RNA-Seq data from publicly available sources, to shortlist potential candidate genes. We found strong evidence for new candidate genes, including Rhbdf2, which association to trabecular bone volume fraction and number was strongly suggested by our analyses. We then examined knockout mice, and validated the causal action of Rhbdf2 on bone mass accrual and microarchitecture.\n\nOur approach revealed new genome-wide QTLs and a series of genes that have never been associated with bone microarchitecture. This study demonstrates for the first time the skeletal role of Rhbdf2 on the physiological remodeling of both the cortical and trabecular bone. This newly assigned function for Rhbdf2 can prove useful in deciphering the predisposing factors of osteoporosis and propose new investigative avenues toward targeted therapeutic solutions.\n\nAuthor summaryIn this study, we used the novel mouse reference population, the Collaborative Cross (CC), to identify new causal genes in the regulation of bone microarchitecture, a critical determinant of bone strength. This approach provides a clear advantage in terms of resolution and dimensionality of the morphometric features (versus humans) and rich allelic diversity (versus classical mouse populations), over current practices of bone-related genome-wide association studies.\n\nOur genome-wide study revealed 5 loci significantly associated with microstructural traits in the cortical and trabecular bone. We found strong evidence for new candidate genes, in particular, Rhbdf2. We then validated the specific role of Rhbdf2 on bone mass accrual and microarchitecture using knockout mice. Importantly, this study is the first demonstration of a physiological role for Rhbdf2.

genetics

Comparison of one-stage and two-stage genome-wide association studies

Linear mixed models are widely used in humans, animals, and plants to conduct genome-wide association studies (GWAS). A characteristic of experimental designs for plants is that experimental units are typically multiple-plant plots of families or lines that are replicated across environments. This structure can present computational challenges to conducting a genome scan on raw (plot-level) data. Two-stage methods have been proposed to reduce the complexity and increase the computational speed of whole-genome scans. The first stage of the analysis fits raw data to a model including environment and line effects, but no individual marker effects. The second stage involves the whole genome scan of marker tests using summary values for each line as the dependent variable. Missing data and unbalanced experimental designs can result in biased estimates of marker association effects from two-stage analyses. In this study, we developed a weighted two-stage analysis to reduce bias and improve power of GWAS while maintaining the computational efficiency of two-stage analyses. Simulation based on real marker data of a diverse panel of maize inbred lines was used to compare power and false discovery rate of the new weighted two-stage method to single-stage and other two-stage analyses and to compare different two-stage models. In the case of severely unbalanced data, only the weighted two-stage GWAS has power and false discovery rate similar to the one-stage analysis. The weighted GWAS method has been implemented in the open-source software TASSEL.

genetics

T-DNA integration is rapid and influenced by the chromatin state of the host genome.

Agrobacterium tumefaciens mediated T-DNA integration is a common tool for plant genome manipulation. However, there is controversy regarding whether T-DNA integration is biased towards genes or randomly distributed throughout the genome. In order to address this question, we performed high-throughput mapping of T-DNA-genome junctions obtained in the absence of selection at several time points after infection. T-DNA-genome junctions were detected as early as 6 hours post-infection. T-DNA distribution was apparently uniform throughout the chromosomes, yet local biases toward AT-rich motifs and T-DNA border sequence micro-homology were detected. Analysis of the epigenetic landscape of integration showed that selected events reported on previously were associated with extremely low methylation and nucleosome occupancy. Conversely, non-selected events from this study showed chromatin marks, such as high nucleosome occupancy and high H3K27me3 that correspond to 3D-interacting heterochromatin islands embedded within euchromatin. Such structures might play a role in capturing and silencing invading T-DNA.

plant biology

Heterogeneous chromatin mobility derived from chromatin states is a determinant of genome organisation in S. cerevisiae

Spatial organisation of the genome is essential for regulating gene activity, yet the mechanisms that shape this three-dimensional organisation in eukaryotes are far from understood. Here, we combine bioinformatic determination of chromatin states during normal growth and heat shock, and computational polymer modelling of genome structure, with quantitative microscopy and Hi-C to demonstrate that differential mobility of yeast chromosome segments leads to spatial self-organisation of the genome. We observe that more than forty percent of chromatin-associated proteins display a poised and heterogeneous distribution along the chromosome, creating a heteropolymer. This distribution changes upon heat shock in a concerted, state-specific manner. Simulating yeast chromosomes as heteropolymers, in which the mobility of each segment depends on its cumulative protein occupancy, results in functionally relevant structures, which match our experimental data. This thermodynamically driven self-organisation achieves spatial clustering of poised genes and mechanistically contributes to the directed relocalisation of active genes to the nuclear periphery upon heat shock.\n\nOne Sentence SummaryUnequal protein occupancy and chromosome segment mobility drive 3D organisation of the genome.

systems biology

The power of a multivariate approach to genome-wide association studies: an example with Drosophila melanogaster wing shape

Due to the complexity of genotype-phenotype relationships, simultaneous analyses of genomic associations with multiple traits will be more powerful and more informative than a series of univariate analyses. In most cases, however, studies of genotype-phenotype relationships have analyzed only one trait at a time, even as the rapid advances in molecular tools have expanded our view of the genotype to include whole genomes. Here, we report the results of a fully integrated multivariate genome-wide association analysis of the shape of the Drosophila melanogaster wing in the Drosophila Genetic Reference Panel. Genotypic effects on wing shape were highly correlated between two different labs. We found 2,396 significant SNPs using a 5% FDR cutoff in the multivariate analyses, but just 4 significant SNPs in univariate analyses of scores on the first 20 principal component axes. A key advantage of multivariate analysis is that the direction of the estimated phenotypic effect is much more informative than a univariate one. Exploiting this feature, we show that the directions of effects were on average replicable in an unrelated panel of inbred lines. Effects of knockdowns of genes implicated in the initial screen were on average more similar than expected under a null model. Association studies that take a phenomic approach in considering many traits simultaneously are an important complement to the power of genomics. Multivariate analyses of such data are more powerful, more informative, and allow the unbiased study of pleiotropy.

genetics

MITE-based drives to transcriptional control of genome host

In a recent past, Transposable Elements (TEs) were referred as selfish genetic components only capable of copying themselves with the aim to increase the odds that will be inherited. Nonetheless, TEs have been initially proposed as positive control elements acting in synergy with the host. Nowadays, it is well known that TE movement into genome host comprise an important evolutionary mechanism capable to produce diverse chromosome rearrangements and thus increase the adaptive fitness. According to as insights into TE functioning are increasing day to day, the manipulation of transposition has raised an interesting possibility to setting the host functions, although the lack of appropriate genome engineering tools has unpaved it. Fortunately, the emergence of genome editing technologies based on programmable nucleases, and especially the arrival of a multipurpose RNA-guided Cas9 endonuclease system, has made it possible to reconsider this challenge. For such purpose, a particular type of transposons referred as Miniature Inverted-repeat Transposable Elements (MITEs) has demonstrated a series of interesting characteristics for designing functional drivers. Here, recent insights into MITE elements and versatile RNA-guided CRISPR/Cas9 genome engineering system are given to outline an effective strategy that allows to deploy the TE potential for control of the host transcriptional activity.

genetics