Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

GAMtools: an automated pipeline for analysis of Genome Architecture Mapping data

Genome Architecture Mapping (GAM) is a recently developed method for mapping chromatin interactions genome-wide. GAM is based on sequencing genomic DNA extracted from thin cryosections of cell nuclei. As a new approach, GAM datasets require specialized analytical tools and approaches. Here we present GAMtools, a pipeline for analysing GAM datasets. GAMtools covers the automated mapping of raw next-generation sequencing data generated by GAM, detection of genomic regions present in each nuclear slice, calculation of quality control metrics, generation of inferred proximity matrices, plotting of heatmaps and detection of genomic features for which chromatin interactions are enriched/depleted.

bioinformatics

The genomic footprint of climate adaptation in Chironomus riparius

The gradual heterogeneity of climatic factors pose varying selection pressures across geographic distances that leave signatures of clinal variation in the genome. Separating signatures of clinal adaptation from signatures of other evolutionary forces, such as demographic processes, genetic drift, and adaptation to non-clinal conditions of the immediate local environment is a major challenge. Here, we examine climate adaptation in five natural populations of the harlequin fly Chironomus riparius sampled along a climatic gradient across Europe. Our study integrates experimental data, individual genome resequencing, Pool-Seq data, and population genetic modelling. Common-garden experiments revealed a positive correlation of population growth rates corresponding to the population origin along the climate gradient, suggesting thermal adaptation on the phenotypic level. Based on a population genomic analysis, we derived empirical estimates of historical demography and migration. We used an FST outlier approach to infer positive selection across the climate gradient, in combination with an environmental association analysis. In total we identified 162 candidate genes as genomic basis of climate adaptation. Enriched functions among these candidate genes involved the apoptotic process and molecular response to heat, as well as functions identified in other studies of climate adaptation in other insects. Our results show that local climate conditions impose strong selection pressures and lead to genomic adaptation despite strong gene flow. Moreover, these results imply that selection to different climatic conditions seems to converge on a functional level, at least between different insect species.

evolutionary biology

Discovering Complete Quasispecies In Bacterial Genomes

Mobile genetic elements can be found in almost all genomes. Possibly the most common non-autonomous mobile genetic elements in bacteria are REPINs that can occur hundreds of times within a genome. The sum of all REPINs within a genome are an evolving populations because they replicate and mutate. We know the exact composition of this population and the sequence of each member of a REPIN population, in contrast to most other biological populations. Here, we model the evolution of REPINs as quasispecies. We fit our quasispecies model to ten different REPIN populations from ten different bacterial strains and estimate duplication rates. We find that our estimated duplication rates range from about 5 x 10-9 to 37 x 10-9 duplications per generation per genome. The small range and the low level of the REPIN duplication rates suggest a universal trade-off between the survival of the REPIN population and the reduction of the mutational load for the host genome. The REPIN populations we investigated also possess features typical of other natural populations. One population shows hallmarks of a population that is going extinct, another population seems to be growing in size and we also see an example of competition between two REPIN populations.

evolutionary biology

HiPiler: Visual Exploration Of Large Genome Interaction Matrices With Interactive Small Multiples

This paper presents an interactive visualization interface--HiPiler--for the exploration and visualization of regions-of-interest in large genome interaction matrices. Genome interaction matrices approximate the physical distance of pairs of regions on the genome to each other and can contain up to 3 million rows and columns with many sparse regions. Regions of interest (ROIs) can be defined, e.g., by sets of adjacent rows and columns, or by specific visual patterns in the matrix. However, traditional matrix aggregation or pan-and-zoom interfaces fail in supporting search, inspection, and comparison of ROIs in such large matrices. In HiPiler, ROIs are first-class objects, represented as thumbnail-like \"snippets\". Snippets can be interactively explored and grouped or laid out automatically in scatterplots, or through dimension reduction methods. Snippets are linked to the entire navigable genome interaction matrix through brushing and linking. The design of HiPiler is based on a series of semi-structured interviews with 10 domain experts involved in the analysis and interpretation of genome interaction matrices. We describe six exploration tasks that are crucial for analysis of interaction matrices and demonstrate how HiPiler supports these tasks. We report on a user study with a series of data exploration sessions with domain experts to assess the usability of HiPiler as well as to demonstrate respective findings in the data.

bioinformatics

Uncovering The Repertoire Of Endogenous Flaviviral Elements In Aedes Mosquito Genomes

Endogenous viral elements derived from non-retroviral RNA viruses were described in various animal genomes. Whether they have a biological function such as host immune protection against related viruses is a field of intense study. Here, we investigated the repertoire of endogenous flaviviral elements (EFVEs) in Aedes mosquitoes, the vectors of arboviruses such as dengue and chikungunya viruses. Previous studies identified three EFVEs from Ae. albopictus and one from Ae. aegypti cell lines. However, in-depth characterization of EFVEs in wild-type mosquito populations and individuals in vivo has not been performed. We detected the full-length DNA sequence of the previously described EFVEs and their respective transcripts in several Ae. albopictus and Ae. aegypti populations from geographically distinct areas. However, EFVE-derived proteins were not detected by mass spectrometry. Using deep sequencing, we detected the production of piRNA-like small RNAs in antisense orientation, targeting the EFVEs and their flanking regions in vivo. The EFVEs were integrated in repetitive regions of the mosquito genomes, and their flanking sequences varied among mosquito populations from different geographical regions. We bioinformatically predicted several new EFVEs from a Vietnamese Ae. albopictus population and observed variation in the occurrence of those elements among mosquito populations. Phylogenetic analysis of an Ae. aegypti EFVE suggested that it integrated prior to the global expansion of the species and subsequently diverged among and within populations. Together, this study revealed substantial structural and nucleotide diversity of flaviviral integrations in Aedes genomes. Unraveling this diversity will help to elucidate the potential biological function of these EFVEs.\n\nImportanceEndogenous viral elements (EVEs) are whole or partial viral sequences integrated in host genomes. Interestingly, some EVEs have important functions for host fitness and antiviral defense. Because mosquitoes also have EVEs in their genomes, we decided to thoroughly characterized them to lay the foundation of the potential use of these EVEs to manipulate the mosquito antiviral response. Here, we focused on EVEs related to the Flavivirus genus, to which dengue and Zika viruses belong, in Aedes mosquito individuals from geographically distinct areas. We showed the existence in vivo of flaviviral EVEs previously identified in mosquito cell lines and we detected new ones. We showed that EVEs have evolved differently in each mosquito population. They produced transcripts and small RNAs, but not proteins, suggesting a function at the RNA level. Our study uncovers the diverse repertoire of flaviviral EVEs in Aedes mosquito populations and suggests a role in the host antiviral system.

microbiology

A Standardized Framework For Representation Of Ancestry Data In Genomics Studies

BackgroundThe accurate description of ancestry is essential to interpret and integrate human genomics data, and to ensure that advances in the field of genomics benefit individuals from all ancestral backgrounds. However, there are no established guidelines for the consistent, unambiguous and standardized description of ancestry. To fill this gap, we provide a framework, designed for the representation of ancestry in GWAS data, but with wider application to studies and resources involving human subjects.\n\nResultHere we describe our framework and its application to the representation of ancestry data in a widely-used publically available genomics resource, the NHGRI-EBI GWAS Catalog. We present the first analyses of GWAS data using our ancestry categories, demonstrating the validity of the framework to facilitate the tracking of ancestry in big data sets. We exhibit the broader relevance and integration potential of our method by its usage to describe the well-established HapMap and 1000 Genomes reference populations. Finally, to encourage adoption, we outline recommendations for authors to implement when describing samples.\n\nConclusionsWhile the known bias towards inclusion of European ancestry individuals in GWA studies persists, African and Hispanic or Latin American ancestry populations contribute a disproportionately high number of associations, suggesting that analyses including these groups may be more effective at identifying new associations. We believe the widespread adoption of our framework will increase standardization of ancestry data, thus enabling improved analysis, interpretation and integration of human genomics data and furthering our understanding of disease.

genetics

Identification And Prioritisation Of Variants In The Short Open-Reading Frame Regions Of The Human Genome

As whole-genome sequencing technologies improve and accurate maps of the entire genome are assembled, short open-reading frames (sORFs) are garnering interest as functionally important regions that were previously overlooked. However, there is a paucity of tools available to investigate variants in sORF regions of the genome. Here we investigate the performance of commonly used tools for variant calling and variant prioritisation in these regions, and present a framework for optimising these processes. First, the performance of four widely used germline variant calling algorithms is systematically compared. Haplotype Caller is found to perform best across the whole genome, but FreeBayes is shown to produce the most accurate variant set in sORF regions. An accurate set of variants is found by taking the intersection of called variants. The potential deleteriousness of each variant is then predicted using a pathogenicity scoring algorithm developed here, called sORF-c. This algorithm uses supervised machine-learning to predict the pathogenicity of each variant, based on a holistic range of functional, conservation-based and region-based scores defined for each variant. By training on a dataset of over 130,000 variants, sORF-c outperforms other comparable pathogenicity scoring algorithms on a test set of variants in sORF regions of the human genome.\n\nList of Abbreviations

genetics

SUMO E3 ligase Mms21 prevents spontaneous DNA damage induced genome rearrangements

Mms21, a subunit of the Smc5/6 complex, possesses an E3 ligase activity for the Small Ubiquitin-like MOdifier (SUMO), which has a major, but poorly understood role in genome maintenance. Here we show mutations that inactivate the E3 ligase activity of Mms21 cause Rad52- and Pol32-dependent break-induced replication (BIR), which specifically requires the Rrm3 DNA helicase. Interestingly, mutations affecting both Mms21 and the Sgs1 helicase, but not sumoylation of Sgs1, cause further accumulation of genome rearrangements, indicating the distinct roles of Mms21 and Sgs1 in suppressing genome rearrangements. Whole genome sequencing further revealed that the Mre11 endonuclease prevents microhomology-mediated translocations and hairpin-mediated inverted duplications in the mms21 mutant. Consistent with the accumulation of endogenous DNA lesions, mms21 cells accumulate spontaneous Ddc2 foci and display a hyper-activated DNA damage checkpoint. Together, these findings support a new paradigm that Mms21 prevents the accumulation of spontaneous DNA lesions that cause diverse genome rearrangements.

genetics

ANNOgesic: A Pipeline To Translate Bacterial/Archaeal RNA-Seq Data Into High-Resolution Genome Annotations

To understand the gene regulation of an organism of interest, a comprehensive genome annotation is essential. While some features, such as coding sequences, can be computationally predicted with high accuracy based purely on the genomic sequence, others, such as promoter elements or non-coding RNAs are harder to detect. RNA-Seq has proven to be an efficient method to identify these genomic features and to improve genome annotations. However, processing and integrating RNA-Seq data in order to generate high-resolution annotations is challenging, time consuming and requires numerous different steps. We have constructed a powerful and modular tool called ANNOgesic that provides the required analyses and simplifies RNA-Seq-based bacterial and archaeal genome annotation. It can integrate data from conventional RNA-Seq and dRNA-Seq, predicts and annotates numerous features, including small non-coding RNAs, with high precision. The software is available under an open source license (ISCL) at https://pypi.org/project/ANNOgesic/.

bioinformatics

Imputation-Based Genomic Coverage Assessments of Current Genotyping Arrays: Illumina HumanCore, OmniExpress, Multi-Ethnic global array and sub-arrays, Global Screening Array, Omni2.5M, Omni5M, and Affymetrix UK Biobank

Genotyping arrays have been widely adopted as an efficient means to interrogate variation across the human genome. Genetic variants may be observed either directly, via genotyping, or indirectly, through linkage disequilibrium with a genotyped variant. The total proportion of genomic variation captured by an array, either directly or indirectly, is referred to as \"genomic coverage.\" Here we use genotype imputation and Phase 3 of the 1000 Genomes Project to assess genomic coverage of several modern genotyping arrays. We find that in general, coverage increases with increasing array density. However, arrays designed to cover specific populations may yield better coverage in those populations compared to denser arrays not tailored to the given population. Ultimately, array choice involves trade-offs between cost, density, and coverage, and our work helps inform investigators weighing these choices and trade-offs.

genetics

Genome Architecture Leads a Bifurcation in Cell Identity

Genome architecture is important in transcriptional regulation and study of its features is a critical part of fully understanding cell identity. Altering cell identity is possible through overexpression of transcription factors (TFs); for example, fibroblasts can be reprogrammed into muscle cells by introducing MYOD1. How TFs dynamically orchestrate genome architecture and transcription as a cell adopts a new identity during reprogramming is not well understood. Here we show that MYOD1-mediated reprogramming of human fibroblasts into the myogenic lineage undergoes a critical transition, which we refer to as a bifurcation point, where cell identity definitively changes. By integrating knowledge of genome-wide dynamical architecture and transcription, we found significant chromatin reorganization prior to transcriptional changes that marked activation of the myogenic program. We also found that the local architectural and transcriptional dynamics of endogenous MYOD1 and MYOG reflected the global genomic bifurcation event. These TFs additionally participate in entrainment of biological rhythms. Understanding the system-level genome dynamics underlying a cell fate decision is a step toward devising more sophisticated reprogramming strategies that could be used in cell therapies.

cell biology

Project MinE: study design and pilot analyses of a large-scale whole-genome sequencing study in amyotrophic lateral sclerosis

The most recent genome-wide association study in amyotrophic lateral sclerosis (ALS) demonstrates a disproportionate contribution from low-frequency variants to genetic susceptibility of disease. We have therefore begun Project MinE, an international collaboration that seeks to analyse whole-genome sequence data of at least 15,000 ALS patients and 7,500 controls. Here, we report on the design of Project MinE and pilot analyses of newly whole-genome sequenced 1,264 ALS patients and 611 controls drawn from the Netherlands. As has become characteristic of sequencing studies, we find an abundance of rare genetic variation (minor allele frequency < 0.1 %), the vast majority of which is absent in public data sets. Principal component analysis reveals local geographical clustering of these variants within The Netherlands. We use the whole-genome sequence data to explore the implications of poor geographical matching of cases and controls in a sequence-based disease study and to investigate how ancestry-matched, externally sequenced controls can induce false positive associations. Also, we have publicly released genome-wide minor allele counts in cases and controls, as well as results from genic burden tests.

genetics

Germline Cas9 Expression Yields Highly Efficient Genome Engineering in a Major Worldwide Disease Vector, Aedes aegypti

The development of CRISPR/Cas9 technologies has dramatically increased the accessibility and efficiency of genome editing in many organisms. In general, in vivo germline expression of Cas9 results in substantially higher activity than embryonic injection. However, no transgenic lines expressing Cas9 have been developed for the major mosquito disease vector Aedes aegypti. Here, we describe the generation of multiple stable, transgenic Ae. aegypti strains expressing Cas9 in the germline, resulting in dramatic improvements in both the consistency and efficiency of genome modifications using CRISPR. Using these strains, we disrupted numerous genes important for normal morphological development, and even generated triple mutants from a single injection. We have also managed to increase the rates of homology directed repair by more than an order of magnitude. Given the exceptional mutagenic efficiency and specificity of the Cas9 strains we built, they can be used for high-throughput reverse genetic screens to help functionally annotate the Ae. aegypti genome. Additionally, these strains represent a first step towards the development of novel population control technologies targeting Ae. aegypti that rely on Cas9-based gene drives.\n\nSignificance StatementAedes aegypti is the principal vector of multiple arboviruses that significantly affect human health including dengue, chikungunya, and zika. Development of tools for efficient genome engineering in this mosquito will not only lay the foundation for the application of novel genetic control strategies that do not rely on insecticides, but will also accelerate basic research on key biological processes involved in disease transmission. Here, we report the development of a transgenic CRISPR approach for rapid gene disruption in this organism. Given their high editing efficiencies, the Cas9 strains we developed can be used to quickly generate novel genome modifications allowing for high-throughput gene targeting, and can possibly facilitate the development of gene drives, thereby accelerating comprehensive functional annotation and development of innovative population control strategies for Ae. aegypti.

bioengineering

CRISPR/Cas9-APEX-mediated proximity labeling enables discovery of proteins associated with a predefined genomic locus in living cells

The activation or repression of a genes expression is primarily controlled by changes in the proteins that occupy its regulatory elements. The most common method to identify proteins associated with genomic loci is chromatin immunoprecipitation (ChIP). While having greatly advanced our understanding of gene expression regulation, ChIP requires specific, high quality, IP-competent antibodies against nominated proteins, which can limit its utility and scope for discovery. Thus, a method able to discover and identify proteins associated with a particular genomic locus within the native cellular context would be extremely valuable. Here, we present a novel technology combining recent advances in chemical biology, genome targeting, and quantitative mass spectrometry to develop genomic locus proteomics, a method able to identify proteins which occupy a specific genomic locus.

biochemistry

Solagigasbacteria: Lone genomic giants among the uncultured bacterial phyla

Recent advances in single-cell genomic and metagenomic techniques have facilitated the discovery of numerous previously unknown, deep branches of the tree of life that lack cultured representatives. Many of these candidate phyla are composed of microorganisms with minimalistic, streamlined genomes lacking some core metabolic pathways, which may contribute to their resistance to growth in pure culture. Here we analyzed single-cell genomes and metagenome bins to show that the \"Candidate phylum SPAM\" represents an interesting exception, by having large genomes (6-8 Mbps), high GC content (66%-71%), and the potential for a versatile, mixotrophic metabolism. We also observed an unusually high genomic heterogeneity among individual SPAM cells in the studied samples. These features may have contributed to the limited recovery of sequences of this candidate phylum in prior metagenomic studies. Based on these observations, we propose renaming SPAM to \"Candidate phylum Solagigasbacteria\". Current evidence suggests that Solagigasbacteria are distributed globally in diverse terrestrial ecosystems, including soils, the rhizosphere, volcanic mud, oil wells, aquifers and the deep subsurface, with no reports from marine environments to date.

microbiology

Gaussian decomposition of high-resolution melt curve derivatives for measuring genome-editing efficiency

We describe a method for measuring genome editing efficiency from in silico analysis of high-resolution melt curve data. The melt curve data derived from amplicons of genome-edited or unmodified target sites were processed to remove the background fluorescent signal emanating from free fluorophore and then corrected for temperature-dependent quenching of fluorescence of double-stranded DNA-bound fluorophore. Corrected data were normalized and numerically differentiated to obtain the first derivatives of the melt curves. These were then mathematically modeled as a sum or superposition of minimal number of Gaussian components. Using Gaussian parameters determined by modeling of melt curve derivatives of unedited samples, we were able to model melt curve derivatives from genetically altered target sites where the mutant population could be accommodated using an additional Gaussian component. From this, the proportion contributed by the mutant component in the target region amplicon could be accurately determined. Mutant component computations compared well with the mutant frequency determination from next generation sequencing data. The results were also consistent with our earlier studies that used difference curve areas from high-resolution melt curves for determining the efficiency of genome-editing reagents. The advantage of the described method is that it does not require calibration curves to estimate proportion of mutants in amplicons of genome-edited target sites.\n\nSignificance StatementGenome editing has been revolutionized by the engineering of molecular scissors that cut DNA at a predetermined location on the chromosome. When these molecular scissors are expressed within cells, these scissors cut the genomic DNA at the designated target site, and the cells respond by repairing the cut sites. This repair process frequently introduces mutations at the target cut site. The more efficient the molecular scissors, the more number of cells in a treated culture dish exhibit these mutations at the cut site. Investigators therefore design several molecular scissors targeting the same region on the chromosome to identify the best ones. We describe a new recipe to measure scissors efficiency in target site cutting.

bioengineering

Prey range and genome evolution of Halobacteriovorax marinus predatory bacteria from an estuary

BackgroundHalobacteriovorax are saltwater-adapted predatory bacteria that attack Gram-negative bacteria and therefore may play an important role in shaping microbial communities. To understand the impact of Halobacteriovorax on ecosystems and develop them as biocontrol agents, it is important to characterize variation in predation phenotypes such as prey range and investigate the forces impacting Halobacteriovorax genome evolution across different phylogenetic distances.\n\nResultsWe isolated H. marinus BE01 from an estuary in Rhode Island using Vibrio from the same site as prey. Small, fast-moving attack phase BE01 cells attach to and invade prey cells, consistent with the intraperiplasmic predation strategy of H. marinus type strain SJ. BE01 is a prey generalist, forming plaques on Vibrio strains from the estuary as well as Pseudomonas from soil and E. coli. Genome analysis revealed that BE01 is very closely related to SJ, with extremely high conservation of gene order and amino acid sequences. Despite this similarity, we identified two regions of gene content difference that likely resulted from horizontal gene transfer. Analysis of modal codon usage frequencies supports the hypothesis that these regions were acquired from bacteria with different codon usage biases compared to Halobacteriovorax. In BE01, one of these regions includes genes associated with mobile genetic elements, such as a transposase not found in SJ and degraded remnants of an integrase occurring as a full-length gene in SJ. The corresponding region in SJ included unique mobile genetic element genes, such as a site-specific recombinase and bacteriophage-related genes not found in BE01. Acquired functions in BE01 include the dnd operon, which encodes a pathway for DNA modification that may protect DNA from nucleases, and a suite of genes involved in membrane synthesis and regulation of gene expression that was likely acquired from another Halobacteriovorax lineage.\n\nConclusionsOur results support previous observations that Halobacteriovorax prey on a broad range of Gram-negative bacteria. Genome analysis suggests strong selective pressure to maintain the genome in the H. marinus lineage represented by BE01 and SJ, although our results also provide further evidence that horizontal gene transfer plays an important role in genome evolution in predatory bacteria.

microbiology

Recovering genomic clusters of secondary metabolites from lakes: a Metagenomics 2.0 approach

BackgroundMetagenomic approaches became increasingly popular in the past decades due to decreasing costs of DNA sequencing and bioinformatics development. So far, however, the recovery of long genes coding for secondary metabolism still represents a big challenge. Often, the quality of metagenome assemblies is poor, especially in environments with a high microbial diversity where sequence coverage is low and complexity of natural communities high. Recently, new and improved algorithms for binning environmental reads and contigs have been developed to overcome such limitations. Some of these algorithms use a similarity detection approach to classify the obtained reads into taxonomical units and to assemble draft genomes. This approach, however, is quite limited since it can classify exclusively sequences similar to those available (and well classified) in the databases.\n\nIn this work, we used draft genomes from Lake Stechlin, north-eastern Germany, recovered by MetaBat, an efficient binning tool that integrates empirical probabilistic distances of genome abundance, and tetranucleotide frequency for accurate metagenome binning. These genomes were screened for secondary metabolism genes, such as polyketide synthases (PKS) and non-ribosomal peptide synthases (NRPS), using the Anti-SMASH and NAPDOS workflows.\n\nResultsWith this approach we were able to identify 243 secondary metabolite clusters from 121 genomes recovered from the lake samples. A total of 18 NRPS, 19 PKS and 3 hybrid PKS/NRPS clusters were found. In addition, it was possible to predict the partial structure of several secondary metabolite clusters allowing for taxonomical classifications and phylogenetic inferences.\n\nConclusionsOur approach revealed a great potential to recover and study secondary metabolites genes from any aquatic ecosystem.

bioinformatics