Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Draft genome of the liver fluke Fasciola gigantica

Fascioliasis is a neglected food-borne disease caused by liver flukes (genus Fasciola) and affects more than 200 million people worldwide. Despite technological advances, little is known about the molecular biology and biochemistry of the fluke. We present the draft genome of Fasciola gigantica for the first time. The assembled draft genome has a size of ~1.04 Gb with an N50 of 129 kb. A total of 20,858 genes were predicted. The de novo repeats identified in the draft genome were 46.85%. In pathway analysis, all the genes of glycolysis, Krebs cycle and fatty acid metabolism were found to be present, but the key genes for fatty acid production in fatty acid biosynthesis were missing. This indicates that the fatty acid required for the survival of the fluke may be acquired from the host bile. The genomic information will provide a comprehensive resource to facilitate the development of novel interventions for fascioliasis control.

bioinformatics

Abnormal LTR retrotransposons formed by the recombination outside the host genome

Long terminal repeat (LTR) retrotransposons are the dominant feature of higher plant genomes, which have a similar life cycle with retrovirus. Previous studies cannot account for all observed complex LTR retrotransposon patterns. In this study, we first identified 63 complex LTR retrotransposons in rice genome, and most of complex elements harbored flanking target-site duplications (TSDs). But these complex elements in which outermost LTRs had not the most highly homologous cant be explained. We propose a new model that the homologous recombination of two new different normal LTR retrotransposon elements in the same family can occur before their integration to the rice genome. The model can explain at least fourteen complex retrotransposons formations. We also find that normal LTR retrotransposons can swap their LTRs to generate abnormal LTR retrotransposons in which LTRs are different because of homologous recombination before their integration to the genome.

genetics

Mapping native R-loops genome-wide using a targeted nuclease approach

R-loops are three-stranded DNA:RNA hybrids that are pervasive in the eukaryotic and prokaryotic genomes and have been implicated in a variety of nuclear processes, including transcription, replication, DNA repair, and chromosome segregation. While R-loops may have physiological roles, the formation of stable, aberrant R-loops has been observed in disease, particularly neurological disorders and cancer. Despite the importance of these structures, methods to assess their distribution in the genome invariably rely on affinity purification, which requires large amounts of input material, is plagued by high level of noise, and is poorly suited to capture dynamic and unstable R-loops. Here, we present a new method that leverages the affinity of RNase H for DNA:RNA hybrids to target micrococcal nuclease to genomic sites that contain R-loops, which are subsequently cleaved, released, and sequenced. Our R-loop mapping method, MapR, is as specific as existing techniques, less prone to recover non-specific repetitive sequences, and more sensitive, allowing for genome-wide coverage with low input material and read numbers, in a fraction of the time.

molecular biology

The complete plastid genomes of four species from Brassicales

Brassicales is a diverse angiosperm order with about 4,700 recognized species. Here, we assembled and described the complete plastid genomes from four species of Brassicales: Capparis urophylla F.Chun (Capparaceae), Carica papaya L. (Caricaceae), Cleome rutidosperma DC. (Cleomaceae), and Moringa oleifera Lam. (Moringaceae), including two plastid genomes newly assembled for two families (Capparaceae and Moringaceae). The four plastid genomes are 159,680 base pairs on average in length and encode 78 protein-coding genes. The genomes each contains a typical structure of a Large Single-Copy (LSC) region and a Small Single-Copy (SSC) region separated by two Inverted Repeat (IR) regions. We performed the maximum-likelihood (ML) phylogenetic analysis using three different data sets of 66 protein-coding genes (ntAll, ntNo3rd and AA). Our phylogenetic results from different dataset are congruent, and are consistent with previous phylogenetic studies of Brassiales.

plant biology

GenPipes: an open-source framework for distributed and scalable genomic analyses

With the decreasing cost of sequencing and the rapid developments in genomics technologies and protocols, the need for validated bioinformatics software that enables efficient large-scale data processing is growing. Here we present GenPipes, a flexible Python-based framework that facilitates the development and deployment of multi-step workflows optimized for High Performance Computing clusters and the cloud. GenPipes already implements 12 validated and scalable pipelines for various genomics applications, including RNA-Seq, ChIP-Seq, DNA-Seq, Methyl-Seq, Hi-C, capture Hi-C, metagenomics and PacBio long read assembly. The software is available under a GPLv3 open source license and is continuously updated to follow recent advances in genomics and bioinformatics. The framework has been already configured on several servers and a docker image is also available to facilitate additional installations. In summary, GenPipes offers genomic researchers a simple method to analyze different types of data, customizable to their needs and resources, as well as the flexibility to create their own workflows.

bioinformatics

Desiccation Does Not Drastically Increase the Accessibility of Foreign DNA to the Nuclear Genomes: Evidence from the Frequency of Endosymbiotic DNA Transfer

Horizontal gene transfer (HGT) is a widely accepted force in the evolution of prokaryotic genomes. However, in eukaryotes, it is still in hot debate. Some bdelloid rotifers that are resistant to extreme desiccation and radiation were reported to have a very high level of HGTs. However, a similar report in another resistant invertebrate, tardigrades, has been mired in controversy. The DNA double-strand breaks (DSBs) induced by prolonged desiccation have been postulated to open the gateway of nuclear genome for foreign DNA integration and thus facilitate the HGT process. If so, the rate of endosymbiotic DNA transfer should also be enhanced. We first surveyed the abundance of nuclear mitochondrial DNAs (NUMTs) and nuclear plastid DNAs (NUPTs) in three groups of eukaryotes that are extremely resistant to desiccation, bdelloid rotifers, A. vaga and A. ricciae, tardigrades, H. dujardini and R. varieornatus, and the resurrection plants, D. hygrometricum and S. tamariscina. Excessive NUMTs or NUPTs have not been detected. Furthermore, we compared nine groups of desiccation-tolerant organisms with their desiccation-sensitive relatives but did not find significant difference in the NUMT/NUPT contents. Desiccation could induce DSBs, but it unlikely dramatically increase the frequency of foreign sequence integration in most eukaryotes. Only in the nuclear genomes enriched in repetitive sequences, the DSBs are predominantly repaired by non-homologous end joining (NHEJ) and desiccation-induced DSBs is possible to enhance the integration of foreign sequences into nuclear genome for some degree.

evolutionary biology

Consequences of breed formation on patterns of genomic diversity and differentiation: the case of highly diverse peripheral Iberian cattle

BackgroundIberian primitive breeds exhibit a remarkable phenotypic diversity over a very limited geographical space. While genomic data are accumulating for most commercial cattle, it is still lacking for these primitive breeds. Whole genome data is key to understand the consequences of historic breed formation and the putative role of earlier admixture events in the observed diversity patterns.\n\nResultsWe sequenced 48 genomes belonging to eight Iberian native breeds and found that the individual breeds are genetically very distinct with FST values ranging from 4 to 16% and have levels of nucleotide diversity similar or larger than those of their European counterparts, namely Jersey and Holstein. All eight breeds display significant gene flow or admixture from African taurine cattle and include mtDNA and Y-chromosome haplotypes from multiple origins. Furthermore, we detected a very low differentiation of chromosome X relative to autosomes within all analyzed taurine breeds, potentially reflecting male-biased gene flow.\n\nConclusionsOur results show that an overall complex history of admixture resulted in unexpectedly high levels of genomic diversity for breeds with seemingly limited geographic ranges that are distantly located from the main domestication center for taurine cattle in the Near East. This is likely to result from a combination of trading traditions and breeding practices in Mediterranean countries. We also found that the levels of differentiation of autosomes vs sex chromosomes across all studied taurine and indicine breeds are likely to have been affected by widespread breeding practices associated with male-biased gene flow.

evolutionary biology

Comparative genomic analysis of the emerging pathogen Streptococcus pseudopneumoniae: novel insights into virulence determinants and identification of a novel species-specific molecular marker

Streptococcus pseudopneumoniae is a close relative of the major human pathogen S. pneumoniae. While initially considered as a commensal species, it has been increasingly associated with lower-respiratory tract infections and high prevalence of antimicrobial resistance (AMR). S. pseudopneumoniae is difficult to identify using traditional typing methods due to similarities with S. pneumoniae and other members of the mitis group (SMG). Using phylogenetic and comparative genomic analyses of SMG genomes, we identified a new molecular marker specific for S. pseudopneumoniae and absent from any other bacterial genome sequenced to date. We found that a large number of known virulence and colonization genes are present in the core S. pseudopneumoniae genome and we reveal the impressive number of known and new surface-exposed proteins encoded by this species. Phylogenetic analyses of S. pseudopneumoniae show that specific clades are associated with allelic variants of core proteins. Resistance to tetracycline and macrolides, the two most common resistances, were encoded by Tn916-like integrating conjugative elements and Mega-2. Overall, we found a tight association of genotypic determinants of AMR as well as phenotypic AMR with a specific lineage of S. pseudopneumoniae. Taken together, our results sheds light on the distribution in S. pseudopneumoniae of genes known to be important during invasive disease and colonization and provide insight into features that could contribute to virulence, colonization and adaptation.\n\nImportanceS. pseudopneumoniae is an overlooked pathogen emerging as the causative agent of lower-respiratory tract infections and associated with chronic obstructive pulmonary disease (COPD) and exacerbation of COPD. However, much remains unknown on its clinical importance and epidemiology, mainly due to the lack of specific means to distinguish it from S. pneumoniae. Here, we provide a new molecular marker entirely specific for S. pseudopneumoniae. Furthermore, our research provides a deep analysis of the presence of virulence and colonization genes, as well as AMR determinants in this species. Our results provide crucial information and pave the way for further studies aiming at understanding the pathogenesis and epidemiology of S. pseudopneumoniae.

microbiology

Comparative study of population genomic approaches for mapping colony-level traits

Social insect colonies exhibit colony-level phenotypes such as social immunity and task coordination, which are the sum of individual phenotypes. Mapping the genetic basis of such phenotypes requires associating the colony-level phenotype with the genotypes in the colony. In this paper, we examine alternative approaches to DNA extraction, library construction and sequencing for genome wide association (GWAS) studies of colony-level traits. We evaluate the accuracy of allele frequency estimation in sequencing a pool of individuals (pool-seq) from each colony in either whole-genome sequencing or reduced representation genomic sequencing. Based on empirical measurement of the experimental noise in sequencing DNA pools, we show that whole-genome pool-seq is more accurate than reduced representation pool-seq. We evaluate the power of the alternative approaches for detecting quantitative trait loci (QTL) of colony-level traits by using simulations that account for an environmental effect on the phenotype. Our results can inform experimental designs and enable optimizing the power of GWAS depending on budget, availability of samples and research goals. We conclude that for a given budget, sequencing un-normalized pools of individuals from each colony achieves greater QTL detection power.

bioinformatics

NucMerge: Genome assembly quality improvement assisted by alternative assemblies and paired-end Illumina reads

BackgroundIn spite of the major breakthroughs in the second-generation sequencing technologies and the developments of a plethora of assemblers over the last ten years, the resulting genome assemblies may still be fragmented and contain errors. It is typical in genome projects with second-generation reads involved to run multiple assemblers with different parameters and choose the best assembly. However, such an approach is always a trade-off between the strengths and weaknesses of the assemblies. To exploit the advantages of different assemblers, an alternative approach that combines the best parts of several assemblies into one may be applied. The existing tools based on such an approach assist in elongation of assembly fragments and/or improvement of assembly accuracy. Though there has been progress with such a strategy, there is still room for improvement of the existing tools.\n\nResultsWe present NucMerge, a tool for improving genome assembly accuracy by incorporating information derived from an alternative assembly and paired-end Illumina reads from the same genome. The tool corrects insertion, deletion, substitution, and inversion errors and locates different inter- and intra-chromosomal rearrangement errors. NucMerge was compared to two existing alternatives, namely Metassembler and GAM-NGS.\n\nConclusionsThe benchmarking results show that NucMerge has generally better performance than the other tools tested, providing accuracy improvement of more assemblies. NucMerge is freely available at https://github.com/uio-bmi/NucMerge under the MPL license.

bioinformatics

Proteomic and Genomic Signatures of Repeat-instability in Cancer and Adjacent Normal Tissues

Repetitive sequences are hotspots of evolution at multiple levels. However, due to technical difficulties involved in their assembly and analysis, the role of repeats in tumor evolution is poorly understood. We developed a rigorous motif-based methodology to quantify variations in the repeat content of proteomes and genomes, directly from proteomic and genomic raw sequence data, and applied it to analyze a wide range of tumors and normal tissues. We identify high similarity between the repeat-instability in tumors and their patient-matched normal tissues, but also tumor-specific signatures, both in protein expression and in the genome, that strongly correlate with cancer progression and robustly predict the tumorigenic state. In a patient, the hierarchy of genomic repeat instability signatures accurately reconstructs tumor evolution, with primary tumors differentiated from metastases. We find an inverse relationship between repeat-instability and point mutation load, within and across patients, and independently of other somatic aberrations. Thus, repeat-instability is a distinct, transient and compensatory adaptive mechanism in tumor evolution.

cancer biology

Pan-Cancer modelling of genomic alterations through gene expression

Cancer is a disease often characterized by the presence of multiple genomic alterations, which trigger altered transcriptional patterns and gene expression, which in turn sustain the processes of tumorigenesis, tumor progression and tumor maintenance. The links between genomic alterations and gene expression profiles can be utilized as the basis to build specific molecular tumorigenic relationships. In this study we perform pan-cancer predictions of the presence of single somatic mutations and copy number variations using machine learning approaches on gene expression profiles. We show that gene expression can be used to predict genomic alterations in every tumor type, where some alterations are more predictable than others. We propose gene aggregation as a tool to improve the accuracy of alteration prediction models from gene expression profiles. Ultimately, we show how this principle can be beneficial in intrinsically noisy datasets, such as those based on single cell sequencing. Author SummaryIn this article we show that transcript abundance can be used to predict the presence or absence of the majority of genomic alterations present in human cancer. We also show how these predictions can be improved by aggregating genes into small networks to counteract the effects of transcript measurement noise.

systems biology

Native CRISPR-Cas mediated in situ genome editing reveals extensive resistance synergy in the clinical multidrug resistant Pseudomonas aeruginosa

Antimicrobial resistance (AMR) is imposing a global public health threat. Despite its importance, resistance characterization in the native background of clinically isolated resistant pathogens is frequently hindered by the lack of genome editing tools in these \"non-model\" strains. Pseudomonas aeruginosa is both a prototypical multidrug resistant (MDR) pathogen and a model species for understanding CRISPR-Cas functions. In this study, we report the successful development of the first native type I-F CRISPR-Cas mediated, one-step genome editing technique in a paradigmatic MDR strain PA154197. The technique is readily applicable in additional type I-F CRISPR-containing, clinical/environmental P. aeruginosa isolates. A two-step In-Del strategy is further developed to edit genomic locus lacking an effective PAM (protospacer adjacent motif) or within an essential gene, which together principally allows any type of non-lethal genomic manipulations in these strains. Exploiting these powerful techniques, a series of reverse mutations are constructed and the key resistant determinants of the MDR PA154197 are elucidated which include over-production of two multidrug efflux pumps MexAB-OprM and MexEF-OprN, and a typical fluoroquinolone (FQ) resistance mutation T83I in the drug target gene gyrA. Characterizing antimicrobial susceptibilities in isogenic strains containing various combinations of single, double, or all three key resistance determinants reveal that i) extensive synergy exists between the target mutation and over-production of efflux pumps, and between the two over-produced tripartite efflux pumps to confer clinically significant FQ resistance; ii) while basal level MexAB-OprM confers resistance only to penicillins, its over-production leads to substantial resistance to all antipseudonmonal {beta}-lactams and additional resistance to FQs; iii) despite the acquisition and over-production of multiple resistant mutations, no obvious evolutionary trade-off of collateral sensitivity is developed in PA154197. Together, these results provide new insights into resistance development in clinical MDR P. aeruginosa strains and demonstrate the great potentials of native CRISPR systems in AMR research.

microbiology

Use of F2 bulks in training sets for genomic prediction of combining ability and hybrid performance

Developing training sets for genomic prediction in hybrid crops requires producing hybrid seed for a large number of entries. In autogamous crop species (e.g., wheat, rice, rapeseed, cotton) this requires elaborate hybridization systems to prevent self-pollination and presents a significant impediment to the implementation of hybrid breeding in general and genomic selection in particular. An alternative to F1 hybrids are bulks of F2 seed from selfed F1 plants (F1:2). Seed production for F1:2 bulks requires no hybridization system because the number of F1 plants needed for producing enough F1:2 seed for multi-environment testing can be generated by hand-pollination. This study evaluated the suitability of F1:2 bulks for use in training sets for genomic prediction of F1 level general combining ability and hybrid performance, under different degrees of divergence between heterotic groups and modes of gene action, using quantitative genetic theory and simulation of a genomic prediction experiment. The simulation, backed by theory, showed that F1:2 training sets are expected to have a lower prediction accuracy relative to F1 training sets, particularly when heterotic groups have strongly diverged. The accuracy penalty, however, was only modest and mostly because of a lower heritability, rather than because of a difference in F1 and F1:2 genetic values. It is concluded that resorting to F1:2 bulks is, in theory at least, a promising approach to remove the significant complication of a hybridization system from the breeding process.

genetics

Dog-wise canine gut metagenome assemblies with reconstructed bacterial genomes and viral candidates

Long-read metagenomic sequencing can improve genome recovery from complex gut microbial communities, yet directly reusable canine gut genome resources remain limited. Here we describe DogMAG, a canine gut metagenome resource based on dog-wise long-read and hybrid assemblies generated by grouping sequencing libraries according to canonical dog identity before assembly. The final dataset comprises 41 assemblies linked to 277 FASTQ records, including 30 Flye long-read-only and 11 OPERA-MS hybrid assemblies. A single integrated BASALT workflow produced 11,276 selected bin/version records, followed by explicit quality-based re-selection of 3,418 medium-quality-or-better metagenome-assembled genome candidates. External dRep dereplication yielded 792 strain-like representatives at 99% average nucleotide identity and 135 species/SGB-like representatives at 95%. GTDB-Tk classified all 792 representatives as Bacteria. Viral screening identified 22,068 geNomad predictions, of which 3,374 Complete, High-quality or Medium-quality viral/proviral candidate rows passed CheckV filtering with contamination [≤]10%. DogMAG provides assemblies, genome and viral candidate sequences, metadata, provenance tables and workflow scripts for reuse, benchmarking and reanalysis.

microbiology

Chromosome-scale comparative sequence analysis unravels molecular mechanisms of genome evolution between two wheat cultivars

BackgroundRecent improvements in DNA sequencing and genome scaffolding have paved the way to generate high-quality de novo assemblies of pseudomolecules representing complete chromosomes of wheat and its wild relatives. These assemblies form the basis to compare the evolutionary dynamics of wheat genomes on a megabase-scale.\n\nResultsHere, we provide a comparative sequence analysis of the ~700-megabase chromosome 2D between two bread wheat genotypes - the old landrace Chinese Spring and the elite Swiss spring wheat line CH Campala Lr22a. There was a high degree of sequence conservation between the two chromosomes. Analysis of large structural variations revealed four large insertions/deletions (InDels) of >100 kb. Based on the molecular signatures at the breakpoints, unequal crossing over and double-strand break repair were identified as the evolutionary mechanisms that caused these InDels. Three of the large InDels affected copy number of NLRs, a gene family involved in plant immunity. Analysis of single nucleotide polymorphism (SNP) density revealed three haploblocks of ~8 Mb, ~9 Mb and ~48 Mb with a 35-fold increased SNP density compared to the rest of the chromosome.\n\nConclusionsThis comparative analysis of two high-quality chromosome assemblies enabled a comprehensive assessment of large structural variations. The insight obtained from this analysis will form the basis of future wheat pan-genome studies.

genomics

Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes

Assessing accuracy of NGS variant calling is immensely facilitated by a robust benchmarking strategy and tools to carry it out in a standard way. Benchmarking variant calls requires careful attention to definitions of performance metrics, sophisticated comparison approaches, and stratification by variant type and genome context. The Global Alliance for Genomics and Health (GA4GH) Benchmarking Team has developed standardized performance metrics and tools for benchmarking germline small variant calls. This team includes representatives from sequencing technology developers, government agencies, academic bioinformatics researchers, clinical laboratories, and commercial technology and bioinformatics developers for whom benchmarking variant calls is essential to their work. Benchmarking variant calls is a challenging problem for many reasons:\n\nO_LIEvaluating variant calls requires complex matching algorithms and standardized counting because the same variant may be represented differently in truth and query callsets.\nC_LIO_LIDefining and interpreting resulting metrics such as precision (aka positive predictive value = TP/(TP+FP)) and recall (aka sensitivity = TP/(TP+FN)) requires standardization to draw robust conclusions about comparative performance for different variant calling methods.\nC_LIO_LIPerformance of NGS methods can vary depending on variant types and genome context; and as a result understanding performance requires meaningful stratification.\nC_LIO_LIHigh-confidence variant calls and regions that can be used as \"truth\" to accurately identify false positives and negatives are difficult to define, and reliable calls for the most challenging regions and variants remain out of reach.\nC_LI\n\nWe have made significant progress on standardizing comparison methods, metric definitions and reporting, as well as developing and using truth sets. Our methods are publicly available on GitHub (https://github.com/ga4gh/benchmarking-tools) and in a web-based app on precisionFDA, which allow users to compare their variant calls against truth sets and to obtain a standardized report on their variant calling performance. Our methods have been piloted in the precisionFDA variant calling challenges to identify the best-in-class variant calling methods within high-confidence regions. Finally, we recommend a set of best practices for using our tools and critically evaluating the results.

genomics

Comprehensive analysis of chromothripsis in 2,658 human cancers using whole-genome sequencing

Chromothripsis is a newly discovered mutational phenomenon involving massive, clustered genomic rearrangements that occurs in cancer and other diseases. Recent studies in cancer suggest that chromothripsis may be far more common than initially inferred from low resolution DNA copy number data. Here, we analyze the patterns of chromothripsis across 2,658 tumors spanning 39 cancer types using whole-genome sequencing data. We find that chromothripsis events are pervasive across cancers, with a frequency of >50% in several cancer types. Whereas canonical chromothripsis profiles display oscillations between two copy number states, a considerable fraction of the events involves multiple chromosomes as well as additional structural alterations. In addition to non-homologous end-joining, we detect signatures of replicative processes and templated insertions. Chromothripsis contributes to oncogene amplification as well as to inactivation of genes such as mismatch-repair related genes. These findings show that chromothripsis is a major process driving genome evolution in human cancer.

genomics