Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Modules Of Co-Occurrence In The Cyanobacterial Pan-Genome

The increasing availability of fully sequenced cyanobacterial genomes opens unprecedented opportunities to investigate the manifold adaptations and functional relationships that determine the genetic content of individual bacterial species. Here, we use comparative genome analysis to investigate the cyanobacterial pan-genome based on 77 strains whose complete genome sequence is available. Our focus is the co-occurrence of likely ortholog genes, denoted as CLOGs. We conjecture that co-occurrence CLOGs is indicative of functional relationships between the respective genes. Going beyond the analysis of pair-wise co-occurrences, we introduce a novel network approach to identify modules of co-occurring ortholog genes. Our results demonstrate that these modules exhibit a high degree of functional coherence and reveal known as well as previously unknown functional relationships. We argue that the high functional coherence observed for the extracted modules is a consequence of the similar-yet-diverse nature of the cyanobacterial phylum. We provide a simple toolbox that facilitates further analysis of our results with respect to specific cyanobacterial genes of interest.

bioinformatics

Migration Harshness Drives Habitat Choice And Local Adaptation In Anadromous Arctic Char: Evidence From Integrating Population Genomics And Acoustic Telemetry

Migration is a ubiquitous life history trait with profound evolutionary and ecological consequences. Recent developments in telemetry and genomics, when combined, can bring significant insights on the migratory ecology of non-model organisms in the wild. Here, we used this integrative approach to document dispersal, gene flow and potential for local adaptation in anadromous Arctic Char from six rivers in the Canadian Arctic. Acoustic telemetry data from 124 tracked individuals indicated asymmetric dispersal, with a large proportion of fish (72%) tagged in three different rivers migrating up the same short river in the fall. Population genomics data from 6,136 SNP markers revealed weak, albeit significant, population differentiation (average pairwise FST = 0.011) and asymmetric dispersal was also revealed by population assignments. Approximate Bayesian Computation simulations suggested the presence of asymmetric gene flow, although in the opposite direction to that observed from the telemetry data, suggesting that dispersal does not necessarily lead to gene flow. These observations suggested that Arctic Char home to their natal river to spawn, but may overwinter in rivers with the shortest migratory route to minimize the costs of migration in non-breeding years. Genome scans and genetic-environment associations identified 90 outlier markers putatively under selection, 23 of which were in or near a gene. Of these, at least four were involved in muscle and cardiac function, consistent with the hypothesis that migratory harshness could drive local adaptation. Our study illustrates the power of integrating genomics and telemetry to study migrations in non-model organisms in logistically challenging environments such as the Arctic.

evolutionary biology

Cancer Genome Interpreter Annotates The Biological And Clinical Relevance Of Tumor Alterations

While tumor genome sequencing has become widely available in clinical and research settings, the interpretation of tumor somatic variants remains an important bottleneck. Most of the alterations observed in tumors, including those in well-known cancer genes, are of uncertain significance. Moreover, the information on tumor genomic alterations shaping the response to existing therapies is fragmented across the literature and several specialized resources. Here we present the Cancer Genome Interpreter (http://www.cancergenomeinterpreter.org), an open access tool that we have implemented to annotate genomic alterations and interpret their possible role in tumorigenesis and in the response to anti-cancer therapies.

cancer biology

Mitigating Mitochondrial Genome Erosion Without Recombination

Mitochondria are ATP-producing organelles of bacterial ancestry that played a key role in the origin and early evolution of complex eukaryotic cells. Most modern eukaryotes transmit mitochondrial genes uniparentally, often without recombination among genetically divergent organelles. While this asymmetric inheritance maintains the efficacy of purifying selection at the level of the cell, the absence of recombination could also make the genome susceptible to Mullers ratchet. How mitochondria escape this irreversible defect accumulation is a fundamental unsolved question. Occasional paternal leakage could in principle promote recombination, but it would also compromise the purifying-selection benefits of uniparental inheritance. We assess this tradeoff using a stochastic population-genetic model. In the absence of recombination, uniparental inheritance of freely segregating genomes mitigates mutational erosion, while paternal leakage exacerbates the ratchet effect. Mitochondrial fusion-fission cycles ensure independent genome segregation, improving purifying selection. Paternal leakage provides opportunity for recombination to slow down the mutation accumulation, but always at a cost of increased steady-state mutation load. Our findings indicate that random segregation of mitochondrial genomes under uniparental inheritance can effectively combat the mutational meltdown, and that homologous recombination under paternal leakage might not be needed.

genetics

RNA-mediated Genomic Arrangements in Mammalian Cells

One of the hallmarks of cancer is the formation of oncogenic fusion genes as a result of chromosomal translocations. Fusion genes are presumed to occur prior to fusion RNA expression. However, studies have reported the presence of fusion RNAs in individuals who were negative for chromosomal translocations. These observations give rise to \"the cart before the horse\" hypothesis, in which fusion RNA precedes the fusion gene and guides the genomic rearrangements that ultimately result in gene fusions. Yet RNA-mediated genomic rearrangement in mammalian cells has never been demonstrated. Here we provide evidence that expression of a chimeric RNA drives formation of a specified gene fusion via genomic rearrangement in mammalian cells. The process is (1) specified by the sequence of chimeric RNA involved, (2) facilitated by physiological hormone levels, (3) permissible regardless of intra-chromosomal (TMPRSS2-ERG) or inter-chromosomal (TMPRSS2-ETV1) fusion, and (4) can occur in normal cells prior to malignant transformation. We demonstrate that, contrary to \"the cart before the horse\" model, it is the antisense rather than sense chimeric RNAs that effectively drive gene fusion, and that this disparity can be explained by transcriptional conflict. Furthermore, we identified an endogenous RNA AZI1 that acts as the initiator RNA to induce TMPRSS2-ERG fusion. RNA-driven gene fusion demonstrated in this report provides important insight in early disease mechanism, and could have fundamental implications in the biology of mammalian genome stability, as well as gene editing technology via mechanisms native to mammalian cells.

cancer biology

Rapid whole genome amplification and sequencing of low cell numbers in a bacteraemia model

Whilst next generation sequencing is frequently used to whole genome sequence bacteria from cultures, its rarely applied directly to clinical samples. Therefore, this study addresses the issue of applying NGS microbial diagnostics directly to blood samples. To demonstrate the potential of direct from blood sequencing a bacteria spiked blood model was developed. Horse blood was spiked with clinical samples of E. coli and S. aureus, and a process developed to isolate bacterial cells whilst removing the majority of host DNA. One sample of each isolate was then amplified using {phi}29 multiple displacement amplification (MDA) and sequenced. The total processing time, from sample to amplified DNA ready for sequencing was 3.5 hours, significantly faster than the 18-hour overnight culture step which is typically required. Both bacteria showed 100% survival through the processing. The direct from sample sequencing resulted in greater than 92% genome coverage of the pathogens whilst limiting the sequencing of host genome (less than 7% of all reads). Analysis of de novo assembled reads allowed accurate genotypic antibiotic resistance prediction. The sample processing is easily applicable to multiple sequencing platforms. Overall this model demonstrates potential to rapidly generate whole genome bacterial data directly from blood.

microbiology

Whole genome sequencing-based detection of antimicrobial resistance and virulence in non-typhoidal Salmonella enterica isolated from wildlife

The aim of this study was to generate a reference set of Salmonella enterica genomes isolated from wildlife from the United States and to determine the antimicrobial resistance and virulence gene profile of the isolates from the genome sequence data. We sequenced the whole genomes of 103 Salmonella isolates sampled between 1988 and 2003 from wildlife and exotic pet cases that were submitted to the Oklahoma Animal Disease Diagnostic Laboratory, Stillwater, Oklahoma. Among 103 isolates, 50.48% were from wild birds, 0.9% was from fish, 24.27% each were from reptiles and mammals. 50.48% isolates showed resistance to at least one antibiotic. Resistance against the aminoglycoside streptomycin was most common while 9 isolates were found to be multi-drug resistant having resistance against more than three antibiotics. Determination of virulence gene profile revealed that the genes belonging to csg operons, the fim genes that encode for type 1 fimbriae and the genes belonging to type III secretion system were predominant among the isolates. The universal presence of fimbrial genes and the genes encoded by pathogenicity islands 1-2 among the isolates we report here indicates that these isolates could potentially cause disease in humans. Therefore, the genomes we report here could be a valuable reference point for future traceback investigations when wildlife is considered to be the potential source of human Salmonellosis.

microbiology

CRISPR-Cpf1 mediates efficient homology-directed repair and temperature-controlled genome editing

Cpf1 is a novel class of CRISPR-Cas DNA endonucleases, with a wide range of activity across different eukaryotic systems. Yet, the underlying determinants of this variability are poorly understood. Here, we demonstrate that LbCpf1, but not AsCpf1, ribonucleoprotein complexes allow efficient mutagenesis in zebrafish and Xenopus. We show that temperature modulates Cpf1 activity by controlling its ability to access genomic DNA. This effect is stronger on AsCpf1, explaining its lower efficiency in ectothermic organisms. We capitalize on this property to show that temporal control of the temperature allows post-translational modulation of Cpf1-mediated genome editing. Finally, we determine that LbCpf1 significantly increases homology-directed repair in zebrafish, improving current approaches for targeted DNA integration in the genome. Together, we provide a molecular understanding of Cpf1 activity in vivo and establish Cpf1 as an efficient and inducible genome engineering tool across ectothermic species.

genetics

Signatures of the evolution of parthenogenesis and cryptobiosis in the genomes of panagrolaimid nematodes

Most animal species reproduce sexually, but parthenogenesis, asexual reproduction of various forms, has arisen repeatedly. Parthenogenetic lineages are usually short lived in evolution; though in some environments parthenogenesis may be advantageous, avoiding the cost of sex. Panagrolaimus nematodes have colonised environments ranging from arid deserts to arctic and antarctic biomes. Many are parthenogenetic, and most have cryptobiotic abilities, being able to survive repeated complete desiccation and freezing. It is not clear which genomic and molecular mechanisms led to the successful establishment of parthenogenesis and the evolution of cryptobiosis in animals in general. At the same time, model systems to study these traits in the laboratory are missing.\n\nWe compared the genomes and transcriptomes of parthenogenetic and sexual Panagrolaimus able to survive crybtobiosis, as well as a non-cryptobiotic Propanogrolaimus species, to identify systems that contribute to these striking abilities. The parthenogens are most probably tripoids originating from hybridisation (allopolyploids). We identified genomic singularities like expansion of gene families, and selection on genes that could be linked to the adaptation to cryptobiosis. All Panagrolaimus have acquired genes through horizontal transfer, some of which are likely to contribute to cryptobiosis. Many genes acting in C. elegans reproduction and development were absent in distant nematode species (including the Panagrolaimids), suggesting molecular pathways cannot directly be transferred from the model system.\n\nThe easily cultured Panagrolaimus nematodes offer a system to study developmental diversity in Nematoda, the molecular evolution of parthenogens, the effects of triploidy on genomes stability, and the origin and biology of cryptobiosis.

evolutionary biology

ShinyGPAS: Interactive genomic prediction accuracy simulator based on deterministic formulas

BackgroundDeterministic formulas highlight the relationships among prediction accuracy and potential factors influencing prediction accuracy prior to performing computationally intensive cross-validation. Visualizing such deterministic formulas in an interactive manner may lead to a better understanding of how genetic factors control prediction accuracy.\n\nResultsThe software to simulate deterministic formulas for genomic prediction accuracy was implemented in R and encapsulated as a web-based Shiny application. ShinyGPAS (Shiny Genomic Prediction Accuracy Simulator) simulates various deterministic formulas and delivers dynamic scatter plots of prediction accuracy vs. genetic factors impacting prediction accuracy, while requiring only mouse navigation in a web browser. ShinyGPAS is available at: https://chikudaisei.shinyapps.io/shinygpas/.\n\nConclusionShinyGPAS is a shiny-based interactive genomic prediction accuracy simulator using deterministic formulas. It can be used for interactively exploring potential factors influencing prediction accuracy in genome-enabled prediction, simulating achievable prediction accuracy prior to genotyping individuals, or supporting in-class teaching. ShinyGPAS is open source software and it is hosted online as a freely available web-based resource with an intuitive graphical user interface.

genetics

Identifying Simultaneous Rearrangements in Cancer Genomes

The traditional view of cancer evolution states that a cancer genome accumulates a sequential ordering of mutations over a long period of time. However, in recent years it has been suggested that a cancer genome may instead undergo a one-time catastrophic event, such as chromothripsis, where a large number of mutations instead occur simultaneously. A number of potential signatures of chromothripsis have been proposed. In this work we provide a rigorous formulation and analysis of the \"ability to walk the derivative chromosome\" signature originally proposed by Korbel and Campbell (2013). In particular, we show that this signature, as originally envisioned, may not always be present in a chromothripsis genome and we provide a precise quantification of under what circumstances it would be present. We also propose a variation on this signature, the H/T alternating fraction, which allows us to overcome some of the limitations of the original signature. We apply our measure to both simulated data and a previously analyzed real cancer dataset and find that the H/T alternating fraction may provide useful signal for distinguishing genomes having acquired mutations simultaneously from those acquired in a sequential fashion. An implementation of the H/T alternating fraction is available at https://bitbucket.org/oesperlab/ht-altfrac.

bioinformatics

Supervised machine learning reveals introgressed loci in the genomes of Drosophila simulans and D. sechellia

Hybridization and gene flow between species appears to be common. Even though it is clear that hybridization is widespread across all surveyed taxonomic groups, the magnitude and consequences of introgression are still largely unknown. Thus it is crucial to develop the statistical machinery required to uncover which genomic regions have recently acquired haplotypes via introgression from a sister population. We developed a novel machine learning framework, called FILET (Finding Introgressed Loci via Extra-Trees) capable of revealing genomic introgression with far greater power than competing methods. FILET works by combining information from a number of population genetic summary statistics, including several new statistics that we introduce, that capture patterns of variation across two populations. We show that FILET is able to identify loci that have experienced gene flow between related species with high accuracy, and in most situations can correctly infer which population was the donor and which was the recipient. Here we describe a data set of outbred diploid Drosophila sechellia genomes, and combine them with data from D. simulans to examine recent introgression between these species using FILET. Although we find that these populations may have split more recently than previously appreciated, FILET confirms that there has indeed been appreciable recent introgression (some of which might have been adaptive) between these species, and reveals that this gene flow is primarily in the direction of D. simulans to D. sechellia.\n\nAUTHOR SUMMARYUnderstanding the extent to which species or diverged populations hybridize in nature is crucially important if we are to understand the speciation process. Accordingly numerous research groups have developed methodology for finding the genetic evidence of such introgression. In this report we develop a supervised machine learning approach for uncovering loci which have introgressed across species boundaries. We show that our method, FILET, has greater accuracy and power than competing methods in discovering introgression, and in addition can detect the directionality associated with the gene flow between species. Using whole genome sequences from Drosophila simulans and Drosophila sechellia we show that FILET discovers quite extensive introgression between these species that has occurred mostly from D. simulans to D. sechellia. Our work highlights the complex process of speciation even within a well-studied system and points to the growing importance of supervised machine learning in population genetics.

evolutionary biology

OligoMiner: A rapid, flexible environment for the design of genome-scale oligonucleotide in situ hybridization probes

Oligonucleotide (oligo)-based fluorescence in situ hybridization (FISH) has emerged as an important tool for the study of chromosome organization and gene expression and has been empowered by the commercial availability of highly complex pools of oligos. However, a dedicated bioinformatic design utility has yet to be created specifically for the purpose of identifying optimal oligo FISH probe sequences on the genome-wide scale. Here, we introduce OligoMiner, a rapid and robust computational pipeline for the genome-scale design of oligo FISH probes that affords the scientist exact control over the parameters of each probe. Our streamlined method uses standard bioinformatic file formats, allowing users to seamlessly integrate existing and new utilities into the pipeline as desired, and introduces a novel method for evaluating the specificity of each probe molecule that connects simulated hybridization energetics to rapidly generated sequence alignments using supervised learning. We demonstrate the scalability of our approach by performing genome-scale probe discovery in numerous model organism genomes and showcase the performance of the resulting probes with both diffraction-limited and single-molecule super-resolution imaging of chromosomal and RNA targets. We anticipate this pipeline will make the FISH probe design process much more accessible and will more broadly facilitate the design of pools of hybridization probes for a variety of applications.

bioinformatics

De Novo Prediction of Human Chromosome Structures: Epigenetic Marking Patterns Encode Genome Architecture

Inside the cell nucleus, genomes fold into organized structures that are characteristic of cell type. Here, we show that this chromatin architecture can be predicted de novo using epigenetic data derived from ChIP-Seq. We exploit the idea that chromosomes encode a one-dimensional sequence of chromatin structural types. Interactions between these chromatin types determine the three-dimensional (3D) structural ensemble of chromosomes through a process similar to phase separation. First, a recurrent neural network is used to infer the relation between the epigenetic marks present at a locus, as assayed by ChIP-Seq, and the genomic compartment in which those loci reside, as measured by DNA-DNA proximity ligation (Hi-C). Next, types inferred from this neural network are used as an input to an energy landscape model for chromatin organization (MiChroM) in order to generate an ensemble of 3D chromosome conformations. After training the model, dubbed MEGABASE (Maximum Entropy Genomic Annotation from Biomarkers Associated to Structural Ensembles), on odd numbered chromosomes, we predict the chromatin type sequences and the subsequent 3D conformational ensembles for the even chromosomes. We validate these structural ensembles by using ChIP-Seq tracks alone to predict Hi-C maps as well as distances measured using 3D FISH experiments. Both sets of experiments support the hypothesis of phase separation being the driving process behind compartmentalization. These findings strongly suggest that epigenetic marking patterns encode sufficient information to determine the global architecture of chromosomes and that de novo structure prediction for whole genomes may be increasingly possible.

biophysics

COGEM: A Toolbox for Computational Genomics in Matlab

MotivationThe Matlab programming language is widely used for both teaching and research in engineering, computer science, and mathematics. Despite its many strengths, it has never been a dominant language in computational genomics or bioinformatics more generally.\n\nResultsHere, we introduce COGEM, a long-term project to develop computational genomics functionality in Matlab. The initial release provides functions for manipulating genomic intervals, stranded or unstranded, with or without numerical data associated. It includes features for both text and binary file input and output, conversion between BAM, BED and BEDGRAPH formats, and numerous functions for manipulating intervals, including shifting, expanding, overlapping, intersecting, unioning, finding nearest intervals, piling up intervals, and performing unary and binary numerical and logical operations on sets of intervals. The toolbox is well-suited to the analysis of high-throughput sequencing data. We demonstrate its functionality by creating a ChIP-seq peak-calling algorithm by chaining together a series of commands, and find it capable of analyzing genome-scale data in reasonable time.\n\nAvailabilityThe current toolbox and reference manual is available as supplementary material, and updated versions will be maintained at www.perkinslab.ca online.

bioinformatics

Integrative analysis of large scale transcriptome data draws a comprehensive landscape of Phaeodactylum tricornutum functional genome and evolutionary origin of diatoms

Diatoms are one of the most successful and ecologically important groups of eukaryotic phytoplankton in the modern ocean. Deciphering their genomes is a key step towards better understanding of their biological innovations, evolutionary origins, and ecological underpinnings. Here, we have used 90 RNA-Seq datasets from different growth conditions combined with published expressed sequence tags and protein sequences from multiple taxa to explore the genome of the model diatom Phaeodactylum tricornutum, and introduce 1,489 novel genes. The new annotation additionally permitted the discovery for the first time of extensive alternative splicing (AS) in diatoms, including intron retention and exon skipping which increases the diversity of transcripts to regulate gene expression in response to nutrient limitations. In addition, we have used up-to-date reference sequence libraries to dissect the taxonomic origins of diatom genomes. We show that the P. tricornutum genome is replete in lineage-specific genes, with up to 47% of the gene models present only possessing orthologues in other stramenopile groups. Finally, we have performed a comprehensive de novo annotation of repetitive elements showing novel classes of TEs such as SINE, MITE, LINE and TRIM/LARD. This work provides a solid foundation for future studies of diatom gene function, evolution and ecology.

bioinformatics

ActiveDriverDB: human disease mutations and genome variation in post-translational modification sites of proteins

Interpretation of genetic variation is required for understanding genotype-phenotype associations, mechanisms of inherited disease, and drivers of cancer. Millions of single nucleotide variants (SNVs) in human genomes are known and thousands are associated with disease. An estimated 20% of disease-associated missense SNVs are located in protein sites of post-translational modifications (PTMs), chemical modifications of amino acids that extend protein function. ActiveDriverDB is a comprehensive human proteo-genomics database that annotates disease mutations and population variants using PTMs. We integrated >385,000 published PTM sites with [~]3.8 million missense SNVs from The Cancer Genome Atlas (TCGA), the ClinVar database of disease genes, and inter-individual variation from human genome sequencing projects. The database includes interaction networks of proteins, upstream enzymes such as kinases, and drugs targeting these enzymes. We also predicted network-rewiring impact of mutations by analyzing gains and losses of kinase-bound sequence motifs. ActiveDriverDB provides detailed visualization, filtering, browsing and searching options for studying PTM-associated SNVs. Users can upload mutation datasets interactively and use our application programming interface for pipelines. Integrative analysis of SNVs and PTMs helps decipher molecular mechanisms of phenotypes and disease, as exemplified by case studies of disease genes TP53, BRCA2 and VHL. The open-source database is available at https://www.ActiveDriverDB.org.

bioinformatics

SeroBA: rapid high-throughput serotyping of Streptococcus pneumoniae from whole genome sequence data

Streptococcus pneumoniae is responsible for 240,000 - 460,000 deaths in children under 5 years of age each year. Accurate identification of pneumococcal serotypes is important for tracking the distribution and evolution of serotypes following the introduction of effective vaccines. Recent efforts have been made to infer serotypes directly from genomic data but current software approaches are limited and do not scale well. Here, we introduce a novel method, SeroBA, which uses a hybrid assembly and mapping approach. We compared SeroBA against real and simulated data and present results on the concordance and computational performance against a validation dataset, the robustness and scalability when analysing a large dataset, and the impact of varying the depth of coverage in the cps locus region on sequence-based serotyping. SeroBA can predict serotypes, by identifying the cps locus, directly from raw whole genome sequencing read data with 98% concordance using a k-mer based method, can process 10,000 samples in just over 1 day using a standard server and can call serotypes at a coverage as low as 10x. SeroBA is implemented in Python3 and is freely available under an open source GPLv3 license from: https://github.com/sanger-pathogens/seroba\n\nDATA SUMMARYO_LIThe reference genome Streptococcus pneumoniae ATCC 700669 is available from National Center for Biotechnology Information (NCBI) with the accession number: FM211187\nC_LIO_LISimulated paired end reads for experiment 2 have been deposited in FigShare: https://doi.org/10.6084/m9.figshare.5086054.v1\nC_LIO_LIAccession numbers for all other experiments are listed in Supplementary Table S1 and Supplementary Table S2.\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTThis article describes SeroBA, a k-mer based method for predicting the serotypes of Streptococcus pneumoniae from Whole Genome Sequencing (WGS) data. SeroBA can identify 92 serotypes and 2 subtypes with constant memory usage and low computational costs. We showed that SeroBA is able to reliably predict serotypes at a depth of coverage as low as 10x and is scalable to large datasets.

bioinformatics