Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Fast and sensitive mapping of error-prone nanopore sequencing reads with GraphMap

Exploiting the power of nanopore sequencing requires the development of new bioinformatics approaches to deal with its specific error characteristics. We present the first nanopore read mapper (GraphMap) that uses a read-funneling paradigm to robustly handle variable error rates and fast graph traversal to align long reads with speed and very high precision (>95%). Evaluation on MinION sequencing datasets against short and long-read mappers indicates that GraphMap increases mapping sensitivity by at least 15-80%. GraphMap alignments are the first to demonstrate consensus calling with <1 error in 100,000 bases, variant calling on the human genome with 76% improvement in sensitivity over the next best mapper (BWA-MEM), precise detection of structural variants from 100bp to 4kbp in length and species and strain-specific identification of pathogens using MinION reads. GraphMap is available open source under the MIT license at https://github.com/isovic/graphmap.

Genomics

Connecting small RNAs and Aging

Website SummarySmall RNAs are important gene regulators of stress response, aging, and many other things. More analysis needs to be done in order to gain a better understanding of these molecules to find connections between small RNAs and important things like aging.\n\nSummarySmall RNAs are a diverse population of gene regulators, but their role in the cell is not fully characterized. Bioinformatics was used to prove their connection with aging and expand current knowledge for these molecules.\n\nAbstractSmall RNAs have a wide range of functions and recent studies have found connections between these molecules and aging pathways. However, the process to systematically characterize this relationship is slow. Prediction tools can be used to expedite this process by finding new genes and pathways that cross talk with each other. Using phylogenetic and systems analysis, connections between small RNAs and aging were proven and new genes that may be related to aging were identified. This type of analysis can be applied to many different pathways in order to fully characterize the role of small RNAs.

Molecular Biology

Structure and conformational states of the bovine mitochondrial ATP synthase by cryo-EM

Adenosine triphosphate (ATP), the chemical energy currency of biology, is synthesized in eukaryotic cells primarily by the mitochondrial ATP synthase. ATP synthases operate by a rotary catalytic mechanism where proton translocation through the membrane-inserted Fo region is coupled to ATP synthesis in the catalytic F1 region via rotation of a central rotor subcomplex. We report here single particle electron cryomicroscopy (cryo-EM) analysis of the bovine mitochondrial ATP synthase. Combining cryo-EM data with bioinformatic analysis allowed us to determine the fold of the a subunit, suggesting a proton translocation path through the Fo region that involves both the a and b subunits. 3D classification of images revealed seven distinct states of the enzyme that show different modes of bending and twisting in the intact ATP synthase. Rotational fluctuations of the c8-ring within the Fo region support a Brownian ratchet mechanism for proton-translocation driven rotation in ATP synthases.

Biochemistry

Clinical metagenomic identification of Balamuthia mandrillaris encephalitis and assembly of the draft genome: the critical need for reference strain sequencing

Primary amoebic meningoencephalitis (PAM) is a rare, often lethal cause of encephalitis, for which early diagnosis and prompt initiation of combination antimicrobials may improve clinical outcomes. In this study, we present the first draft assembly of the Balamuthia mandrillaris genome recovered from a rare survivor of PAM, in total comprising 49 Mb of sequence. Comparative analysis of the mitochondrial genome and high-copy number genes from 6 additional Balamuthia mandrillaris strains demonstrated remarkable sequence variation, with the closest homologs corresponding to other amoebae, hydroids, algae, slime molds, and peat moss. We also describe the use of unbiased metagenomic next-generation sequencing (NGS) and SURPI bioinformatics analysis to diagnose an ultimately fatal case of Balamuthia mandrillaris encephalitis in a 15-year old girl. Real-time NGS testing of a hospital day 6 CSF sample detected Balamuthia on the basis of high-quality hits to 16S and 18S ribosomal RNA sequences present in the National Center for Biotechnology Information (NCBI) nt reference database. Retrospective analysis of a day 1 CSF sample revealed that more timely identification of Balamuthia by metagenomic NGS, potentially resulting in a better outcome, would have required availability of the complete genome sequence. These results underscore the diverse evolutionary origins underpinning this eukaryotic pathogen, and the critical importance of whole-genome reference sequences for microbial detection by NGS.

Genomics

Development of EST-SSR annotated database in olive (Olea europaea)

Olive tree (Olea europaea L.) is one of the most important oil producing crops in the world and the genetic identification of several genotypes by using molecular markers is the first step in its breeding programs. A set of 1,801 well-informative EST-SSR primers targeting specific Olive genes included in different biological processes and pathways were generated using 11,215 Olive EST sequences acquired from the NCBI database. Our bioinformatics analytical procedure showed that 8295 SSR motifs were detected which belonged to different motif types with occurrences of 77.6%, 11.84%, 8.62%, 0.84%, 0.77% and 0.29% for Mononucleotide, trinucleotide, dinucleotide, hexanucleotide, pentanucleotide and tetranucleotide respectively. The appearance of the AAG/CTT repeat was highly represented in trinucleotide and the representation of AG/CT was high in dinucleotide repeats. Results obtained from functional annotation of olives EST sequences targeted with our primers set indicated that 78.5% of these sequences having homology with known proteins, while 4.2% was homologous to hypothetical, predicted, unnamed or uncharacterized proteins and the 17.3% sequences did not possess homology with any known proteins. Our EST-SSR primer set cover a total of 92 biological pathways such as carbohydrate metabolism pathway, energy metabolism& carbon fixation in photosynthetic organism pathway including 11 pathways associated with lipid metabolism. A twenty five randomly selected primers were applied to 9 Egyptian cultivated olive accessions to test its amplification and polymorphism detection efficacy. All tested primers were successfully amplified and only 10 exhibited detectable polymorphism.

Genomics

Tying down loose ends in the Chlamydomonas genome

The Chlamydomonas genome has been sequenced, assembled and annotated to produce a rich resource for genetics and molecular biology in this well-studied model organism. The annotated genome is very rich in open reading frames upstream of the annotated coding sequence ( uORFs): almost three quarters of the assigned transcripts have at least one uORF, and frequently more than one. This is problematic with respect to the standard scanning model for eukaryotic translation initiation. These uORFs can be grouped into three classes: class 1, initiating in-frame with the coding sequence (cds) (thus providing a potential in-frame N-terminal extension); class 2, initiating in the 5UT and terminating out-of-frame in the cds; and class 3, initiating and terminating within the 5UT. Multiple bioinformatics criteria (including analysis of Kozak consensus sequence agreement and BLASTP comparisons to the closely related Volvox genome, and statistical comparison to cds and to random-sequence controls) indicate that of ~4000 class 1 uORFs, approximately half are likely in vivo translation initiation sites. The proposed resulting N-terminal extensions in many cases will sharply alter the predicted biochemical properties of the encoded proteins. These results suggest significant modifications in ~2000 of the ~20,000 transcript models with respect to translation initiation and encoded peptides. In contrast, class 2 uORFs may be subject to purifying selection, and the existent ones (surviving selection) are likely inefficiently translated. Class 3 uORFs are remarkably similar to random sequence expectations with respect to size, number and composition and therefore may be largely selectively neutral; their very high abundance (found in more than half of transcripts, frequently with multiple uORFs per transcript) nevertheless suggests the possibility of translational regulation on a wide scale.

Genomics

Phylogenomic Reconstruction Supports Supercontinent Origins for Leishmania

Leishmania, a genus of parasites transmitted to human hosts and mammalian/reptilian reservoirs by an insect vector, is the causative agent of the human disease complex leishmaniasis. The evolutionary relationships within the genus Leishmania and its origins are the source of ongoing debate, reflected in conflicting phylogenetic and biogeographic reconstructions. This study employs a recently described bioinformatics method, SISRS, to identify over 200,000 informative sites across the genome from newly sequenced and publicly available Leishmania data. This dataset is used to reconstruct the evolutionary relationships of this genus. Additionally, we constructed a large multi-gene dataset; we used this dataset to reconstruct the phylogeny and estimate divergence dates for species. We conclude that the genus Leishmania evolved at least 90-100 million years ago. Our results support the hypothesis that Leishmania clades separated prior to, and during, the breakup of Gondwana. Additionally, we confirm that reptile-infecting Leishmania are derived from mammalian forms, and that the species that infect porcupines and sloths form a clade long separated from other species. We also firmly place the guinea-pig infecting species, L. enrietti, the globally dispersed L. siamensis, and the newly identified Australia species from kangaroos as sibling species whose distribution arises from the ancient connection between Australia, Antarctica, and South America.

Evolutionary Biology

Patching holes in the Chlamydomonas genome

The Chlamydomonas genome has been sequenced, assembled and annotated to produce a rich resource for genetics and molecular biology in this well-studied model organism. However, the current reference genome contains ~1000 blocks of unknown sequence ( N-islands), which are frequently placed in introns of annotated gene models. We developed a strategy, using careful bioinformatics analysis of short-sequence cDNA and genomic DNA reads, to search for previously unknown exons hidden within such blocks, and determine the sequence and exon/intron boundaries of such exons. These methods are based on assembly and alignment completely independent of prior reference assembly or reference annotation. Our evidence indicates that ~one-quarter of the annotated intronic N-islands actually contain hidden exons. For most of these our algorithm recovers full exonic sequence with associated splice junctions and exon-adjacent intron sequence, that can be joined to the reference genome assembly and annotated transcript models. These new exons represent de novo sequence generally present nowhere in the assembled genome, and the added sequence can be shown in many cases to greatly improve evolutionary conservation of the predicted encoded peptides. At the same time, our results confirm the purely intronic status for a substantial majority of N-islands annotated as intronic in the reference annotated genome, increasing confidence in this valuable resource.

Genomics

Methylome analysis reveals dysregulated developmental and viral pathways in breast cancer

BackgroundBreast cancer (BC) ranks among the most common cancers in Sudan and worldwide with hefty toll on female health and human resources. Recent studies have uncovered a common BC signature characterized by low frequency of oncogenic mutations and high frequency of epigenetic silencing of major BC tumor suppressor genes. Therefore, we conducted a genome-wide methylome study to characterize aberrant DNA methylation in breast cancer.\n\nResultsDifferential methylation analysis between primary tumor samples and normal samples from healthy adjacent tissues yielded 20188 differentially methylated positions (DMPs), which is further divided into 13633 hypermethylated sites corresponding to 5339 genes and 6555 hypomethylated sites corresponding to 2811 genes. Moreover, bioinformatics analysis revealed epigenetic dysregulation of major developmental pathways including hippo signaling pathway. We also uncovered many clues to a possible role for EBV infection in BC\n\nConclusionOur results clearly show the utility of epigenetic assays in interrogating breast cancer tumorigenesis, and pinpointing specific developmental and viral pathways dysregulation that might serve as potential biomarkers or targets for therapeutic interventions.

Genomics

Divorcing strain classification from species names

Confusion about strain classification and nomenclature permeates modern microbiology. Although taxonomists have traditionally acted as gatekeepers of order, the numbers of and speed at which new strains are identified has outpaced the opportunity for professional classification for many lineages. Furthermore, the growth of bioinformatics and database fueled investigations have placed metadata curation in the hands of researchers with little taxonomic experience. Here I describe practical challenges facing modern microbial taxonomy, provide an overview of complexities of classification for environmentally ubiquitous taxa like Pseudomonas syringae, and emphasize that classification and nomenclature need not be the one in the same. A move toward implementation of relational classification schemes based on inherent properties of whole genomes could provide sorely needed continuity in how strains are referenced across manuscripts and data sets.

Microbiology

gEVAL - A web based browser for evaluating genome assemblies

MotivationFor most research approaches, genome analyses are dependent on the existence of a high quality genome reference assembly. However, the local accuracy of an assembly remains difficult to assess and improve. The gEVAL browser allows the user to interrogate an assembly in any region of the genome by comparing it to different datasets and evaluating the concordance. These analyses include: a wide variety of sequence alignments, comparative analyses of multiple genome assemblies, and consistency with optical and other physical maps. gEVAL highlights allelic variations, regions of low complexity, abnormal coverage, and potential sequence and assembly errors, and offers strategies for improvement. While gEVAL focuses primarily on sequence integrity, it can also display arbitrary annotation including Ensembl or TrackHub sources. We provide gEVAL web sites for many human, mouse, zebrafish and chicken assemblies to support the Genome Reference Consortium, and gEVAL is also downloadable to enable its use for any organism and assembly.\n\nAvailabilityWeb Browser: http://geval.sanger.ac.uk, Plugin: http://wchow.github.io/wtsi-geval-plugin.\n\nContactkj2@sanger.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics

An open library of human kinase domain constructs for automated bacterial expression

Kinases play a critical role in many cellular signaling pathways and are dysregulated in a number of diseases, such as cancer, diabetes, and neurodegeneration. Since the FDA approval of imatinib in 2001, therapeutics targeting kinases now account for roughly 50% of current cancer drug discovery efforts. The ability to explore human kinase biochemistry, biophysics, and structural biology in the laboratory is essential to making rapid progress in understanding kinase regulation, designing selective inhibitors, and studying the emergence of drug resistance. While insect and mammalian expression systems are frequently used for the expression of human kinases, bacterial expression systems are superior in terms of simplicity and cost-effectiveness but have historically struggled with human kinase expression. Following the discovery that phosphatase coexpression could produce high yields of Src and Abl kinase domains in bacterial expression systems, we have generated a library of 52 His-tagged human kinase domain constructs that express above 2 {micro}g/mL culture in a simple automated bacterial expression system utilizing phosphatase coexpression (YopH for Tyr kinases, Lambda for Ser/Thr kinases). Here, we report a structural bioinformatics approach to identify kinase domain constructs previously expressed in bacteria likely to express well in a simple high-throughput protocol, experiments demonstrating our simple construct selection strategy selects constructs with good expression yields in a test of 84 potential kinase domain boundaries for Abl, and yields from a high-throughput expression screen of 96 human kinase constructs. Using a fluorescence-based thermostability assay and a fluorescent ATP-competitive inhibitor, we show that the highest-expressing kinases are folded and have well-formed ATP binding sites. We also demonstrate how the resulting expressing constructs can be used for the biophysical and biochemical study of clinical mutations by engineering a panel of 48 Src mutations and 46 Abl mutations via single-primer mutagenesis and screening the resulting library for expression yields. The wild-type kinase construct library is available publicly via Addgene, and should prove to be of high utility for experiments focused on drug discovery and the emergence of drug resistance.

Biochemistry

DNA from dust: comparative genomics of large DNA viruses in field surveillance samples

Mareks disease (MD) is a lymphoproliferative disease of chickens caused by airborne gallid herpesvirus type 2 (GaHV-2, aka MDV-1). Mature virions are formed in the feather follicle epithelium cells of infected chickens from which the virus is shed as fine particles of skin and feather debris, or poultry dust. Poultry dust is the major source of virus transmission between birds in agricultural settings. Despite both clinical and laboratory data that show increased virulence in field isolates of MDV-1 over the last 40 years, we do not yet understand the genetic basis of MDV-1 pathogenicity. Our present knowledge on genome-wide variation in the MDV-1 genome comes exclusively from laboratory-grown isolates. MDV-1 isolates tend to lose virulence with increasing passage number in vitro, raising concerns about their ability to accurately reflect virus in the field. The ability to rapidly and directly sequence field isolates of MDV-1 is critical to understanding the genetic basis of rising virulence in circulating wild strains. Here we present the first complete genomes of uncultured, field-isolated MDV-1. These five consensus genomes were derived directly from poultry dust or single chicken feather follicles without passage in cell culture. These sources represent the shed material that is transmitted to new hosts, vs. the virus produced by a point source in one animal. We developed a new procedure to extract and enrich viral DNA, while reducing host and environmental contamination. DNA was sequenced using Illumina MiSeq high-throughput approaches and processed through a recently described bioinformatics workflow for de novo assembly and curation of herpesvirus genomes. We comprehensively compared these genomes to one another and also to previously described MDV-1 genomes. The field-isolated genomes had remarkably high DNA identity when compared to one another, with few variant proteins between them. In an analysis of genetic distance, the five new field genomes grouped separately from all previously described genomes. Each consensus genome was also assessed to determine the level of polymorphisms within each sample, which revealed that MDV-1 exists in the wild as a polymorphic population. By tracking a new polymorphic locus in ICP4 over time, we found that MDV-1 genomes can evolve in short period of time. Together these approaches advance our ability to assess MDV-1 variation within and between hosts, over time, and during adaptation to changing conditions.

Microbiology

Assemblytics: a web analytics tool for the detection of assembly-based variants

SummaryAssemblytics is a web app for detecting and analyzing structural variants from a de novo genome assembly aligned to a reference genome. It incorporates a unique anchor filtering approach to increase robustness to repetitive elements, and identifies six classes of variants based on their distinct alignment signatures. Assemblytics can be applied both to comparing aberrant genomes, such as human cancers, to a reference, or to identify differences between related species. Multiple interactive visualizations enable in-depth explorations of the genomic distributions of variants.\n\nAvailability and Implementationhttp://qb.cshl.edu/assemblytics, https://github.com/marianattestad/assemblytics\n\nContact: mnattest@cshl.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics

Functional metagenomics reveals novel β-galactosidases not predictable from gene sequences

A soil metagenomic library carried in pJC8 (an IncP cosmid) was used for functional complementation for {beta}-galactosidase activity in both -Proteobacteria (Sinorhizobium meliloti) and{gamma} -Proteobacteria (Escherichia coli). One {beta}-galactosidase, encoded by overlapping clones selected in both hosts, was identified as a member of glycoside hydrolase family 2. ORFs obviously encoding possible {beta}-galactosidases were not identified in 19 other clones that were only able to complement S. meliloti. Based on low sequence similarity to known glycoside hydrolases but not {beta}-galactosidases, three ORFs were examined further. Biochemical analysis confirmed that all encoded {beta}-galactosidase activity. Bioinformatic and structural modeling implied that Lac161_ORF10 protein represented a novel enzyme family with a five-bladed propeller glycoside hydrolase domain.

Microbiology

Characterization of sterol synthesis in bacteria

Sterols are essential components of eukaryotic cells whose biosynthesis and function in eukaryotes has been studied extensively. Sterols are also recognized as the diagenetic precursors of steranes preserved in sedimentary rocks where they can function as geological proxies for eukaryotic organisms and/or aerobic metabolisms and environments. However, production of these lipids is not restricted to the eukaryotic domain as a few bacterial species also synthesize sterols. Phylogenomic studies have identified genes encoding homologs of sterol biosynthesis proteins in the genomes of several additional species, indicating that sterol production may be more widespread in the bacterial domain than previously thought. Although the occurrence of sterol synthesis genes in a genome indicates the potential for sterol production, it provides neither conclusive evidence of sterol synthesis nor information about the composition and abundance of basic and modified sterols that are actually being produced. Here, we coupled bioinformatics with lipid analyses to investigate the scope of bacterial sterol production. We identified oxidosqualene cyclase (Osc), which catalyzes the initial cyclization of oxidosqualene to the basic sterol structure, in 34 bacterial genomes from 5 phyla (Bacteroidetes, Cyanobacteria, Planctomycetes, Proteobacteria and Verrucomicrobia) and in 176 metagenomes. Our data indicate that bacterial sterol synthesis likely occurs in diverse organisms and environments and also provides evidence that there are as yet uncultured groups of bacterial sterol producers. Phylogenetic analysis of bacterial and eukaryotic Osc sequences revealed two potential lineages of the sterol pathway in bacteria indicating a complex evolutionary history of sterol synthesis in this domain. We characterized the lipids produced by Osc-containing bacteria and found that we could generally predict the ability to synthesize sterols. However, predicting the final modified sterol based on our current knowledge of bacterial sterol synthesis was difficult. Some bacteria produced demethylated and saturated sterol products even though they lacked homologs of the eukaryotic proteins required for these modifications emphasizing that several aspects of bacterial sterol synthesis are still completely unknown. It is possible that bacteria have evolved distinct proteins for catalyzing sterol modifications and this could have significant implications for our understanding of the evolutionary history of this ancient biosynthetic pathway.

Microbiology

Analytical considerations for comparative transcriptomics of wild organisms.

Comparative transcriptomics can now be conducted on organisms in natural settings, which has greatly enhanced understanding of genome-environment interactions. However, important data handling and quality control challenges remain, particularly when working with non-model species outside of a controlled laboratory environment. Here, we demonstrate the utility and potential pitfalls of comparative transcriptomics of wild organisms, with an example from three cyprinid fish species (Teleostei:Cypriniformes). We present computational solutions for processing, annotating and summarizing comparative transcriptome data for assessing genome-environment interactions across species. The resulting bioinformatics pipeline addresses the following points: (1) the potential importance of \"essential genes\", (2) the influence of microbiomes and other exogenous DNA, (3) potentially novel, species-specific genes, and (4) genomic rearrangements (e.g., whole genome duplication). Quantitative consideration of these points contributes to a firmer foundation for future comparative work across distantly related taxa for a variety of sub-disciplines, including stress and immune response, community ecology, ecotoxicology, and climate change.

Genomics

Optimizing multiplex CRISPR/Cas9-based genome editing for wheat

BackgroundCRISPR/Cas9-based genome editing holds great promise to accelerate the development of new crop varieties by providing a powerful tool to modify the genomic regions controlling major agronomic traits. To diversify the set of tools available for wheat genome engineering, we have established a tRNA-based multiplex gene editing strategy for hexaploid wheat.\n\nResultsThe functionality of the various CRISPR/Cas9 components was assessed using the transient expression in the wheat protoplasts followed by next-generation sequencing (NGS) of the targeted genomic regions. The efficiency of wheat codon-optimized Cas9 for targeted gene editing in wheat was validated. Multiple single guide RNAs (gRNAs) were evaluated for the ability to edit the homoeologous copies of four genes affecting some important agronomic traits in wheat. Low correspondence was found between the gRNA efficiency predicted bioinformatically and that assessed in the transient expression assay. A multiplex gene editing construct with several gRNA-tRNA units under the control of a single promoter for the RNA polymerase III generated indels at the targets sites with the efficiency comparable to that obtained for a single gRNA construct.\n\nConclusionsBy integrating the protoplast transformation assay with multiplexed NGS, it is possible to perform fast functional screens for a large number of gRNAs and to optimize constructs for effective editing of multiple independent targets in the wheat genome. The multiplexing capacity of the tandemly arrayed tRNA-gRNA construct is well suited for the simultaneous editing of the redundant gene copies in the allopolyploid genomes or genomic regions beneficially affecting multiple agronomic traits. A polycistronic gene construct that can be quickly assembled using the Golden Gate reaction along with the wheat codon optimized Cas9 will further expand the set of tools available for engineering the wheat genome.

Genomics