Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

SupportNet: a novel incremental learning framework through deep learning and support data

MotivationIn most biological data sets, the amount of data is regularly growing and the number of classes is continuously increasing. To deal with the new data from the new classes, one approach is to train a classification model, e.g., a deep learning model, from scratch based on both old and new data. This approach is highly computationally costly and the extracted features are likely very different from the ones extracted by the model trained on the old data alone, which leads to poor model robustness. Another approach is to fine tune the trained model from the old data on the new data. However, this approach often does not have the ability to learn new knowledge without forgetting the previously learned knowledge, which is known as the catastrophic forgetting problem. To our knowledge, this problem has not been studied in the field of bioinformatics despite its existence in many bioinformatic problems.\n\nResultsHere we propose a novel method, SupportNet, to solve the catastrophic forgetting problem efficiently and effectively. SupportNet combines the strength of deep learning and support vector machine (SVM), where SVM is used to identify the support data from the old data, which are fed to the deep learning model together with the new data for further training so that the model can review the essential information of the old data when learning the new information. Two powerful consolidation regularizers are applied to ensure the robustness of the learned model. Comprehensive experiments on various tasks, including enzyme function prediction, subcellular structure classification and breast tumor classification, show that SupportNet drastically outperforms the state-of-the-art incremental learning methods and reaches similar performance as the deep learning model trained from scratch on both old and new data.\n\nAvailabilityOur program is accessible at: https://github.com/lykaust15/SupportNet.

bioinformatics

Floating search methodology for combining classification models for site recognition in DNA sequences

Recognition of the functional sites of genes, such as translation initiation sites, donor and acceptor splice sites and stop codons, is a relevant part of many current problems in bioinformatics. Recognition of the functional sites of genes is also a fundamental step in gene structure predictions in the most powerful programs. The best approaches to this type of recognition use sophisticated classifiers, such as support vector machines. However, with the rapid accumulation of sequence data, methods for combining many sources of evidence are necessary as it is unlikely that a single classifier can solve this type of problem with the best possible performance.\n\nA major issue is that the number of possible models to combine is large and the use of all of these models is impractical. In this paper, we present a framework that is based on floating search for combining as many classifiers as needed for the recognition of any functional sites of a gene. The methodology can be used for the recognition of translation initiation sites, donor and acceptor splice sites and stop codons. Furthermore, we can combine any number of classifiers that are trained on any species. The method is also scalable to large datasets, as is shown in experiments in which the whole human genome is used. The method is also applicable to other recognition tasks.\n\nWe present experiments on the recognition of these four functional sites in the human genome, which is used as the target genome, and use another 20 species as sources of evidence. The proposed methodology shows significant improvement over state-of-the-art methods for use in a thorough evaluation process. The proposed method is also able to improve heuristic selection of species to be used as sources of evidence as the search finds the most useful datasets.\n\nAuthor summaryIn this paper we present a methodology for combining many sources of information to recognize some of the most important functional sites in a genomic sequence. The functional sites of the sequences, such as, translation start sites, translation initiation sites, acceptor and donor splice sites and stop codons, play a very relevant role in many Bioinformatics tasks. Their accurate recognition is an important task by itself and also as part of gene structure prediction programs.\n\nOur approach uses a methodology usually termed in Computer Science as \"floating search\". This is a powerful heuristics applicable when the cost of evaluating each possible solution is high. The methodology is applied to the recognition of four different functional sites in the human genome using as additional sources of evidence the annotated genomes of other twenty different species.\n\nThe results show an advantage of the proposed method and also challenge the standard assumption of using only genomes not very close and not very far from the human to improve the recognition of functional sites in the human genome.

bioinformatics

RBV: Read balance validator, a tool for prioritising copy number variations in germline conditions

BackgroundThe popularisation and decreased cost of genome resequencing has resulted in an increased use in molecular diagnostics. While there are a number of established and high quality bioinfomatic tools for identifying small genetic variants including single nucleotide variants and indels, currently there is no established standard for the detection of copy number variants (CNVs) from sequence data. The requirement for CNV detection from high throughput sequencing has resulted in the development of a large number of software packages. These tools typically utilise the sequence data characteristics: read depth, split reads, read pairs, and assembly-based techniques. However the additional source of information from read balance, defined as relative proportion of reads of each allele at each position, has been underutilised in the existing applications.\n\nResultsWe present Read Balance Validator (RBV), a bioinformatic tool which uses read balance for prioritisation and validation of putative CNVs. The software simultaneously interrogates nominated regions for the presence of deletions or multiplications, and can differentiate larger CNVs from diploid regions. Additionally, the utility of RBV to test for inheritance of CNVs is demonstrated in this report.\n\nConclusionsRBV is a CNV validation and prioritisation bioinformatic tool for both genome and exome sequencing available as a python package from https://github.com/whitneywhitford/RBV

bioinformatics

Characterizing molecular flexibility by combining lRMSD measures

The root mean square deviation (RMSD) and the least RMSD are two widely used similarity measures in structural bioinformatics. Yet, they stem from global comparisons, possibly obliterating locally conserved motifs. We correct these limitations with the so-called combined RMSD, which mixes independent lRMSD measures, each computed with its own rigid motion. The combined RMSD can be used to compare (quaternary) structures based on motifs defined from the sequence (domains, SSE), or to compare structures based on structural motifs yielded by local structural alignment methods.\n\nWe illustrate the benefits of combined RMSD over the usual RMSD on three problems, namely (i) the analysis of conformational changes based on combined RMSD of rigid structural motifs (case study: a class II fusion protein), (ii) the calculation of structural phylogenies (case study: class II fusion proteins), and (iii) the assignment of quaternary structures for hemoglobin. Using these, we argue that the combined RMSD is a tool a choice to perform positive and negative discrimination of degree of freedom, with applications to the design of move sets and collective coordinates.\n\nCombined RMSD are available within the Structural Bioinformatics Library (http://sbl.inria.fr).

bioinformatics

The Integrated Rapid Infectious Disease Analysis (IRIDA) Platform

Whole genome sequencing (WGS) is a powerful tool for public health infectious disease investigations owing to its higher resolution, greater efficiency, and cost-effectiveness over traditional genotyping methods. Implementation of WGS in routine public health microbiology laboratories is impeded by a lack of user-friendly automated and semi-automated pipelines, restrictive jurisdictional data sharing policies, and the proliferation of non-interoperable analytical and reporting systems. To address these issues, we developed the Integrated Rapid Infectious Disease Analysis (IRIDA) platform (irida.ca), a user-friendly, decentralized, open-source bioinformatics and analytical web platform to support real-time infectious disease outbreak investigations using WGS data. Instances can be independently installed on local high-performance computing infrastructure, enabling private and secure data management and analyses according to organizational policies and governance. IRIDAs data management capabilities enable secure upload, storage and sharing of all WGS data and metadata. The core platform currently includes pipelines for quality control, assembly, annotation, variant detection, phylogenetic analysis, in silico serotyping, multi-locus sequence typing, and genome distance calculation. Analysis pipeline results can be visualized within the platform through dynamic line lists and integrated phylogenomic clustering for research and discovery, and for enhancing decision-making support and hypothesis generation in epidemiological investigations. Communication and data exchange between instances are provided through customizable access controls. IRIDA complements centralized systems, empowering local analytics and visualizations for genomics-based microbial pathogen investigations. IRIDA is currently transforming the Canadian public health ecosystem and is freely available at https://github.com/phac-nml/irida and www.irida.ca.\n\nImpact StatementWhole genome sequencing (WGS) is revolutionizing infectious disease analysis and surveillance due to its cost effectiveness, utility, and improved analytical power. To date, no \"one-size-fits-all\" genomics platform has been universally adopted, owing to differences in national (and regional) health information systems, data sharing policies, computational infrastructures, lack of interoperability and prohibitive costs. The Integrated Rapid Infectious Disease Analysis (IRIDA) platform is a user-friendly, decentralized, open-source bioinformatics and analytical web platform developed to support real-time infectious disease outbreak investigations using WGS data. IRIDA empowers public health, regulatory and clinical microbiology laboratory personnel to better incorporate WGS technology into routine operations by shielding them from the computational and analytical complexities of big data genomics. IRIDA is now routinely used as part of a validated suite of tools to support outbreak investigations in Canada. While IRIDA was designed to serve the needs of the Canadian public health system, it is generally applicable to any public health and multi-jurisdictional environment. IRIDA enables localized analyses but provides mechanisms and standard outputs to enable data sharing. This approach can help overcome pervasive challenges in real-time global infectious disease surveillance, investigation and control, resulting in faster responses, and ultimately, better public health outcomes.\n\nDATA SUMMARYO_LIData used to generate some of the figures in this manuscript can be found in the NCBI BioProject PRJNA305824.\nC_LI

bioinformatics

GranatumX: A community engaging and flexible software environment for single-cell analysis

We present GranatumX, a next-generation software environment for single-cell data analysis. GranatumX is inspired by the interactive web tool Granatum. It enables biologists to access the latest single-cell bioinformatics methods in a web-based graphical environment. It also offers software developers the opportunity to rapidly promote their own tools with others in customizable pipelines. The architecture of GranatumX allows for easy inclusion of plugin modules, named Gboxes, that wrap around bioinformatics tools written in various programming languages and on various platforms. GranatumX can be run on the cloud or private servers and generate reproducible results. It is a community-engaging, flexible, and evolving software ecosystem for scRNA-Seq analysis, connecting developers with bench scientists. GranatumX is freely accessible at http://garmiregroup.org/granatumx/app.

bioinformatics

rCASC: reproducible Classification Analysis of Single Cell sequencing data

SummarySingle-cell RNA sequencing has emerged as an essential tool to investigate cellular heterogeneity, and highlighting cell sub-population specific signatures. Nowadays, dedicated and user-friendly bioinformatics workflows are required to exploit the deconvolution of single-cells transcriptome. Furthermore, there is a growing need of bioinformatics workflows granting both functional, i.e. saving information about data and analysis parameters, and computation reproducibility, i.e. storing the real image of the computation environment. Here, we present rCASC a modular RNAseq analysis workflow allowing data analysis from counts generation to cell sub-population signatures identification, granting both functional and computation reproducibility.\n\nAvailability and ImplementationrCASC is part of the reproducible bioinfomatics project. rCASC is a docker based application controlled by a R package available at https://github.com/kendomaniac/rCASC.\n\nSupplementary informationSupplementary data are available at rCASC github

bioinformatics

A targeted subgenomic approach for phylogenomics based on microfluidic PCR and high throughput sequencing

Advances in high-throughput sequencing (HTS) have allowed researchers to obtain large amounts of biological sequence information at speeds and costs unimaginable only a decade ago. Phylogenetics, and the study of evolution in general, is quickly migrating towards using HTS to generate larger and more complex molecular datasets. In this paper, we present a method that utilizes microfluidic PCR and HTS to generate large amounts of sequence data suitable for phylogenetic analyses. The approach uses a Fluidigm microfluidic PCR array and two sets of PCR primers to simultaneously amplify 48 target regions across 48 samples, incorporating sample-specific barcodes and HTS adapters (2,304 unique amplicons per microfluidic array). The final product is a pooled set of amplicons ready to be sequenced, and thus, there is no need to construct separate, costly genomic libraries for each sample. Further, we present a bioinformatics pipeline to process the raw HTS reads to either generate consensus sequences (with or without ambiguities) for every locus in every sample or--more importantly--recover the separate alleles from heterozygous target regions in each sample. This is important because it adds allelic information that is well suited for coalescent-based phylogenetic analyses that are becoming very common in conservation and evolutionary biology. To test our subgenomic method and bioinformatics pipeline, we sequenced 576 samples across 96 target regions belonging to the South American clade of the genus Bartsia L. in the plant family Orobanchaceae. After sequencing cleanup and alignment, the experiment resulted in [~]25,300bp across 486 samples for a set of 48 primer pairs targeting the plastome, and [~]13,500bp for 363 samples for a set of primers targeting regions in the nuclear genome. Finally, we constructed a combined concatenated matrix from all 96 primer combinations, resulting in a combined aligned length of [~]40,500bp for 349 samples.

Evolutionary Biology

What’s in my pot? Real-time species identification on the MinION™

Whole genome sequencing on next-generation instruments provides an unbiased way to identify the organisms present in complex metagenomic samples. However, the time-to-result can be protracted because of fixed-time sequencing runs and cumbersome bioinformatics workflows. This limits the utility of the approach in settings where rapid species identification is crucial, such as in the quality control of food-chain components, or in during an outbreak of an infectious disease. Here we present Whats in my Pot? (WIMP), a laboratory and analysis workflow in which, starting with an unprocessed sample, sequence data is generated and bacteria, viruses and fungi present in the sample are classified to subspecies and strain level in a quantitative manner, without prior knowledge of the sample composition, in approximately 3.5 hours. This workflow relies on the combination of Oxford Nanopore Technologies MinION sensing device with a real-time species identification bioinformatics application.

Genomics

Progressive lengthening of 3′ untranslated regions of mRNAs by alternative cleavage and polyadenylation in cellular senescence of mouse embryonic fibroblasts

BackgroundCellular senescence has historically been viewed as an irreversible cell cycle arrest that acts to prevent cancer. Recent discoveries demonstrated that cellular senescence also played a vital role in normal embryonic development, tissue renewal and senescence-related diseases. Alternative cleavage and polyadenylation (APA) is an important layer of post-transcriptional regulation, which has been found playing an essential role in development, activation of immune cells and cancer progression. However, the role of APA in the process of cellular senescence remains unclear.\n\nMaterials and MethodsWe applied high-throughput paired-end polyadenylation sequencing (PA-seq) and strand-specific RNA-seq sequencing technologies, combined systematic bioinformatics analyses and experimental validation to investigate APA regulation in different passages of mouse embryonic fibroblasts (MEFs) and in aortic vascular smooth muscle cells of rats (VSMCs) with different ages.\n\nResultsBased on PA-seq, we found that genes in senescent cells tended to use distal pA sites and an independent bioinformatics analysis for RNA-seq drew the same conclusion. In consistent with these global results, both the number of genes significantly preferred to use distal pAs in senescent MEFs and VSMCs were significantly higher than genes tended to use proximal pAs. Interestingly, the expression levels of genes preferred to use distal pAs in senescent MFEs and VSMCs tended to decrease, while genes with single pAs did not show such trend. More importantly, genes preferred to use distal pAs in senescent MFEs and VSMCs were both enriched in common senescence-related pathways, including ubiqutin mediated proteolysis, regulation of actin cytoskeleton, cell cycle and wnt signaling pathway. By cis-elements analyses, we found that the longer 3' UTRs of the genes tended to use distal pAs progressively can introduce more conserved binding sites of senescence-related miRNAs and RBPs. Furthermore, 375 genes with progressive 3' UTR lengthening during MEF senescence tended to use more strong and conserved polyadenylation signal (PAS) around distal pA sites and this was accompanied the observation that expression level of core factors involved in cleavage and polyadenylation complex was decreased.\n\nConclusionsOur finding that genes preferred distal pAs in senescent mouse and rat cells provide new insights for aging cells posttranscriptional gene regulation in the view of alternative polyadenylation given senescence response was thought to be a tumor suppression mechanism and more genes tended to use proximal pAs in cancer cells. In short, APA was a hidden layer of post-transcriptional gene expression regulation involved in cellular senescence.

Genomics

Computational Pan-Genomics: Status, Promises and Challenges

Many disciplines, from human genetics and oncology to plant breeding, microbiology and virology, commonly face the challenge of analyzing rapidly increasing numbers of genomes. In case of Homo sapiens, the number of sequenced genomes will approach hundreds of thousands in the next few years. Simply scaling up established bioinformatics pipelines will not be sufficient for leveraging the full potential of such rich genomic datasets. Instead, novel, qualitatively different computational methods and paradigms are needed. We will witness the rapid extension of computational pan-genomics, a new sub-area of research in computational biology. In this paper, we generalize existing definitions and understand a pan-genome as any collection of genomic sequences to be analyzed jointly or to be used as a reference. We examine already available approaches to construct and use pan-genomes, discuss the potential benefits of future technologies and methodologies, and review open challenges from the vantage point of the above-mentioned biological disciplines. As a prominent example for a computational paradigm shift, we particularly highlight the transition from the representation of reference genomes as strings to representations as graphs. We outline how this and other challenges from different application domains translate into common computational problems, point out relevant bioinformatics techniques and identify open problems in computer science. With this review, we aim to increase awareness that a joint approach to computational pan-genomics can help address many of the problems currently faced in various domains.

Genomics

A framework to interpret short tandem repeat variation in humans

Identifying regions of the genome that are depleted of mutations can reveal potentially deleterious variants. Short tandem repeats (STRs), also known as microsatellites, are among the largest contributors of de novo mutations in humans and are implicated in a variety of human disorders. However, because of the challenges STRs pose to bioinformatics tools, per-locus studies of STR mutations have been limited to highly ascertained panels of several dozen loci. Here, we harnessed bioinformatics tools and a novel analytical framework to estimate mutation parameters for each STR in the human genome by correlating STR genotypes with local sequence heterozygosity. We applied our method to obtain robust estimates of the impact of local sequence features on mutation parameters and used this to create a framework for measuring constraint at STRs by comparing observed vs. expected mutation rates. Constraint scores identified known pathogenic variants with early onset effects. Our constraint metrics will provide a valuable tool for prioritizing pathogenic STRs in medical genetics studies.

genetics

rAmpSeq: Using repetitive sequences for robust genotyping

Repetitive sequences have been used for DNA fingerprinting and genotyping for more than a quarter century. Now, with our knowledge of whole genome sequences, repetitive sequences can be used to identify polymorphisms that can be mapped and scored in a systematic manner. We have developed a simple, robust platform for designing primers, PCR amplification, and high throughput cloning that allows hundreds to thousands of markers to be scored for less than $5 per sample. Conserved regions were used to design PCR primers for amplifying thousands of middle repetitive regions of the maize (Zea mays ssp. mays) genome. Bioinformatic scans were then used to identify DNA sequence polymorphisms in the low copy intervening sequences. When used in conjunction with simple DNA preps, optimized PCR conditions, high multiplex Illumina indexing and a bioinformatic marker calling platform tailored for repetitive sequences, this methodology provides a cost effective genotyping strategy for large-scale genomic selection projects. We show detailed results from four maize primer sets that produced between 1,335-3,225 good coverage loci with 1056 that segregated appropriately in a bi-parental family. This approach could have wide applicability to breeding and conservation biology, where hundreds of thousands of samples need to be genotyped for very minimal cost.

genomics

Qinichelins, novel catecholate-hydroxamate siderophores synthesized via a multiplexed convergent biosynthesis pathway

The explosive increase in genome sequencing and the advances in bioinformatic tools have revolutionized the rationale for natural product discovery from actinomycetes. In particular, this has revealed that actinomycete genomes contain numerous orphan gene clusters that have the potential to specify many yet unknown bioactive specialized metabolites, representing a huge unexploited pool of chemical diversity. Here, we describe the discovery of a novel group of catecholate-hydroxamate siderophores termed qinichelins (2-5) from Streptomyces sp. MBT76. Correlation between the metabolite levels and the protein expression profiles identified the biosynthetic gene cluster (BGC; named qch) most likely responsible for qinichelin biosynthesis. The structure of the molecules was elucidated by bioinformatics, mass spectrometry and NMR. Synthesis of the qinichelins requires the interplay between four gene clusters, for its synthesis and for precursor supply. This biosynthetic complexity provides new insights into the challenges scientists face when applying synthetic biology approaches for natural product discovery.\n\nPride repository reviewer account details:\n\nURL: https://www.ebi.ac.uk/pride/archive/login\n\nProject accession: PXD006577\n\nUsername: reviewer35793@ebi.ac.uk\n\nPassword: 3H0iM1FK

microbiology

Translated Blast of L Polymerase as a Hit for Novel Arenaviruses Species

Many pathogenic viruses can transmit between human and animals as zoonotic viruses and cause dangerous diseases with obvious clinical signs globally. However, the world deals seriously with these viruses when the viruses infected either human or animals especially if the infection were confirmed that classified as zoonotic (Lal et al. 2005). There are many viruses distribution in many countries around the world including Ebolavirus, Marburgvirus, SARS and MERS coronaviruses, Hendra, Nipah and arenavirus haemorrhagic fever viruses were categorized as zoonotic RNA viruses that cause epidemic in some regions such as African countries (Fichet-Calvet & Rogers 2009) (Ehichioya et al. 2010). Consequently, structural bioinformatics of virus protein like L polymerase of arenaviruses was used for monitor the future outbreak that could be happens by new species of viruses. At this research, significant similarities with hemorrhagic fever viruses including arenaviruses were found on GenBank database. Translated blast (tBLASTn) available on https://blast.ncbi.nlm.nih.gov/Blast.cgi was used for searching translated nucleotide databases using a protein query of arenavirus L polymerase (McGinnis & Madden 2004). At this research, the new and archival metazoan transcriptome sequence data of the new TSA species that available on NCBI was used for identification with arenaviruses genes. Therefore, structure bioinformatics was utilized for better understanding and predication the evolution and natural history of the pools of uncharacterized virus on Genbank database that have led to emerging haemorrhagic fever in near future around the world.

microbiology

Evidence-Based Design and Evaluation of a Whole Genome Sequencing Clinical Report for the Reference Microbiology Laboratory

BackgroundMicrobial genome sequencing is now being routinely used in many clinical and public health laboratories. Understanding how to report complex genomic test results to stakeholders who may have varying familiarity with genomics - including clinicians, laboratorians, epidemiologists, and researchers - is critical to the successful and sustainable implementation of this new technology; however, there are no evidence-based guidelines for designing such a report in the pathogen genomics domain. Here, we describe an iterative, human-centered approach to creating a report template for communicating tuberculosis (TB) genomic test results.\n\nMethodsWe used Design Study Methodology - a human centered multi-stage approach drawn from the information visualization domain - to redesign an existing clinical report. We used expert consults and an online questionnaire to discover various stakeholders needs around the types of data and tasks related to TB that they encounter in their daily workflow. We also evaluated their perceptions of and familiarity with genomic data, as well as its utility at various clinical decision points. These data shaped the design of multiple prototype reports that were compared against the existing report through a second online survey, with the resulting qualitative and quantitative data informing the final, redesigned, report.\n\nResultsWe recruited 78 participants, 65 of whom were clinicians, nurses, laboratorians, researchers, and epidemiologists involved in TB diagnosis, treatment, and/or surveillance. Our first survey indicated that participants were largely enthusiastic about genomic data, with the majority agreeing on its utility for certain TB diagnosis and treatment tasks and many reporting some confidence in their ability to interpret this type of data (between 58.8% and 94.1%, depending on the specific data type). When we compared our four prototype reports against the existing design, we found that for the majority (86.7%) of design comparisons, participants preferred the alternative prototype designs over the existing version, and that both clinicians and non-clinicians expressed similar design preferences. Participants articulated clearer design preferences when asked to compare individual design elements versus entire reports. Both the quantitative and qualitative data informed the design of a revised report, which is available online as a LaTeX template.\n\nConclusionsWe show how a human-centered design approach integrating quantitative and qualitative feedback can be used to design an alternative report for representing complex microbial genomic data. We suggest experimental and design guidelines to inform future design studies in the bioinformatics and microbial genomics domains, and suggest that this type of mixed-methods study is important to facilitate the successful translation of pathogen genomics in the clinic, not only for clinical reports but also more complex bioinformatics data visualization software.

genomics

HiMAP: robust Phylogenomics from Highly Multiplexed Amplicon sequencing

High-throughput sequencing has fundamentally changed how molecular phylogenetic datasets are assembled, and phylogenomic datasets commonly contain 50-100-fold more loci than those generated using traditional Sanger-based approaches. Here, we demonstrate a new approach for building phylogenomic datasets using single tube, highly multiplexed amplicon sequencing, which we name HiMAP (Highly Multiplexed Amplicon-based Phylogenomics), and present bioinformatic pipelines for locus selection based on genomic and transcriptomic data resources and post-sequencing consensus calling and alignment. This method is inexpensive and amenable to sequencing a large number (hundreds) of taxa simultaneously, requires minimal hands-on time at the bench (<1/2 day), and data analysis can be accomplished without the need for read mapping or assembly. We demonstrate this approach by sequencing 878 amplicons in single reactions for 82 species of tephritid fruit flies across seven genera (384 individuals), including some of the most economically-important agricultural insect pests. The resulting dataset (>150,000 bp concatenated alignment) contained >40,000 phylogenetically informative characters, and although some discordance was observed between analyses, it provided unparalleled resolution of many phylogenetic relationships in this group. Most notably, we found high support for the generic status of Zeugodacus and the sister relationship between Dacus and Zeugodacus. We discuss HiMAP, with regard to its molecular and bioinformatic strengths, and the insight the resulting dataset provides into relationships of this diverse insect group.

genomics

Coupling MALDI-TOF mass spectrometry protein and specialized metabolite analyses to rapidly discriminate bacterial function

For decades, researchers have lacked the ability to rapidly correlate microbial identity with bacterial metabolism. Since specialized metabolites are critical to bacterial function and survival in the environment, we designed a data acquisition and bioinformatics technique (IDBac) that utilizes in situ matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) to analyze protein and specialized metabolite spectra of single bacterial colonies from agar plates. We demonstrated the power of our approach by discriminating between two Bacillus subtilis colonies in under 30 minutes, which differ by a single genomic mutation, solely on the basis of their differential ability to produce cyclic peptide antibiotics surfactin and plipastatin. Next, we employed our IDBac technique to detect subtle intra-species differences in the production of metal scavenging acyl-desferrioxamines in a group of eight freshwater Micromonospora isolates that share >99% sequence similarity in the 16S rRNA gene. Finally, we employed our method to simultaneously extract protein and specialized metabolite MS profiles from unidentified species of Lake Michigan sponge-associated bacteria cultivated on an agar plate. In just 3 hours, we created hierarchical protein MS groupings of 11 environmental isolates (10 MS replicates each, for a total of 110 samples) that accurately mirrored phylogenetic groupings. We further distinguished isolates within these groupings, which share nearly identical 16S rRNA gene sequence identity, based on inter- and intra-species differences in specialized metabolite production. To our knowledge, IDBac is the first attempt to couple in situ MS analyses of protein content and specialized metabolite production to allow the distinction of closely related bacterial colonies.\n\nSignificanceMass spectrometry is a powerful technique that has been used to identify bacteria via protein content, and to assess bacterial function in an environment via analysis of specialized metabolites. However, until now these analyses have operated independently, and this has resulted in the inability to rapidly connect bacterial phylogenetic identity with patterns of specialized metabolism. To bridge this gap, we designed a MALDI-TOF mass spectrometry data acquisition and bioinformatics pipeline (IDBac) to discriminate both intact protein and specialized metabolite spectra directly from bacterial cells grown on agar. To our knowledge, this is the first technique that organizes bacteria into highly similar phylogenetic groups and allows for comparison of metabolic differences of hundreds of isolates in just a few hours.

microbiology