Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

STAPLER: a simple tool for creating, managing and parallelizing common high-throughput sequencing workflows

STAPLER is a command line program intended for creating, managing and parallelizing bioinformatics workflows. Considerable emphasis has been placed on the ease of adoption and use by effortless installation, simple definition of workflows and quick-start tutorials. Custom workflows can be defined in an easy, modular way allowing the user to choose the desired input data, analysis tools and parameters with a simple parameter file. STAPLER then generates shell scripts that execute the workflow on a personal computer or in a supercomputing environment. Log files are generated to ensure that experimental results can be reproduced, and features are provided for validating run success and allowing rerunning parts of workflow if necessary. STAPLER is freely available on the web at https://github.com/tyrmi/STAPLER, implemented in Python 2 and supported on any UNIX or UNIX-like platform.

bioinformatics

A computational framework for systematic exploration of biosynthetic diversity from large-scale genomic data

Genome mining has become a key technology to explore and exploit natural product diversity through the identification and analysis of biosynthetic gene clusters (BGCs). Initially, this was performed on a single-genome basis; currently, the process is being scaled up to large-scale mining of pan-genomes of entire genera, complete strain collections and metagenomic datasets from which thousands of bacterial genomes can be extracted at once. However, no bioinformatic framework is currently available for the effective analysis of datasets of this size and complexity. Here, we provide a streamlined computational workflow, tightly integrated with antiSMASH and MIBiG, that consists of two new software tools, BiG-SCAPE and CORASON. BiG-SCAPE facilitates rapid calculation and interactive visual exploration of BGC sequence similarity networks, grouping gene clusters at multiple hierarchical levels, and includes a glocal alignment mode that accurately groups both complete and fragmented BGCs. CORASON employs a phylogenomic approach to elucidate the detailed evolutionary relationships between gene clusters by computing high-resolution multi-locus phylogenies of all BGCs within and across gene cluster families (GCFs), and allows researchers to comprehensively identify all genomic contexts in which particular biosynthetic gene cassettes are found. We validate BiG-SCAPE by correlating its GCF output to metabolomic data across 403 actinobacterial strains. Furthermore, we demonstrate the discovery potential of the platform by using CORASON to comprehensively map the phylogenetic diversity of the large detoxin/rimosamide gene cluster clan, prioritizing three new detoxin families for subsequent characterization of six new analogs using isotopic labeling and analysis of tandem mass spectrometric data.

bioinformatics

Multiple polyvalency provided by intrinsically disordered segments is a key feature of postsynaptic scaffold proteins

The human postsynaptic density is an elaborate network comprising thousands of proteins, playing a vital role in the molecular events of learning and the formation of memory. Despite our growing knowledge of specific proteins and their interactions, atomic-level details of their full three-dimensional structure and their rearrangements are mostly elusive. Advancements in structural bioinformatics enabled us to depict the characteristic features of proteins involved in different processes aiding neurotransmission. We show that postsynaptic protein-protein interactions are mediated through the delicate balance of intrinsically disordered regions and folded domains, and this duality is also imprinted in the amino acid sequence. We introduce Diversity of Potential Interactions (DPI), a structure and regulation based descriptor to assess the diversity of interactions. Our approach reveals that the postsynaptic proteome has its own characteristic features and these properties reliably discriminate them from other proteins of the human proteome. Our results suggest that postsynaptic proteins are especially susceptible to forming diverse interactions with each other, which might be key in the reorganization of the PSD in molecular processes related to learning and memory.

bioinformatics

A comprehensive analysis of the usability and archival stability of omics computational tools and resources

Developing new software tools for analysis of large-scale biological data is a key component of advancing modern biomedical research. Scientific reproduction of published findings requires running computational tools on data generated by such studies, yet little attention is presently allocated to the installability and archival stability of computational software tools. Scientific journals require data and code sharing, but none currently require authors to guarantee the continuing functionality of newly published tools. We have estimated the archival stability of computational biology software tools by performing an empirical analysis of the internet presence for 36,702 omics software resources published from 2005 to 2017. We found that almost 28% of all resources are currently not accessible through URLs published in the paper they first appeared in. Among the 98 software tools selected for our installability test, 51% were deemed \"easy to install,\" and 28% of the tools failed to be installed at all due to problems in the implementation. Moreover, for papers introducing new software, we found that the number of citations significantly increased when authors provided an easy installation process. We propose for incorporation into journal policy several practical solutions for increasing the widespread installability and archival stability of published bioinformatics software.

bioinformatics

Virulence in a Pseudomonas syringae Strain with a Small Repertoire of Predicted Effectors

Both type III effector proteins and non-ribosomal peptide toxins play important roles for Pseudomonas syringae pathogenicity in host plants, but whether and how these virulence pathways interact to promote infection remains unclear. Genomic evidence from one clade of P. syringae suggests a tradeoff between the total number of type III effector proteins and presence of syringomycin, syringopeptin, and syringolin A toxins. Here we report the complete genome sequence from P. syringae CC1557, which contains the lowest number of known type III effectors to date and has also acquired genes similar to sequences encoding syringomycin pathways from other strains. We demonstrate that this strain is pathogenic on Nicotiana benthamiana and that both the type III secretion system and a new type III effector family, hopBJ1, contribute to virulence. We further demonstrate that virulence activity of HopBJ1 is dependent on similar catalytic sites as the E. coli CNF1 toxin. Taken together, our results provide additional support for a negative correlation between type III effector repertoires and the potential to produce syringomycin-like toxins while also highlighting how genomic synteny and bioinformatics can be used to identify and characterize novel virulence proteins.

Microbiology

The Drosophila Genome Nexus: a population genomic resource of 605 Drosophila melanogaster genomes, including 197 genomes from a single ancestral range population

Hundreds of wild-derived D. melanogaster genomes have been published, but rigorous comparisons across data sets are precluded by differences in alignment methodology. The most common approach to reference-based genome assembly is a single round of alignment followed by quality filtering and variant detection. We evaluated variations and extensions of this approach, and settled on an assembly strategy that utilizes two alignment programs and incorporates both SNPs and short indels to construct an updated reference for a second round of mapping prior to final variant detection. Utilizing this approach, we reassembled published D. melanogaster population genomic data sets (previous DPGP releases and the DGRP freeze 2.0), and added unpublished genomes from several sub-Saharan populations. Most notably, we present aligned data from phase 3 of the Drosophila Population Genomics Project (DPGP3), which provides 197 genomes from a single ancestral range population of D. melanogaster (from Zambia). The large sample size, high genetic diversity, and potentially simpler demographic history of the DPGP3 sample will make this a highly valuable resource for fundamental population genetic research. The complete set of assemblies described here, termed the Drosophila Genome Nexus, presently comprises 605 consistently aligned genomes, and is publicly available in multiple formats with supporting documentation and bioinformatic tools. This resource will greatly facilitate population genomic analysis in this model species by reducing the methodological differences between data sets.

Genomics

Phinch: An interactive, exploratory data visualization framework for –Omic datasets

Using environmental sequencing approaches, we now have the ability to deeply characterize biodiversity and biogeographic patterns in understudied, uncultured microbial taxa (investigations of bacteria, archaea, and microscopic eukaryotes using 454/Illumina sequencing platforms). However, the sheer volume of data produced from these new technologies requires fundamentally different approaches and new paradigms for effective data analysis. Scientific visualization represents an innovative method towards tackling the current bottleneck in bioinformatic workflows. In addition to giving researchers a unique approach for exploring large datasets, it stands to empower biologists with the ability to conduct powerful analyses without requiring a deep level of computational knowledge. Here we present Phinch, an interactive, browser-based visualization framework that can be used to explore and analyze biological patterns in high-throughput -Omic datasets. This project takes advantage of standard file formats from computational pipelines in order to bridge the gap between biological software (e.g. microbial ecology pipelines) and existing data visualization capabilities (harnessing the flexibility and scalability of technologies such as HTML5).

Genomics

Microbial community composition and diversity via 16S rRNA gene amplicons: evaluating the Illumina platform

As new sequencing technologies become cheaper and older ones disappear, laboratories switch vendors and platforms. Validating the new setups is a crucial part of conducting rigorous scientific research. Here we report on the reliability and biases of performing bacterial 16S rRNA gene amplicon paired-end sequencing on the MiSeq Illumina platform. We designed a protocol using 50 barcode pairs to run samples in parallel and coded a pipeline to process the data. Sequencing the same sediment sample in 248 replicates as well as 70 samples from alkaline soda lakes, we evaluated the performance of the method with regards to estimates of alpha and beta diversity.\n\nUsing different purification and DNA quantification procedures we always found up to 5-fold differences in the yield of sequences between individually barcodes samples. Using either a one-step or a two-step PCR preparation resulted in significantly different estimates in both alpha and beta diversity. Comparing with a previous method based on 454 pyrosequencing, we found that our Illumina protocol performed in a similar manner - with the exception for evenness estimates where correspondence between the methods was low.\n\nWe further quantified the data loss at every processing step eventually accumulating to 50% of the raw reads. When evaluating different OTU clustering methods, we observed a stark contrast between the results of QIIME with default settings and the more recent UPARSE algorithm when it comes to the number of OTUs generated. Still, overall trends in alpha and beta diversity corresponded highly using both clustering methods.\n\nOur procedure performed well considering the precisions of alpha and beta diversity estimates, with insignificant effects of individual barcodes. Comparative analyses suggest that 454 and Illumina sequence data can be combined if the same PCR protocol and bioinformatic workflows are used for describing patterns in richness, beta-diversity and taxonomic composition.

Molecular Biology

The genetic architecture of local adaptation I: The genomic landscape of foxtail pine (Pinus balfouriana Grev. & Balf.) as revealed from a high-density linkage map

Explaining the origin and evolutionary dynamics of the genetic architecture of adaptation is a major research goal of evolutionary genetics. Despite controversy surrounding success of the attempts to accomplish this goal, a full understanding of adaptive genetic variation necessitates knowledge about the genomic location and patterns of dispersion for the genetic components affecting fitness-related phenotypic traits. Even with advances in next generation sequencing technologies, the production of full genome sequences for non-model species is often cost prohibitive, especially for tree species such as pines where genome size often exceeds 20 to 30 Gbp. We address this need by constructing a dense linkage map for foxtail pine (Pinus balfouriana Grev. & Balf.), with the ultimate goal of uncovering and explaining the origin and evolutionary dynamics of adaptive genetic variation in natural populations of this forest tree species. We utilized megagametophyte arrays (n = 76-95 megagametophytes/tree) from four maternal trees in combination with double-digestion restriction site associated DNA sequencing (ddRADseq) to produce a consensus linkage map covering 98.58% of the foxtail pine genome, which was estimated to be 1276 cM in length (95% CI: 1174 cM to 1378 cM). A novel bioinformatic approach using iterative rounds of marker ordering and imputation was employed to produce single-tree linkage maps (507-17 066 contigs/map; lengths: 1037.40- 1572.80 cM). These linkage maps were collinear across maternal trees, with highly correlated marker orderings (Spearmans{rho} > 0.95). A consensus linkage map derived from these single-tree linkage maps contained 12 linkage groups along which 20 655 contigs were non-randomly distributed across 901 unique positions (n = 23 contigs/position), with an average spacing of 1.34 cM between adjacent positions. Of the 20 655 contigs positioned on the consensus linkage map, 5627 had enough sequence similarity to contigs contained within the most recent build of the loblolly pine (P. taeda L.) genome to identify them as putative homologs containing both genic and non-genic loci. Importantly, all 901 unique positions on the consensus linkage map had at least one contig with putative homology to loblolly pine. When combined with the other biological signals that predominate in our data (e.g., correlations of recombination fractions across single trees), we show that dense linkage maps for non-model forest tree species can be efficiently constructed using next generation sequencing technologies. We subsequently discuss the usefulness of these maps as community-wide resources and as tools with which to test hypotheses about the genetic architecture of adaptation.

Evolutionary Biology

What to compare and how: comparative transcriptomics for Evo-Devo

Evolutionary developmental biology has grown historically from the capacity to relate patterns of evolution in anatomy to patterns of evolution of expression of specific genes, whether between very distantly related species, or very closely related species or populations. Scaling up such studies by taking advantage of modern transcriptomics brings promising improvements, allowing us to estimate the overall impact and molecular mechanisms of convergence, constraint or innovation in anatomy and development. But it also presents major challenges, including the computational definitions of anatomical homology and of organ function, the criteria for the comparison of developmental stages, the annotation of transcriptomics data to proper anatomical and developmental terms, and the statistical methods to compare transcriptomic data between species to highlight significant conservation or changes. In this article, we review these challenges, and the ongoing efforts to address them, which are emerging from bioinformatics work on ontologies, evolutionary statistics, and data curation, with a focus on their implementation in the context of the development of our database Bgee (http://bgee.org).

Evolutionary Biology

Widespread polycistronic transcripts in mushroom-forming fungi revealed by single-molecule long-read mRNA sequencing

Genes in prokaryotic genomes are often arranged into clusters and co-transcribed into polycistronic RNAs. Isolated examples of polycistronic RNAs were also reported in some eukaryotes but their presence was generally considered rare. Here we developed a long-read sequencing strategy to identify polycistronic transcripts in several mushroom forming fungal species including Plicaturopsis crispa, Phanerochaete chrysosporium, Trametes versicolor and Gloeophyllum trabeum1. We found genome-wide prevalence of polycistronic transcription in these Agaricomycetes, and it involves up to 8% of the transcribed genes. Unlike polycistronic mRNAs in prokaryotes, these co-transcribed genes are also independently transcribed, and upstream transcription may interfere downstream transcription. Further comparative genomic analysis indicates that polycistronic transcription is likely a feature unique to these fungi. In addition, we also systematically demonstrated that short-read assembly is insufficient for mRNA isoform discovery, especially for isoform-rich loci. In summary, our study revealed, for the first time, the genome prevalence of polycistronic transcription in a subset of fungi. Futhermore, our long-read sequencing approach combined with bioinformatics pipeline is a generic powerful tool for precise characterization of complex transcriptomes.

Genomics

Pollen feeding proteomics: salivary proteins of the passion flower butterfly, Heliconius melpomene

While most adult Lepidoptera use flower nectar as their primary food source, butterflies in the genus Heliconius have evolved the novel ability to acquire amino acids from consuming pollen. Heliconius butterflies collect pollen on their proboscis, moisten the pollen with saliva, and use a combination of mechanical disruption and chemical degradation to release free amino acids that are subsequently re-ingested in the saliva. Little is known about the molecular mechanisms of this complex pollen feeding adaptation. Here we report an initial shotgun proteomic analysis of saliva from Heliconius melpomene. Results from liquid-chromatography tandem mass-spectrometry confidently identified 31 salivary proteins, most of which contained predicted signal peptides, consistent with extracellular secretion. Further bioinformatic annotation of these salivary proteins indicated the presence of four distinct functional classes: proteolysis (10 proteins), carbohydrate hydrolysis (5), immunity (6), and \"housekeeping\"(4). Additionally, six proteins could not be functionally annotated beyond containing a predicted signal sequence. The presence of several salivary proteases is consistent with previous demonstrations that Heliconius saliva has proteolytic capacity. It is likely these proteins play a key role in generating free amino acids during pollen digestion. The identification of proteins functioning in carbohydrate hydrolysis is consistent with Heliconius butterflies consuming nectar, like other lepidopterans, as well as pollen. Immune-related proteins in saliva are also expected, given that ingestion of pathogens is a very likely route to infection. The few \"housekeeping\" proteins are likely not true salivary proteins and reflect a modest level of contamination that occurred during saliva collection. Among the unannotated proteins were two sets of paralogs, each seemingly the result of a relatively recent tandem duplication. These results offer a first glimpse into the molecular foundation of Heliconius pollen feeding and provide a substantial advance towards comprehensively understanding this striking evolutionary novelty.

Molecular Biology

An homology- and coevolution-consistent structural model of bacterial copper-tolerance protein CopM supports function as a ‘metal sponge’ and suggests regions for metal-dependent interactions with other proteins

Copper is essential for life but toxic, therefore all organisms control tightly its intracellular abundance. Bacteria have indeed whole operons devoted to copper resistance, with genes that code for efflux pumps, oxidases, etc. Recently, the CopM protein of the CopMRS operon was described as a novel important element for copper tolerance in Synechocystis. This protein consists of a domain of unknown function, and was proposed to act as a periplasmic/extracellular copper binder. This work describes a bioinformatic study of CopM including structural models based on homology modeling and on residue coevolution, to help expand on its recent biochemical characterization. The protein is predicted to be periplasmic but membrane-anchored, not secreted. Two disordered regions are predicted, both possibly involved in protein-protein interactions. The 3D models disclose a 4-helix bundle with several potential copper-binding sites, most of them largely buried inside the bundle lumen. Some of the predicted copper-binding sites involve residues from the disordered regions, suggesting they could gain structure upon copper binding and thus possibly modulate the interactions they mediate. All models are provided as PDB files in the Supporting Information and can be visualized online at http://lucianoabriata.altervista.org/modelshome.html Note (January 2017): Recent X-ray structures of apo, copper- and silver-bound CopM are < 3[A] RMSD away from the models, and reveal metal-dependent structural flexibility (Zhao et al Acta Crystallogr D Struct Biol. 2016)

Biochemistry

A variant in TAF1 is associated with a new syndrome with severe intellectual disability and characteristic dysmorphic features

We describe the discovery of a new genetic syndrome, RykDax syndrome, driven by a whole genome sequencing (WGS) study of one family from Utah with two affected male brothers, presenting with severe intellectual disability (ID), a characteristic intergluteal crease, and very distinctive facial features including a broad, upturned nose, sagging cheeks, downward sloping palpebral fissures, prominent periorbital ridges, deep-set eyes, relative hypertelorism, thin upper lip, a high-arched palate, prominent ears with thickened helices, and a pointed chin. This Caucasian family was recruited from Utah, USA. Illumina-based WGS was performed on 10 members of this family, with additional Complete Genomics-based WGS performed on the nuclear portion of the family (mother, father and the two affected males). Using WGS datasets from 10 members of this family, we can increase the reliability of the biological inferences with an integrative bioinformatic pipeline. In combination with insights from clinical evaluations and medical diagnostic analyses, these DNA sequencing data were used in the study of three plausible genetic disease models that might uncover genetic contribution to the syndrome. We found a 2 to 5-fold difference in the number of variants detected as being relevant for various disease models when using different sets of sequencing data and analysis pipelines. We de-rived greater accuracy when more pipelines were used in conjunction with data encompassing a larger portion of the family, with the number of putative de-novo mutations being reduced by 80%, due to false negative calls in the parents. The boys carry a maternally inherited mis-sense variant in a X-chromosomal gene TAF1, which we consider as disease relevant. TAF1 is the largest subunit of the general transcription factor IID (TFIID) multi-protein complex, and our results implicate mutations in TAF1 as playing a critical role in the development of this new intellectual disability syndrome.

Genetics

Changes in postural syntax characterize sensory modulation and natural variation of C. elegans locomotion

Locomotion is driven by shape changes coordinated by the nervous system through time; thus, enumerating an animals complete repertoire of shape transitions would provide a basis for a comprehensive understanding of locomotor behaviour. Here we introduce a discrete representation of behaviour in the nematode C. elegans. At each point in time, the worms posture is approximated by its closest matching template from a set of 90 postures and locomotion is represented as sequences of postures. The frequency distribution of postural sequences is heavy-tailed with a core of frequent behaviours and a much larger set of rarely used behaviours. Responses to optogenetic and environmental stimuli can be quantified as changes in postural syntax: worms show different preferences for different sequences of postures drawn from the same set of templates. A discrete representation of behaviour will enable the use of methods developed for other kinds of discrete data in bioinformatics and language processing to be harnessed for the study of behaviour.\n\nAuthor SummaryTechnology for recording neural activity is advancing rapidly and whole-brain imaging with single neuron resolution has already been demonstrated for smaller animals. To interpret such complex neural recordings, we need comprehensive characterizations of behaviour, which is the principal output of the brain. Animal tracking can increasingly be performed automatically but an outstanding challenge is finding ways to represent these behavioural data. We have focused on the movement of the nematode worm C. elegans to develop a quantitative representation of behaviour as a series of distinct postures. Each posture is analogous to a word in a language and so we can directly count the number of phrases that makes up the C. elegans behavioural repertoire. C. elegans has a very small nervous system but we find that its behavioural repertoire is still complex. As with human languages, there is a large number of possible phrases, but most are rarely used. When comparing different populations of worms or worms in different environments, we find that the difference between their behaviour is due to a subset of their entire repertoire. In the language analogy, these would correspond to idiomatic phrases that distinguish groups of speakers. A quantitative understanding of the nature of behavioural variation will inform research on the function and evolution of neural circuits.

Animal Behavior and Cognition

Building Genomic Analysis Pipelines in a Hackathon Setting with Bioinformatician Teams: DNA-seq, Epigenomics, Metagenomics and RNA-seq

We assembled teams of genomics professionals to assess whether we could rapidly develop pipelines to answer biological questions commonly asked by biologists and others new to bioinformatics by facilitating analysis of high-throughput sequencing data. In January 2015, teams were assembled on the National Institutes of Health (NIH) campus to address questions in the DNA-seq, epigenomics, metagenomics and RNA-seq subfields of genomics. The only two rules for this hackathon were that either the data used were housed at the National Center for Biotechnology Information (NCBI) or would be submitted there by a participant in the next six months, and that all software going into the pipeline was open-source or open-use. Questions proposed by organizers, as well as suggested tools and approaches, were distributed to participants a few days before the event and were refined during the event. Pipelines were published on GitHub, a web service providing publicly available, free-usage tiers for collaborative software development (https://github.com/features/). The code was published at https://github.com/DCGenomics/ with separate repositories for each team, starting with hackathon_v001.

Genomics

Shifts in diversification rates linked to biogeographic movement into new areas, an example of disparate continental distributions and a recent radiation in the Andes

AcknowledgementsWe would like to thank S. Mathews for kindly sharing genomic DNA for some taxa used in this study. J. Sullivan, L. Harmon, E. Roalson, J. Beaulieu, B. Moore, N. Nurk, M. Pennell, T. Peterson, and two anonymous reviewers for helpful suggestions or comments on the manuscript. J. Beaulieu, L. Harmon, C. Blair and the University of Idaho Institute for Bioinformatics and Evolutionary Studies (NIH/NCRR P20RR16448 and P20RR016454) for computational aid. Funding for this work was provided by NSF DEB-1210895 to DCT for SUC, NSF DEB-1253463 to DCT, and Graduate Student Research Grants to SUC from the Botanical Society of America (BSA), the Society of Systematic Biologists (SSB), the American Society of Plant Taxonomists (ASPT), and the University of Idaho Stillinger Herbarium Expedition Funds.\n\nPremise of the studyClade specific bursts in diversification are often associated with the evolution of key innovations. However, in groups with no obvious morphological innovations, observed upticks in diversification rates have also been attributed to the colonization of a new geographic environment. In this study, we explore the systematics, diversification dynamics, and historical biogeography of the plant clade Rhinantheae in the Orobanchaceae, with a special focus on the Andean clade of the genus Bartsia L..\n\nMethodsWe sampled taxa from across Rhinantheae, including a representative sample of Andean Bartsia species. Using standard phylogenetic methods, we reconstructed evolutionary relationships, inferred divergence times among the clades of Rhinantheae, elucidated their biogeographic history, and investigated diversification dynamics.\n\nKey resultsWe confirmed that the South American Bartsia species form a highly supported monophyletic group. The median crown age of Rhinantheae was determined to be ca. 30 Ma, and Europe played an important role in the biogeographic history of the lineages. South America was first reconstructed in the biogeographic analyses around 9 Ma, and with a median age of 2.59 Ma, this clade shows a significant uptick in diversification.\n\nConclusionsIncreased net diversification of the South American clade corresponds with biogeographic movement into the New World. This happened at a time when the Andes were reaching the necessary elevation to host an alpine environment. Although a specific route could not be identified with certainty, we provide plausible hypotheses to how the group colonized the New World.

Evolutionary Biology

Rapid metagenomic identification of viral pathogens in clinical samples by real-time nanopore sequencing analysis

We report unbiased metagenomic detection of chikungunya virus (CHIKV), Ebola virus (EBOV), and hepatitis C virus (HCV) from four human blood samples by MinION nanopore sequencing coupled to a newly developed, web-based pipeline for real-time bioinformatics analysis on a computational server or laptop (MetaPORE). At titers ranging from 107-108 copies per milliliter, reads to EBOV from two patients with acute hemorrhagic fever and CHIKV from an asymptomatic blood donor were detected within 4 to 10 minutes of data acquisition, while lower titer HCV virus (1x105 copies per milliliter) was detected within 40 minutes. Analysis of mapped nanopore reads alone, despite an average individual error rate of 24% [range 8-49%], permitted identification of the correct viral strain in all 4 isolates, and 90% of the genome of CHIKV was recovered with >98% accuracy. Using nanopore sequencing, metagenomic detection of viral pathogens directly from clinical samples was performed within an unprecedented <6 hours sample-to-answer turnaround time and in a timeframe amenable for actionable clinical and public health diagnostics.

Genomics