Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Scalable Design of Paired CRISPR Guide RNAs for Genomic Deletion

Using CRISPR/Cas9, diverse genomic elements may be studied in their endogenous context. Pairs of single guide RNAs (sgRNAs) are used to delete regulatory elements and small RNA genes, while longer RNAs can be silenced through promoter deletion. We here present CRISPETa, a bioinformatic pipeline for flexible and scalable paired sgRNA design based on an empirical scoring model. Multiple sgRNA pairs are returned for each target. Any number of targets can be analyzed in parallel, making CRISPETa equally appropriate for studies of individual elements, or complex library screens. Fast run-times are achieved using a precomputed off-target database. sgRNA pair designs are output in a convenient format for visualisation and oligonucleotide ordering. We present a series of pre-designed, high-coverage library designs for entire classes of non-coding elements in human, mouse, zebrafish, Drosophila and C. elegans. Using an improved version of the DECKO deletion vector, together with a quantitative deletion assay, we test CRISPETa designs by deleting an enhancer and exonic fragment of the MALAT1 oncogene. These achieve efficiencies of [≥]50%, resulting in production of mutant RNA. CRISPETa will be useful for researchers seeking to harness CRISPR for targeted genomic deletion, in a variety of model organisms, from single-target to high-throughput scales.

Genomics

Suitability of different mapping algorithms for genome-wide polymorphism scans with Pool-Seq data

The cost-effectiveness of sequencing pools of individuals (Pool-Seq) provides the basis for the popularity and wide-spread use of this method for many research questions, ranging from unravelling the genetic basis of complex traits to the clonal evolution of cancer cells. Because the accuracy of Pool-Seq could be affected by many potential sources of error, several studies determined, for example, the influence of the sequencing technology, the library preparation protocol, and mapping parameters. Nevertheless, the impact of the mapping tools has not yet been evaluated. Using simulated and real Pool-Seq data, we demonstrate a substantial impact of the mapping tools leading to characteristic false positives in genome-wide scans. The problem of false positives was particularly pronounced when data with different read lengths and insert sizes were compared. Out of 14 evaluated algorithms novoalign, bwa mem and clc4 are most suitable for mapping Pool-Seq data. Nevertheless, no single algorithm is sufficient for avoiding all false positives. We show that the intersection of the results of two mapping algorithms provides a simple, yet effective strategy to eliminate false positives. We propose that the implementation of a consistent Pool-seq bioinformatics pipeline building on the recommendations of this study can substantially increase the reliability of Pool-Seq results, in particular when libraries generated with different protocols are being compared.

Genomics

Phages rarely encode antibiotic resistance genes: a cautionary tale for virome analysis

Antibiotic resistance genes (ARG) are pervasive in gut microbiota, but it remains unclear how often ARG are transferred, particularly to pathogens. Traditionally, ARG spread is attributed to horizontal transfer mediated either by DNA transformation, bacterial conjugation or generalized transduction. However, recent viral metagenome (virome) analyses suggest that ARG are frequently carried by phages, which is inconsistent with the traditional view that phage genomes rarely encode ARG. Here we used exploratory and conservative bioinformatic strategies found in the literature to detect ARG in phage genomes, and experimentally assessed a subset of ARG predicted using exploratory thresholds. ARG abundances in 1,181 phage genomes were vastly over-estimated using exploratory thresholds (421 predicted vs 2 known), due to low similarities and matches to protein unrelated to antibiotic resistance. Consistent with this, 4 ARG predicted using exploratory thresholds were experimentally evaluated and failed to confer antibiotic resistance in Escherichia coli. Re-analysis of available human-or mouse-associated viromes for ARG and their genomic context suggested that bona fide ARG attributed to phages in viromes were previously over-estimated. These findings provide guidance for documentation of ARG in viromes, and re-assert that ARG are rarely encoded in phages.

Ecology

Comparing the Statistical Fate of Paralogous and Orthologous Sequences

Since several decades, sequence alignment is a widely used tool in bioinformatics. For instance, finding homologous sequences with known function in large databases is used to get insight into the function of non-annotated genomic regions. Very efficient tools, like BLAST have been developed to identify and rank possible homologous sequences. To estimate the significance of the homology, the ranking of alignment scores takes a background model for random sequences into account. Using this model one can estimate the probability to find two exactly matching subsequences by chance in two unrelated sequences. The corresponding probability for two homologous sequences is much higher allowing to identify them. Here we focus on the distribution of lengths of exact sequence matches in protein coding regions pairs of evolutionary distant genomes. We show that this distribution exhibits a power-law tail with exponent = --5. Developing a simple model of sequence evolution by substitutions and segmental duplications, we show analytically that paralogous and orthologous gene pairs contribute differently to this distribution. Our model explains the differences observed in the comparison of coding and non-coding parts of genomes, thus providing with a better understanding of statistical properties of genomic sequences and their evolution.

Evolutionary Biology

Steady at the wheel: conservative sex and the benefits of bacterial transformation

Many bacteria are highly sexual, but the reasons for their promiscuity remain obscure. Did bacterial sex evolve to maximize diversity and facilitate adaptation in a changing world, or does it instead help to retain the bacterial functions that work right now? In other words, is bacterial sex innovative or conservative? Our aim in this review is to integrate experimental, bioinformatic and theoretical studies to critically evaluate these alternatives, with a main focus on natural genetic transformation, the bacterial equivalent of eukaryotic sexual reproduction. First, we provide a general overview of several hypotheses that have been put forward to explain the evolution of transformation. Next, we synthesize a large body of evidence highlighting the numerous passive and active barriers to transformation that have evolved to protect bacteria from foreign DNA, thereby increasing the likelihood that transformation takes place among clonemates. Our critical review of the existing literature provides support for the view that bacterial transformation is maintained as a means of genomic conservation that provides direct benefits to both individual bacterial cells and to transformable bacterial populations. We examine the generality of this view across bacteria and contrast this explanation with the different evolutionary roles proposed to maintain sex in eukaryotes.

Evolutionary Biology

Persistent activation of interlinked Th2-airway epithelial gene networks in sputum-derived cells from aeroallergen-sensitized symptomatic atopic asthmatics

RationaleAtopic asthma is a persistent disease characterized by intermittent wheeze and progressive loss of lung function. The disease is thought to be driven primarily by chronic aeroallergen-induced Th2-associated airways inflammation. However, the vast majority of atopics do not develop asthma-related wheeze, despite ongoing exposure to aeroallergens to which they are strongly sensitized, indicating that additional pathogenic mechanism(s) operate in conjunction with Th2 immunity to drive asthma pathogenesis.\n\nObjectivesEmploy systems level analyses to identify inflammation-associated gene networks operative at baseline in sputum-derived RNA from house dust mite-sensitized (HDMs) subjects with/without wheezing history; identify networks characteristic of the ongoing asthmatic state. All subjects resided in the constitutively-HDMhigh Perth environment.\n\nMethodsGenome wide expression profiling by RNASeq followed by gene coexpression network analysis.\n\nMeasurements/ResultsHDMs-nonwheezers displayed baseline gene expression in sputum including IL-5, IL-13 and CCL17. HDMs-wheezers showed equivalent expression of these classical Th2-effector genes but their overall baseline sputum signatures were more complex, comprising hundreds of Th2-associated and epithelial-associated genes, networked into two separate coexpression modules. The first module was connected by the hubs EGFR, ERBB2, CDH1 and IL-13. The second module was associated with CDHR3, and contained genes that control mucociliary clearance.\n\nConclusionsOur findings provide new insight into the inflammatory mechanisms operative at baseline in the airway mucosal microenvironment in atopic asthmatics undergoing natural perennial aeroallergen exposure. The molecular mechanism(s) that determine susceptibility to asthma amongst these subjects involve interactions between Th2-and epithelial function-associated genes within a complex co-expression network, which is not operative in equivalently sensitized/exposed atopic non-asthmatics.\n\nFundingThis study was funded by the Asthma Foundation WA, the Department of Health WA, and the NHMRC. AB is funded by a BrightSpark Foundation McCusker Fellowship. GLH is a NHMRC Fellow. AG is supported by the McCusker Charitable Foundation Bioinformatics Centre. ACJ is a recipient of an Australian Postgraduate Award and a Top-Up Award from the University of Western Australia.

Immunology

Pherotype polymorphism in Streptococcus pneumoniae and its effects on population structure and recombination

Natural transformation in the Gram-positive pathogen Streptococcus pneumoniae occurs when cells become \"competent\", a state that is induced in response to high extracellular concentrations of a secreted peptide signal called CSP (Competence Stimulating Peptide) encoded by the comC locus. Two main CSP signal types (pherotypes) are known to dominate the pherotype diversity across strains. Using thousands of fully sequenced pneumococcal genomes, we confirm that pneumococcal populations are highly genetically structured and that there is significant variation among diverged populations in pherotype frequencies; most carry only a single pherotype. Moreover, we find that the relative frequencies of the two dominant pherotypes significantly vary within a small range across geographical sites. It has been variously proposed that pherotypes either promote genetic exchange among cells expressing the same pherotype, or conversely that they promote recombination between strains bearing different pherotypes. We distinguish these hypotheses using a bioinformatics approach by estimating recombination frequencies within and between pherotypes across 4,089 full genomes. Despite underlying population structure, we observe extensive recombination between populations; additionally, we found significantly higher rates of genetic exchange between strains expressing different pherotypes than among isolates carrying the same pherotype. Our results indicate that pherotypes do not restrict, and marginally facilitate, recombination between strains. Furthermore, our results suggest that the CSP balanced polymorphism does not causally underlie population differentiation. Therefore, when strains carrying different pherotypes encounter one another during co-colonization, genetic exchange can freely occur.

Evolutionary Biology

Evaluating Mendelian nephrotic syndrome genes for evidence of risk alleles or oligogenicity that explain heritability

BackgroundMore than 30 genes can harbor rare exonic variants sufficient to cause nephrotic syndrome (NS), and the number of genes implicated in monogenic NS continues to grow. However, outside the first year of life, the majority of affected patients, particularly in ancestrally mixed populations, do not have a known monogenic form of NS. Even in those children classified with a monogenic form of NS, there is phenotypic heterogeneity. Thus, we have only discovered a fraction of the heritability of NS - the underlying genetic factors contributing to phenotypic variation. Part of the \"missing heritability\" for NS has been posited to be explained by patients harboring coding variants across one or more previously implicated NS genes, insufficient to cause NS in a classical Mendelian manner, but that nonetheless impact protein function enough to cause disease. However, systematic evaluation in patients with NS for rare or low-frequency risk alleles within single genes, or in combination across genes (\"oligogenicity\"), has not been reported.\n\nObjectiveTo determine whether, as compared to a reference population, patients with NS have either a significantly increased burden of protein-altering variants (\"risk-alleles\"), or unique combination of them (\"oligogenicity\"), in a set of 21 genes implicated in Mendelian forms of NS.\n\nMethodsIn 303 patients with NS enrolled in the Nephrotic Syndrome Study Network (NEPTUNE), we performed targeted amplification paired with next-generation sequencing of 21 genes implicated in monogenic NS. We created a high-quality variant call set and compared it to a variant call set of the same genes in a reference population composed of 2535 individuals from Phase 3 of 1000 Genomes Project. We created both a \"stringent\" and \"relaxed\" pathogenicity filtering pipeline, applied them to both cohorts, and computed the (1) burden of variants in the entire gene set per cohort, (2) burden of variants in the entire gene set per individual, (3) burden of variants within a single gene per cohort, and (4) unique combinations of variants across two or more genes per cohort.\n\nResultsWith few exceptions when using the relaxed filter, and which are likely the result of confounding by population stratification, NS patients did not have significantly increased burden of variants in Mendelian NS genes in comparison to a reference cohort, nor was there any evidence of oligogenicity. This was true when using both the relaxed and stringent variant pathogenicity filter.\n\nConclusionIn our study, the burden or particular combinations of low-frequency or rare protein altering variants in previously implicated Mendelian NS genes cohort does not significantly differ between North American patients with NS and a reference population. Studies in larger independent cohorts or meta-analyses are needed to assess generalizability of our discoveries and also address whether there is in fact small but significant enrichment of risk alleles or oligogenicity in NS cases undetectable with this current sample size. It is still possible that rare protein altering variants in these genes, insufficient to cause Mendelian disease, still contribute to NS as risk alleles and/or via oligogenicity. However, we suggest that more accurate bioinformatic analyses and the incorporation of functional assays would be necessary to identify bona fide instances of this form of genetic architecture as a contributor to the heritability of NS.

Genetics

Nuclear pore-like structures in a compartmentalized bacterium

Planctomycetes are distinguished from other Bacteria by compartmentalization of cells via internal membranes, interpretation of which has been subject to recent debate regarding potential relations to Gram-negative cell structure. In our interpretation of the available data, the planctomycete Gemmata obscuriglobus contains a nuclear body compartment, and thus possesses a type of cell organization with parallels to the eukaryote nucleus. Here we show that pore-like structures occur in internal membranes of G.obscuriglobus and that they have elements structurally similar to eukaryote nuclear pores, including a basket, ring-spoke structure, and eight-fold rotational symmetry. Bioinformatic analysis of proteomic data reveals that some of the G. obscuriglobus proteins associated with pore-containing membranes possess structural domains found in eukaryote nuclear pore complexes. Moreover, immuno-gold labelling demonstrates localization of one such protein, containing a {beta}-propeller domain, specifically to the G. obscuriglobus pore-like structures. Finding bacterial pores within internal cell membranes and with structural similarities to eukaryote nuclear pore complexes raises the dual possibilities of either hitherto undetected homology or stunning evolutionary convergence.

Microbiology

SiLiCO: A Simulator of Long Read Sequencing in PacBio and Oxford Nanopore

SummaryLong read sequencing platforms, which include the widely used Pacific Biosciences (PacBio) platform and the emerging Oxford Nanopore platform, aim to produce sequence fragments in excess of 15-20 kilobases, and have proved advantageous in the identification of structural variants and easing genome assembly. However, long read sequencing remains relatively expensive and error prone, and failed sequencing runs represent a significant problem for genomics core facilities. To quantitatively assess the underlying mechanics of sequencing failure, it is essential to have highly reproducible and controllable reference data sets to which sequencing results can be compared. Here, we present SiLiCO, the first in silico simulation tool to generate standardized sequencing results from both of the leading long read sequencing platforms.\n\nAvailabilitySiLiCO is an open source package written in Python. It is freely available at https://www.github.com/ethanagbaker/SiLiCO under the GNU GPL 3.0 license.\n\nContact \n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics

Nanopore DNA Sequencing and Genome Assembly on the International Space Station

The emergence of nanopore-based sequencers greatly expands the reach of sequencing into low-resource field environments, enabling in situ molecular analysis. In this work, we evaluated the performance of the MinION DNA sequencer (Oxford Nanopore Technologies) in-flight on the International Space Station (ISS), and benchmarked its performance off-Earth against the MinION, Illumina MiSeq, and PacBio RS II sequencing platforms in terrestrial laboratories. Samples contained mixtures of genomic DNA extracted from lambda bacteriophage, Escherichia coli (strain K12) and Mus musculus (BALB/c). The in-flight sequencing experiments generated more than 80,000 total reads with mean 2D accuracies of 85 - 90%, mean 1D accuracies of 75 - 80%, and median read lengths of approximately 6,000 bases. We were able to construct directed assemblies of the ~4.7 Mb E. coli genome, ~48.5 kb lambda genome, and a representative M. musculus sequence (the ~16.3 kb mitochondrial genome), at 100%, 100%, and 96.7% pairwise identity, respectively, and de novo assemblies of the lambda and E. coli genomes generated solely from nanopore reads yielded 100% and 99.8% genome coverage, respectively, at 100% and 98.5% pairwise identity. Across all surveyed metrics (base quality, throughput, stays/base, skips/base), no observable decrease in MinION performance was observed while sequencing DNA in space. Simulated runs of in-flight nanopore data using an automated bioinformatic pipeline and cloud or laptop based genomic assembly demonstrated the feasibility of real-time sequencing analysis and direct microbial identification in space. Applications of sequencing for space exploration include infectious disease diagnosis, environmental monitoring, evaluating biological responses to spaceflight, and even potentially the detection of extraterrestrial life on other planetary bodies.

Genomics

A dual-strategy expression screen for candidate connectivity labels in the developing thalamus

The thalamus or \"inner chamber\" of the brain is divided into ~30 discrete nuclei, with highly specific patterns of afferent and efferent connectivity. To identify genes that may direct these patterns of connectivity, we used two strategies. First, we used a bioinformatics pipeline to survey the predicted proteomes of nematode, fruitfly, mouse and human for extracellular proteins containing any of a list of motifs found in known guidance or connectivity molecules. Second, we performed clustering analyses on the Allen Developing Mouse Brain Atlas data to identify genes encoding surface proteins expressed with temporal profiles similar to known guidance or connectivity molecules. In both cases, we then screened the resultant genes for selective expression patterns in the developing thalamus. These approaches identified 82 candidate connectivity labels in the developing thalamus. These molecules include many members of the Ephrin, Eph-receptor, cadherin, protocadherin, semaphorin, plexin, Odz/teneurin, Neto, cerebellin, calsyntenin and Netrin-G families, as well as diverse members of the immunoglobulin (Ig) and leucine-rich receptor (LRR) superfamilies, receptor tyrosine kinases and phosphatases, a variety of growth factors and receptors, and a large number of miscellaneous membrane-associated or secreted proteins not previously implicated in axonal guidance or neuronal connectivity. The diversity of their expression patterns indicates that thalamic nuclei are highly differentiated from each other, with each one displaying a unique repertoire of these molecules, consistent with a combinatorial logic to the specification of thalamic connectivity.

Neuroscience

Reproducibility and replicability of rodent phenotyping in preclinical studies

The scientific community is increasingly concerned with cases of published \"discoveries\" that are not replicated in further studies. The field of mouse behavioral phenotyping was one of the first to raise this concern, and to relate it to other complicated methodological issues: the complex interaction between genotype and environment; the definitions of behavioral constructs; and the use of the mouse as a model animal for human health and disease mechanisms. In January 2015, researchers from various disciplines including genetics, behavior genetics, neuroscience, ethology, statistics and bioinformatics gathered in Tel Aviv University to discuss these issues. The general consent presented here was that the issue is prevalent and of concern, and should be addressed at the statistical, methodological and policy levels, but is not so severe as to call into question the validity and the usefulness of model organisms as a whole. Well-organized community efforts, coupled with improved data and metadata sharing, were agreed by all to have a key role to play in identifying specific problems and promoting effective solutions. As replicability is related to validity and may also affect generalizability and translation of findings, the implications of the present discussion reach far beyond the issue of replicability of mouse phenotypes but may be highly relevant throughout biomedical research.

Scientific Communication and Education

Genetic,transcriptome, proteomic and epidemiological evidence for blood brain barrier disruption and polymicrobial brain invasion as determinant factors in Alzheimers disease.

Multiple pathogens have been detected in Alzheimers disease (AD) brains. A bioinformatics approach was used to assess relationships between pathogens and AD genes (GWAS), the AD hippocampal transcriptome and plaque or tangle proteins. Host/pathogen interactomes (C.albicans, C.Neoformans, Bornavirus, B.Burgdorferri, cytomegalovirus, Ebola virus, HSV-1, HERV-W, HIV-1, Epstein-Barr, hepatitis C, influenza, C.Pneumoniae, P.Gingivalis, H.Pylori, T.Gondii, T.Cruzi) significantly overlap with misregulated AD hippocampal genes, with plaque and tangle proteins and, except Bornavirus, Ebola and HERV-W, with AD genes. Upregulated AD hippocampal genes match those upregulated by multiple bacteria, viruses, fungi or protozoa in immunocompetent blood cells. AD genes are enriched in bone marrow and immune locations and in GWAS datasets reflecting pathogen diversity, suggesting selection for pathogen resistance. The age of AD patients implies resistance to infections afflicting the younger. APOE4 protects against malaria and hepatitis C, and immune/inflammatory gain of function applies to APOE4, CR1, TREM2 and presenilin variants. 30/78 AD genes are expressed in the blood brain barrier (BBB), which is disrupted by AD risk factors (ageing, alcohol, aluminium, concussion, cerebral hypoperfusion, diabetes, homocysteine, hypercholesterolaemia, hypertension, obesity, pesticides, pollution, physical inactivity, sleep disruption and smoking). The BBB and AD benefit from statins, NSAIDs, oestrogen, melatonin and the Mediterranean diet. Polymicrobial involvement is supported by the upregulation of pathogen sensors/defenders (bacterial, fungal, viral) in the AD brain, blood or CSF. Cerebral pathogen invasion permitted by BBB inadequacy, activating a hyper-efficient immune/inflammatory system, betaamyloid and other antimicrobial defence may be responsible for AD which may respond to antibiotic, antifungal or antiviral therapy.

Neuroscience

Aer is a family of Energy-Taxis Receptors in Pseudomonas which also involves Aer-2 and CttP

Chemotaxis allows bacteria to sense gradients in their environment and respond by directing their swimming. Aer is a receptor that, instead of responding to a specific chemoattractant, allows bacteria to sense cellular energy levels and move towards favourable environments. In Pseudomonas, the number of apparent Aer homologs differs between the only two species it had been characterized in, P. aeruginosa and P. putida. Here we combined bioinformatic approaches with deletional mutagenesis in P. pseudoalcaligenes KF707 to further characterize Aer. It was determined that the number of Aer homologs varies between 0-4 throughout the Pseudomonas genus, and they were phylogenetically classified into 5 subgroups. We also used sequence analysis to show that these homologous receptors differ in their HAMP signal transduction domains. Genetic analysis also indicated that some Aer homologs have likely been subject to horizontal transfer. P. pseudoalcaligenes KF707 was unique among species for having three Aer homologs as well as the receptors CttP and McpB. Phenotypic characterization in this species showed the most prevalent homolog of Aer was key, but not essential for energy-taxis. This study demonstrates that energy-taxis in Pseudomonas varies between species and provides a new naming convention and associated phylogenetic details for Aer chemoreceptors.

microbiology

Searching for the chromatin determinants of human hematopoiesis

Hematopoiesis is one of the best characterized biological systems but the connection between chromatin changes and lineage differentiation is not yet well understood. We have developed a bioinformatic workflow to generate a chromatin space that allows to classify forty-two human healthy blood epigenomes from the BLUEPRINT, NIH ROADMAP and ENCODE consortia by their cell type. This approach let us to distinguish different cells types based on their epigenomic profiles, thus recapitulating important aspects of human hematopoiesis. The analysis of the orthogonal dimension of the chromatin space identify 32,662 chromatin determinant regions (CDRs), genomic regions with different epigenetic characteristics between the cell types. Functional analysis revealed that these regions are linked with cell identities. The inclusion of leukemia epigenomes in the healthy hematological chromatin sample space gives us insights on the healthy cell types that are more epigenetically similar to the disease samples. Further analysis of tumoral epigenetic alterations in hematopoietic CDRs points to sets of genes that are tightly regulated in leukemic transformations and commonly mutated in other tumors. Our method provides an analytical approach to study the relationship between epigenomic changes and cell lineage differentiation. Method availability: https://github.com/david-juan/ChromDet

genomics

16S rRNA Gene Sequencing as a Clinical Diagnostic Aid for Gastrointestinal-related Conditions

Accurate detection of the microorganisms underlying gut dysbiosis in the patient is critical to initiate the appropriate treatment. However, most clinical microbiology techniques used to detect gut bacteria were developed over a century ago and rely on culture-based approaches that are often laborious, unreliable, and subjective. Further, culturing does not scale well for multiple targets and detects only a minority of the microorganisms in the human gastrointestinal tract. Here we present a clinical test for gut microorganisms based on targeted sequencing of the prokaryotic 16S rRNA gene. We tested 46 clinical prokaryotic targets in the human gut, 28 of which can be identified by a bioinformatics pipeline that includes sequence analysis and taxonomic annotation. Using microbiome samples from a cohort of 897 healthy individuals, we established a reference range defining clinically relevant relative levels for each of the 28 targets. Our assay accurately quantified all 28 targets and correctly reflected 38/38 verification samples of real and synthetic stool material containing known pathogens. Thus, we have established a new test to interrogate microbiome composition and diversity, which will improve patient diagnosis, treatment and monitoring. More broadly, our test will facilitate epidemiological studies of the microbiome as it relates to overall human health and disease.

microbiology

The megabase-sized fungal genome of Rhizoctonia solani assembled from nanopore reads only.

The ability to quickly obtain accurate genome sequences of eukaryotic pathogens at low costs provides a tremendous opportunity to identify novel targets for therapeutics, develop pesticides with increased target specificity and breed for resistance in food crops. Here, we present the first report of the ~54 MB eukaryotic genome sequence of Rhizoctonia solani, an important pathogenic fungal species of maize, using nanopore technology. Moreover, we show that optimizing the strategy for wet-lab procedures aimed to isolate high quality and ultra-pure high molecular weight (HMW) DNA results in increased read length distribution and thereby allowing generation of the most contiguous genome assembly for R. solani to date. We further determined sequencing accuracy and compared the assembly to short-read technologies. With the current sequencing technology and bioinformatics tool set, we are able to deliver an eukaryotic fungal genome at low cost within a week. With further improvements of the sequencing technology and increased throughput of the PromethION sequencer we aim to generate near-finished assemblies of large and repetitive plant genomes and cost-efficiently perform de novo sequencing of large collections of microbial pathogens and the microbial communities that surround our crops.

genomics