Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A novel core genome approach to enable prospective and dynamic monitoring of infectious outbreaks

Whole-genome sequencing is increasingly adopted in clinical settings to identify pathogen transmissions. Currently, such studies are performed largely retrospectively, but to be actionable they need to be carried out prospectively, in which samples are continuously added and compared to previous samples. To enable prospective pathogen comparison, genomic relatedness metrics based on single nucleotide differences must be consistent across time, efficient to compute and reliable for a large variety of samples. The choice of genomic regions to compare, i.e., the core genome, is critical to obtain a good metric.\n\nWe propose a novel core genome method that selects conserved sequences in the reference genome by comparing its k-mer content to that of publicly available genome assemblies. The conserved-sequence genome is sample set-independent, which enables prospective pathogen monitoring. Based on clinical data sets of 3436 S. aureus, 1362 K. pneumoniae and 348 E. faecium samples, we show that the conserved-sequence genome disambiguates same-patient samples better than a core genome consisting of conserved genes. The conserved-sequence genome confirms outbreak samples with high accuracy: in a set of 2335 S. aureus samples, it correctly identifies 44 out of 45 outbreak samples, whereas the conserved gene method confirms 38 out of 45 outbreak samples.

bioinformatics

How to make a rodent giant: Genomic basis and tradeoffs of gigantism in the capybara, the world’s largest rodent

Gigantism is the result of one lineage within a clade evolving extremely large body size relative to its small-bodied ancestors, a phenomenon observed numerous times in animals. Theory predicts that the evolution of giants should be constrained by two tradeoffs. First, because body size is negatively correlated with population size, purifying selection is expected to be less efficient in species of large body size, leading to a genome-wide elevation of the ratio of non-synonymous to synonymous substitution rates (dN/dS) or mutation load. Second, gigantism is achieved through higher number of cells and higher rates of cell proliferation, thus increasing the likelihood of cancer. However, the incidence of cancer in gigantic animals is lower than the theoretical expectation, a phenomenon referred to as Petos Paradox. To explore the genetic basis of gigantism in rodents and uncover genomic signatures of gigantism-related tradeoffs, we sequenced the genome of the capybara, the worlds largest living rodent. We found that dN/dS is elevated genome wide in the capybara, relative to other rodents, implying a higher mutation load. Conversely, a genome-wide scan for adaptive protein evolution in the capybara highlighted several genes involved in growth regulation by the insulin/insulin-like growth factor signaling (IIS) pathway. Capybara-specific gene-family expansions included a putative novel anticancer adaptation that involves T cell-mediated tumor suppression, offering a potential resolution to Petos Paradox in this lineage. Gene interaction network analyses also revealed that size regulators function simultaneously as growth factors and oncogenes, creating an evolutionary conflict. Based on our findings, we hypothesize that gigantism in the capybara likely involved three evolutionary steps: 1) Increase in body size by cell proliferation through the ISS pathway, 2) coupled evolution of growth-regulatory and cancer-suppression mechanisms, possibly driven by intragenomic conflict, and 3) establishment of the T cell-mediated tumor suppression pathway as an anticancer adaptation. Interestingly, increased mutation load appears to be an inevitable outcome of an increase in body size.\n\nAuthor SummaryThe existence of gigantic animals presents an evolutionary puzzle. Larger animals have more cells and undergo exponentially more cell divisions, thus, they should have enormous rates of cancer. Moreover, large animals also have smaller populations making them vulnerable to extinction. So, how do gigantic animals such as elephants and blue whales protect themselves from cancer, and what are the consequences of evolving a large size on the genetic health of a species? To address these questions we sequenced the genome of the capybara, the worlds largest rodent, and performed comparative genomic analyses to identify the genes and pathways involved in growth regulation and cancer suppression. We found that the insulin-signaling pathway was involved in the evolution of gigantism in the capybara. We also found a putative novel anticancer mechanism mediated by the detection of tumors by T-cells, offering a potential solution to how capybaras mitigated the tradeoff imposed by cancer. Furthermore, we show that capybara genome harbors a higher proportion of slightly deleterious mutations relative to all other rodent genomes. Overall, this study provides insights at the genomic level into the evolution of a complex and extreme phenotype, and offers a detailed picture of how the evolution of a giant body size in the capybara has shaped its genome.

evolutionary biology

Exon capture optimization in large-genome amphibians

BackgroundGathering genomic-scale data efficiently is challenging for non-model species with large, complex genomes. Transcriptome sequencing is accessible for even large-genome organisms, and sequence capture probes can be designed from such mRNA sequences to enrich and sequence exonic regions. Maximizing enrichment efficiency is important to reduce sequencing costs, but, relatively little data exist for exon capture experiments in large-genome non-model organisms. Here, we conducted a replicated factorial experiment to explore the effects of several modifications to standard protocols that might increase sequence capture efficiency for large-genome amphibians.\n\nMethodsWe enriched 53 genomic libraries from salamanders for a custom set of 8,706 exons under differing conditions. Libraries were prepared using pools of DNA from 3 different salamanders with approximately 30 gigabase genomes: California tiger salamander (Ambystoma californiense), barred tiger salamander (Ambystoma mavortium), and an F1 hybrid between the two. We enriched libraries using different amounts of c0t-1 blocker, individual input DNA, and total reaction DNA. Enriched libraries were sequenced with 150 bp paired-end reads on an Illumina HiSeq 2500, and the efficiency of target enrichment was quantified using unique read mapping rates and average depth across targets. The different enrichment treatments were evaluated to determine if c0t-1 and input DNA significantly impact enrichment efficiency in large-genome amphibians.\n\nResultsIncreasing the amounts of c0t-1 and individual input DNA both reduce the rates of PCR duplication. This reduction led to an increase in the percentage of unique reads mapping to target sequences, essentially doubling overall efficiency of the target capture from 10.4% to nearly 19.9%. We also found that post-enrichment DNA concentrations and qPCR enrichment verification were useful for predicting the success of enrichment.\n\nConclusionsIncreasing the amount of individual sample input DNA and the amount of c0t-1 blocker both increased the efficiency of target capture in large-genome salamanders. By reducing PCR duplication rates, the number of unique reads mapping to targets increased, making target capture experiments more efficient and affordable. Our results indicate that target capture protocols can be modified to efficiently screen large-genome vertebrate taxa including amphibians.

Genomics

Extensive sequencing of seven human genomes to characterize benchmark reference materials

The Genome in a Bottle Consortium, hosted by the National Institute of Standards and Technology (NIST) is creating reference materials and data for human genome sequencing, as well as methods for genome comparison and benchmarking. Here, we describe a large, diverse set of sequencing data for seven human genomes; five are current or candidate NIST Reference Materials. The pilot genome, NA12878, has been released as NIST RM 8398. We also describe data from two Personal Genome Project trios, one of Ashkenazim Jewish ancestry and one of Chinese ancestry. The data come from 12 technologies: BioNano Genomics, Complete Genomics paired-end and LFR, Ion Proton exome, Oxford Nanopore, Pacific Biosciences, SOLiD, 10X Genomics GemCode WGS, and Illumina exome and WGS paired-end, mate-pair, and synthetic long reads. Cell lines, DNA, and data from these individuals are publicly available. Therefore, we expect these data to be useful for revealing novel information about the human genome and improving sequencing technologies, SNP, indel, and structural variant calling, and de novo assembly.

Genomics

Major improvements to the Heliconius melpomene genome assembly used to confirm 10 chromosome fusion events in 6 million years of butterfly evolution

The Heliconius butterflies are a widely studied adaptive radiation of 46 species spread across Central and South America, several of which are known to hybridise in the wild. Here, we present a substantially improved assembly of the Heliconius melpomene genome, developed using novel methods that should be applicable to improving other genome assemblies produced using short read sequencing. Firstly, we whole genome sequenced a pedigree to produce a linkage map incorporating 99% of the genome. Secondly, we incorporated haplotype scaffolds extensively to produce a more complete haploid version of the draft genome. Thirdly, we incorporated ~20x coverage of Pacific Biosciences sequencing and scaffolded the haploid genome using an assembly of this long read sequence. These improvements result in a genome of 795 scaffolds, 275 Mb in length, with an L50 of 2.1 Mb, an N50 of 34 and with 99% of the genome placed and 84% anchored on chromosomes. We use the new genome assembly to confirm that the Heliconius genome underwent 10 chromosome fusions since the split with its sister genus Eueides, over a period of about 6 million years.

Genomics

Geographic cline analysis as a tool for studying genome-wide variation: a case study of pollinator-mediated divergence in a monkeyflower

A major goal of speciation research is to reveal the genomic signatures that accompany the speciation process. Genome scans are routinely used to explore genome-wide variation and identify highly differentiated loci that may contribute to ecological divergence, but they do not incorporate spatial, phenotypic, or environmental data that might enhance outlier detection. Geographic cline analysis provides a potential framework for integrating diverse forms of data in a spatially-explicit framework, but it has not been used to study genome-wide patterns of divergence. Aided by a first-draft genome assembly, we combine an FCT scan and geographic cline analysis to characterize patterns of genome-wide divergence between divergent pollination ecotypes of Mimulus aurantiacus. FCT analysis of 58,872 SNPs generated via RADseq revealed little ecotypic differentiation (mean FCT = 0.041), though a small number of loci were moderately to highly diverged. Consistent with our previous results from the gene MaMyb2, which contributes to differences in flower color, 130 loci have cline shapes that recapitulate the spatial pattern of trait divergence, suggesting that they reside in or near the genomic regions that contribute to pollinator isolation. In the narrow hybrid zone between the ecotypes, extensive admixture among individuals and low linkage disequlibrium between markers indicate that outlier loci are scattered throughout the genome, rather than being restricted to one or a few regions. In addition to revealing the genomic consequences of ecological divergence in this system, we discuss how geographic cline analysis is a powerful but under-utilized framework for studying genome-wide patterns of divergence.

Genomics

Evidence-Based Design and Evaluation of a Whole Genome Sequencing Clinical Report for the Reference Microbiology Laboratory

BackgroundMicrobial genome sequencing is now being routinely used in many clinical and public health laboratories. Understanding how to report complex genomic test results to stakeholders who may have varying familiarity with genomics - including clinicians, laboratorians, epidemiologists, and researchers - is critical to the successful and sustainable implementation of this new technology; however, there are no evidence-based guidelines for designing such a report in the pathogen genomics domain. Here, we describe an iterative, human-centered approach to creating a report template for communicating tuberculosis (TB) genomic test results.\n\nMethodsWe used Design Study Methodology - a human centered multi-stage approach drawn from the information visualization domain - to redesign an existing clinical report. We used expert consults and an online questionnaire to discover various stakeholders needs around the types of data and tasks related to TB that they encounter in their daily workflow. We also evaluated their perceptions of and familiarity with genomic data, as well as its utility at various clinical decision points. These data shaped the design of multiple prototype reports that were compared against the existing report through a second online survey, with the resulting qualitative and quantitative data informing the final, redesigned, report.\n\nResultsWe recruited 78 participants, 65 of whom were clinicians, nurses, laboratorians, researchers, and epidemiologists involved in TB diagnosis, treatment, and/or surveillance. Our first survey indicated that participants were largely enthusiastic about genomic data, with the majority agreeing on its utility for certain TB diagnosis and treatment tasks and many reporting some confidence in their ability to interpret this type of data (between 58.8% and 94.1%, depending on the specific data type). When we compared our four prototype reports against the existing design, we found that for the majority (86.7%) of design comparisons, participants preferred the alternative prototype designs over the existing version, and that both clinicians and non-clinicians expressed similar design preferences. Participants articulated clearer design preferences when asked to compare individual design elements versus entire reports. Both the quantitative and qualitative data informed the design of a revised report, which is available online as a LaTeX template.\n\nConclusionsWe show how a human-centered design approach integrating quantitative and qualitative feedback can be used to design an alternative report for representing complex microbial genomic data. We suggest experimental and design guidelines to inform future design studies in the bioinformatics and microbial genomics domains, and suggest that this type of mixed-methods study is important to facilitate the successful translation of pathogen genomics in the clinic, not only for clinical reports but also more complex bioinformatics data visualization software.

genomics

Higher-order inter-chromosomal hubs shape 3-dimensional genome organization in the nucleus

Eukaryotic genomes are packaged into a 3-dimensional structure in the nucleus of each cell. There are currently two distinct views of genome organization that are derived from different technologies. The first view, derived from genome-wide proximity ligation methods (e.g. Hi-C), suggests that genome organization is largely organized around chromosomes. The second view, derived from in situ imaging, suggests a central role for nuclear bodies. Yet, because microscopy and proximity-ligation methods measure different aspects of genome organization, these two views remain poorly reconciled and our overall understanding of how genomic DNA is organized within the nucleus remains incomplete. Here, we develop Split-Pool Recognition of Interactions by Tag Extension (SPRITE), which moves away from proximity-ligation and enables genome-wide detection of higher-order DNA interactions within the nucleus. Using SPRITE, we recapitulate known genome structures identified by Hi-C and show that the contact frequencies measured by SPRITE strongly correlate with the 3-dimensional distances measured by microscopy. In addition to known structures, SPRITE identifies two major hubs of inter-chromosomal interactions that are spatially arranged around the nucleolus and nuclear speckles, respectively. We find that the majority of genomic regions exhibit preferential spatial association relative to one of these nuclear bodies, with regions that are highly transcribed by RNA Polymerase II organizing around nuclear speckles and transcriptionally inactive and centromere-proximal regions organizing around the nucleolus. Together, our results reconcile the two distinct pictures of nuclear structure and demonstrate that nuclear bodies act as inter-chromosomal hubs that shape the overall 3-dimensional packaging of genomic DNA in the nucleus.

genomics

The genome of the biting midge Culicoides sonorensis and gene expression analyses of vector competence for Bluetongue virus

BackgroundThe use of the new genomic technologies has led to major advances in control of several arboviruses of medical importance such as Dengue. However, the development of tools and resources available for vectors of non-zoonotic arboviruses remains neglected. Biting midges of the genus Culicoides transmit some of the most important arboviruses of wildlife and livestock worldwide, with a global impact on economic productivity, health and welfare. The absence of a suitable reference genome has hindered genomic analyses to date in this important genus of vectors. In the present study, the genome of Culicoides sonorensis, a vector of bluetongue virus (BTV) in the USA, has been sequenced to provide the first reference genome for these vectors. In this study, we also report the use of the reference genome to perform initial transcriptomic analyses of vector competence for BTV.\n\nResultsOur analyses reveal that the genome is 197.4 Mb, assembled in 7,974 scaffolds. Its annotation using the transcriptomic data generated in this study and in a previous study has identified 15,629 genes. Gene expression analyses of C. sonorensis females infected with BTV performed in this study revealed 165 genes that were differentially expressed between vector competent and refractory females. Two candidate genes, glutathione S-transferase (gst) and the antiviral helicase ski2, previously recognized as involved in vector competence for BTV in C. sonorensis (gst) and repressing dsRNA virus propagation (ski2), were confirmed in this study.\n\nConclusionsThe reference genome of C. sonorensis has enabled preliminary analyses of the gene expression profiles of vector competent and refractory individuals. The genome and transcriptomes generated in this study provide suitable tools for future research on arbovirus transmission. These provide a significant resource for these vector lineage, which diverged from other major Dipteran vector families over 200 million years ago. The genome will be a valuable source of comparative data for other important Dipteran vector families including mosquitoes (Culicidae) and sandflies (Psychodidae), and yield potential targets for transgenic modification in vector control and functional studies.

genomics

Hybrid de novo assembly of the draft genome of the freshwater mussel Venustaconcha ellipsiformis (Bivalvia: Unionida).

Freshwater mussels (Bivalvia: Unionida) serve an important role as aquatic ecosystem engineers but are one of the most critically imperilled groups of animals. Here, we used a combination of sequencing strategies to assemble and annotate a draft genome of Venustaconcha ellipsiformis, which will serve as a valuable genomic resource given the ecological value and unique \"doubly uniparental inheritance\" mode of mitochondrial DNA transmission of freshwater mussels. The genome described here was obtained by combining high coverage short reads (65X genome coverage of Illumina paired-end and 11X genome coverage of mate-pairs sequences) with low coverage Pacific Biosciences long reads (0.3X genome coverage). Briefly, the final scaffold assembly accounted for a total size of 1.54Gb (366,926 scaffolds, N50 = 6.5Kb, with 2.3% of \"N\" nucleotides), representing 86% of the predicted genome size of 1.80Gb, while over one third of the genome (37.5%) consisted of repeated elements and more than 85% of the core eukaryotic genes were recovered. Given the repeated genetic bottlenecks of V. ellipsiformis populations as a result of glaciations events, heterozygosity was also found to be remarkably low (0.6%), in contrast to most other sequenced bivalve species. Finally, we reassembled the full mitochondrial genome and found six polymorphic sites with respect to the previously published reference. This resource opens the way to comparative genomics studies to identify genes related to the unique adaptations of freshwater mussels and their distinctive mitochondrial inheritance mechanism.

genomics

High-quality genome assemblies of 15 Drosophila species generated using Nanopore sequencing

The Drosophila genus is a unique group containing a wide range of species that occupy diverse ecosystems. In addition to the most widely studied species, Drosophila melanogaster, many other members in this genus also possess a well-developed set of genetic tools. Indeed, high-quality genomes exist for several species within the genus, facilitating studies of the function and evolution of cis-regulatory regions and proteins by allowing comparisons across at least 50 million years of evolution. Yet, the available genomes still fail to capture much of the substantial genetic diversity within the Drosophila genus. We have therefore tested protocols to rapidly and inexpensively sequence and assemble the genome from any Drosophila species using single-molecule sequencing technology from Oxford Nanopore. Here, we use this technology to present high-quality genome assemblies of 15 Drosophila species: 10 of the 12 originally sequenced Drosophila species (ananassae, erecta, mojavensis, persimilis, pseudoobscura, sechellia, simulans, virilis, willistoni, and yakuba), four additional species that had previously reported assemblies (biarmipes, bipectinata, eugracilis, and mauritiana), and one novel assembly (triauraria). Genomes were generated from an average of 29x depth-of-coverage data that after assembly resulted in an average contig N50 of 4.4 Mb. Subsequent alignment of contigs from the published reference genomes demonstrates that our assemblies could be used to close over 60% of the gaps present in the currently published reference genomes. Importantly, the materials and reagents cost for each genome was approximately $1,000 (USD). This study demonstrates the power and cost-effectiveness of long-read sequencing for genome assembly in Drosophila and provides a framework for the affordable sequencing and assembly of additional Drosophila genomes.

genomics

GET_PHYLOMARKERS, a software package to select optimal orthologous clusters for phylogenomics and inferring pan-genome phylogenies, used for a critical geno-taxonomic revision of the genus Stenotrophomonas

The massive accumulation of genome-sequences in public databases promoted the proliferation of genome-level phylogenetic analyses in many areas of biological research. However, due to diverse evolutionary and genetic processes, many loci have undesirable properties for phylogenetic reconstruction. These, if undetected, can result in erroneous or biased estimates, particularly when estimating species trees from concatenated datasets. To deal with these problems, we developed GET_PHYLOMARKERS, a pipeline designed to identify high-quality markers to estimate robust genome phylogenies from the orthologous clusters, or the pan-genome matrix (PGM), computed by GET_HOMOLOGUES. In the first context, a set of sequential filters are applied to exclude recombinant alignments and those producing anomalous or poorly resolved trees. Multiple sequence alignments and maximum likelihood (ML) phylogenies are computed in parallel on multi-core computers. A ML species tree is estimated from the concatenated set of top-ranking alignments at the DNA or protein levels, using either FastTree or IQ-TREE (IQT). The latter is used by default due to its superior performance revealed in an extensive benchmark analysis. In addition, parsimony and ML phylogenies can be estimated from the PGM.\n\nWe demonstrate the practical utility of the software by analyzing 170 Stenotrophomonas genome sequences available in RefSeq and 10 new complete genomes of environmental S. maltophilia complex (Smc) isolates reported herein. A combination of core-genome and PGM analyses was used to revise the molecular systematics of the genus. An unsupervised learning approach that uses a goodness of clustering statistic identified 20 groups within the Smc at a core-genome average nucleotide identity of 95.9% that are perfectly consistent with strongly supported clades on the core- and pan-genome trees. In addition, we identified 14 misclassified RefSeq genome sequences, 12 of them labeled as S. maltophilia, demonstrating the broad utility of the software for phylogenomics and geno-taxonomic studies. The code, a detailed manual and tutorials are freely available for Linux/UNIX servers under the GNU GPLv3 license at https://github.com/vinuesa/get_phylomarkers. A docker image bundling GET_PHYLOMARKERS with GET_HOMOLOGUES is available at https://hub.docker.com/r/csicunam/get_homologues/, which can be easily run on any platform.

genomics

Genomic resources for Goniozus legneri, Aleochara bilineata and Paykullia maculata, representing three independent origins of the parasitoid lifestyle in insects

Parasitoid insects are important model systems for a multitude of biological research topics and widely used as biological control agents against insect pests. While the parasitoid lifestyle has evolved numerous times in different insect groups, research has focused almost exclusively on Hymenoptera from the parasitica clade. The genomes of several members of this group have been sequenced, but no genomic resources are available from any of the other, independent evolutionary origins of the parasitoid lifestyle. Our aim here was to develop genomic resources for three parasitoid insects outside the parasitica. We present draft genome assemblies for Goniozus legneri, a parasitoid Hymenopteran more closely related to the non-parasitoid wasps and bees than to the parasitica wasps, the Coleopteran parasitoid Aleochara bilineata and the Dipteran parasitoid Paykullia maculata. The genome assemblies are fragmented, but complete in terms of gene content. We also provide preliminary structural annotations. We anticipate that these genomic resources will be valuable for testing the generality of findings obtained from parasitica wasps in future comparative studies.\n\nData availabilityThe Whole Genome Shotgun projects have been deposited at DDBJ/ENA/GenBank under the accessions NCVS00000000 (G. legneri), NBZA00000000 (A. bilineata) and NDXZ00000000 (P. maculata). The versions described in this paper are versions NCVS01000000, NBZA01000000 and NDXZ01000000, respectively. Mapped reads and genome annotations are available through http://parasitoids.labs.vu.nl/parasitoids/. This website also includes genome browsers and viroblast instances for each genome.

genomics

Assembly of chloroplast genomes with long- and short-read data: a comparison of approaches using Eucalyptus pauciflora as a test case

BackgroundChloroplasts are organelles that conduct photosynthesis in plant and algal cells. Chloroplast genomes code for around 130 genes, and the information they contain is widely used in agriculture and studies of evolution and ecology. Correctly assembling complete chloroplast genomes can be challenging because the chloroplast genome contains a pair of long inverted repeats (10-30 kb). The advent of long-read sequencing technologies should alleviate this problem by providing sufficient information to completely span the inverted repeat regions. Yet, long-reads tend to have higher error rates than short-reads, and relatively little is known about the best way to combine long- and short-reads to obtain the most accurate chloroplast genome assemblies. Using Eucalyptus pauciflora, the snow gum, as a test case, we evaluated the effect of multiple parameters, such as different coverage of long (Oxford nanopore) and short (Illumina) reads, different long-read lengths, different assembly pipelines, and different genome polishing steps, with a view to determining the most accurate and efficient approach to chloroplast genome assembly.\n\nResultsHybrid assemblies combining at least 20x coverage of both long-reads and short-reads generated a single contig spanning the entire chloroplast genome with few or no detectable errors. Short-read-only assemblies generated three contigs representing the long single copy, short single copy and inverted repeat regions of the chloroplast genome. These contigs contained few single-base errors but tended to exclude several bases at the beginning or end of each contig. Long-read-only assemblies tended to create multiple contigs with a much higher single-base error rate, even after polishing. The chloroplast genome of Eucalyptus pauciflora is 159,942 bp, contains 131 genes of known function, and confirms the phylogenetic position of Eucalyptus pauciflora as a close relative of Eucalyptus regnans.\n\nConclusionsOur results suggest that very accurate assemblies of chloroplast genomes can be achieved using a combination of at least 20x coverage of long- and short-reads respectively, provided that the long-reads contain at least ~5x coverage of reads longer than the inverted repeat region. We show that further increases in coverage give little or no improvement in accuracy, and that hybrid assemblies are more accurate than long-read-only or short-read-only assemblies.

genomics

Comparative Genomic Analysis of Bacillus thuringiensis Reveals Molecular Adaptation to Copper Tolerance

Bacillus thuringiensis is a type of Gram positive and rod shaped bacterium that is found in a wide range of habitats. Despite the intensive studies conducted on this bacterium, most of the information available are related to its pathogenic characteristics, with only a limited number of publications mentioning its ability to survive in extreme environments. Recently, a B. thuringiensis MCMY1 strain was successfully isolated from a copper contaminated site in Mamut Copper Mine, Sabah. This study aimed to conduct a comparative genomic analysis by using the genome sequence of MCMY1 strain published in GenBank (PRJNA374601) as a target genome for comparison with other available B. thuringiensis genomes at the GenBank. Whole genome alignment, Fragment all-against-all comparison analysis, phylogenetic reconstruction and specific copper genes comparison were applied to all forty-five B. thuringiensis genomes to reveal the molecular adaptation to copper tolerance. The comparative results indicated that B. thuringiensis MCMY1 strain is closely related to strain Bt407 and strain IS5056. This strain harbors almost all available copper genes annotated from the forty-five B. thuringiensis genomes, except for the gene for Magnesium and cobalt efflux protein (CorC) which plays an indirect role in reducing the oxidative stress that caused by copper and other metal ions. Furthermore, the findings also showed that the Copper resistance gene family, CopABCDZ and its repressor (CsoR) are conserved in almost all sequenced genomes but the presence of the genes for Cytoplasmic copper homeostasis protein (CutC) and CorC across the sample genomes are highly inconsonant. The variation of these genes across the B. thuringiensis genomes suggests that each strain may have adapted to their specific ecological niche. However, further investigations will be need to support this preliminary hypothesis.

genomics

Evaluation of strategies for the assembly of diverse bacterial genomes using MinION long-read sequencing

BackgroundShort-read sequencing technologies have made microbial genome sequencing cheap and accessible. However, closing genomes is often costly and assembling short reads from genomes that are repetitive and/or have extreme %GC content remains challenging. Long-read, single-molecule sequencing technologies such as the Oxford Nanopore MinION have the potential to overcome these difficulties, although the best approach for harnessing their potential remains poorly evaluated.\n\nResultsWe sequenced nine bacterial genomes spanning a wide range of GC contents using Illumina MiSeq and Oxford Nanopore MinION sequencing technologies to determine the advantages of each approach, both individually and combined. Assemblies using only MiSeq reads were highly accurate but lacked contiguity, a deficiency that was partially overcome by adding MinION reads to these assemblies. Even more contiguous genome assemblies were generated by using MinION reads for initial assembly, but these were more error-prone and required further polishing. This was especially pronounced when Illumina libraries were biased, as was the case for our strains with both high and low GC content. Increased genome contiguity dramatically improved the annotation of insertion sequences and secondary metabolite biosynthetic gene clusters, likely because long-reads can disambiguate these highly repetitive but biologically important genomic regions.\n\nConclusionsGenome assembly using short-reads is challenged by repetitive sequences and extreme GC contents. Our results indicate that these difficulties can be largely overcome by using single-molecule, long-read sequencing technologies such as the Oxford Nanopore MinION. Using MinION reads for assembly followed by polishing with Illumina reads generated the most contiguous genomes and enabled the accurate annotation of important but difficult to sequence genomic features such as insertion sequences and secondary metabolite biosynthetic gene clusters. The combination of Oxford Nanopore and Illumina sequencing is cost effective and dramatically advances studies of microbial evolution and genome-driven drug discovery.

genomics

A Computational Framework To Assess Genome-Wide Distribution Of Polymorphic Human Endogenous Retrovirus-K In Human Populations

Human Endogenous Retrovirus type K (HERV-K) is the only HERV known to be insertionally polymorphic. It is possible that HERV-Ks contribute to human disease because people differ in both number and genomic location of these retroviruses. Indeed viral transcripts, proteins, and antibody against HERV-K are detected in cancers, auto-immune, and neurodegenerative diseases. However, attempts to link a polymorphic HERV-K with any disease have been frustrated in part because population frequency of HERV-K provirus at each site is lacking and it is challenging to identify closely related elements such as HERV-K from short read sequence data. We present an integrated and computationally robust approach that uses whole genome short read data to determine the occupation status at all sites reported to contain a HERV-K provirus. Our method estimates the proportion of fixed length genomic sequence (k-mers) from whole genome sequence data matching a reference set of k-mers unique to each HERV-K loci and applies mixture model-based clustering to account for low depth sequence data. Our analysis of 1000 Genomes Project Data (KGP) reveals numerous differences among the five KGP super-populations in the frequency of individual and co-occurring HERV-K proviruses; we provide a visualization tool to easily depict the prevalence of any combination of HERV-K among KGP populations. Further, the genome burden of polymorphic HERV-K is variable in humans, with East Asian (EAS) individuals having the fewest integration sites. Our study identifies population-specific sequence variation for several HERV-K proviruses. We expect these resources will advance research on HERV-K contributions to human diseases.\n\nAuthor summaryHuman Endogenous Retrovirus type K (HERV-K) is the youngest of retrovirus families in the human genome and is the only group that is polymorphic; a HERV-K can be present in one individual but absent from others. HERV-Ks could contribute to disease risk but establishing a link of a polymorphic HERV-K to a specific disease has been difficult. We develop an easy to use method that reveals the considerable variation existing among global populations in the frequency of individual and co-occurring polymorphic HERV-K, and in the total number of HERV-K that any individual has in their genome. Our study provides a global reference set of HERV-K genomic diversity and tools needed to determine the genomic landscape of HERV-K in any patient population.

genomics

Comparative analysis of Streptococcus genomes

BackgroundGenome sequencing of multiple strains demonstrated high variability in gene content even in closely related strains of the same species and created a newly emerged object for genomic analysis, the pan-genome, that is, the complete set of genes observed in a given species or a higher level taxon. Here we analysed the pan-genome structure and the genome evolution of 25 strains of Streptococcus suis, 50 strains of Streptococcus pyogenes and 28 strains of Streptococcus pneumoniae.\n\nResultsFractions of the pan-genome, unique, periphery, and universal genes differ in size, functional composition, the level of nucleotide substitutions, and predisposition to horizontal gene transfer and genomic rearrangements. The density of substitutions in intergenic regions appears to be correlated with selection acting on adjacent genes, implying that more conserved genes tend to have more conserved regulatory regions. The total pan-genome of the genus is open, but only due to strain-specific genes, whereas other pan-genome fractions reach saturation. The strain-specific fraction is enriched with mobile elements and hypothetical proteins, but also contains a number of candidate virulence-related genes, so it may have a strong impact on adaptability and pathogenicity.\n\nAbout 7% of single-copy periphery genes have been found in different syntenic regions. More than a half of these genes are rare in all Streptococcus species; others are rare in at least one species. We have identified the set of genes with phylogenies inconsistent with species and non-conserved location in the chromosome; these genes are candidates for horizontal transfer between species.\n\nAn inversion of length 15 kB found in four independent branches of S. pneumoniae has breakpoints formed by genes encoding a surface antigen protein (PhtD). The observed parallelism may indicate the action of an antigen variation mechanism.\n\nConclusionsMembers of the genus Streptococcus have a highly dynamic, open pan-genome, that potentially confers them with the ability to adapt to changing environmental conditions, i.e. antibiotic resistance or transmission between different hosts. Hence, understanding of genome evolution is important for the identification of potential pathogens and design of drugs and vaccines.

genomics