Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Fast and global detection of periodic sequence repeats in large genomic resources

Periodically repeating DNA and protein elements are involved in various important biological events including genomic evolution, gene regulation, protein complex formation, and immunity. Notably, the currently used genome editing tools such as ZFNs, TALENs, and CRISPRs are also all associated with periodically repeating biomolecules of natural organisms. Despite the biological importance of periodically repeating sequences and the expectation that new genome editing modules could be discovered from such periodical repeats, no software that globally detects such structured elements in large genomic resources in a high-throughput and unsupervised manner has been developed. Here, we developed new software, SPADE (Search for Patterned DNA Elements), that exhaustively explores periodic DNA and protein repeats from large-scale genomic datasets based on k-mer periodicity evaluation. SPADE precisely captured reported genome-editing-associated sequences and other protein families involving repeating domains with significantly better performance than the other software designed for limited sets of repetitive biomolecular sequences.

bioinformatics

Whole genome analysis of local Kenyan and global sequences unravels the epidemiological and molecular evolutionary dynamics of RSV genotype ON1 strains

The respiratory syncytial virus (RSV) group A variant with the 72-nucleotide duplication in the G gene, genotype ON1, was first detected in Kilifi in 2012 and has almost completely replaced previously circulating genotype GA2 strains. This replacement suggests some fitness advantage of ON1 over the GA2 viruses, and might be accompanied by important genomic substitutions in ON1 viruses. Close observation of such a new virus introduction over time provides an opportunity to better understand the transmission and evolutionary dynamics of the pathogen. We have generated and analyzed 184 RSV-A whole genome sequences (WGS) from Kilifi (Kenya) collected between 2011 and 2016, the first ON1 genomes from Africa and the largest collection globally from a single location. Phylogenetic analysis indicates that RSV-A transmission into this coastal Kenya location is characterized by multiple introductions of viral lineages from diverse origins but with varied success in local transmission. We identify signature amino acid substitutions between ON1 and GA2 viruses within genes encoding the surface proteins (G, F), polymerase (L) and matrix M2-1 proteins, some of which were identified as positively selected, and thereby provide an enhanced picture of RSV-A diversity. Furthermore, five of the eleven RSV open reading frames (ORF) (i.e. G, F, L, N and P), analyzed separately, formed distinct phylogenetic clusters for the two genotypes. This might suggest that coding regions outside of the most frequently studied G ORF play a role in the adaptation of RSV to host populations with the alternative possibility that some of the substitutions are nothing more than genetic hitchhikers. Our analysis provides insight into the epidemiological processes that define RSV spread, highlights the genetic substitutions that characterize emerging strains, and demonstrates the utility of large-scale WGS in molecular epidemiological studies.\n\nAuthor summaryRespiratory syncytial virus (RSV) is the leading viral cause of severe pneumonia and bronchiolitis among infants and children globally. No vaccine exists to date. The high genetic variability of this RNA virus, characterized by group (A or B), genotype (within group) and variant (within genotype) replacement in populations, may pose a challenge to effective vaccine design by enabling immune response escape. To date most sequence data exists for the highly variable G gene encoding the RSV attachment protein, and there is little globally-sampled RSV genomic data to provide a fine resolution of the epidemiology and evolutionary dynamics of the pathogen. Here we use long-term RSV surveillance in coastal Kenya to track the introduction, spread and evolution of a new RSV genotype known as ON1 (having a 72-nucleotide duplication in the G gene). We present a set of 184 RSV-A whole genomes, including 176 of RSV ON1 (the first from Africa), describe patterns of local ON1 spread and show genome-wide changes between the two major RSV-A genotypes that may define the pathogens adaptation to the host. These findings have implications for vaccine design and improved understanding of RSV epidemiology and evolution.

epidemiology

Genomics of microgeographic adaptation in the hyperdominant Amazonian tree Eperua falcata Aubl. (Fabaceae).

Plant populations can undergo very localized adaptation, allowing widely distributed populations to adapt to divergent habitats in spite of recurrent gene flow. Neotropical trees - whose large and undisturbed populations often span a variety of environmental conditions and local habitats - are particularly good models to study this process. Here, we carried out a genome scan for selection through whole-genome sequencing of pools of populations, sampled according to a replicated sampling design, to evaluate microgeographic adaptation in the hyperdominant Amazonian tree Eperua falcata Aubl. (Fabaceae). A high-coverage genomic resource of [~]250 Mb was assembled de novo and annotated, leading to 32,789 predicted genes. 97,062 bi-allelic SNPs were detected over 25,803 contigs, and a custom Bayesian model was implemented to uncover candidate genomic targets of divergent selection. A set of 290 divergence outlier SNPs was detected at the regional scale (between study sites), while 185 SNPs located in the vicinity of 106 protein-coding genes were detected as replicated outliers between microhabitats within regions. Outlier genomic regions are involved in a variety of physiological processes, including plant responses to stress (e.g., oxidative stress, hypoxia and metal toxicity) and biotic interactions. Together with evidence suggesting microgeographic divergence in functional traits, the discovery of genomic targets of microgeographic adaptation in the Neotropics is consistent with the hypothesis that local adaptation is a key driver of ecological diversification, operating across multiple spatial scales, from large- (i.e. regional) to microgeographic- (i.e. landscape) scales.

evolutionary biology

Quality Control and Integration of Genotypes from Two Calling Pipelines for Whole Genome Sequence Data in the Alzheimer’s Disease Sequencing Project

The Alzheimers Disease Sequencing Project (ADSP) performed whole genome sequencing (WGS) of 584 subjects from 111 multiplex families at three sequencing centers. Genotype calling of single nucleotide variants (SNVs) and insertion-deletion variants (indels) was performed centrally using GATK-HaplotypeCaller and Atlas V2. The ADSP Quality Control (QC) Working Group applied QC protocols to project-level variant call format files (VCFs) from each pipeline, and developed and implemented a novel protocol, termed \"consensus calling,\" to combine genotype calls from both pipelines into a single high-quality set. QC was applied to autosomal bi-allelic SNVs and indels, and included pipeline-recommended QC filters, variant-level QC, and sample-level QC. Low-quality variants or genotypes were excluded, and sample outliers were noted. Quality was assessed by examining Mendelian inconsistencies (MIs) among 67 parent-offspring pairs, and MIs were used to establish additional genotype-specific filters for GATK calls. After QC, 578 subjects remained. Pipeline-specific QC excluded ~12.0% of GATK and 14.5% of Atlas SNVs. Between pipelines, ~91% of SNV genotypes across all QCed variants were concordant; 4.23% and 4.56% of genotypes were exclusive to Atlas or GATK, respectively; the remaining ~0.01% of discordant genotypes were excluded. For indels, variant-level QC excluded ~36.8% of GATK and 35.3% of Atlas indels. Between pipelines, ~55.6% of indel genotypes were concordant; while 10.3% and 28.3% were exclusive to Atlas or GATK, respectively; and ~0.29% of discordant genotypes were. The final WGS consensus dataset contains 27,896,774 SNVs and 3,133,926 indels and is publicly available.\n\nAbbreviationsAD, Alzheimers disease; QC, Quality Control; LSSAC, Large-Scale Sequencing and Analysis Center; Broad, Broad Institute Genomics Service; Baylor, Baylor College of Medicine Human Genome Sequencing Center; WashU, Washington University-St. Louis McDonnell Genome Institute; WGS, whole genome sequencing; WES, whole exome sequencing; indel, insertion-deletion variants; VCF, variant control format; MI, Mendelian inconsistency; MC, Mendelian consistency; GWAS, genome-wide association study; VR, referent allele read depth; DP, overall read depth; MS, mapping score; GQ, genotype quality score; Ti/Tv, Transition/Transversion; CS, concordance code

genetics

Marine sponges as Chloroflexi hot-spots: Genomic insights and high resolution visualization of an abundant and diverse symbiotic clade

Chloroflexi represent a widespread, yet enigmatic bacterial phylum. Meta-and single cell genomics were performed to shed light on the functional gene repertoire of Chloroflexi symbionts from the HMA sponge Aplysina aerophoba. Eighteen draft genomes were reconstructed and placed into phylogenetic context of which six were investigated in detail. Common genomic features of Chloroflexi sponge symbionts were related to central energy and carbon converting pathways, amino acid and fatty acid metabolism and respiration. Clade specific metabolic features included a massively expanded genomic repertoire for carbohydrate degradation in Anaerolineae and Caldilineae genomes, and amino acid utilization as nutrient source by SAR202. While Anaerolineae and Caldilineae import cofactors and vitamins, SAR202 genomes harbor genes encoding for co-factor biosynthesis. A number of features relevant to symbiosis were further identified, including CRISPRs-Cas systems, eukaryote-like repeat proteins and secondary metabolite gene clusters. Chloroflexi symbionts were visualized in the sponge extracellular matrix at ultrastructural resolution by FISH-CLEM method. Chloroflexi cells were generally rod-shaped and about 1 m in length, albeit displayed different and characteristic cellular morphotypes per each class. The extensive potential for carbohydrate degradation has been reported previously for Ca. Poribacteria and SAUL, typical symbionts of HMA sponges, and we propose here that HMA sponge symbionts collectively engage in degradation of dissolved organic matter, both labile and recalcitrant. Thus sponge microbes may not only provide nutrients to the sponge host, but also contribute to DOM re-cycling and primary productivity in reef ecosystems via a pathway termed the \"sponge loop\".

microbiology

Genome Functional Annotation using Deep Convolutional Neural Network

Deep neural network application is today a skyrocketing field in many disciplinary domains. In genomics the development of deep neural networks is expected to revolutionize current practice. Several approaches relying on convolutional neural networks have been developed to associate short genomic sequences with a functional role such as promoters, enhancers or protein binding sites along genomes. These approaches rely on the generation of sequences batches with known annotations for learning purpose. While they show good performance to predict annotations from a test subset of these batches, they usually perform poorly when applied genome-wide.\n\nIn this study, we address this issue and propose an optimal strategy to train convolutional neural networks for this specific application. We use as a case study transcription start sites and show that a model trained on one organism can be used to predict transcription start sites in a different specie. This cross-species application of convolutional neural networks trained with genomic sequence data provides a new technique to annotate any genome from previously existing annotations in related species. It also provides a way to determine whether the sequence patterns recognized by chromatin associated proteins in different species are conserved or not.

bioinformatics

An Integrative Boosting Approach for Predicting Survival Time With Multiple Genomics Platforms

Recent technological advances have made it possible to collect multiple types of genomics data on the same set of patients. It is of great interest to integrate multiple genomics data types together for predicting disease outcomes. We propose a variable selection method, termed Integrative Boosting (I-Boost), that makes proper use of all available clinical and genomics data in predicting individual patient survival time. Through simulation studies and applications to data sets from The Cancer Genome Atlas, we demonstrate that I-Boost provides substantially higher prediction accuracy than existing variable selection methods. Using I-Boost, we show that (1) the integration of multiple genomics platforms with clinical variables significantly improves the prediction accuracy for survival time over the use of clinical variables alone; (2) gene expression values are typically more prognostic of survival time than other genomics data types; and (3) gene modules/signatures are at least as prognostic as the collection of individual gene expression data.

bioinformatics

Parallel sexual and parasexual population genomic structure in Trypanosoma cruzi

Genetic exchange and hybridization in parasitic organisms is fundamental to the exploitation of new hosts and host populations. Variable mating frequency often coincides with strong metapopulation structure, where patchy selection or demography may favor different reproductive modes. Evidence for genetic exchange in Trypanosoma cruzi over the last 30 years has been limited and inconclusive. The reproductive modes of other medically important trypanosomatids are better established, although little is known about their variability on a spatio-temporal scale. Targeting a contemporary focus of T. cruzi transmission in southern Ecuador, we present compelling evidence from 45 sequenced genomes that T. cruzi (discrete typing unit I) maintains sexual populations alongside others that represent clonal bursts of parasexual origin. Strains from one site exhibit genome-wide Hardy-Weinberg equilibrium and intra-chromosomal linkage decay consistent with meiotic reproduction. Strains collected from adjacent areas (>6 km) show excess heterozygosity, near-identical haplo-segments, common mitochondrial sequences and levels of aneuploidy incompatible with Mendelian sex. Certain individuals exhibit trisomy in as many as fifteen chromosomes. Others present fewer, yet shared, aneuploidies reminiscent of mitotic genome erosion and parasexual genetic exchange. Genomic and intra-genomic phylogenetics as well as haplotype co-ancestry analyses indicate a clear break in gene-flow between these distinct populations, despite the fact that they occasionally co-occur in vectors and hosts. We propose biological explanations for the fine-scale disconnectivity we observe and discuss the epidemiological consequences of flexible reproductive modes and their genomic architecture for this medically important parasite.

evolutionary biology

Analyses of cancer data in the Genomic Data Commons Data Portal with new functionalities in the TCGAbiolinks R/Bioconductor package

The advent of Next Generation Sequencing (NGS) technologies has opened new perspectives in deciphering the genetic mechanisms underlying complex diseases. Nowadays, the amount of genomic data is massive and substantial efforts and new tools are required to unveil the information hidden in the data.\n\nThe Genomic Data Commons (GDC) Data Portal is a large data collection platform that includes different genomic studies included the ones from The Cancer Genome Atlas (TCGA) and the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) initiatives, accounting for more than 40 tumor types originating from nearly 30000 patients. Such platforms, although very attractive, must make sure the stored data are easily accessible and adequately harmonized. Moreover, they have the primary focus on the data storage in a unique place, and they do not provide a comprehensive toolkit for analyses and interpretation of the data. To fulfill this urgent need, comprehensive but easily accessible computational methods for integrative analyses of genomic data without renouncing a robust statistical and theoretical framework are needed. In this context, the R/Bioconductor package TCGAbiolinks was developed, offering a variety of bioinformatics functionalities. Here we introduce new features and enhancements of TCGAbiolinks in terms of i) more accurate and flexible pipelines for differential expression analyses, ii) different methods for tumor purity estimation and filtering, iii) integration of normal samples from the Genotype-Tissue-Expression (GTEx) platform iv) support for other genomics datasets, here exemplified by the TARGET data.\n\nEvidence has shown that accounting for tumor purity is essential in the study of tumorigenesis, as these factors promote confounding behavior regarding differential expression analysis. Henceforth, we implemented these filtering procedures in TCGAbiolinks. Moreover, a limitation of some of the TCGA datasets is the unavailability or paucity of corresponding normal samples. We thus integrated into TCGAbiolinks the possibility to use normal samples from the Genotype-Tissue Expression (GTEx) project, which is another large-scale repository cataloging gene expression from healthy individuals. The new functionalities are available in the TCGABiolinks v 2.8 and higher released in Bioconductor version 3.7.

bioinformatics

The functional genomic circuitry of human glioblastoma stem cells

SummarySuccessful glioblastoma (GBM) therapies have remained elusive due to limitations in understanding mechanisms of growth and survival of the tumorigenic population. Using CRISPR-Cas9 approaches in patient-derived GBM stem cells to interrogate function of the coding genome, we identify diverse actionable pathways responsible for growth that reveal the gene-essential circuitry of GBM stemness. In particular, we describe the Sox developmental transcription factor family; H3K79 methylation by DOT1L; and ufmylation stress responsiveness programs as essential for GBM stemness. Additionally, we find mechanisms of temozolomide resistance and sensitivity that could lead to combination strategies with this standard of care treatment. By reaching beyond static genome analysis of bulk tumors, with a genome wide functional approach, we dive deep into a broad range of biological processes to provide new understanding of GBM growth and treatment resistance.\n\nSignificanceGlioblastoma (GBM) remains an incurable disease despite an increasingly thorough depth of knowledge of the genomic and epigenomic alterations of bulk tumors. Evidence from multiple approaches support that GBM reflects an aberrant developmental hierarchy, with GBM stem cells (GSCs), fueling tumor growth and invasion. The properties of this tumor subpopulation may also in part explain treatment resistance and disease recurrence. Unfortunately, we still have a limited knowledge of the molecular circuitry of these cells and progress has been slow as we have not been able, until recently, to interrogate function at the genome-wide scale. Here, using parallel genome-wide CRISPR-Cas9 screens, we identify the essential genes for GSC growth. Further, by screening in the presence of low and high dose temozolomide, we identify mechanisms of drug resistance and sensitivity. These functional screens in patient derived cells reveal new aspects of GBM biology and identify a diversity of actionable targets such as genes governing stem cell traits, epigenome regulation and the response to stress stimuli.

cancer biology

Cicada endosymbionts have tRNAs that are correctly processed despite having genomes that do not encode all of the tRNA processing machinery

Gene loss and genome reduction are defining characteristics of nutritional endosymbiotic bacteria. In extreme cases, even essential genes related to core cellular processes such as replication, transcription, and translation are lost from endosymbiont genomes. Computational predictions on the genomes of the two bacterial symbionts of the cicada Diceroprocta semicincta, \"Candidatus Hodgkinia cicadicola\" (Alphaproteobacteria) and \"Ca. Sulcia muelleri\" (Betaproteobacteria), find only 26 and 16 tRNA, and 15 and 10 aminoacyl tRNA synthetase genes, respectively. Furthermore, the original \"Ca. Hodgkinia\" genome annotation is missing several essential genes involved in tRNA processing, such as RNase P and CCA tRNA nucleotidyltransferase, as well as several RNA editing enzymes required for tRNA maturation. How \"Ca. Sulcia\" and \"Ca. Hodgkinia\" preform basic translation-related processes without these genes remains unknown. Here, by sequencing eukaryotic mRNA and total small RNA, we show that the limited tRNA set predicted by computational annotation of \"Ca. Sulcia\" and \"Ca. Hodgkinia\" is likely correct. Furthermore, we show that despite the absence of genes encoding tRNA processing activities in the symbiont genomes, symbiont tRNAs have correctly processed 5 and 3 ends, and seem to undergo nucleotide modification. Surprisingly, we find that most \"Ca. Hodgkinia\"and \"Ca. Sulcia\" tRNAs exist as tRNA halves. Finally, and in contrast with other related insects, we show that cicadas have experienced little horizontal gene transfer that might complement the activities missing from the endosymbiont genomes. We conclude that \"Ca. Sulcia\" and \"Ca. Hodgkinia\" tRNAs likely function in bacterial translation, but require host-encoded enzymes to do so.

evolutionary biology

Population assignment from cancer genome profiling data

For a variety of human malignancies, incidence, treatment efficacy and overall prognosis show considerable variation between different populations and ethnic groups. Disentangling the effects related to particular population backgrounds can help in both understanding cancer biology and in tailoring therapeutic interventions. Because self-reported or inferred patient data can be incomplete or misleading due to migration and genomic admixture, a data-driven ancestry estimation should be preferred. While algorithms to analyze ancestry structure from healthy individuals have been developed, an easy-to-use tool to assign population groups based on genotyping data from SNP profiles is still missing and benchmarking for the validity of population assignment strategy for aberrant cancer genomes was not tested.\n\nWe benchmarked the consistency and accuracy of cross-platform population assignment. We also demonstrated its high accuracy to process unaltered as well as cancer genomes. Despite widespread and extensive somatic mutations of cancer profiling data, population assignment consistency between germline and highly mutated samples from cancer patients reached of 97% and 92% for assignment into 5 and 26 populations re-spectively. Comparison of our benchmarked results with self-reported meta-data estimated a matching rate between 88% to 92%. Despite a relatively high matching rate, the ethnicity labels indicated in meta-data are vague compared to the standardized output from our tool.\n\nWe have developed a bioinformatics tool to assign the populations from genome profiling data and validated its performance in healthy as well as aberrant cancer genomes. It is ready-to-use for genotyping data from nine commercial SNP array platforms or sequencing data. This tool is effective to scrutinize the population structure in cancer genomes and provides better measure to integrate genotyping data from various platforms instead of self-reported information. It will facilitate research on interplay between ethnicity related genetic background and molecular patterns in cancer entities and disentangling possible hereditary contributions.\n\nThe docker image of the tool is provided in DockerHub as \"baudisgroup/snp2pop\".

bioinformatics

Examination of Australian Streptococcus suis Isolates From Clinically Affected Pigs in a Global Context and the Genomic Characterisation of ST1 as a Predictor of Virulence

Streptococcus suis is a major zoonotic pathogen that causes severe disease in both humans and pigs. In this study, we investigated S. suis from 148 cases of clinical disease in pigs from 46 pig herds over a period of seven years. These isolates underwent whole genome sequencing, genome analysis and antimicrobial susceptibility testing. Genome sequence data of Australian isolates was compared at the core genome level to clinical isolates from overseas. Results demonstrated eight predominant multi-locus sequence types and two major cps gene types (cps2 and 3). At the core genome level Australian isolates clustered predominantly within one large clade consisting of isolates from the UK, Canada and North America. In particular, serotype 2 MLST25 strains were very closely associated with Canadian and North American strains. A very small proportion of Australian swine isolates (5%) were phylogenetically associated with south-east Asian and UK isolates, many of which were classified as causing systemic disease, and derived from cases of human and swine disease. In addition, we show that ST1 clones carry a constellation of putative virulence genes not present in other Australian STs, and that this is mirrored in overseas ST1 clones. Based on this dataset we provide a comprehensive outline of the current S. suis clones associated with disease in Australian pigs and their global context, and discuss the implications this has on antimicrobial therapy, potential vaccine candidates and public health.\n\nImportanceIn this study, we examine in detail, the genomic characteristics of 148 Streptococcus suis isolates from clinically diseased Australian pigs. We report the antimicrobial susceptibility profiles, virulence gene analysis and relationship to isolates from other regions of the world. We also demonstrate that ST1 clones, regardless of serotype, carry a large array of putative virulence genes while maintaining a small total gene content. This compilation of data has major ramifications for vaccine development, and refines the understanding of the distribution of various strains of this potentially-fatal zoonotic agent in the global pig industry

microbiology

SonHi-C: a set of non-procedural approaches for predicting 3D genome organization from Hi-C data

1BackgroundMany computational methods have been developed that leverage the results from biological experiments (such as Hi-C) to infer the 3D organization of the genome. Formally, this is referred to as the 3D genome reconstruction problem (3D-GRP). None of the existing methods for solving the 3D-GRP have utilized a non-procedural programming approach (such as constraint programming or integer programming) despite the established advantages and successful applications of such approaches for predicting the 3D structure of other biomolecules. Our objective was to develop a set of mathematical models and corresponding non-procedural implementations for solving the 3D-GRP to realize the same advantages.\n\nResultsWe present a set of non-procedural approaches for predicting 3D genome organization from Hi-C data (collectively referred to as SonHi-C and pronounced \"sonic\"). Specifically, this set is comprised of three mathematical models based on constraint programming (CP), graph matching (GM) and integer programming (IP). All of the mathematical models were implemented using non-procedural languages and tested with Hi-C data from Schizosaccharomyces pombe (fission yeast). The CP implementation could not optimally solve the problem posed by the fission yeast data after several days of execution time. The GM and IP implementations were able to predict a 3D model of the fission yeast genome in 1.088 and 294.44 seconds, respectively. These 3D models were then biologically validated through literature search which verified that the predictions were able to recapitulate key documented features of the yeast genome.\n\nConclusionsOverall, the mathematical models and programs developed here demonstrate the power of non-procedural programming and graph theoretic techniques for quickly and accurately modelling the 3D genome from Hi-C data. Additionally, they highlight the practical differences observed when differing non-procedural approaches are utilized to solve the 3D-GRP.

bioinformatics

A Multiplexed DNA FISH strategy for Assessing Genome Architecture in C. elegans

Eukaryotic DNA is highly organized within nuclei and this genomic organization is important for genome function. Fluorescent in situ hybridization (FISH) approaches allow the 3D architecture of genomes to be visualized. Scalable FISH technologies, which can be applied to whole animals, are needed to help unravel how genomic architecture regulates, or is regulated by, development, growth, reproduction, and aging. Here, we describe a multiplexed DNA FISH Oligopaint library that targets the entire C. elegans genome at chromosome, three megabase, and 500 kb scales. We describe a hybridization strategy that provides flexibility to DNA FISH experiments by coupling a single primary probe synthesis reaction to dye conjugated detection oligos via bridge oligos, eliminating the time and cost typically associated with labeling probe sets for individual DNA FISH experiments. The approach allows visualization of genome organization at varying scales in all/most cells across all stages of development in an intact animal model system.

cell biology

A draft reference genome sequence for Scutellaria baicalensis Georgi

Scutellaria baicalensis Georgi is an important medicinal plant used worldwide. Information about the genome of this species is important for scientists studying the metabolic pathways that synthesise the bioactive compounds in this plant. Here, we report a draft reference genome sequence for S. baicalensis obtained by a combination of Illumina and PacBio sequencing, which was assembled using 10 X Genomics and Hi-C technologies. We assembled 386.63 Mb of the 408.14 Mb genome, amounting to about 94.73% of the total genome size, and the sequences were anchored onto 9 pseudochromosomes with a super-N50 of 33.2 Mb. The reference genome sequence of S. baicalensis offers an important foundation for understanding the biosynthetic pathways for bioactive compounds in this medicinal plant and for its improvement through molecular breeding.

plant biology

Comparative genomics and divergence time estimation of the anaerobic fungi in herbivorous mammals

The anaerobic gut fungi (AGF) or Neocallimastigomycota inhabit the rumen and alimentary tract of herbivorous mammals, where they play an important role in the degradation of plant fiber. Comparative genomic and phylogenomic analysis of the AGF has long been hampered by their fastidious growth pattern as well as their large and AT-biased genomes. We sequenced 21 AGF transcriptomes and combined them with 5 available genome sequences of AGF taxa to explore their evolutionary relationships, time their divergence, and characterize patterns of gene gain/loss associated with their evolution. We estimate that the most recent common ancestor of the AGF diverged 66 ({+/-}10) million years ago, a timeframe that coincides with the evolution of grasses (Poaceae), as well as the mammalian transition from insectivory to herbivory. The concordance of these independently estimated ages of AGF evolution, grasses evolution, and mammalian transition to herbivory suggest that AGF have been important in shaping the success of mammalian herbivory transition by improving the efficiency of energy acquisition from recalcitrant plant materials. Comparative genomics identified multiple lineage-specific genes and protein domains in the AGF, two of which were acquired from an animal host (galectin) and rumen gut bacteria (carbohydrate-binding domain) via horizontal gene transfer (HGT). Four of the bacterial derived "Cthe_2159" genes in AGF genomes also encode eukaryotic Pfam domains ("Atrophin-1", "eIF-3_zeta", "Nop14", and "TPH") indicating possible gene fusion events after the acquisition of "Cthe_2159" domain. A third AGF domain, plant-like polysaccharide lyase N-terminal domain ("Rhamnogal_lyase"), represents the first report from fungi that potentially aids AGF to degrade pectin. Analysis of genomic and transcriptomic sequences confirmed the presence and expression of these lineage-specific genes in nearly all AGF clades supporting the hypothesis that these laterally acquired and novel genes in fungi are likely functional. These genetic elements may contribute to the exceptional abilities of AGF to degrade plant biomass and enable metabolism of the rumen microbes and animal hosts.

evolutionary biology

Metagenomic profiling of ticks: identification of novel rickettsial genomes and detection of tick-borne canine parvovirus

BackgroundAcross the world, ticks act as vectors of human and animal pathogens. Ticks rely on bacterial endosymbionts, which often share close and complex evolutionary links with tick-borne pathogens. As the prevalence, diversity and virulence potential of tick-borne agents remain poorly understood, there is a pressing need for microbial surveillance of ticks as potential disease vectors.\n\nMethodology/Principal FindingsWe developed a two-stage protocol that includes 16S-amplicon screening of pooled samples of hard ticks collected from dogs, sheep and camels in Palestine, followed by shotgun metagenomics on individual ticks to detect and characterise tick-borne pathogens and endosymbionts. Two ticks isolated from sheep yielded an abundance of reads from the genus Rickettsia, which were assembled into draft genomes. One of the resulting genomes was highly similar to Rickettsia massiliae strain MTU5. Analysis of signature genes showed that the other represents the first genome sequence of the potential pathogen Candidatus Rickettsia barbariae. Ticks from a dog and a sheep yielded draft genome sequences of strains of the Coxiella-like endosymbiont Candidatus Coxeilla mudrowiae. A sheep tick yielded sequences from the sheep pathogen Anaplasma ovis, while Hyalomma ticks from camels yielded sequences belonging to Francisella-like endosymbionts. From the metagenome of a dog tick from Jericho, we generated a genome sequence of a canine parvovirus.\n\nSignificanceHere, we have shown how a cost-effective two-stage protocol can be used to detect and characterise tick-borne pathogens and endosymbionts. In recovering genome sequences from an unexpected pathogen (canine parvovirus) and a previously unsequenced pathogen (Candidatus Rickettsia barbariae), we demonstrate the open-ended nature of metagenomics. We also provide evidence that ticks can carry canine parvovirus, raising the possibility that ticks might contribute to the spread of this troublesome virus.\n\nAuthor SummaryWe have shown how DNA sequencing can be used to detect and characterise potentially pathogenic microorganisms carried by ticks. We surveyed hard ticks collected from domesticated animals across the West Bank territory of Palestine. All the ticks came from species that are also capable of feeding on humans. We detected several important pathogens, including two species of Rickettsia, the sheep pathogen Anaplasma ovis and canine parvovirus. These findings highlight the importance of hard ticks and the hazards they present for human and animal health in Palestine and the opportunities presented by high-throughput sequencing and bioinformatics analyses of DNA sequences in this setting.

microbiology