Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

MEBS, a software platform to evaluate large (meta)genomic collections according to their metabolic machinery: unraveling the sulfur cycle

BACKGROUNDThe increasing number of metagenomic and genomic sequences has dramatically improved our understanding of microbial diversity, yet our ability to infer metabolic capabilities in such datasets remains challenging.\n\nFINDINGSWe describe the Multigenomic Entropy Based Score pipeline (MEBS), a software platform designed to evaluate, compare and infer complex metabolic pathways in large omic datasets, including entire biogeochemical cycles. MEBS is open source and available through https://github.com/eead-csic-compbio/metagenome_Pfam_score. To demonstrate its use we modeled the sulfur cycle by exhaustively curating the molecular and ecological elements involved (compounds, genes, metabolic pathways and microbial taxa). This information was reduced to a collection of 112 characteristic Pfam protein domains and a list of complete-sequenced sulfur genomes. Using the mathematical framework of relative entropy (H), we quantitatively measured the enrichment of these domains among sulfur genomes. The entropy of each domain was used to both: build up a final score that indicates whether a (meta)genomic sample contains the metabolic machinery of interest and to propose marker domains in metagenomic sequences such as DsrC (PF04358). MEBS was benchmarked with a dataset of 2,107 non-redundant microbial genomes from RefSeq and 935 metagenomes from MG-RAST. Its performance, reproducibility, and robustness were evaluated using several approaches, including random sampling, linear regression models, Receiver Operator Characteristic plots and the Area Under the Curve metric (AUC). Our results support the broad applicability of this algorithm to accurately classify (AUC=0.985) hard to culture genomes (e.g., Candidatus Desulforudis audaxviator), previously characterized ones and metagenomic environments such as hydrothermal vents, or deep-sea sediment.\n\nCONCLUSIONSOur benchmark indicates that an entropy-based score can capture the metabolic machinery of interest and be used to efficiently classify large genomic and metagenomic datasets, including uncultivated/unexplored taxa

bioinformatics

Sweeping genomic remodeling through repeated selection of alternatively adapted haplotypes occurs in the first decades after marine stickleback colonize new freshwater ponds

Heterogeneous genetic divergence can accumulate across the genome when populations adapt to different habitats while still exchanging alleles. How long does diversification take and how much of the genome is affected? When divergence occurs in parallel from standing genetic variation, how often are the same haplotypes used? We explore these questions using RAD-seq genotyping data, and show that broad-scale genomic re-patterning, fueled by standing variation, can emerge in just dozens of generations in replicate natural populations of threespine stickleback fish (Gasterosteus aculeatus). After the catastrophic 1964 Alaskan earthquake, marine stickleback colonized newly created ponds on seismically uplifted islands. We find that freshwater fish in these young ponds differ from their marine ancestors across the same genomic segments previously shown to have diverged in much older lake populations. Outside of these core divergent regions the genome shows no population structure across the ocean-freshwater divide, consistent with strong local selection acting in alternative environments on stickleback populations still connected by significant gene flow.Reinforcing this inference, a majority of divergent haplotypes that are at high frequency in ponds are shared across independent freshwater populations and are detectable, at low frequency, in the sea even across great geographic distances. Building upon previous work in this model species for population genomics, our data suggest that a long history of divergent selection and gene flow across stickleback in oceanic and freshwater habitats has created balanced polymorphism in large genomic blocks of alternatively adapted DNA sequences, ultimately stoking - and potentially channeling - rapid, parallel evolution.

evolutionary biology

Somatic inactivating PTPRJ mutations and dysregulated pathways identified in canine melanoma by integrated comparative genomic analysis

Canine malignant melanoma, a significant cause of mortality in domestic dogs, is a powerful comparative model for human melanoma, but little is known about its genetic etiology. We mapped the genomic landscape of canine melanoma through multi-platform analysis of 37 tumors (31 mucosal, 3 acral, 2 cutaneous, and 1 uveal) and 17 matching constitutional samples including long- and short-insert whole genome sequencing, RNA sequencing, array comparative genomic hybridization, single nucleotide polymorphism array, and targeted Sanger sequencing analyses. We identified novel predominantly truncating mutations in the putative tumor suppressor gene PTPRJ in 19% of cases. No BRAF mutations were detected, but activating RAS mutations (24% of cases) occurred in conserved hotspots in all cutaneous and acral and 13% of mucosal subtypes. MDM2 amplifications (24%) and TP53 mutations (19%) were mutually exclusive. Additional low-frequency recurrent alterations were observed amidst low point mutation rates, an absence of ultraviolet light mutational signatures, and an abundance of copy number and structural alterations. Mutations that modulate cell proliferation and cell cycle control were common and highlight therapeutic axes such as MEK and MDM2 inhibition. This mutational landscape resembles that seen in BRAF wild-type and sun-shielded human melanoma subtypes. Overall, these data inform biological comparisons between canine and human melanoma while suggesting actionable targets in both species.\n\nAUTHOR SUMMARYMelanoma, an aggressive cancer arising from transformed melanocytes, commonly occurs in pet dogs. Unlike human melanoma, which most often occurs in sun-exposed cutaneous skin, canine melanoma typically arises in sun-shielded oral mucosa. Clinical features of canine melanoma resemble those of human melanoma, particularly the less common sun-shielded human subtypes. However, whereas the genomic basis of diverse human melanoma subtypes is well understood, canine melanoma genomics remain poorly defined. Similarly, although diverse new treatments for human melanoma based on a biologic disease understanding have recently shown dramatic improvements in outcomes for these patients, treatments for canine melanoma are limited and outcomes remain universally poor. Detailing the genomic basis of canine melanoma thus provides untapped potential for improving the lives of pet dogs while also helping to establish canine melanoma as a comparative model system for informing human melanoma biology and treatment. In order to better define the genomic landscape of canine melanoma, we performed multi-platform characterization of 37 tumors. Our integrated analysis confirms that these tumors commonly contain mutations in canine orthologs of human cancer genes such as RAS, MDM2, and TP53 as well mutational patterns that share important similarities with human melanoma subtypes. We have also found a new putative cancer gene, PTPRJ, frequently mutated in canine melanoma. These data will guide additional biologic and therapeutic studies in canine melanoma while framing the utility of comparative studies of canine and human cancers more broadly.

cancer biology

Evidence of independent acquisition and adaption of ultra-small bacteria to human hosts across the highly diverse yet reduced genomes of the phylum Saccharibacteria

Recently, we discovered that a member of the Saccharibacteria/TM7 phylum (strain TM7x) isolated from the human oral cavity, has an ultra-small cell size (200-300nm), a highly reduced genome (705 Kbp) with limited de novo biosynthetic capabilities, and a very novel lifestyle as an obligate epibiont on the surface of another bacterium 1. There has been considerable interest in uncultivated phyla, particularly those that are now classified as the proposed candidate phyla radiation (CPR) reported to include 35 or more phyla and are estimated to make up nearly 15% of the domain Bacteria. Most members of the larger CPR group share genomic properties with Saccharibacteria including reduced genomes (<1Mbp) and lack of biosynthetic capabilities, yet to date, strain TM7x represents the only member of the CPR that has been cultivated and is one of only three CPR routinely detected in the human body. Through small subunit ribosomal RNA (SSU rRNA) gene surveys, members of the Saccharibacteria phylum are reported in many environments as well as within a diversity of host species and have been shown to increase dramatically in human oral and gut diseases. With a single copy of the 16S rRNA gene resolved on a few limited genomes, their absolute abundance is most often underestimated and their potential role in disease pathogenesis is therefore underappreciated. Despite being an obligate parasite dependent on other bacteria, six groups (G1-G6) are recognized using SSU rRNA gene phylogeny in the oral cavity alone. At present, only genomes from the G1 group, which includes related and remarkably syntenic environmental and human oral associated representatives1, have been uncovered to date. In this study we systematically captured the spectrum of known diversity in this phylum by reconstructing completely novel Class level genomes belonging to groups G3, G6 and G5 through cultivation enrichment and/or metagenomic binning from humans and mammalian rumen. Additional genomes for representatives of G1 were also obtained from modern oral plaque and ancient dental calculus. Comparative analysis revealed remarkable divergence in the host-associated members across this phylum. Within the human oral cavity alone, variation in as much as 70% of the genes from nearest oral clade (AAI 50%) as well as wide GC content variation is evident in these newly captured divergent members (G3, G5 and G6) with no environmental relatives. Comparative analyses suggest independent episodes of transmission of these TM7 groups into humans and convergent evolution of several key functions during adaptation within hosts. In addition, we provide evidence from in vivo collected samples that each of these major groups are ultra-small in size and are found attached to larger cells.

microbiology

THEA: A novel approach to gene identification in phage genomes

MotivationCurrently there are no tools specifically designed for annotating genes in phages. Several tools are available that have been adapted to run on phage genomes, but due to their underlying design they are unable to capture the full complexity of phage genomes. Phages have adapted their genomes to be extremely compact, having adjacent genes that overlap, and genes completely inside of other longer genes. This non-delineated genome structure makes it difficult for gene prediction using the currently available gene annotators. Here we present THEA (The Algorithm), a novel method for gene calling specifically designed for phage genomes. While the compact nature of genes in phages is a problem for current gene annotators, we exploit this property by treating a phage genome as a network of paths: where open reading frames are favorable, and overlaps and gaps are less favorable, but still possible. We represent this network of connections as a weighted graph, and use graph theory to find the optimal path.\n\nResultsWe compare THEA to other gene callers by annotating a set of 2,133 complete phage genomes from GenBank, using THEA and the three most popular gene callers. We found that the four programs agree on 82% of the total predicted genes, with THEA predicting significantly more genes than the other three. We searched for these extra genes in both GenBanks non-redundant protein database and sequence read archive, and found that they are present at levels that suggest that these are functional protein coding genes.\n\nAvailability and ImplementationThe source code and all files can be found at: https://github.com/deprekate/THEA\n\nContactKatelyn McNair: deprekate@gmail.com

bioinformatics

Crossbrowse: A versatile genome browser for visualizing comparative experimental data

The recent beyond-exponential growth in diverse collections of deep sequencing datasets creates enormous opportunities for discovery, concomitant with new challenges for displaying and interpreting these data. Notably, the availability of scores of whole genome sequences in multiple species clades enables comparative studies of functional elements. However, current genome browsers do not permit effective visualization of multigenome experimental data. Here, we present CrossBrowse, a standalone desktop application for displaying and browsing cross-species genomic datasets. We utilize data standards and graphic representation of popular browsers, and incorporate an intuitive graphical visualization of genome synteny that facilitates and drives human interrogation of comparative data. Our platform permits users with minimal informatics capacity to select arbitrary sets of genomes for display, upload and configure multiple datasets, and interact with vertebrate-sized genomic datasets in real-time. We illustrate the utility of CrossBrowse with interrogation of comparative invertebrate and mammalian datasets that provide insights into diverse aspects of transcriptional and post-transcriptional regulation. Of note, we show examplars of both preservation and divergence of functional elements that cannot be inferred from sequence alignments alone. Moreover, we demonstrate how inspection of primary data using CrossBrowse exposes an artifact in a typical strategy for assigning species-specific functional elements, and drives the implementation of an improved computational strategy. We anticipate that CrossBrowse will greatly foster user-based discovery within multispecies genomic datasets, and inform their bioinformatic interpretation.

bioinformatics

Genome-specific histories of divergence and introgression between an allopolyploid unisexual salamander lineage and two sexual species

Quantifying genetic introgression between sexual species and polyploid lineages traditionally thought to be asexual is an important step in understanding what factors drive the longevity of putatively asexual groups. However, the presence of multiple distinct subgenomes within a single lineage provides a significant logistical challenge to evaluating the origin of genetic variation in most polyploids. Here, we capitalize on three recent innovations--variation generated from ultraconserved elements (UCEs), bioinformatic techniques for assessing variation in polyploids, and model-based methods for evaluating historical gene flow--to measure the extent and tempo of introgression over the evolutionary history of an allopolyploid lineage of all-female salamanders and two ancestral sexual species. We first analyzed variation from more than a thousand UCEs using a reference mapping method developed for polyploids to infer subgenome specific patterns of variation in the all-female lineage. We then used PHRAPL to choose between sets of historical models that reflected different patterns of introgression and divergence between the genomes of the parental species and the same genomes found within the polyploids. Our analyses support a scenario in which the genomes sampled in unisexuals salamanders were present in the lineage [~]3.4 million years ago, followed by an extended period of divergence from their parental species. Recent secondary introgression has occurred at different times between each sexual species and their representative genomes within the unisexuals during the last 500,000 years. Sustained introgression of sexual genomes into the unisexual lineage has been the defining characteristic of their reproductive mode, but this study provides the first evidence that unisexual genomes have also undergone long periods of divergence without introgression. Unlike other unisexual, sperm-dependent taxa in which introgression is rare, the alternating periods of divergence and introgression between unisexual salamanders and their sexual relatives could reveal the scenarios in which the influx of novel genomic material is favored and potentially explain why these salamanders are among the oldest described unisexual animals.

evolutionary biology

An ultra-dense haploid genetic map for evaluating the highly fragmented genome assembly of Norway spruce (Picea abies)

Norway spruce (Picea abies (L.) Karst.) is a conifer species of substanital economic and ecological importance. In common with most conifers, the P. abies genome is very large ([~]20 Gbp) and contains a high fraction of repetitive DNA. The current P. abies genome assembly (v1.0) covers approximately 60% of the total genome size but is highly fragmented, consisting of >10 million scaffolds. The genome annotation contains 66,632 gene models that are at least partially validated (www.congenie.org), however, the fragmented nature of the assembly means that there is currently little information available on how these genes are physically distributed over the 12 P. abies chromosomes. By creating an ultra-dense genetic linkage map, we anchored and ordered scaffolds into linkage groups, which complements the fine-scale information available in assembly contigs. Our ultra-dense haploid consensus genetic map consists of 21,056 markers derived from 14,336 scaffolds that contain 17,079 gene models (25.6% of the validated gene models) that we have anchored to the 12 linkage groups. We used data from three independent component maps, as well as comparisons with previously published Picea maps to evaluate the accuracy and marker ordering of the linkage groups. We demonstrate that approximately 3.8% of the anchored scaffolds and 1.6% of the gene models covered by the consensus map have likely assembly errors as they contain genetic markers that map to different regions within or between linkage groups. We further evaluate the utility of the genetic map for the conifer research community by using an independent data set of unrelated individuals to assess genome-wide variation in genetic diversity using the genomic regions anchored to linkage groups. The results show that our map is sufficiently dense to enable detailed evolutionary analyses across the P. abies genome.

genetics

Development of a joint evolutionary model for the genome and the epigenome

BackgroundInterspecies epigenome comparisons yielded functional information that cannot be revealed by genome comparison alone, begging for theoretical advances that enable principled analysis approaches. Whereas probabilistic genome evolution models provided theoretical foundation to comparative genomics studies, it remains challenging to extend DNA evolution models to epigenomes.\n\nResultsWe present an effort to develop ab initio evolution models for epigenomes, by explicitly expressing the joint probability of multispecies DNA sequences and histone modifications on homologous genomic regions. This joint probability is modeled as a mixture of four components representing four evolutionary hypotheses, namely dependence and independence of interspecies epigenomic variations to sequence mutations and to sequence insertions and deletions (indels). For model fitting, we implemented a maximum likelihood method by coupling downhill simplex algorithm with dynamic programming. Based on likelihood comparisons, the model can be used to infer whether interspecies epigenomic variations depend on mutation or indels in local genomic sequences. We applied this model to analyze DNase hypersensitive regions and spermatid H3K4me3 ChIP-seq data from human and rhesus macaque. Approximately 5.5% of homologous regions in the genomes exhibited H3K4me3 modification in either species, among which approximately 67% homologous regions exhibited sequence-dependent interspecies H3K4me3 variations. Mutations accounted for less sequence-dependent H3K4me3 variations than indels. Among transposon-mediated indels, ERV1 insertions and L1 insertions were most strongly associated with H3K4me3 gains and losses, respectively.\n\nConclusionThis work initiates a class of probabilistic evolution models that jointly model the genomes and the epigenomes, thus helps to bring evolutionary principles to comparative epigenomic studies.

bioinformatics

Illuminating the microbiome’s dark matter: a functional genomic toolkit for the study of human gut Actinobacteria

Despite the remarkable evolutionary and metabolic diversity found within the human microbiome, the vast majority of mechanistic studies focus on two phyla: the Bacteroidetes and the Proteobacteria. Generalizable tools for studying the other phyla are urgently needed in order to transition microbiome research from a descriptive to a mechanistic discipline. Here, we focus on the Coriobacteriia class within the Actinobacteria phylum, detected in the distal gut of 90% of adult individuals around the world, which have been associated with both chronic and infectious disease, and play a key role in the metabolism of pharmaceutical, dietary, and endogenous compounds. We established, sequenced, and annotated a strain collection spanning 14 genera, 8 decades, and 3 continents, with a focus on Eggerthella lenta. Genome-wide alignments revealed inconsistencies in the taxonomy of the Coriobacteriia for which amendments have been proposed. Re-sequencing of the E. lenta type strain from multiple culture collections and our laboratory stock allowed us to identify errors in the finished genome and to identify point mutations associated with antibiotic resistance. Analysis of 24 E. lenta genomes revealed an \"open\" pan-genome suggesting we still have not fully sampled the genetic and metabolic diversity within this bacterial species. Consistent with the requirement for arginine during in vitro growth, the core E. lenta genome included the arginine dihydrolase pathway. Surprisingly, glycolysis and the citric acid cycle was also conserved in E. lenta despite the lack of evidence for carbohydrate utilization. We identified a species-specific marker gene and validated a multiplexed quantitative PCR assay for simultaneous detection of E. lenta and specific genes of interest from stool samples. Finally, we demonstrated the utility of comparative genomics for linking variable genes to strain-specific phenotypes, including antibiotic resistance and drug metabolism. To facilitate the continued functional genomic analysis of the Coriobacteriia, we have deposited the full collection of strains in DSMZ and have written a general software tool (ElenMatchR) that can be readily applied to novel phenotypic traits of interest. Together, these tools provide a first step towards a molecular understanding of the many neglected but clinically-relevant members of the human gut microbiome.

microbiology

Using Core Genome Alignments to Assign Bacterial Species

With the exponential increase in the number of bacterial taxa with genome sequence data, a new standardized method is needed to assign bacterial species designations using genomic data that is consistent with the classically-obtained taxonomy. This is particularly acute for unculturable obligate intracellular bacteria like those in the Rickettsiales, where classical methods like DNA-DNA hybridization cannot be used to define species. Within the Rickettsiales, species designations have been applied inconsistently, often obfuscating the relationship between organisms and the context for experimental results. In this study, we generated core genome alignments for a wide range of genera with classically defined species, including Arcobacter, Caulobacter, Erwinia, Neisseria, Polaribacter, Ralstonia, Thermus, as well as genera within the Rickettsiales including Rickettsia, Orientia, Ehrlichia, Neoehrlichia, Anaplasma, eorickettsia, and Wolbachia. A core genome alignment sequence identity (CGASI) threshold of 96.8% was found to maximize the prediction of classically-defined species. Using the CGASI cutoff, the Wolbachia genus can be delineated into species that differ from the currently used supergroup designations, while the Rickettsia genus is delineated into nine species, as opposed to the current 27 species. Additionally, we find that core genome alignments cannot be constructed between genomes belonging to different genera, establishing a bacterial genus cutoff that suggests the need to create new genera from the Anaplasma and Neorickettsia. By using core genome alignments to assign taxonomic designations, we aim to provide a high-resolution, robust method for bacterial nomenclature that is aligned with classically-obtained results.

microbiology

The genomic landscape of recombination rate variation in Chlamydomonas reinhardtii reveals effects of linked selection

Recombination confers a major evolutionary advantage by breaking up linkage disequilibrium (LD) between harmful and beneficial mutations and facilitating selection. Here, we use genome-wide patterns of LD to infer fine-scale recombination rate variation in the genome of the model green alga Chlamydomonas reinhardtii and estimate rates of LD decay across the entire genome. We observe recombination rate variation of up to two orders of magnitude, finding evidence of recombination hotspots playing a role in the genome. Recombination rate is highest just upstream of genic regions, suggesting the preferential targeting of recombination breakpoints in promoter regions. Furthermore, we observe a positive correlation between GC content and recombination rate, suggesting a role for GC-biased gene conversion or selection on base composition within the GC-rich genome of C. reinhardtii. We also find a positive relationship between nucleotide diversity and recombination, consistent with widespread influence of linked selection in the genome. Finally, we use estimates of the effective rate of recombination to calculate the rate of sex that occurs in natural populations of this important model microbe, estimating a sexual cycle roughly every 770 generations. We argue that the relatively infrequent rate of sex and large effective population size creates an population genetic environment that increases the influence of linked selection on the genome.

evolutionary biology

Divergent selection and drift shape the genomes of two avian sister species spanning a saline-freshwater ecotone

The role of species divergence due to ecologically-based divergent selection - or ecological speciation - in generating and maintaining biodiversity is a central question in evolutionary biology. Comparison of the genomes of phylogenetically related taxa spanning a selective habitat gradient enables discovery of divergent signatures of selection and thereby provides valuable insight into the role of divergent ecological selection in speciation. Tidal marsh ecosystems provide tractable opportunities for studying organisms adaptations to selective pressures that underlie ecological divergence. Sharp environmental gradients across the saline-freshwater ecotone within tidal marshes present extreme adaptive challenges to terrestrial vertebrates. Here we sequence 20 whole genomes of two avian sister species endemic to tidal marshes - the Saltmarsh Sparrow (Ammodramus caudacutus) and Nelsons Sparrow (A. nelsoni) - to evaluate the influence of selective and demographic processes in shaping genome-wide patterns of divergence. Genome-wide divergence between these two recently diverged sister species was notably high (genome-wide FST = 0.32). Against a background of high genome-wide divergence, regions of elevated divergence were widespread throughout the genome, as opposed to focused within islands of differentiation. These patterns may be the result of genetic drift acting during past tidal march colonization events in addition to divergent selection to different environments. We identified several candidate genes that exhibited elevated divergence between Saltmarsh and Nelsons sparrows, including genes linked to osmotic regulation, circadian rhythm, and plumage melanism - all putative candidates linked to adaptation to tidal marsh environments. These findings provide new insights into the roles of divergent selection and genetic drift in generating and maintaining biodiversity.

evolutionary biology

Seed Genome Hypomethylated Regions Are Enriched In Transcription Factor Genes

The precise mechanisms that control gene activity during seed development remain largely unknown. Previously, we showed that several genes essential for seed development, including those encoding storage proteins, fatty acid biosynthesis enzymes, and transcriptional regulators, such as ABI3 and FUS3, are located within hypomethylated regions of the soybean genome. These hypomethylated regions are similar to the DNA methylation valleys (DMVs), or canyons, found in mammalian cells. Here, we address the question of the extent to which DMVs are present within seed genomes, and what role they might play in seed development. We scanned soybean and Arabidopsis seed genomes from post-fertilization through dormancy and germination for regions that contain < 5% or < 0.4% bulk methylation in CG-, CHG-, and CHH-contexts over all developmental stages. We found that DMVs represent extensive portions of seed genomes, range in size from 5 to 76 kb, are scattered throughout all chromosomes, and are hypomethylated throughout the plant life cycle. Significantly, DMVs are enriched greatly in transcription factor genes, and other developmental genes, that play critical roles in seed formation. Many DMV genes are regulated with respect to seed stage, region, and tissue - and contain H3K4me3, H3K27me3, or bivalent marks that fluctuate during development. Our results indicate that DMVs are a unique regulatory feature of both plant and animal genomes, and that a large number of seed genes are regulated in the absence of methylation changes during development - probably by the action of specific transcription factors and epigenetic events at the chromatin level.\n\nSignificanceWe scanned soybean and Arabidopsis seed genomes for hypomethylated regions, or DNA Methylation Valleys (DMVs), present in mammalian cells. A significant fraction of seed genomes contain DMV regions that have < 5% bulk DNA methylation, or, in many cases, no detectable DNA methylation. Methylation levels of seed DMVs do not vary detectably during seed development with respect to time, region, and tissue, and are present prior to fertilization. Seed DMVs are enriched in transcription factor genes and other genes critical for seed development, and are also decorated with histone marks that fluctuate with developmental stage, resembling in significant ways their animal counterparts. We conclude that many genes playing important roles in seed formation are regulated in the absence of detectable DNA methylation events, and suggest that selective action of transcriptional activators and repressors, as well as chromatin epigenetic events play important roles in making a seed - particularly embryo formation.

plant biology

TGFam-Finder: An optimal solution for target-gene family annotation in eukaryotic genomes

Whole genome annotation errors that omit essential protein-coding genes hinder further research. We developed Target Gene Family Finder (TGFam-Finder), an optimal tool for structural annotation of protein-coding genes containing target domain(s) of interest in eukaryotic genomes. Large-scale re-annotation of 100 publicly available eukaryotic genomes led to the discovery of essential genes that were missed in previous annotations. An average of 117 (346%) and 148 (45%) additional FAR1 and NLR genes were newly identified in 50 plant genomes. Furthermore, 117 (47%) additional C2H2 zinc finger genes were detected in 50 animal genomes including human and mouse. Accuracy of the newly annotated genes was validated by RT-PCR and cDNA sequencing in human, mouse and rice. In the human genome, 26 newly annotated genes were identical with known functional genes. TGFam-Finder along with the new gene models provide an optimized platform for unbiased functional and comparative genomics and comprehensive evolutionary study in eukaryotes.

bioinformatics

Using machine learning to predict antimicrobial minimum inhibitory concentrations and associated genomic features for nontyphoidal Salmonella

Nontyphoidal Salmonella species are the leading bacterial cause of food-borne disease in the United States. Whole genome sequences and paired antimicrobial susceptibility data are available for Salmonella strains because of surveillance efforts from public health agencies. In this study, a collection of 5,278 nontyphoidal Salmonella genomes, collected over 15 years in the United States, were used to generate XGBoost-based machine learning models for predicting minimum inhibitory concentrations (MICs) for 15 antibiotics. The MIC prediction models have average accuracies between 95-96% within {+/-} 1 two-fold dilution factor and can predict MICs with no a priori information about the underlying gene content or resistance phenotypes of the strains. By selecting diverse genomes for training sets, we show that highly accurate MIC prediction models can be generated with fewer than 500 genomes. We also show that our approach for predicting MICs is stable over time despite annual fluctuations in antimicrobial resistance gene content in the sampled genomes. Finally, using feature selection, we explore the important genomic regions identified by the models for predicting MICs. To date, this is one of the largest MIC modeling studies to be published. Our strategy for developing whole genome sequence-based models for surveillance and clinical diagnostics can be readily applied to other important human pathogens.

bioinformatics

AlleleHMM: a data-driven method to identify allele-specific differences in distributed functional genomic marks

How DNA sequence variation influences gene expression remains poorly understood. Diploid organisms have two homologous copies of their DNA sequence in the same nucleus, providing a rich source of information about how genetic variation affects a wealth of biochemical processes. However, few computational methods have been developed to discover allele-specific differences in functional genomic data. Existing methods either treat each SNP independently, limiting statistical power, or combine SNPs across gene annotations, preventing the discovery of allele specific differences in unexpected genomic regions. Here we introduce AlleleHMM, a new computational method to identify blocks of neighboring SNPs that share similar allele-specific differences in mark abundance. AlleleHMM uses a hidden Markov model to divide the genome among three hidden states based on allele frequencies in genomic data: a symmetric state (state S) which shows no difference between alleles, and regions with a higher signal on the maternal (state M) or paternal (state P) allele. AlleleHMM substantially outperformed naive methods using both simulated and real genomic data, particularly when input data had realistic levels of overdispersion. Using PRO-seq data, AlleleHMM identified thousands of allele specific blocks of transcription in both coding and non-coding genomic regions. AlleleHMM is a powerful tool for discovering allele-specific regions in functional genomic datasets.

bioinformatics

Expanded Analysis of the Pantoea stewartii subsp. stewartii DC283 Complete Genome Reveals Plasmid-borne Virulence Factors

Pantoea stewartii subsp. stewartii, a Gram-negative proteobacterium, causes Stewarts wilt disease in corn. Bacterial transmission to plants occurs primarily via the corn flea beetle insect vector, which is native to North America. P. stewartii DC283 is the wild-type reference strain most used to study pathogenesis. Previously the complete genome of P. stewartii was released. Here, the method whereby the genome was assembled is described in greater detail. Data from a matepair library preparation with 3.5 kilobase insert size and high-throughput sequencing from the MiSeq Illumina platform, together with the available incomplete genome sequence of AHIE00000000.1 (containing 65 contigs) was used. This work resulted in the complete assembly of one circular chromosome, ten circular plasmids and one linear phage from P. stewartii DC283. A high number of sequences encoding repetitive transposases (> 400) were found in the complete genome. The separation of plasmids from genomic DNA revealed that two Type III secretion systems in P. stewartii DC283 are located on two separate mega-plasmids. Interestingly, the assembly identified a previously unknown 66-kb region in a location interior to a contig in the previous reference genome. Overall, a novel approach was successfully utilized to fully assemble a prokaryotic genome that contains large numbers of repetitive sequences and multiple plasmids, which resulted in some interesting biological findings.

bioinformatics