Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Genome-wide scans of selection highlight the impact of biotic and abiotic constraints in natural populations of the model grass Brachypodium distachyon.

Grasses are essential plants for ecosystem functioning. Quantifying the selective pressures that act on natural variation in grass species is therefore essential regarding biodiversity maintenance. In this study, we investigate the selection pressures that act on two distinct populations of the grass model Brachypodium distachyon without prior knowledge about the traits under selection. We took advantage of whole-genome sequencing data produced for 44 natural accessions of B. distachyon and used complementary genome-wide scans of selection (GWSS) methods to detect genomic regions under balancing and positive selection. We show that selection is shaping genetic diversity at multiple temporal and spatial scales in this species and affects different genomic regions across the two populations. Gene Ontology annotation of candidate genes reveals that pathogens may constitute important factors of positive and balancing selection in Brachypodium distachyon. We eventually cross-validated our results with QTL data available for leaf-rust resistance in this species and demonstrate that, when paired with classical trait mapping, GWSS can help pinpointing candidate genes for further molecular validation. Thanks to a near-base perfect reference genome and the large collection of freely available natural accessions collected across its natural range, B. distachyon appears as a prime system for studies in ecology, population genomics and evolutionary biology.

evolutionary biology

Integration and analysis of CPTAC proteomics data in the context of cancer genomics in the cBioPortal

The Clinical Proteomic Tumor Analysis Consortium (CPTAC) has produced extensive mass spectrometry based proteomics data for selected breast, colon and ovarian tumors from The Cancer Genome Atlas (TCGA). We have incorporated the CPTAC proteomics data into the cBioPotal to support easy exploration and integrative analysis of these proteomic datasets in the context of the clinical and genomics data from the same tumors. cBioPortal is an open source platform for exploring, visualizing, and analyzing multi-dimensional cancer genomics and clinical data. The public instance of the cBioPortal (http://cbioportal.org/) hosts more than 100 cancer genomics studies including all of the data from TCGA. Its biologist-friendly interface provides many rich analysis features, including a graphical summary of gene-level data across multiple platforms, correlation analysis between genes or other data types, survival analysis, and network visualization. Here, we present the integration of the CPTAC mass spectrometry based proteomics data into the cBioPortal, consisting of 77 breast, 95 colorectal, and 174 ovarian tumors that already have been profiled by TCGA for mutations, copy number alterations, gene expression, and DNA methylation. As a result, the CPTAC data can now be easily explored and analyzed in the cBioPortal in the context of clinical and genomics data. By integrating CPTAC data into cBioPortal, limitations of TCGA proteomics array data can be overcome while also providing a user-friendly web interface, a web API and an R client to query the mass spectrometry data together with genomic, epigenomic, and clinical data.

cancer biology

Mind your gaps: Overlooking assembly gaps confounds statistical testing in genome analysis

BackgroundThe difficulties associated with sequencing and assembling some regions of the DNA sequence result in gaps in the reference genomes that are typically represented as stretches of Ns. Although the presence of assembly gaps causes a slight reduction in the mapping rate in many experimental settings, that does not invalidate the typical statistical testing comparing read count distributions across experimental conditions. However, we hypothesize that not handling assembly gaps in the null model may confound statistical testing of co-localization of genomic features.\n\nResultsFirst, we performed a series of explorative analyses to understand whether and how the public genomic tracks intersect the assembly gaps track (hg19). The findings rightly confirm that the genomic regions in public genomic tracks intersect very little with assembly gaps and the intersection was observed only at the beginning and end regions of the assembly gaps rather than covering the whole gap sizes. Further, we simulated a set of query and reference genomic tracks in a way that nullified any dependence between them to test our hypothesis that not avoiding assembly gaps in the null model would result in spurious inflation of statistical significance. We then contrasted the distributions of test statistics and p-values of Monte Carlo simulation-based permutation tests that either avoided or not avoided assembly gaps in the null model when testing for significant co-localization between a pair of query and reference tracks. We observed that the statistical tests that did not account for the assembly gaps in the null model resulted in a distribution of the test statistic that is shifted to the right and a distribu tion of p-values that is shifted to the left (leading to inflated significance).\n\nConclusionOur results shows that not accounting for assembly gaps in statistical testing of co-localization analysis may lead to false positives and over-optimistic findings.

bioinformatics

Genome-wide selection scans integrated with association mapping reveal mechanisms of physiological adaptation across a salinity gradient in killifish

Adaptive divergence between marine and freshwater environments is important in generating phyletic diversity within fishes, but the genetic basis of adaptation to freshwater habitats remains poorly understood. Available approaches to detect adaptive loci include genome scans for selection, but these can be difficult to interpret because of incomplete knowledge of the connection between genotype and phenotype. In contrast, genome wide association studies (GWAS) are powerful tools for linking genotype to phenotype, but offer limited insight into the evolutionary forces shaping variation. Here, we combine GWAS and selection scans to identify loci important in the adaptation of complex physiological traits to freshwater environments. We focused on freshwater (FW)-native and brackish water (BW)-native populations of the Atlantic killifish (Fundulus heteroclitus) as well as a population that is a natural admixture of these two populations. We measured phenotypes for multiple physiological traits that differ between populations and that may contribute to adaptation across osmotic niches (salinity tolerance, hypoxia tolerance, metabolic rate, and body shape) and used a reduced representation approach for genome-wide genotyping. Our results show patterns of population divergence in physiological capabilities that are consistent with local adaptation. Selection scans between BW-native and FW-native populations identified genomic regions that presumably aect fitness between BW and FW environments, while GWAS revealed loci that contribute to variation for each physiological trait. There was substantial overlap in the genomic regions putatively under selection and loci associated with the measured physiological traits, suggesting that these phenotypes are important for adaptive divergence between BW and FW environments. Our analysis also implicates candidate genes likely involved in physiological capabilities, some of which validate a priori hypotheses. Together, these data provide insight into the mechanisms that enable diversification of fishes across osmotic boundaries.\n\nAuthor SummaryIdentifying the genes that underlie adaptation is important for understanding the evolutionary process, but this is technically challenging. We bring multiple lines of evidence to bear for identifying genes that underlie adaptive divergence. Specifically, we integrate genotype-phenotype association mapping with genome-wide scans for signatures of natural selection to reveal genes that underlie phenotypic variation and that are adaptive in populations of killifish that are diverging between marine and freshwater environments. Because adaptation is likely manifest in multiple physiological traits, we focus on hypoxia tolerance, salinity tolerance, and metabolic rate; traits that are divergent between marine and freshwater populations. We show that each of these phenotypes is evolving by natural selection between environments; genetic variants that contribute to variation in these physiological traits tend to be evolving by natural selection between marine and freshwater populations. Furthermore, one of our top candidate genes provides a mechanistic explanation for previous hypotheses that suggest the adaptive importance of cellular tight junctions. Together, these data demonstrate a powerful approach to identify genes involved in adaptation and help to reveal the mechanisms enabling transitions of fishes across osmotic boundaries.

evolutionary biology

Exploratory re-encoding of Yellow Fever Virus genome: analysis of in vitro and in vivo replicative phenotypes

Virus attenuation by genome re-encoding is a pioneering approach for generating effective live-attenuated vaccine candidates. Its core principle is to introduce a large number of synonymous substitutions into the viral genome to produce stable attenuation of the targeted virus. Introduction of large numbers of mutations has also been shown to maintain stability of the attenuated phenotype by lowering the risk of reversion and recombination of re-encoded genomes. Identifying mutations with low fitness cost is pivotal as this increases the number that can be introduced and generates more stable and attenuated viruses. Here, we sought to identify mutations with low deleterious impact on the in vivo replication and virulence of yellow fever virus (YFV). Following comparative bioinformatic analyses of flaviviral genomes, we categorized synonymous transition mutations according to their impact on CpG/UpA composition and secondary RNA structures. We then designed 17 re-encoded viruses with 100-400 synonymous mutations in the NS2A-to-NS4B coding region of YFV Asibi and Ap7M (hamster-adapted) genomes. Each virus contained a panel of synonymous mutations designed according to the above categorisation criteria. The replication and fitness characteristics of parent and re-encoded viruses were compared in vitro using cell culture competition experiments. In vivo laboratory hamster models were also used to compare relative virulence and immunogenicity characteristics. Most of the re-encoded strains showed no decrease in replicative fitness in vitro. However, they showed reduced virulence and, in some instances, decreased replicative fitness in vivo. Importantly, the most attenuated of the re-encoded strains induced robust, protective immunity in hamsters following challenge with Ap7M, a virulent virus. Overall, the introduction of transitions with no or a marginal increase in the number of CpG/UpA dinucleotides had the mildest impact on YFV replication and virulence in vivo. Thus, this strategy can be incorporated in procedures for the finely tuned creation of substantially re-encoded viral genomes.

microbiology

Kaptive Web: user-friendly capsule and lipopolysaccharide serotype prediction for Klebsiella genomes

As whole genome sequencing becomes an established component of the microbiologists toolbox, it is imperative that researchers, clinical microbiologists and public health professionals have access to genomic analysis tools for rapid extraction of epidemiologically and clinically relevant information. For the gram-negative hospital pathogens such as Klebsiella pneumoniae, initial efforts have focused on detection and surveillance of antimicrobial resistance genes and clones. However, with the resurgence of interest in alternative infection control strategies targeting Klebsiella surface polysaccharides, the ability to extract information about these antigens is increasingly important.\n\nHere we present Kaptive Web, an online tool for rapid typing of Klebsiella K and O loci, which encode the polysaccharide capsule and lipopolysaccharide O antigen, respectively. Kaptive Web enables users to upload and analyse genome assemblies in a web browser. Results can be downloaded in tabular format or explored in detail via the graphical interface, making it accessible for users at all levels of computational expertise.\n\nWe demonstrate Kaptive Webs utility by analysis of >500 K. pneumoniae genomes. We identify extensive K and O locus diversity among 201 genomes belonging to the carbapenemase- associated clonal group 258 (25 K and six O loci). Characterisation of a further 309 genomes indicates that such diversity is common among the multi-drug resistant clones and that these loci represent useful epidemiological markers for strain subtyping. These findings reinforce the need for rapid, reliable and accessible typing methods such as Kaptive Web.\n\nKaptive Web is available for use at kaptive.holtlab.net and source code is available at github.com/kelwyres/Kaptive-Web.

microbiology

Evolutionary insights into Bean common mosaic necrosis virus and Cowpea aphid borne mosaic virus using global isolates and thirteen new near complete genomes from Kenya

Plant viral diseases are one of the major limitations in legume production within sub Saharan Africa (SSA), as they account for up to 100 % in production losses within smallholder farms. In this study, field surveys were conducted in the western highlands of Kenya with viral symptomatic leaf samples collected. Subsequently, next-generation sequencing was carried out. The main aim was to gain insights into the selection pressure and evolutionary relationships of Bean common mosaic necrosis virus (BCMNV) and Cowpea aphid-borne mosaic virus (CABMV), within symptomatic common beans and cowpeas. Eleven near complete genomes of BCMNV and two for CABMV sequences were obtained from SSA. Bayesian phylogenomic analysis and tests for differential selection pressure within sites and across tree branches of the viral genomes was carried out. Three distinct well-supported clades were identified across the whole genome tree, and were in agreement with individual gene trees. Selection pressure analysis within sites and across phylogenetic branches suggested both viruses were evolving independently, but under strong purifying selection, with a slow evolutionary rate. These findings provide valuable insights on the evolution of BCMNV and CABMV genomes and their relationship to other viral genomes globally. These results will contribute greatly to the knowledge gap surrounding the phylogenomic relationship of these viruses, particularly for CABMV, for which there are few genome sequences available, and support the current breeding efforts towards resistance for BCMNV and CABMV.

evolutionary biology

Prediction of Optimal Growth Temperature using only Genome Derived Features

Optimal growth temperature is a fundamental characteristic of all living organisms. Knowledge of this temperature is central to the study the organism, the thermal stability and temperature dependent activity of its genes, and the bioprospecting of its genome for thermally adapted proteins. While high throughput sequencing methods have dramatically increased the availability of genomic information, the growth temperatures of the source organisms are often unknown. This limits the study and technological application of these species and their genomes. Here, we present a novel method for the prediction of growth temperatures of prokaryotes using only genomic sequences. By applying the reverse ecology principle that an organisms genome includes identifiable adaptations to its native environment, we can predict a species optimal growth temperature with an accuracy of 4.69 {degrees}C root-mean-square error and a correlation coefficient of 0.908. The accuracy can be further improved for specific taxonomic clades or by excluding psychrophiles. This method provides a valuable tool for the rapid calculation of organism growth temperature when only the genome sequence is known.

bioinformatics

Understanding the factors that shape patterns of nucleotide diversity in the house mouse genome

A major goal of population genetics has been to determine the extent to which selection at linked sites influences patterns of neutral nucleotide diversity in the genome. Multiple lines of evidence suggest that diversity is influenced by both positive and negative selection. For example, in many species there are troughs in diversity surrounding functional genomic elements, consistent with the action of either background selection (BGS) or selective sweeps. In this study, we investigated the causes of the diversity troughs that are observed in the wild house mouse genome. Using the unfolded site frequency spectrum (uSFS), we estimated the strength and frequencies of deleterious and advantageous mutations occurring in different functional elements in the genome. We then used these estimates to parameterize forward-in-time simulations of chromosomes, using realistic distributions of functional elements and recombination rate variation in order to determine if selection at linked sites can explain the observed patterns of nucleotide diversity. The simulations suggest that BGS alone cannot explain the dips in diversity around either exons or conserved non-coding elements (CNEs). A combination of BGS and selective sweeps, however, can explain the troughs in diversity around CNEs. This is not the case for protein-coding exons, where observed dips in diversity cannot be explained by parameter estimates obtained from the uSFS. We discuss the extent to which our results provide evidence of sweeps playing a role in shaping patterns of nucleotide diversity and the limitations of using the uSFS for obtaining inferences of the frequency and effects of advantageous mutations.\n\nAuthor SummaryWe present a study examining the causes of variation in nucleotide diversity across the mouse genome. The status of mice as a model organism in the life sciences makes them an excellent model system for studying molecular evolution in mammals. In our study, we analyse how natural selection acting on new mutations can affect levels of nucleotide diversity through the processes of background selection and selective sweeps. To perform our analyses, we first estimated the rate and strengths of selected mutations from a sample of wild mice and then use our estimates in realistic population genetic simulations. Analysing simulations, we find that both harmful and beneficial mutations are required to explain patterns of nucleotide diversity in regions of the genome close to gene regulatory elements. For protein-coding genes, however, our approach is not able to fully explain observed patterns and we think that this is because there are strongly advantageous mutations that occur in protein-coding genes that we were not able to detect.

evolutionary biology

The mitochondrial genomes of the mesozoans Intoshia linei, Dicyema sp., and Dicyema japonicum

The Dicyemida and Orthonectida are two groups of tiny, simple, vermiform parasites that have historically been united in a group named the Mesozoa. Both Dicyemida and Orthonectida have just two cell layers and appear to lack any defined tissues. They were initially thought to be evolutionary intermediates between protozoans and metazoans but more recent analyses indicate that they are protostomian metazoans that have undergone secondary simplification from a complex ancestor. Here we describe the first almost complete mitochondrial genome sequence from an orthonectid, Intoshia linei, and describe nine and eight mitochondrial protein-coding genes from Dicyema sp. and Dicyema japonicum, respectively. The 14,247 base pair long I. linei sequence has typical metazoan gene content, but is exceptionally AT-rich, and has a divergent gene order compared to other metazoans. The data we present from the Dicyemida provide very limited support for the suggestion that dicyemid mitochondrial genes are found on discrete mini-circles, as opposed to the large circular mitochondrial genomes that are typical across the Metazoa. The cox1 gene from dicyemid species has a series of conserved in-frame deletions that is unique to this lineage. Using cox1 genes from across the genus Dicyema, we report the first internal phylogeny of this group.\n\nKey FindingsO_LIWe report the first almost-complete mitochondrial genome from an orthonectid parasite, Intoshia linei, including 12 protein-coding genes; 20 tRNAs and putative sequences for large and small subunit rRNAs. We find that the I. linei mitochondrial genome is exceptionally AT-rich and has a novel gene order compared to other published metazoan mitochondrial genomes. These findings are indicative of the rapid rate of evolution that has occurred in the I. linei mitochondrial genome.\nC_LIO_LIWe also report nine and eight protein-coding genes, respectively, from the dicyemid species Dicyema sp. and Dicyema japonicum, and use the cox1 genes from both species for phylogenetic inference of the internal phylogeny of the dicyemids.\nC_LIO_LIWe find that the cox1 gene from dicyemids has a series of four conserved in-frame deletions which appear to be unique to this group.\nC_LI

evolutionary biology

Deep genome annotation of the opportunistic human pathogen Streptococcus pneumoniae D39

A precise understanding of the genomic organization into transcriptional units and their regulation is essential for our comprehension of opportunistic human pathogens and how they cause disease. Using single-molecule real-time (PacBio) sequencing we unambiguously determined the genome sequence of Streptococcus pneumoniae strain D39 and revealed several inversions previously undetected by short-read sequencing. Significantly, a chromosomal inversion results in antigenic variation of PhtD, an important surface-exposed virulence factor. We generated a new genome annotation using automated tools, followed by manual curation, reflecting the current knowledge in the field. By combining sequence-driven terminator prediction, deep paired-end transcriptome sequencing and enrichment of primary transcripts by Cappable-Seq, we mapped 1,015 transcriptional start sites and 748 termination sites. Using this new genomic map, we identified several new small RNAs (sRNAs), riboswitches (including twelve previously misidentified as sRNAs), and antisense RNAs. In total, we annotated 92 new protein-encoding genes, 39 sRNAs and 165 pseudogenes, bringing the S. pneumoniae D39 repertoire to 2,151 genetic elements. We report operon structures and observed that 9% of operons lack a 5-UTR. The genome data is accessible in an online resource called PneumoBrowse (https://veeninglab.com/pneumobrowse) providing one of the most complete inventories of a bacterial genome to date. PneumoBrowse will accelerate pneumococcal research and the development of new prevention and treatment strategies.

microbiology

Graph Peak Caller: calling ChIP-Seq Peaks on Graph-based Reference Genomes

Graph-based representations are considered to be the future for reference genomes, as they allow integrated representation of the steadily increasing data on individual variation. Currently available tools allow de novo assembly of graph-based reference genomes, alignment of new read sets to the graph representation as well as certain analyses like variant calling and haplotyping. We here present a first method for calling ChIP-Seq peaks on read data aligned to a graph-based reference genome. The method is a graph generalization of the peak caller MACS2, and is implemented in an open source tool, Graph Peak Caller. By using the existing tool vg to build a pan-genome of Arabidopsis thaliana, we validate our approach by showing that Graph Peak Caller with a pan-genome reference graph can trace variants within peaks that are not part of the linear reference genome, and find peaks that in general are more motif-enriched than those found by MACS2.

bioinformatics

Analysis of genome-wide differentiation between native and introduced populations of the cupped oysters Crassostrea gigas and Crassostrea angulata

The Pacific cupped oyster is genetically subdivided into two sister taxa, Crassostrea gigas and C. angulata, which are in contact in the north-western Pacific. The nature and origin of their genetic and taxonomic differentiation remains controversial due the lack of known reproductive barriers and morphologic similarity. In particular, whether ecological and/or intrinsic isolating mechanisms participate to species divergence remains unknown. The recent co-introduction of both taxa into Europe offers a unique opportunity to test how genetic differentiation maintains under new environmental and demographic conditions. We generated a pseudo-chromosome assembly of the Pacific oyster genome using a combination of BAC-end sequencing and scaffold anchoring to a new high-density linkage map. We characterized genome-wide differentiation between C. angulata and C. gigas in both their native and introduced ranges, and showed that gene flow between species has been facilitated by their recent co-introductions in Europe. Nevertheless, patterns of genomic divergence between species remain highly similar in Asia and Europe, suggesting that the environmental transition caused by the co-introduction of the two species did not affect the genomic architecture of their partial reproductive isolation. Increased genetic differentiation was preferentially found in regions of low recombination. Using historical demographic inference, we show that the heterogeneity of differentiation across the genome is well explained by a scenario whereby recent gene flow has eroded past differentiation at different rates across the genome after a period of geographical isolation. Our results thus support the view that low-recombining regions help in maintaining intrinsic genetic differences between the two species.

evolutionary biology

Creation and multi-omics characterization of a genomically hybrid strain in the nitrogen-fixing symbiotic bacterium Sinorhizobium meliloti

Many bacteria, often associated with eukaryotic hosts and of relevance for biotechnological applications, harbour a multipartite genome composed by more than one replicon. Biotechnologically relevant phenotypes are often encoded by genes residing on the secondary replicons. A synthetic biology approach to developing enhanced strains for biotechnological purposes could therefore involve merging pieces or entire replicons from multiple strains into a single genome. Here we report the creation of a genomic hybrid strain in a model multipartite genome species, the plant-symbiotic bacterium Sinorhizobium meliloti. In particular, we moved the secondary replicon pSymA (accounting for nearly 20% of total genome content) from a donor S. meliloti strain to an acceptor strain. The cis-hybrid strain was screened for a panel of complex phenotypes (carbon/nitrogen utilization phenotypes, intra- and extra-cellular metabolomes, symbiosis, and various microbiological tests). Additionally, metabolic network reconstruction and constraint-based modelling were employed for in silico prediction of metabolic flux reorganization. Phenotypes of the cis-hybrid strain were in good agreement with those of both parental strains. Interestingly, the symbiotic phenotype showed a marked cultivar-specific improvement with the cis-hybrid strains compared to both parental strains. These results provide a proof-of-principle for the feasibility of genome-wide replicon-based remodelling of bacterial strains for improved biotechnological applications in precision agriculture.

synthetic biology

Core Genome Multi Locus Sequence Typing and Single Nucleotide Polymorphism Analysis in the Epidemiology of Brucella melitensis Infections

The use of whole genome sequencing (WGS) using next generation sequencing (NGS) technology has become a widely accepted method for microbiology laboratories in the application of molecular typing for outbreak tracing and genomic epidemiology. Several studies demonstrated the usefulness of WGS data analysis through Single Nucleotide Polymorphism (SNP) calling from a reference sequence analysis for Brucella melitensis, whereas gene-by-gene comparison through core-genome Multilocus Sequence Typing (cgMLST) has not been explored so far. The current study developed an allele-based method cgMLST and compared its performance to the genome-wide SNP approach and the traditional MLVA on a defined sample collection. The dataset comprised of 37 epidemiologically linked animal cases of brucellosis as well as 71 epidemiologically unrelated human and animal isolates collected in Italy. The cgMLST scheme generated in this study contained 2,687 targets of the B. melitensis 16M reference genome (75.4% of the complete genome). We established the potential criteria necessary for inclusion of an isolate into a brucellosis outbreak cluster to be [≤]4 loci in the cgMLST and [≤]10 in WGS SNP analysis. CgMLST and SNP analysis provided much higher phylogenetic distance resolution than MLVA, particularly for strains belonging to the same lineage thus allowing diverse and unrelated genotypes to be identified with greater confidence. The application of this cgMLST scheme to the characterization of B. melitensis strains provided insights into the epidemiology of this pathogen and it is a candidate to be a benchmark tool for outbreak investigations in human and animal brucellosis.

microbiology

Three new genome assemblies support a rapid radiation in Musa acuminata (wild banana)

Edible bananas result from interspecific hybridization between Musa acuminata and Musa balbisiana, as well as among subspecies in M. acuminata. Four particular M. acuminata subspecies have been proposed as the main contributors of edible bananas, all of which radiated in a short period of time in southeastern Asia. Clarifying the evolution of these lineages at a whole-genome scale is therefore an important step toward understanding the domestication and diversification of this crop. This study reports the de novo genome assembly and gene annotation of a representative genotype from three different subspecies of M. acuminata. These data are combined with the previously published genome of the fourth subspecies to investigate phylogenetic relationships and genome evolution. Analyses of shared and unique gene families reveal that the four subspecies are quite homogenous, with a core genome representing at least 50% of all genes and very few M. acuminata species-specific gene families. Multiple alignments indicate high sequence identity between homologous single copy-genes, supporting the close relationships of these lineages. Interestingly, phylogenomic analyses demonstrate high levels of gene tree discordance, due to both incomplete lineage sorting and introgression. This pattern suggests rapid radiation within Musa acuminata subspecies that occurred after the divergence with M. balbisiana. Introgression between M. a. ssp. malaccensis and M. a. ssp. burmannica was detected across a substantial portion of the genome, though multiple approaches to resolve the subspecies tree converged on the same topology. To support future evolutionary and functional analyses, we introduce the PanMusa database, which enables researchers to exploration of individual gene families and trees.

evolutionary biology

Machine learning based prediction of functional capabilities in metagenomically assembled microbial genomes

The increasing popularity of genome resolved meta genomics - the binning of genomes of potentially uncultured organisms direct from the environmental DNA - has resulted in a deluge of draft genomes. There is a pressing need to develop methods to interpret this data. Here, we used machine learning to predict functional and metabolic traits of microbes from their genomes. We collated an extensive database of 84 phenotypic traits associated with 9407 prokaryotic genomes and trained different machine learning models on this data. We found that a lasso logistic regression based on the frequency of gene orthologs had the best combination of functional prediction performance and interpretability. This model was able to classify 65 phenotypic traits with greater than 90

microbiology

Parallel Evolution of Key Genomic Features and Cellular Bioenergetics Across the Marine Radiation of a Bacterial Phylum

Diverse bacterial and archaeal lineages drive biogeochemical cycles in the global ocean, but the evolutionary processes that have shaped their genomic properties and physiological capabilities remain obscure. Here we track the genome evolution of the globally-abundant marine bacterial phylum Marinimicrobia across its diversification into modern marine environments and demonstrate that extant lineages have repeatedly switched between epipelagic and mesopelagic habitats. Moreover, we show that these habitat transitions have been accompanied by repeated and fundamental shifts in genomic organization, cellular bioenergetics, and metabolic modalities. Lineages present in epipelagic niches independently acquired genes necessary for phototrophy and environmental stress mitigation, and their genomes convergently evolved key features associated with genome streamlining. Conversely, lineages residing in mesopelagic waters independently acquired nitrate respiratory machinery and a variety of cytochromes, consistent with the use of alternative terminal electron acceptors in oxygen minimum zones (OMZs). Further, while surface water clades have retained an ancestral Na+-pumping respiratory complex, deep water lineages have largely replaced this complex with a canonical H+-pumping respiratory complex I, potentially due to the increased efficiency of the latter together with more energy-limiting environments deep in the oceans interior. These parallel evolutionary trends across disparate clades suggest that the evolution of key features of genomic organization and cellular bioenergetics in abundant marine lineages may in some ways be predictable and driven largely by environmental conditions and nutrient dynamics.

microbiology