Search bioRxiv⌕ Search

Biology subjects

Wynne, J. H.

Publications and source records attributed to Wynne, J. H..

3 recordsLinked to original sources

When seeps give ANME-SRB the cold shoulder: putative role of denitrification mediated methane oxidation in an Antarctic Cold Seep

Antarctica represents a significant, unresolved, and unstable source of methane to the atmosphere. To advance our understanding of the biological filter of methane in Antarctica, here we identify the taxa and functional genes present during methane oxidation in an Antarctic Methane Seep. Methane oxidation was present in all sediments, including in a non seep control site. Using 16S rRNA analysis alongside metagenomics, we found that ANaerobic MEthane oxidizing (ANME) archaea coupled to Sulfate-Reducing Bacteria (SRB), documented as the most important marine methane sink in other locations, were not present. Instead, we observed the presence of denitrification-dependent methane oxidizers, including the anaerobic genus Candidatus Methylomirabilis, alongside the nitrate reducing archaea Candidatus Methanoperedens through short-read metagenomic classification. In addition, we note the presence of multiple aerobic methanotrophs, with a particularly high abundance of the Methylobacter, Methylomonas, and Methyloprofundus genera. Our results support denitrification-mediated methane oxidation and aerobic methanotrophy as the primary potential methane sinks in the Ross Sea. The widespread methane oxidation, including in control sediment, combined with the possibility of anaerobic methane oxidation linked to denitrification rather than sulfate reduction highlights the ubiquity and uniqueness of the Antarctic methane cycle.

ecology↗

Metagenomic contextualization of proteins with state space models

Since the early adoption of metagenomics (the culture-free sequencing of microbial community genomes) in 2011, sequence data has increased over 500-fold across ecosystems. This surge in data has outpaced reliable taxonomic and functional annotation, with over half of sequences lacking confident functional assignment. These unknown sequences limit our understanding of microbial processes central to planetary health and human health. Recent advances in genomic language modeling have made progress in the interpretation of metagenomics datasets. Most state-of-the-art models rely on transformer architectures, which limit the maximum sequence length and therefore capture only a fraction of assembled metagenomic sequences due to the quadratic scaling of attention. This prevents training and inference on sequences with broad context, including multiple coding and non-coding regions. To overcome this limitation, we propose leveraging new model architectures that scale linearly with sequence length, making them more suitable for modeling longer metagenomic sequences. Here, we introduce Nammu, a mixed-modality Mamba-based foundation model with 167M parameters trained on the OpenMetaGenomic (OMG) corpus. Nammu is a bidirectional encoder trained with a 20K context length using a curriculum strategy, first on 64M protein sequences and then on 32M mixed-modality metagenomic contigs. We compared Nammu to gLM2, a mixed-modality transformer also trained on OMG using 37% more tokens, using taxonomy inference on a marine dataset from the Critical Assessment of Metagenome Interpretation (CAMI). Nammu outperforms gLM2 at every taxonomic level. We further assessed function via KEGG Orthology prediction in deep-sea metagenome-assembled genomes, where Nammu outperforms gLM2 (150M). These results demonstrate improved performance.

bioinformatics↗

BioGeoFormer: A deep learning approach to classify unknown genes associated with critical biogeochemical cycles

Remote functional annotation continues to impede progress in microbial ecology, as alignment-based approaches still leave over one-third of microbial sequences functionally unresolved. In contrast, pre-trained natural-language-processing approaches have shown strong potential for inferring functions from diverse biological sequences, and here we introduce a protein language modeling approach allowing us to classify sequences into 37 defined key pathway categories involved in 4 major biogeochemical cycles (methane, sulfur, nitrogen and phosphorus cycles). To do so, we fine-tuned ESM2-8m using databases curated for biogeochemical cycling pathways. Our resultant BioGeochemical cycling transFormer (BioGeoFormer or BGF) was high-performing on validation and test sets, producing embeddings that exhibit an ability to infer protein function at a metabolic pathway level. BGF was applied to a dataset of metagenome-assembled genomes (MAGs) constructed from methane-fueled, deep-sea, "cold seep" environments to demonstrate its utility in contrast to current informatics approaches. We employed multiple gene assignments to identify gene function within these MAGs. A total of 1.05M genes were assigned biogeochemical functions, with BGF alone suggesting putative ecosystem roles for 0.49M (46%) of these at a confidence of 85% or greater; these genes were classified as unknown by the other approaches. Across the pathways of interest, BGF identified 6 times as many genes, on average, as Hidden Markov models (HMMs) as well as alignment-based approaches across the various pathways. BGF provides a novel tool that is capable of informing process-based hypotheses in diverse systems, highlighting cryptic proteins most notably linked to methane, nitrogen, and phosphorus cycling while uncovering the mysteries within microbial dark matter. Author summaryWhen investigating the function of microbes in the environment, scientists are often left with vast amounts of genes or proteins where no knowledge about their function is available. This represents a huge amount of information left to be discovered in many fields of biology. One recent approach that has shown significant potential in further understanding unknown proteins are protein language models, which are deep-learning methods leveraging large datasets to understand the language of proteins. We aimed to apply protein language modeling to further understand the function of proteins as they relate to large-scale environmental transformation of nutrients and carbon. Specifically, we designed our approach to further understand the metabolism of microbes that affect methane, nitrogen, phosphorus, and sulfur, all elements that are highly impactful to the planets function and health. We used our new approach on a deep-sea microbiology dataset, and showed the methods utility in further understanding the function of proteins and their impact on the environment. Overall, we found our method is an important new tool in the toolset of environmental scientists working to better understand the function of microbes and their proteins.

microbiology↗