Search bioRxiv⌕ Search

Biology subjects

Coughlan, K.

Publications and source records attributed to Coughlan, K..

2 recordsLinked to original sources

PNGaseA-mediated N-glycan stripping from peptides by infant-derived Bifidobacterium bifidum

N-glycans are highly common sources of nutrition for human colonic-dwelling bacteria. These microbes have evolved a several methods to remove N-glycans from proteins; herein we describe the biochemical and structural characterisation of one such enzyme, a PNGaseA superfamily member produced by the infant-associated Bifidobacterium bifidum LMG13195. This PNGase was demonstrated to elicit activity against a wide variety of N-glycan structures yet exhibited a high preference for N-glycans attached to a peptide rather than to a denatured or native protein. This unusual specificity highlights how bacterial species tune their enzymology to different types of substrates. The structural characterisation of this PNGase reveals how its structure determines this specificity while being the first structure presented from the PNGaseA superfamily, revealing a unique ten-stand {beta}-sheet cradling a canonical PNGase catalytic module.

biochemistry↗

Alignment-free Bacterial Taxonomy Classification with Genomic Language Models

Advances in natural language processing, including the ability to process long sequences, have paved the way for the development of Genomic Language Models (gLM). This study evaluates the feasibility of four models for bacterial classification using 16S rRNA sequences and demonstrates that gLM embeddings can be applied to effectively classify sequences at the species level, matching or outperforming the accuracy of established bioinformatics tools like BLAST+ and VSEARCH. We adopt cosine similarity as a computationally efficient metric, enabling classification orders of magnitude faster than current methods, and show that it carries biologically relevant signals. In addition, we demonstrate how sequence embeddings can be used to identify mislabeled sequences. Our findings place gLM embeddings as a promising alternative to traditional alignment-based methods, especially in large-scale applications such as metataxonomic assignments. Despite its wide potential, key challenges remain, including the sensitivity of embeddings to sequences of different lengths.

bioinformatics↗