Search bioRxiv⌕ Search

Biology subjects

Atkinson, T.

Publications and source records attributed to Atkinson, T..

4 recordsLinked to original sources

AbBFN2: A flexible antibody foundation model based on Bayesian Flow Networks

Antibody engineering is marked by diverse data and desiderata, making it a prime candidate for multi-objective design, but is commonly tackled as a series of sequential optimisation tasks. Here, we present AbBFN2, a generative antibody foundation model trained on paired antibody sequences as well as genetic and biophysical metadata. This is achieved using the Bayesian Flow Network paradigm, which allows unified modelling of diverse data sources and flexible conditional generation at inference time. By virtue of its rich set of features and architectural flexibility, AbBFN2 can be adapted to a number of tasks commonly tackled by individual models, consolidating traditional computational pipelines into a single step. We demonstrate the adaptability of AbBFN2 using sequence inpainting, humanisation, biophysical property optimisation, and conditional de novo library generation of antibodies with rare attributes as example tasks. By removing the need for task-specific training, we hope that AbBFN2 will accelerate machine learning-based antibody design and development workflows.

bioinformatics↗

Protein Sequence Modelling with Bayesian Flow Networks

Exploring the vast and largely uncharted territory of amino acid sequences is crucial for understanding complex protein functions and the engineering of novel therapeutic proteins. Whilst generative machine learning has advanced protein sequence modelling, no existing approach is proficient for both unconditional and conditional generation. In this work, we propose that Bayesian Flow Networks (BFNs), a recently introduced framework for generative modelling, can address these challenges. We present ProtBFN, a 650M parameter model trained on protein sequences curated from UniProtKB, which generates natural-like, diverse, structurally coherent, and novel protein sequences, significantly outperforming leading autoregressive and discrete diffusion models. Further, we fine-tune ProtBFN on heavy chains from the Observed Antibody Space (OAS) to obtain an antibody-specific model, AbBFN, which we use to evaluate zero-shot conditional generation capabilities. AbBFN is found to be competitive with, or better than, antibody-specific BERT-style models, when applied to predicting individual framework or complimentary determining regions (CDR).

bioinformatics↗

Contrasting Sequence with Structure: Pre-training Graph Representations with PLMs

Understanding protein function is vital for drug discovery, disease diagnosis, and protein engineering. While Protein Language Models (PLMs) pre-trained on vast protein sequence datasets have achieved remarkable success, equivalent Protein Structure Models (PSMs) remain underrepresented. We attribute this to the relative lack of high-confidence structural data and suitable pre-training objectives. In this context, we introduce BioCLIP, a contrastive learning framework that pre-trains PSMs by leveraging PLMs, generating meaningful per-residue and per-chain structural representations. When evaluated on tasks such as protein-protein interaction, Gene Ontology annotation, and Enzyme Commission number prediction, BioCLIP-trained PSMs consistently outperform models trained from scratch and further enhance performance when merged with sequence embeddings. Notably, BioCLIP approaches, or exceeds, specialized methods across all benchmarks using its singular pre-trained design. Our work addresses the challenges of obtaining quality structural data and designing self-supervised objectives, setting the stage for more comprehensive models of protein function. Source code is publicly available2.

bioinformatics↗

An efficient method for high molecular weight bacterial DNA extraction suitable for shotgun metagenomics from skin swabs

The human skin microbiome represents a variety of complex microbial ecosystems that play a key role in host health. Molecular methods to study these communities have been developed but have been largely limited to low-throughput quantification and short amplicon sequencing, providing limited functional information about the communities present. Shotgun metagenomic sequencing has emerged as a preferred method for microbiome studies as it provides more comprehensive information about the species/strains present in a niche and the genes they encode. However, the relatively low bacterial biomass of skin, in comparison to other areas such as the gut microbiome, makes obtaining sufficient DNA for shotgun metagenomic sequencing challenging. Here we describe an optimised high-throughput method for extraction of high molecular weight DNA suitable for shotgun metagenomic sequencing. We validated the performance of the extraction method, and analysis pipeline on skin swabs collected from both adults and babies. The pipeline effectively characterised the bacterial skin microbiota with a cost and throughput suitable for larger longitudinal sets of samples. Application of this method will allow greater insights into community compositions and functional capabilities of the skin microbiome. Impact StatementDetermining the functional capabilities of microbial communities within different human microbiomes is important to understand their impacts on health. Extraction of sufficient DNA is challenging, especially from low biomass samples, such as skin swabs suitable for shotgun metagenomics, which is needed for taxonomic resolution and functional information. Here we describe an optimised DNA extraction method that produces enough DNA from skin swabs, suitable for shotgun metagenomics, and demonstrate it can be used to effectively characterise the skin microbiota. This method will allow future studies to identify taxonomic and functional changes in the skin microbiota which is needed to develop interventions to improve and maintain skin health. Data SummaryAll sequence data and codes can be accessed at: NCBI Bio Project ID: PRJNA937622 DOI: https://github.com/quadram-institute-bioscience/coronahit_guppy DOI: https://github.com/ilianaserghiou/Serghiou-et-al.-2023-Codes

genomics↗