Search bioRxiv⌕ Search

Biology subjects

Sisson, Z.

Publications and source records attributed to Sisson, Z..

3 recordsLinked to original sources

Unified sampling framework and experimental benchmarking of sequence- and structure-based protein models

Generative models are increasingly used for protein design, but the lack of standardized evaluation frameworks limits comparison across model classes and hinders translation to experimental success. Here, we introduce a unified sampling and benchmarking framework that enables controlled sequence generation across alignment, protein language, and structure-based models, and apply it to Tobacco etch virus (TEV) protease. Across hundreds of thousands of designed sequences, different models explore distinct regions of sequence space with no clear computational selection metrics to assess enzymatic function. Experimental evaluation reveals large differences in functional outcomes, ranging from non-functional variants to sequences with 9-fold higher activity than wildtype. Machine learning-designed libraries achieve a 39.32% hit rate (percentage of variants matching or exceeding wildtype activity) compared to 6.06% for an error-prone PCR baseline. Structure-based models perform best overall, with hit rates of 74.4% and 66.8% for ESM-IF1 and ProteinMPNN, respectively. Commonly used selection metrics do not strongly correlate with experimental activity, highlighting a gap between in silico evaluation and enzyme function. Together, these results establish a generalizable framework for benchmarking generative protein models and demonstrate the necessity of experimental validation for guiding model development and sequence prioritization.

bioinformatics↗

GROQ-seq Datasets Across Transcription Factors (LacI, RamR, VanR), T7 RNA Polymerase and TEV Protease

Predicting any proteins function from its sequence alone would be a significant breakthrough in molecular biology. Although machine learning approaches have sought to tackle this, their limited generalizability reflects the absence of sufficiently large, open, diverse, and unified datasets. To address this data gap, we developed a high-throughput experimental platform called GROQ-seq (Growth-based Quantitative Sequencing). In GROQ-seq, a proteins function can be linked to a sequencing-based readout that enables scalable characterization of large variant libraries in Escherichia coli. Here, we present pilot datasets demonstrating its performance across three distinct protein function classes: transcription factors, polymerases, and proteases. The objective of this report is to present the datasets and to provide users with a clear and transparent characterization of their properties, including both the strengths and limitations.

bioengineering↗

Proxi-RIMS-seq2 applied to native microbiomes uncovers hundreds of known and novel m5C methyltransferase specificities.

Methylation patterns in bacteria can be used to study Restriction-Modification (RM) or other defense systems with novel properties. While m4C and m6A methylation is well characterized mainly through PacBio sequencing, the landscape of m5C methylation is under-characterized. To bridge this gap, we performed RIMS-seq2 on microbiomes composed of resolved assemblies of distinct genomes through proximity ligation. This high-throughput approach enables the identification of m5C methylated motifs and links them to cognate methyltransferases directly on native microbiomes without the need to isolate bacterial strains. Methylation patterns can also be identified on viral DNA and compared to host DNA, strengthening evidence for virus-host interaction. Applied to three different microbiomes, the method unveils over 1900 motifs that were deposited in REBASE. The motifs include a novel 8-base recognition site (CATm5CGATG) that was experimentally validated by characterizing its cognate methyltransferase. Our findings suggest that microbiomes harbor arrays of untapped m5C methyltransferase specificities, providing insights to bacterial biology and biotechnological applications.

genomics↗