Search bioRxiv⌕ Search

Biology subjects

Pashkova, L.

Publications and source records attributed to Pashkova, L..

2 recordsLinked to original sources

ProteusAI: An Open-Source and User-Friendly Platform for Machine Learning-Guided Protein Design and Engineering

AO_SCPLOWBSTRACTC_SCPLOWProtein design and engineering are crucial for advancements in biotechnology, medicine, and sustainability. Machine learning (ML) models are used to design or enhance protein properties such as stability, catalytic activity, and selectivity. However, many existing ML tools require specialized expertise or lack open-source availability, limiting broader use and further development. To address this, we developed ProteusAI, a user-friendly and open-source ML platform to streamline protein engineering and design tasks. ProteusAI offers modules to support researchers in various stages of the design-build-test-learn (DBTL) cycle, including protein discovery, structure-based design, zero-shot predictions, and ML-guided directed evolution (MLDE). Our benchmarking results demonstrate ProteusAIs efficiency in improving proteins and enyzmes within a few DBTL-cycle iterations. ProteusAI democratizes access to ML-guided protein engineering and is freely available for academic and commercial use. Future work aims to expand and integrate novel methods in computational protein and enzyme design to further develop ProteusAI.

bioinformatics↗

PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining

The exponential growth of microbial genome data presents unprecedented opportunities for mining the potential of microorganisms. The burgeoning field of pangenomics offers a framework for extracting insights from this big biological data. Recent advances in microbial pangenomic research have generated substantial data and literature, yielding valuable knowledge across diverse microbial species. PanKB (pankb.org), a knowledgebase designed for microbial pangenomics research and biotechnological applications, was built to capitalize on this wealth of information. PanKB currently includes 51 pangenomes on 8 industrially relevant microbial families, comprising 8, 402 genomes, over 500, 000 genes, and over 7M mutations. To describe this data, PanKB implements four main components: 1) Interactive pangenomic analytics to facilitate exploration, intuition, and potential discoveries; 2) Alleleomic analytics, a pangenomic- scale analysis of variants, providing insights into intra-species sequence variation and potential mutations for applications; 3) A global search function enabling broad and deep investigations across pangenomes to power research and bioengineering workflows; 4) A bibliome of 833 open- access pangenomic papers and an interface with an LLM that can answer in-depth questions using their knowledge. PanKB empowers researchers and bioengineers to harness the full potential of microbial pangenomics and serves as a valuable resource bridging the gap between pangenomic data and practical applications. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=134 SRC="FIGDIR/small/608241v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@4d3daaorg.highwire.dtl.DTLVardef@10b6d49org.highwire.dtl.DTLVardef@133da4forg.highwire.dtl.DTLVardef@141ae7e_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗