Search bioRxiv⌕ Search

Biology subjects

Kim, L. M.

Publications and source records attributed to Kim, L. M..

4 recordsLinked to original sources

Predicting the protein interaction landscape of a mycobacterial pathogen

High-dimensional phenotypic screens of bacterial loss-of-function mutant libraries have determined gene-gene connections and specific phenotypes for thousands of bacterial genes in many species, but deciphering the underlying mechanisms remains decidedly low-throughput. Here, we demonstrate the utility of proteome-wide AI-based protein-protein interaction (PPI) predictions for overcoming this gap by using pooled-AlphaFold3 to assess all [~]1.3 million possible pairwise interactions in the proteome of Mycobacterium leprae. We identify [~]2,000 strong and intermediate PPIs that underlie a significant fraction of phenotypes and gene-gene connections observed in large-scale chemical genomics screens from Mycobacterium tuberculosis, Mycobacterium smegmatis, and Corynebacterium glutamicum. This combined approach predicts specific functions for dozens of previously uncharacterized core, conserved, and essential mycobacterial proteins. We highlight new information derived from the study, including insights into mycobacterial envelope assembly, peptidoglycan remodeling, and new modulators of the central dogma enzymes RNA polymerase and DNA gyrase. These data establish combined pooled-AlphaFold3 PPI prediction and high-throughput genomics approach as the gold standard for large-scale characterization of protein function.

systems biology↗

The phenotypic landscape of the model firmicute Bacillus subtilis

Firmicutes are gram-positive bacteria with important roles in human health, disease, and industry. However, more than a quarter of genes in the model firmicute Bacillus subtilis remain completely uncharacterized, including numerous core phylum-specific genes. Here, we design a compact pooled CRISPRi library targeting all protein-coding genes in B. subtilis and test its growth in [~]150 stress conditions. Using data from this screen as a hypothesis generator, we perform targeted experiments, revealing that the conserved essential firmicute protein YneF is part of the SRP co-translational protein secretion pathway. We also demonstrate that ECF-transporters play a previously unknown but broadly conserved role in cell wall homeostasis, perform an unbiased analysis of amino acid crossfeeding, and make additional discoveries about bacterial competition. In addition to these major contributions to our understanding of B. subtilis biology (and gram-positive firmicutes in general), this work provides a rich dataset that will nucleate future studies of uncharacterized genes and presents a framework for accessible full-genome functional genomic screens in other bacteria. SIGNIFICANCELarge-scale chemical genomics screens facilitate the characterization of genes by generating phenotypic data across a library of gene mutants. Here, using CRISPRi, we designed a compact pooled library targeting all genes in B. subtilis and screened growth of the library in [~]150 conditions. Importantly, we used this dataset to expand B. subtilis biology on levels ranging from molecular pathways to bacterial communities. We discovered a role for an essential firmicute gene in protein secretion, implicated two new players in cell wall homeostasis, and unraveled fundamental factors driving competition and cooperation in the soil. This rich dataset will serve as a hypothesis generator, expanding our set of bacterial phenotypic data and driving future experiments in B. subtilis and other gram-positive firmicutes.

microbiology↗

The phenotypic landscape of the mycobacterial cell

The Mycobacteriales are an order of diverse bacteria that thrive in many environmental and host-associated niches. Because the most notorious member of this clade, Mycobacterium tuberculosis, is a major human pathogen, research on Mycobacteriales has focused on pathogenesis, and, as a consequence, many fundamental aspects of Mycobacterial biology remain understudied. Here, we address this gap by performing a genome-wide CRISPRi chemical genomics screen using a diverse set of >35 antibiotics, detergents, and other anti-microbials predominantly targeting the cell envelope of Mycobacterium smegmatis, a saprophytic model Mycobacterium. We highlight new information derived from this screen, including the identification of novel functions for previously uncharacterized conserved and essential genes (in mycolic acid and arabinogalactan synthesis), the discovery of a new drug scaffold/protein target pair, and insights into the mechanism of action of two commonly used antibiotics. These data are also a valuable resource for the mycobacterial research community, as they provide thousands of novel phenotypes for uncharacterized genes and meaningful phenotypic correlations between annotated and uncharacterized genes.

microbiology↗

Predicting the protein interaction landscape of a free-living bacterium with pooled-AlphaFold3

Accurate prediction of protein complex structures by AlphaFold3 and similar programs has been used to predict the presence of protein-protein interactions (PPIs), but this technique has never been applied to an entire genome due to onerous computational requirements and questionable utility. Here we present pooled-PPI prediction, a technique that dramatically improves the accuracy of genome-scale screens compared to a paired approach while simultaneously reducing inference time ([~]2-fold) and the number of jobs ([~]100-fold). We use this technique to predict the structure of all 113,050 pairwise PPIs in Mycoplasma genitalium using only 2,027 AlphaFold3 jobs. This unbiased and comprehensive dataset was highly predictive of known interactions, revealed a previously unappreciated but widespread size bias in AlphaFold interface scores, correctly identified protein-protein interfaces in macromolecular complexes, and uncovered new biology in M. genitalium. This work establishes pooled-PPI prediction as a highly scalable method for uncovering protein-protein interactions and a powerful addition to the functional genomics toolkit.

systems biology↗