Search bioRxiv⌕ Search

Biology subjects

Catoiu, E. A.

Publications and source records attributed to Catoiu, E. A..

4 recordsLinked to original sources

QSProteome: A Community-Driven Interactive Platform for Large-Scale Exploration and Evaluation of Predicted Protein Complex Structures.

QSProteome (https://QSProteome.org) is a community-scale platform for modeling, evaluating, and refining quaternary protein structure. The resource hosts 35,528 unique modeled assemblies spanning over 42,000 genes, covering nearly all curated complexes in BioCyc and ComplexPortal databases. Each model is displayed on an interactive page with 3D visualization, chain-level confidence metrics, structural alignments, functional annotations, and automated stoichiometry checks against database expectations. A cloud-based server supports continuous user uploads and automated processing pipelines--enabling submission and validation of >54,000 models within 14 weeks. To promote iterative refinement, QSProteome includes a gamified re-curation workflow that transitions users from training modules into live curation, enabling the community-led assessment of 1,547 ABC transporter complexes. Together, these components form a dynamic, scalable infrastructure for proteome-scale structural biology. By unifying modeling, validation, and annotation in a reusable, searchable, and community-extensible framework, QSProteome enables proteome-scale structure accessibility and reuse--powering discovery, annotation, and collaborative refinement across the structural biology community.

bioinformatics↗

Establishing comprehensive quaternary structural proteomes from genome sequence

A critical body of knowledge has developed through advances in protein microscopy, protein-fold modeling, structural biology software, availability of sequenced bacterial genomes, large-scale mutation databases, and genome-scale models. Based on these recent advances, we develop a computational framework that; i) identifies the oligomeric structural proteome encoded by an organisms genome from available structural resources; ii) maps multi-strain alleleomic variation, resulting in the structural proteome for a species; and iii) calculates the 3D orientation of proteins across subcellular compartments with residue-level precision. Using the platform, we; iv) compute the quaternary E. coli K-12 MG1655 structural proteome; v) use a dataset of 12,000 mutations to build Random Forest classifiers that can predict the severity of mutations; and, in combination with a genome-scale model that computes proteome allocation, vi) obtain the spatial allocation of the E. coli proteome. Thus, in conjunction with relevant datasets and increasingly accurate computational models, we can now annotate quaternary structural proteomes, at genome-scale, to obtain a molecular-level understanding of whole-cell functions. SignificanceAdvancements in experimental and computational methods have revealed the shapes of multi-subunit proteins. The absence of a unified platform that maps actionable datatypes onto these increasingly accurate structures creates a barrier to structural analyses, especially at the genome-scale. Here, we describe QSPACE, a computational annotation platform that evaluates existing resources to identify the best-available structure for each protein in a users query, maps the 3D location of actionable datatypes (e.g., active sites, published mutations) onto the selected structures, and uses third-party APIs to determine the subcellular compartment of all amino acids of a protein. As proof-of-concept, we deployed QSPACE to generate the quaternary structural proteome of E. coli MG1655 and demonstrate two use-cases involving large-scale mutant analysis and genome-scale modelling.

bioinformatics↗

Laboratory evolution reveals transcriptional mechanisms underlying thermal adaptation of Escherichia coli

Adaptive laboratory evolution (ALE) is able to generate microbial strains which exhibit extreme phenotypes, revealing fundamental biological adaptation mechanisms. Here, we use ALE to evolve Escherichia coli strains that grow at temperatures as high as 45.3{degrees}C, a temperature lethal to wild type cells. The strains adopted a hypermutator phenotype and employed multiple systems-level adaptations that made global analysis of the DNA mutations difficult. Given the challenge at the genomic level, we were motivated to uncover high temperature tolerance adaptation mechanisms at the transcriptomic level. We employed independently modulated gene set (iModulon) analysis to reveal five transcriptional mechanisms underlying growth at high temperatures. These mechanisms were connected to acquired mutations, changes in transcriptome composition, sensory inputs, phenotypes, and protein structures. They are: (i) downregulation of general stress responses while upregulating the specific heat stress responses; (ii) upregulation of flagellar basal bodies without upregulating motility, and upregulation fimbriae; (iii) shift toward anaerobic metabolism, (iv) shift in regulation of iron uptake away from siderophore production, and (v) upregulation of yjfIJKL, a novel heat tolerance operon which we characterized using AlphaFold. iModulons associated with these five mechanisms explain nearly half of all variance in the gene expression in the adapted strains. These thermotolerance strategies reveal that optimal coordination of known stress responses and metabolism can be achieved with a small number of regulatory mutations, and may suggest a new role for large protein export systems. ALE with transcriptomic characterization is a productive approach for elucidating and interpreting adaptation to otherwise lethal stresses.

systems biology↗

Laboratory-acquired mutations fall outside the wild-type alleleome of Escherichia coli

Inexpensive DNA sequencing has led to a rapidly increasing number of whole genome sequences in the public domain. Natural sequence variation can now be assessed across a large number of sequenced strains of a bacterial species, resulting in the definition of the wild-type alleleome (the collection of alleles for every gene found in the species). Concurrently, laboratory evolution emerged as a new approach to address biological questions and to develop new phenotypic traits, and a large number of laboratory acquired mutations can be found in databases. The availability of this large-scale sequence variation data now allows for a detailed comparison of mutations fixed in natural versus laboratory evolutions. Such comparison shows that laboratory-acquired mutations are rarely found in the wild-type alleleome of Escherichia coli. The E. coli alleleome is highly conserved as most of the sequence variation is concentrated in about 2% of the coding region. We find that there are typically two alternate amino acids coded for in the variable locations, and switches between the two are found in the data sets. Finally, we find that adaptive laboratory mutations, unlike wild-type mutations, do not utilize the redundancy built into the genetic code: they are less likely to be synonymous and rely on changing a single nucleotide in a codon. However, the uniqueness of mutations fixed in laboratory evolutions bodes well for synthetic biology by revealing novel exploitable sequence space untouched by natural evolution.

bioengineering↗