Search bioRxiv⌕ Search

bioRxiv · 10.1101/2022.09.27.509819

COVIDpro: Database for mining protein dysregulation in patients with COVID-19

Abstract

BackgroundThe ongoing pandemic of the coronavirus disease 2019 (COVID-19) caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) still has limited treatment options partially due to our incomplete understanding of the molecular dysregulations of the COVID-19 patients. We aimed to generate a repository and data analysis tools to examine the modulated proteins underlying COVID-19 patients for the discovery of potential therapeutic targets and diagnostic biomarkers. MethodsWe built a web server containing proteomic expression data from COVID-19 patients with a toolset for user-friendly data analysis and visualization. The web resource covers expert-curated proteomic data from COVID-19 patients published before May 2022. The data were collected from ProteomeXchange and from select publications via PubMed searches and aggregated into a comprehensive dataset. Protein expression by disease subgroups across projects was compared by examining differentially expressed proteins. We also visualize differentially expressed pathways and proteins. Moreover, circulating proteins that differentiated severe cases were nominated as predictive biomarkers. FindingsWe built and maintain a web server COVIDpro (https://www.guomics.com/covidPro/) containing proteomics data generated by 41 original studies from 32 hospitals worldwide, with data from 3077 patients covering 19 types of clinical specimens, the majority from plasma and sera. 53 protein expression matrices were collected, for a total of 5434 samples and 14,403 unique proteins. Our analyses showed that the lipopolysaccharide-binding protein, as identified in the majority of the studies, was highly expressed in the blood samples of patients with severe disease. A panel of significantly dysregulated proteins was identified to separate patients with severe disease from non-severe disease. Classification of severe disease based on these proteomic signatures on five test sets reached a mean AUC of 0.87 and ACC of 0.80. InterpretationCOVIDpro is an online database with an integrated analysis toolkit. It is a unique and valuable resource for testing hypotheses and identifying proteins or pathways that could be targeted by new treatments of COVID-19 patients. FundingNational Key R&D Program of China: Key PDPM technologies (2021YFA1301602, 2021YFA1301601, 2021YFA1301603), Zhejiang Provincial Natural Science Foundation for Distinguished Young Scholars (LR19C050001), Hangzhou Agriculture and Society Advancement Program (20190101A04), National Natural Science Foundation of China (81972492) and National Science Fund for Young Scholars (21904107), National Resource for Network Biology (NRNB) from the National Institute of General Medical Sciences (NIGMS-P41 GM103504) Research in contextO_ST_ABSEvidence before this studyC_ST_ABSAlthough an increasing number of therapies against COVID-19 are being developed, they are still insufficient, especially with the rise of new variants of concern. This is partially due to our incomplete understanding of the diseases mechanisms. As data have been collected worldwide, several questions are now worth addressing via meta-analyses. Most COVID-19 drugs function by targeting or affecting proteins. Effectiveness and resistance to therapeutics can be effectively assessed via protein measurements. Empowered by mass spectrometry-based proteomics, protein expression has been characterized in a variety of patient specimens, including body fluids (e.g., serum, plasma, urea) and tissue (i.e., formalin-fixed and paraffin-embedded (FFPE)). We expert-curated proteomic expression data from COVID-19 patients published before May 2022, from the largest proteomic data repository ProteomeXhange as well as from literature search engines. Using this resource, a COVID-19 proteome meta-analysis could provide useful insights into the mechanisms of the disease and identify new potential drug targets. Added value of this studyWe integrated many published datasets from patients with COVID-19 from 11 nations, with over 3000 patients and more than 5434 proteome measurements. We collected these datasets in an online database, and generated a toolbox to easily explore, analyze, and visualize the data. Next, we used the database and its associated toolbox to identify new proteins of diagnostic and therapeutic value for COVID-19 treatment. In particular, we identified a set of significantly dysregulated proteins for distinguishing severe from non-severe patients using serum samples. Implications of all the available evidenceCOVIDpro will support the navigation and analysis of patterns of dysregulated proteins in various COVID-19 clinical specimens for identification and verification of protein biomarkers and potential therapeutic targets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhang, F., Luna, A., Tan, T., Chen, Y., Sander, C., Guo, T.. 2022-09-28. COVIDpro: Database for mining protein dysregulation in patients with COVID-19. https://doi.org/10.1101/2022.09.27.509819

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

INFORME: coupling information-theoretic experimental design with nonlinear mixed-effects modeling for efficient observation scheduling

Mathematical models of treatment response can inform individualized therapy, but their calibration often requires longitudinal measurements that are costly, burdensome, and collected on fixed schedules. Such schedules may be inefficient, over-sampling patients whose response is already well characterized while delaying informative measurements for those whose model parameters remain uncertain. We present INFORME (INFORmation-theoretic design with Mixed Effects), a framework that combines Bayesian information-theoretic experimental design with nonlinear mixed-effects modeling to adaptively select each patients next measurement time. Population and response-subgroup parameter distributions learned from an existing cohort provide informative priors, allowing candidate measurement times to be ranked by their expected reduction in patient-specific parameter uncertainty. As observations accumulate, priors can be updated to reflect the response subgroup most consistent with the patients data. We evaluate INFORME in two radiotherapy datasets: 150 synthetic tumor volume trajectories from a hybrid cellular automaton model of prostate cancer spheroids (HD1) and longitudinal tumor volumes from 39 patients with head-and-neck cancer (HD2). In HD1, population priors allowed omission of both pretreatment scans, while adaptive scheduling reduced the protocol from nine scans to three or four, with the response group identified from a single post-treatment scan on day 27. In HD2, the adaptive schedule used three scans instead of six and improved prediction by delaying the first on-treatment scan from week 1 to week 2, avoiding transient dynamics that produced false-positive and false-negative response projections. Across both datasets, the adaptive schedules used a mean of 2.7 scans in stead of seven and advanced completion of the patient-specific prediction by a mean of 15.5 days (95% CI, 6.7-24.3) relative to the equidistant protocol, while treatment duration remained unchanged. INFORME therefore reduces measurement burden and accelerates patient-specific prediction by concentrating observations at times that are most informative for model calibration.

systems biology↗

Sobetirome, a thyroid hormone receptor beta agonist, is a potential therapeutic agent for pulmonary fibrosis

Idiopathic pulmonary fibrosis (IPF) is a progressive and fatal disease with limited treatment options. Our group previously identified the antifibrotic potential of thyroid hormone, triiodothyronine (T3); however, clinical translation of thyroid hormone therapy is limited by its systemic adverse effects. In this study, we investigate whether sobetirome, a selective and well tolerated thyroid hormone receptor beta (THRB) agonist, offers antifibrotic benefits of thyroid hormone while minimizing systemic toxicity. Our study reveals that sobetirome, administered via intraperitoneal or inhalational routes, effectively mitigates bleomycin-induced pulmonary fibrosis in mice, with no evidence of toxicity. We identified that sobetirome restores mitochondrial homeostasis via activating the THRB-PPARGC1a axis. This protects alveolar type II epithelial cells from injury-induced apoptosis while selectively inducing apoptosis and metabolic reprogramming in apoptosis resistant IPF fibroblasts. Cell-specific deletion of Ppargc1a in either alveolar epithelial cells or fibroblasts abolishes sobetirome-mediated protection, establishing PPARGC1a as an essential mediator of therapeutic response. Importantly, sobetirome reverses fibrosis-associated transcriptional programs in human IPF lung tissue, reducing expression of key fibrosis-associated genes, including collagen I alpha 1 (COL1A1), collagen III alpha 1 (COL3A1), periostin (POSTN), cathepsin K (CTSK), and Chitinase 3 Like 1 (CHI3L1), while promoting extracellular matrix remodeling, epithelial restoration, and tissue homeostasis. Collectively, our findings identify THRB activation as a novel metabolic strategy for reversing pulmonary fibrosis. Across complementary in vitro, in vivo, and human ex vivo models, sobetirome restores mitochondrial function, modulates apoptotic pathways in pathogenic cells, and promotes fibrosis resolution, highlighting its potential as a lung-targeted therapeutic approach for IPF and other fibrotic lung diseases.

systems biology↗

Mechanistic modeling of bacterial translation initiation across growth conditions

Translation frequency in bacteria depends on how ribosomes, mRNAs, and initiation factors are allocated across growth conditions. Here, we developed a mechanistic ODE-based model of Escherichia coli translation that represents initiation, elongation, termination, and coupled auxiliary processes. Growth-dependent abundances were derived from physiological relationships and reprocessed omics data, and simulated outputs were compared with translation-frequency and active-ribosome references. The model predicts a continuous shift from complex-formation-limited toward ribosome-limited behavior as growth increases. This shift is characterized by a decline in free-ribosome abundance, whereas initiation-factor pools remain largely unbound and do not become depleted in parallel. Together with the implemented IF-dependent kinetic term, this preserved availability provides a model-internal route through which productive initiation can be maintained despite increasing ribosome utilization. Consistently, transcript-wide ribosome loading remains below its theoretical maximum, while COG-level simulations reveal distinct sector-specific translation-frequency trajectories. The study therefore provides a resource-allocation framework for interpreting how mRNA--ribosome interactions shape bacterial translation across growth conditions.

systems biology↗