Search bioRxiv⌕ Search

Biology subjects

Beder, T.

Publications and source records attributed to Beder, T..

3 recordsLinked to original sources

Machine learning on large scale perturbation screens for SARS-CoV-2 host factors identifies β-catenin/CBP inhibitor PRI-724 as a potent antiviral

Expanding antiviral treatment options against SARS-CoV-2 remains crucial as the virus evolves rapidly and drug resistant strains have emerged. Broad spectrum host-directed antivirals (HDA) are promising therapeutic options, however the robust identification of relevant host factors by CRISPR/Cas9 or RNA interference screens remains challenging due to low consistency in the resulting hits. To address this issue, we employed machine learning based on experimental data from knockout screens and a drug screen. As gold standard, we assembled perturbed genes reducing virus replication or protecting the host cells. The machines based their predictions on features describing cellular localization, protein domains, annotated gene sets from Gene Ontology, gene and protein sequences, and experimental data from proteomics, phospho-proteomics, protein interaction and transcriptomic profiles of SARS-CoV-2 infected cells. The models reached a remarkable performance with a balanced accuracy of 0.82 (knockout based classifier) and 0.71 (drugs screen based classifier), suggesting patterns of intrinsic data consistency. The predicted host dependency factors were enriched in sets of genes particularly coding for development, morphogenesis, and neural related processes. Focusing on development and morphogenesis-associated gene sets, we found {beta}-catenin to be central and selected PRI-724, a canonical {beta}-catenin/CBP disruptor, as a potential HDA. PRI-724 limited infection with SARS-CoV-2 variants, SARS-CoV-1, MERS-CoV and IAV in different cell line models. We detected a concentration-dependent reduction in CPE development, viral RNA replication, and infectious virus production in SARS-CoV-2 and SARS-CoV-1-infected cells. Independent of virus infection, PRI-724 treatment caused cell cycle deregulation which substantiates its potential as a broad spectrum antiviral. Our proposed machine learning concept may support focusing and accelerating the discovery of host dependency factors and the design of antiviral therapies. Authors summaryDrug resistance to pathogens is a well-known phenomenon which was also observed for SARS-CoV-2. Given the gradually increasing evolutionary pressure on the virus by herd immunity, we attempted to enlarge the available antiviral repertoire by focusing on host proteins that are usurped by viruses. The identification of such proteins was followed within several high throughput screens in which genes are knocked out individually. But, so far, these efforts led to very different results. Machine learning helps to identify common patterns and normalizes independent studies to their individual designs. With such an approach, we identified genes that are indispensable during embryonic development, i.e., when cells are programmed for their specific destiny. Shortlisting the hits revealed {beta}-catenin, a central player during development, and PRI-724, which inhibits the interaction of {beta}-catenin with cAMP responsive element binding (CREB) binding protein (CBP). In our work, we confirmed that the disruption of this interaction impedes virus replication and production. In A549-AT cells treated with PRI-724, we observed cell cycle deregulation which might contribute to the inhibition of virus infection, however the exact underlying mechanisms needs further investigation.

microbiology↗

The gene expression classifier ALLCatchR identifies B-precursor ALL subtypes and underlying developmental trajectories across age

Current classifications (WHO-HAEM5 / ICC) define up to 26 molecular B-cell precursor acute lymphoblastic leukemia (BCP-ALL) disease subtypes, which are defined by genomic driver aberrations and corresponding gene expression signatures. Identification of driver aberrations by RNA-Seq is well established, while systematic approaches for gene expression analysis are less advanced. Therefore, we developed ALLCatchR, a machine learning based classifier using RNA-Seq expression data to allocate BCP-ALL samples to 21 defined molecular subtypes. Trained on n=1,869 transcriptome profiles with established subtype definitions (4 cohorts; 55% pediatric / 45% adult), ALLCatchR allowed subtype allocation in 3 independent hold-out cohorts (n=1,018; 75% pediatric / 25% adult) with 95.7% accuracy (averaged sensitivity across subtypes: 91.1% / specificity: 99.8%). High confidence predictions were achieved in 84.6% of samples with 99.7% accuracy. Only 1.2% of samples remained unclassified. ALLCatchR outperformed existing tools and identified novel candidates in previously unassigned samples. We established a novel RNA-Seq reference of human B-lymphopoiesis. Implementation in ALLCatchR enabled projection of BCP-ALL samples to this trajectory, which identified shared patterns of proximity of BCP-ALL subtypes to normal lymphopoiesis stages. ALLCatchR sustains RNA-Seq routine application in BCP-ALL diagnostics with systematic gene expression analysis for accurate subtype allocations and novel insights into underlying developmental trajectories.

cancer biology↗

Identifying essential genes across eukaryotes by machine learning

Identifying essential genes on a genome scale is resource intensive and has been performed for only a few eukaryotes. For less studied organisms essentiality might be predicted by gene homology. However, this approach cannot be applied to non-conserved genes. Additionally, divergent essentiality information is obtained from studying single cells or whole, multi-cellular organisms, and particularly when derived from human cell line screens and human population studies. We employed machine learning across six model eukaryotes and 60,381 genes, using 41,635 features derived from sequence, gene functions and network topology. Within a leave-one-organism-out cross-validation, the classifiers showed a high generalizability with an average accuracy close to 80% in the left-out species. As a case study, we applied the method to Tribolium castaneum and validated predictions experimentally yielding similar performance. Finally, using the classifier based on the studied model organisms enabled linking the essentiality information of human cell line screens and population studies.

bioinformatics↗