Search bioRxiv⌕ Search

Biology subjects

Anderson, O.

Publications and source records attributed to Anderson, O..

3 recordsLinked to original sources

Evaluating Multiomics Integration Architectures for Training With Structured Missingness

Multimodal bioinformatics datasets are increasingly common in biomedical research, for tasks such as cancer subtyping and outcome prediction. It is feasible that data from a given patient, or even all data from a given institution, does not have coverage across all modalities; the available data is contingent on both the assay choice at the institution alongside technical aspects and associated drop-out. Consequently, algorithms for machine learning models must be tolerant to structured data missingness (or occurrence) for entire modalities when training. In this paper, we compare general strategies for training multimodal models in the context of structured modality missingness, employing suitable strategies for the stage of modality integration: early by concatenation of features, intermediate by max pooling of latent features, and late by aggregating model predictions probabilistically. We evaluate our strategies on a real-world bioinformatics dataset for the task of breast cancer subtyping, constructing a range of structured missingness scenarios. We highlight that, despite their inability to learn cross-modality interactions, late integration models outperform against early and intermediate integration strategies across a range of scenarios according to the level and nature of missingness. Logistic regression models, although simple, also outperform neural networks within the same settings. Fundamentally, we show that understanding the structure of missingness within a dataset is necessary when selecting a method of integration, and that simple models and approaches should not be dismissed when working with structured missingness.

bioinformatics↗

GRPhIN: Graphlet Characterization of Regulatory and Physical Interaction Networks

Graphs are powerful tools for modeling and analyzing molecular interaction networks. Graphs typically represent either undirected physical interactions or directed regulatory relationships, which can obscure a particular proteins functional context. Graphlets can describe local topologies and patterns within graphs, and combining physical and regulatory interactions offer new graphlet configurations that can provide biological insights. We present GRPhIN, a tool for characterizing graphlets and protein roles within graphlets in mixed physical and regulatory interaction networks. We describe the graphlets of mixed networks in B. subtilis, C. elegans, D. melanogaster, D. rerio, and S. cerevisiae and examine local topologies of proteins and sub- networks related to the oxidative stress response pathway. We found a number of graphlets that were abundant in all species, specific node positions (orbits) within graphlets that were over-represented in stress-associated proteins, and rarely-occurring graphlets that were over-represented in oxidative stress subnetworks. These results showcase the potential for using graphlets in mixed physical and regulatory interaction networks to identify new patterns beyond a single interaction type.

systems biology↗

ProteinWeaver: A Webtool to Visualize Ontology-Annotated Protein Networks

Molecular interaction networks are a vital tool for studying biological systems. While many tools exist that visualize a protein or a pathway within a network, no tool provides the ability for a researcher to consider a proteins position in a network in the context of a specific biological process or pathway. We developed ProteinWeaver, a web-based tool designed to visualize and analyze non-human protein interaction networks by integrating known biological functions. ProteinWeaver provides users with an intuitive interface to situate a user-specified protein in a user-provided biological context (as a Gene Ontology term) in five model organisms. Protein-Weaver also reports the presence of physical and regulatory network motifs within the queried subnetwork and statistics about the proteins distance to the biological process or pathway within the network. These insights can help researchers generate testable hypotheses about the proteins potential role in the process or pathway under study. Two cell biology case studies demonstrate ProteinWeavers potential to generate hypotheses from the queried subnetworks. ProteinWeaver is available at https://proteinweaver.reedcompbio.org/.

systems biology↗