Search bioRxiv⌕ Search

Biology subjects

Melendez, C.

Publications and source records attributed to Melendez, C..

3 recordsLinked to original sources

Unified imputation of missing data modalities and features in multi-omic data via shared representation learning

Multi-omic studies promise a more comprehensive view of biological systems by jointly measuring multiple molecular layers. In practice, however, such datasets are rarely complete: entire molecular modalities may be missing for many samples, and observed modalities often contain substantial feature-level missingness. Existing imputation approaches typically address only one of these two problems, relying either on feature-level imputation within a single modality or on pairwise translation models that cannot accommodate arbitrary combinations of missing modalities. We present MIMIR, a deep learning framework for unified multi-omic imputation of bulk data that addresses both missing modalities and missing values through shared representation learning. MIMIR first learns modality-specific representations using masked autoencoders and then projects these representations into a common latent space, enabling reconstruction from any subset of observed modalities. Evaluated on pan-cancer multi-omic data from The Cancer Genome Atlas, MIMIR consistently outperforms baseline methods across a range of missing-modality and missing-value scenarios, including missing completely at random and missing not at random settings. Analysis of the learned shared space reveals structured cross-modal dependencies that explain modality-specific differences in imputation accuracy, with transcriptional and epigenetic modalities forming a strongly aligned core and copy number variation contributing more distinct signal. Together, these results demonstrate that shared representation learning provides an effective and flexible foundation for multi-omic imputation under heterogeneous patterns of missingness.

bioinformatics↗

Accounting for digestion enzyme bias in Casanovo

A key parameter of any proteomics mass spectrometry experiment is the identity of the enzyme that is used to digest proteins in the sample into peptides. The Casanovo de novo sequencing model was trained using data that was generated with trypsin digestion; consequently, the model prefers to predict peptides that end with the amino acids "K" or "R." This bias is desirable when the Casanovo is used to analyze data that was also generated using trypsin but can be problematic if the data was generated using some other digestion enzyme. In this work, we modify Casanovo to take as input the identify of the digestion enzyme, alongside each observed spectrum. We then train Casanovo with data generated using several different restriction enzymes, and we demonstrate that the resulting model successfully learns to capture enzyme-specific behavior. However, we find, surprisingly, that this new model does not yield a significant improvement in sequencing accuracy relative to a model trained without the enzyme information but using the same training set. This observation may have important implications for future attempts to make use of experimental metadata in de novo sequencing models.

bioinformatics↗

In Vitro Antibacterial Activity of Dinuclear Thiolato-Bridged Ruthenium(II)-Arene Compounds

The antibacterial activity of 22 thiolato-bridged dinuclear ruthenium(II)-arene compounds was assessed in vitro against Escherichia coli, Streptococcus pneumoniae and Staphylococcus aureus. None of the compounds efficiently inhibited the growth of the three E. coli strains tested and only compound 5 exhibited a medium activity against this bacterium (MIC (minimum inhibitory concentration) of 25 M). However, a significant antibacterial activity was observed against S. pneumoniae, with MIC values ranging from 1.3 to 2.6 M for compounds 1-3, 5 and 6. Similarly, compounds 2, 5-7 and 20-22 had MIC values ranging from 2.5 to 5 M against S. aureus. The tested diruthenium compounds have a bactericidal effect significantly faster than that of penicillin. Fluorescence microscopy assays performed on S. aureus using the BODIPY-tagged diruthenium complex 15 showed that this type of metal compound enter the bacteria and do not accumulate in the cell wall of gram-positive bacteria. Cellular internalization was further confirmed by inductively coupled plasma mass spectrometry (ICP-MS) experiments. The nature of the substituents anchored on the bridging thiols and the compounds molecular weight appear to significantly influence the antibacterial activity. Thus, if overall a decrease of the bactericidal effect with the increase of compounds molecular weight is observed, however the complexes bearing larger benzo-fused lactam substituents had low MIC values. This first antibacterial activity screening demonstrated that the thiolato-diruthenium compounds exhibit promising activity against S. aureus and S. pneumoniae and deserve to be considered for further studies.

pharmacology and toxicology↗