Search bioRxiv⌕ Search

Biology subjects

Metz, T.

Publications and source records attributed to Metz, T..

6 recordsLinked to original sources

Reference-free compound identification using computational prediction of molecular properties and multi-dimensional spectrometric measurements: a fentanyl case study

Mass spectrometry is used to identify chemicals to which humans are exposed, but it cannot directly determine molecular structures. Instead, structures are inferred by matching experimental spectra to libraries of spectra constructed from analyses of pure reference compounds. However, the chemical space of human exposures far exceeds the amount of experimental library spectra. Here, we evaluate a reference-free strategy for confident identification of unknown molecules. Using fentanyl as a case study, we created a suspect library of over 1 billion computationally predicted fentanyl analogs and predicted molecular properties through machine learning, molecular dynamics, and density functional theory. Multi-dimensional spectra from a blinded analysis of a mock fentanyl tablet were matched with the predicted library, yielding an average of three candidate structures per measured analog, with six exact identifications. This work emphasizes the promise of reference-free molecular measurements for assessing human exposure by merging computational predictions with high-dimensional measurements.

scientific communication and education↗

A comprehensive reference database to support untargeted metabolomics in Pseuudomonas putida

Pseudomonas putida strain KT2440 is a crucial model organism for synthetic biology and bioengineering applications, yet there currently exists no comprehensive metabolomics database comparable to those available for other model organisms. This gap hinders the use of untargeted metabolomics for exploratory analyses in this system. We developed the P. putida metabolome reference database (PPMDB v1) to address this limitation by consolidating metabolite information from multiple sources and expanding coverage through computational predictions. The database was constructed by curating metabolites from BioCyc, BiGG, and other literature sources, then computationally expanding this collection using BioTransformer environmental transformation predictions to generate additional predicted metabolites. We enhanced the databases utility for molecular annotation in metabolomics studies by incorporating analytical properties including collision cross-sections, tandem mass spectra, and gas-phase infrared spectra. These analytical properties were gathered from existing measurement data or predicted using computational tools. We further augmented the database through inclusion of reaction information and pathway annotations, facilitating biological interpretation of metabolomics data. This publicly available resource fills a critical gap in P. putida research infrastructure, supporting metabolite annotation and biological interpretation in untargeted metabolomics studies and enabling in-depth exploratory analyses of this important synthetic biology platform at the molecular level. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=110 SRC="FIGDIR/small/713193v1_ufig1.gif" ALT="Figure 1"> View larger version (26K): org.highwire.dtl.DTLVardef@c8828forg.highwire.dtl.DTLVardef@1f3a5c5org.highwire.dtl.DTLVardef@1084535org.highwire.dtl.DTLVardef@1f7ca4a_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

A Benchmarking Study of Feature Screening Approaches Across Omics Classification Settings

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime. Author SummaryA common goal when analyzing high throughput omics data is to identify key biomolecules which may be important in complex biological systems. This type of large-scale analysis is critical, as it can significantly reduce the set of total biomolecules which are considered for more detailed, targeted analysis. Additionally, biomolecules which are highly predictive of specific phenotypes are useful markers to guide treatment. For instance, identifying which biomolecules are predictive of type 1 diabetes could improve early diagnosis, which may improve patient prognosis. However, modern instrumentation can often detect thousands or tens of thousands of biomolecules, and studies typically have a disproportionately limited number of samples. Many measured biomolecules, or features, are often noisy and uninformative, and overcoming this noise to identify an informative set of biomolecules can be challenging for machine learning (ML) models. Fortunately, there are many feature selection strategies which can intelligently reduce the number of "uninformative" biomolecules an ML model needs to overcome. In this work, sure screening, which is a class of feature selection methods that retains the true feature set under certain conditions, is evaluated in the context of omics data analyses. The pros and cons of various sure screening methods, including software availability, are summarized, and their use in the larger feature selection literature is contextualized. Additionally, we benchmark performance by applying a suite of sure screening approaches on several real omics datasets generated to understand the progression of type 1 diabetes.

bioinformatics↗

Dysregulation of lung epithelial cell homeostasis and immunity contributes to Middle East Respiratory Syndrome coronavirus disease severity

Coronaviruses (CoV) emerge suddenly from animal reservoirs to cause novel diseases in new hosts. Discovered in 2012, Middle East respiratory syndrome coronavirus (MERS-CoV) is endemic in camels in the Middle East and is continually causing local outbreaks and epidemics. While all three newly emerging human CoV from past 20 years (SARS-CoV, SARS-CoV-2, MERS-CoV) cause respiratory disease, each CoV has unique host interactions that drive differential pathogeneses. To better understand the virus and host interactions driving lethal MERS-CoV infection, we performed a longitudinal multi-omics analysis of sublethal and lethal MERS-CoV infection in mice. Significant differences were observed in body weight loss, virus titers and acute lung injury among lethal and sub-lethal virus doses. Virus induced apoptosis of type I and II alveolar epithelial cells suggest that loss or dysregulation of these key cell populations was a major driver of severe disease. Omics analysis suggested differential pathogenesis was multi-factorial with clear differences among innate and adaptive immune pathways as well as those that regulate lung epithelial homeostasis. Infection of mice lacking functional T and B-cells showed that adaptive immunity was important in controlling viral replication but also increased pathogenesis. In summary, we provide a high-resolution host response atlas for MERS-CoV infection and disease severity. Multi-omics studies of viral pathogenesis offer a unique opportunity to not only better understand the molecular mechanisms of disease but also to identify genes and pathways that can be exploited for therapeutic intervention all of which is important for our future pandemic preparedness. ImportanceEmerging coronaviruses like SARS-CoV, SARS-CoV-2 and MERS-CoV cause a range of disease outcomes in humans from asymptomatic, moderate and severe respiratory disease which can progress to death but the factors causing these disparate outcomes remain unclear. Understanding host responses to mild and life-threatening infection provides insight into virus-host networks within and across organ systems that contribute to disease outcomes. We used multi-omics approaches to comprehensively define the host response to moderate and severe MERS-CoV infection. Severe respiratory disease was associated with dysregulation of the immune response. Key lung epithelial cell populations that are essential for lung function get infected and die. Mice lacking key immune cell populations experienced greater virus replication but decreased disease severity implicating the immune system in both protective and pathogenic roles in the response to MERS-CoV. These data could be utilized to design new therapeutic strategies targeting specific pathways that contribute to severe disease.

microbiology↗

Associations between SARS-CoV-2 Infection or COVID-19 Vaccination and Human Milk Composition: A Multi-Omics Approach

The risk of contracting SARS-CoV-2 via human milk-feeding is virtually non-existent. Adverse effects of COVID-19 vaccination for lactating individuals are not different from the general population, and no evidence has been found that their infants exhibit adverse effects. Yet, there remains substantial hesitation among this population globally regarding the safety of these vaccines. Herein we aimed to determine if compositional changes in milk occur following infection or vaccination, including any evidence of vaccine components. Using an extensive multi-omics approach, we found that compared to unvaccinated individuals SARS-CoV-2 infection was associated with significant compositional differences in 67 proteins, 385 lipids, and 13 metabolites. In contrast, COVID-19 vaccination was not associated with any changes in lipids or metabolites, although it was associated with changes in 13 or fewer proteins. Compositional changes in milk differed by vaccine. Changes following vaccination were greatest after 1-6 hours for the mRNA-based Moderna vaccine (8 changed proteins), 3 days for the mRNA-based Pfizer (4 changed proteins), and adenovirus-based Johnson and Johnson (13 changed proteins) vaccines. Proteins that changed after both natural infection and Johnson and Johnson vaccine were associated mainly with systemic inflammatory responses. In addition, no vaccine components were detected in any milk sample. Together, our data provide evidence of only minimal changes in milk composition due to COVID-19 vaccination, with much greater changes after natural SARS-CoV-2 infection. IMPORTANCEThe impact of the observed changes in global milk composition on infant health remain unknown. These findings emphasize the importance of vaccinating the lactating population against COVID-19, as compositional changes in milk were found to be far less evident after vaccination compared to SARS-CoV-2 infection. Importantly, vaccine components were not detected in milk after vaccination.

molecular biology↗

Shifts from non-obligate generalists to obligate specialists in simulations of mutualistic network assembly

Understanding ecosystem recovery after perturbation is crucial for ecosystem conservation. Mutualisms contribute key functions for plants such as pollination and seed dispersal. We modelled the assembly of mutualistic networks based on trait matching between plants and their animal partners that have different degrees of specialization on plant traits. Additionally, we addressed the role of non-obligate animal mutualists, including facultative mutualists or non-resident species that have their main resources outside the target site. Our computer simulations show that non-obligate animals facilitate network assembly during the early stages, furthering colonization by an increase in niche space and reduced competition. While non-obligate and generalist animals provide most of the fitness benefits to plants in the early stages of the assembly, obligate and specialist animals dominate at the end of the assembly. Our results thus demonstrate the combined occurrence of shifts from diet, trait, and habitat generalists to more specialised animals.

ecology↗