Search bioRxiv⌕ Search

Biology subjects

Dauda, K. A.

Publications and source records attributed to Dauda, K. A..

4 recordsLinked to original sources

Exploring the Flexible Penalization of Bayesian Survival Analysis Using Beta Process Prior for Baseline Hazard

High-dimensional data has significantly captured the interest of many researchers, particularly in the context of variable selection. However, when dealing with time-to-event data in survival analysis, where censoring is a key consideration, progress in addressing this complex problem has remained somewhat limited. More-over, in microarray research, it is common to identify groupings of genes involved in the same biological pathways. These gene groupings frequently collaborate and operate as a unified entity. Therefore, this study is motivated to adopt the idea of a Penalized semi-parametric Bayesian Cox (PSBC) model through elastic-net and group lasso penalty functions (PSBC-EN-G and PSBC-GL-G) to incorporate the grouping structure of the covariates (genes) and optimally perform variable selection. The proposed methods assign a beta prior process to the cumulative baseline hazard function (PSBC-EN-B and PSBC-GL-B), instead of the gamma prior process used in existing methods (PSBC-EN-G and PSBC-GL-G). Three real-life datasets and simulation scenarios were considered to compare and validate the efficiency of the modified methods with existing techniques, using Bayesian Information Criteria (BIC). The results of the simulated studies provided empirical evidence that the proposed methods performed better than the existing methods across a wide range of data scenarios. Similarly, the results of the real-life study showed that the proposed methods revealed a substantial improvement over the existing techniques in terms of feature selection and grouping behavior.

bioengineering↗

Clustering large-scale biomedical data to model dynamic accumulation processes in disease progression and anti-microbial resistance evolution

Accumulation modelling uses machine learning to discover the dynamics by which systems acquire discrete features over time. Many systems of biomedical interest show such dynamics: from bacteria acquiring resistances to sets of drugs, to patients acquiring symptoms during the course of progressive disease. Existing approaches for accumulation modelling are typically limited either in the number of features they consider or their ability to characterise interactions between these features - a limitation for the large-scale genetic and/or phenotypic datasets often found in modern biomedical applications. Here, we demonstrate how clustering can make such large-scale datasets tractable for powerful accumulation modelling approaches. Clustering resolves issues of sparsity and high dimensionality in datasets but complicates the intepretation of the inferred dynamics, especially if observations are not independent. Focussing on hypercubic hidden Markov models (HyperHMM), we introduce several approaches for interpreting, estimating, and bounding the results of the dynamics in these cases and show how biomedical insight can be gained in such cases. We demonstrate this Cluster-based HyperHMM (CHyperHMM) pipeline for synthetic data, clinical data on disease progression in severe malaria, and genomic data for anti-microbial resistance evolution in Klebsiella pneumoniae, reflecting two global health threats.

bioinformatics↗

HyperTraPS-CT: Inference and prediction for accumulation pathways with flexible data and model structures

Accumulation processes, where many potentially coupled features are acquired over time, occur throughout the sciences, from evolutionary biology to disease progression, and particularly in the study of cancer progression. Existing methods for learning the dynamics of such systems typically assume limited (often pairwise) relationships between feature subsets, cross-sectional or untimed observations, small feature sets, or discrete orderings of events. Here we introduce HyperTraPS-CT (Hypercubic Transition Path Sampling in Continuous Time) to compute posterior distributions on continuous-time dynamics of many, arbitrarily coupled, traits in unrestricted state spaces, accounting for uncertainty in observations and their timings. We demonstrate the capacity of HyperTraPS-CT to deal with cross-sectional, longitudinal, and phylogenetic data, which may have no, uncertain, or precisely specified sampling times. HyperTraPS-CT allows positive and negative interactions between arbitrary subsets of features (not limited to pairwise interactions), supporting Bayesian and maximum-likelihood inference approaches to identify these interactions, consequent pathways, and predictions of future and unobserved features. We also introduce a range of visualisations for the inferred outputs of these processes and demonstrate model selection and regularisation for feature interactions. We apply this approach to case studies on the accumulation of mutations in cancer progression and the acquisition of anti-microbial resistance genes in tuberculosis, demonstrating its flexibility and capacity to produce predictions aligned with applied priorities.

bioinformatics↗

A New Generalized Gamma-Weibull Distribution with Applications toTime-to-event Data

In this research, a new class of probability distributions referred to as Generalized Gamma Weibull (GGW) distributions was introduced within the context of parametric survival analysis. This distribution represents a modification of the gamma Weibull distribution and offers valuable insights, particularly when dealing with highly skewed lifetime data. The study extensively examined the mathematical characteristics of these distributions, encompassing hazard functions, moments, quantile functions, and order statistics. Furthermore, the research delved into parameter estimation methods for these newly proposed distributions, employing the maximum likelihood technique, Fisher Information (FI), and deriving asymptotic confidence intervals for both censored and uncensored scenarios. To illustrate the practical utility of these proposed distributions, the study applied them to analyze two sets of real-life survival data and two sets of real-life data, resulting in a total of four distinct datasets. To gauge the effectiveness of the GGW distributions in comparison to existing methods such as Generalized Weibull and Generalized gamma (G-Weibull and G-Gamma) distributions, the research employed statistical indices including the Akaike Information Criterion (AIC), Corrected Akaike Information Criterion (CAIC), and Bayesian Information Criterion (BIC). The outcomes of this comparative analysis demonstrated the superior performance of the newly introduced GGW distributions (AIC=338.6313, BIC=346.2794, and CAIC=339.5202) when contrasted with the existing methods (G-Weibull: AIC=376.1946, BIC=381.9307, and CAIC=376.5424) across all three criteria, thereby highlighting the enhanced suitability of GGW distributions for modeling and analyzing skewed lifetime data.

cancer biology↗