Search bioRxiv⌕ Search

Biology subjects

Werner, C. M.

Publications and source records attributed to Werner, C. M..

2 recordsLinked to original sources

Interpretable and predictive models to harness the life science data revolution

The proliferation of high-dimensional data in ecology and evolutionary biology raises the promise of statistical and machine learning models that are highly predictive and interpretable. However, high-dimensional data are commonly burdened with an inherent trade-off: in-sample prediction of outcomes will improve as additional variables are included in the model, but this may come at the cost of poor predictive accuracy and limited generalizability for future or unsampled observations (out-of-sample prediction). To confront this problem of overfitting, sparse models can focus on key variables by correctly placing low weight on unimportant variables. We competed nine methods to quantify their performance in variable selection and prediction using simulated data with different sample sizes, numbers of variables, and strengths of effects. Overfitting was typical for many methods and simulation scenarios. Despite this, in-sample and out-of-sample prediction converged on the true predictive target for simulations with more observations, larger causal effects, and fewer variables. Accurate variable selection to support process-based understanding will be unattainable for many realistic sampling schemes in ecology and evolution. We use our analyses to characterize data attributes for which statistical learning is possible, and illustrate how some sparse methods can achieve predictive accuracy while mitigating and learning the extent of overfitting.

genomics↗

Disentangling key species interactions in diverse and heterogeneous communities: A Bayesian sparse modeling approach

1Modeling species interactions in diverse communities traditionally requires a prohibitively large number of species-interaction coefficients, especially when considering environmental dependence of parameters. We implemented Bayesian variable selection via sparsity-inducing priors on non-linear species abundance models to determine which species-interactions should be retained and which can be represented as an average heterospecific interaction term, reducing the number of model parameters. We evaluated model performance using simulated communities, computing out-of-sample predictive accuracy and parameter recovery across different input sample sizes. We applied our method to a diverse empirical community, allowing us to disentangle the direct role of environmental gradients on species intrinsic growth rates from indirect effects via competitive interactions. We also identified a few neighboring species from the diverse community that had non-generic interactions with our focal species. This sparse modeling approach facilitates exploration of species-interactions in diverse communities while maintaining a manageable number of parameters.

ecology↗