Search bioRxivSearch

bioRxiv · 10.1101/350017

The intrinsic predictability of ecological time series and its potential to guide forecasting

Abstract

Successfully predicting the future states of systems that are complex, stochastic and potentially chaotic is a major challenge. Model forecasting error (FE) is the usual measure of success; however model predictions provide no insights into the potential for improvement. In short, the realized predictability of a specific model is uninformative about whether the system is inherently predictable or whether the chosen model is a poor match for the system and our observations thereof. Ideally, model proficiency would be judged with respect to the systems intrinsic predictability - the highest achievable predictability given the degree to which system dynamics are the result of deterministic v. stochastic processes. Intrinsic predictability may be quantified with permutation entropy (PE), a model-free, information-theoretic measure of the complexity of a time series. By means of simulations we show that a correlation exists between estimated PE and FE and show how stochasticity, process error, and chaotic dynamics affect the relationship. This relationship is verified for a dataset of 461 empirical ecological time series. We show how deviations from the expected PE-FE relationship are related to covariates of data quality and the nonlinearity of ecological dynamics.\n\nThese results demonstrate a theoretically-grounded basis for a model-free evaluation of a systems intrinsic predictability. Identifying the gap between the intrinsic and realized predictability of time series will enable researchers to understand whether forecasting proficiency is limited by the quality and quantity of their data or the ability of the chosen forecasting model to explain the data. Intrinsic predictability also provides a model-free baseline of forecasting proficiency against which modeling efforts can be evaluated.\n\nGlossaryActive information: The amount of information that is available to forecasting models (redundant information minus lost information; Fig. 1).\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=45 SRC=\"FIGDIR/small/350017_fig1a.gif\" ALT=\"Figure 1A\">\nView larger version (9K):\norg.highwire.dtl.DTLVardef@934d9eorg.highwire.dtl.DTLVardef@ccdc10org.highwire.dtl.DTLVardef@1839ed9org.highwire.dtl.DTLVardef@31bd70_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1A.C_FLOATNO The total information content of an observation of a system at a given state in time, St, is depicted by filled circles with past states (St-1 and St-2) represented by shades of grey, i) lack of overlap between past and present states illustrating a case where no information is transmitted from past states (i.e. a purely stochastic system), with low redundancy and high Shannon entropy rate, ii) intermediate overlap indicating a case when some information is transferred from past to present (i.e. a deterministic system strongly driven by stochastic forcing), with intermediate redundancy and Shannon entropy rate, iii) large overlap indicating a case when the current state is mostly determined by the previous state (i.e. a highly deterministic system), with high redundancy and low Shannon entropy rate. Note that both the redundancy and Shannon entropy rate of a system are intrinsic properties of the system and will only change if the system itself changes.\n\nC_FIG Forecasting error (FE): A measure of the discrepancy between a models forecasts and the observed dynamics of a system. Common measures of forecast error are root mean squared error and mean absolute error.\n\nEntropy: Measures the average amount of information in the outcome of a stochastic process.\n\nInformation: Any entity that provides answers and resolves uncertainty about a process. When information is calculated using logarithms to the base two (i.e. information in bits), it is the minimum number of yes/no questions required, on average, to determine the identity of the symbol (Jost 2006). The information in an observation consists of information inherited from the past (redundant information), and of new information.\n\nIntrinsic predictability: the maximum achievable predictability of a system (Beckage et al. 2011).\n\nLost information: The part of the redundant information lost due to measurement or sampling error, or transformations of the data (Fig. 1).\n\nNew information, Shannon entropy rate: The Shannon entropy rate quantifies the average amount of information per observation in a time series that is unrelated to the past, i.e., the new information (Fig. 1).\n\nNonlinearity: When the deterministic processes governing system dynamics depend on the state of the system.\n\nPermutation entropy (PE): permutation entropy is a measure of the complexity of a time series (Bandt & Pompe, 2002) that is negatively correlated with a systems predictability (Garland et al. 2015). Permutation entropy quantifies the combined new and lost information. PE is scaled to range between a minimum of 0 and a maximum of 1.\n\nRealized predictability: the achieved predictability of a system from a given forecasting model.\n\nRedundant information: The information inherited from the past, and thus the maximum amount of information available for use in forecasting (Fig. 1).\n\nSymbols, words, permutations: symbols are simply the smallest unit in a formal language such as the letters in the English alphabet i.e., {\"A\", \"B\",..., \"Z\"}. In information theory the alphabet is more abstract, such as elements in the set {\"up\", \"down\"} or {\"1\", \"2\", \"3\"}. Words, of length m refer to concatenations of the symbols (e.g., up-down-down) in a set. Permutations are the possible orderings of symbols in a set. In this manuscript, the words are the permutations that arise from the numerical ordering of m data points in a time series.\n\nWeighted permutation entropy (WPE): a modification of permutation entropy (Fadlallah et al., 2013) that distinguishes between small-scale, noise-driven variation and large-scale, system-driven variation by considering the magnitudes of changes in addition to the rank-order patterns of PE.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Pennekamp, F., Iles, A., Garland, J., Brennan, G., Brose, U., Gaedke, U., Jacob, U., Kratina, P., Matthews, B., Munch, S., Novak, M., Palamara, G. M., Rall, B., Rosenbaum, B., Tabi, A., Ward, C., Williams, R., Ye, H., Petchey, O.. 2018-06-19. The intrinsic predictability of ecological time series and its potential to guide forecasting. https://doi.org/10.1101/350017

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Beyond Single-Metric Assessments: Uncovering Masked Butterfly Declines via Multi-Scalar Analysis in Central Alberta

1. This study analyzed 21 years (2000-2025) of butterfly count data from Central Alberta, integrated with intensive 5-year (2021-2025) high-resolution intra-seasonal sampling. 2. Long-term macro-scale analysis revealed a significant decline in Shannon Diversity, a change that remained obscured when relying solely on traditional metrics of species richness and evenness. 3. This diversity decline was primarily driven by the severe, long-term collapse of the native Common Ringlet (Coenonympha tullia). 4. Four other dominant species--Cabbage White (Pieris rapae), Clouded Sulphur (Colias eriphyle), European Skipper (Thymelicus lineola), and Common Wood Nymph (Cercyonis pegala)--maintained long-term population stability, though their abundances were significantly constrained by extreme winter minimum temperatures and rapid spring warming. 5. High-resolution intra-seasonal analysis (2021-2025) demonstrated that community indices and species-specific abundances were strongly limited by daily weather, particularly wind velocity and temperature. 6. These findings illustrate that while traditional metrics like richness and evenness are fundamental to community ecology, they provide incomplete insights when applied in isolation; they are most effective when utilized as part of a complementary, multi-scalar framework. 7. This study highlights the necessity of coupling multi-decadal historical datasets with high-frequency, fine-scale sampling to accurately identify the mechanisms of community turnover that simpler metrics may overlook. 8. The results underscore the critical importance of standardized citizen science monitoring in quantifying environmental impacts and establishing conservation priorities for terrestrial insect groups.

ecology

From concentration to export: resource contrasts and bee traits shape pollinator spillover to crops

Floral plantings can either concentrate bees or export them to adjacent crops, yet the ecological conditions influencing these outcomes remain unclear. Here, we develop a mathematical model as proof of concept for our previous integrative hypothesis: concentrator and exporter outcomes can arise as alternative, context-dependent outcomes of the same underlying resource-selection process. Using bees as a model and focusing specifically on spillover from floral plantings to crops, we identified resource-specific thresholds separating concentration- and export-favoring conditions. Our model translates differences in relative patch attractiveness into context-dependent concentration and export outcomes and generates resource-specific, testable predictions about the conditions favoring pollinator movement into crops. In our simulations, the concentrator-exporter transition occurred at a lower flowering-intensity contrast than at pollen or nectar contrasts, which suggests that flowering intensity may provide an initial cue for bee movement, whereas nectar and pollen rewards refine or sustain bee responses once crops are perceived as attractive. Spillover thresholds differed among resource contrasts, whereas response steepness varied across bee-trait and community scenarios. Under the model's trait-sensitivity formulation, predicted spillover probability responded more strongly to flowering contrast for specialists than for generalists; colony size amplified this response, whereas bee richness dampened it. Together, these patterns show how flowering and resource contrasts interact with bee traits and community context to shape predicted spillover. Our results confirm that the concentrator and exporter hypotheses can be understood as context-dependent outcomes of the same ecological process rather than as mutually exclusive alternatives. Experimental tests of the predicted thresholds conducted in the field could reveal when and where floral plantings are most likely to promote bee spillover to crops, potentially supporting crop pollination.

ecology

A Computational Re-evaluation of Spatial Trials for Zoonotic Tuberculosis Control: Model Misspecification, Diagnostic Miss-classification, and the Illusion of Wildlife Culling Efficacy

1. Wildlife reservoir management frequently relies on the Randomised Badger Culling Trial's (RBCT) trade-off hypothesis, which posits that reductions in cattle herd infections are offset by a perturbation effect driven by disrupted host dispersal. This paper evaluates the computational and epidemiological robustness of this historical trial, which serves as the foundational empirical experiment guiding zoonotic tuberculosis (Mycobacterium bovis) control policies. 2. Using generalized linear mixed models with a generalized Poisson error distribution to explicitly address historical data overdispersion, this study contrasts traditional parametric inference against exact cluster-constrained permutation tests across distinct operational definitions of disease incidence. 3. Non-parametric diagnostics reveal that previously reported treatment and perturbation effects render as statistical artifacts under exact non-parametric permutation. Inside culling zones, parametric significance fails to withstand exact permutation verification due to extreme data leverage in localized cluster blocks. 4. Crucially, when diagnostic misclassification biases are eliminated by analysing total reactor datasets, all apparent culling effects disappear, and information criteria overwhelmingly favour nested null architectures. Unconfirmed reactors likely represent true biological infections missed by low-sensitivity post-mortem macro-necropsy, proving that host removal tracks observation noise rather than genuine zoonotic transmission pathways. 5. Finally, empirical scaling conducted in this study identifies a novel mathematical saturation effect, demonstrating that this sub-linear scaling is an operational artifact of unmodelled herd-level disease recurrence over time. 6. Policy implications. Because current zoonotic tuberculosis intervention frameworks are built upon a structurally misspecified statistical model, they have driven large-scale veterinary policies resulting in substantial, unevidenced ecological and economic interventions while failing to provide genuine public health, animal health, or disease control benefits.

ecology