Search bioRxivSearch

Biology subjects

Novak, M.

Publications and source records attributed to Novak, M..

8 recordsLinked to original sources

The intrinsic predictability of ecological time series and its potential to guide forecasting

Successfully predicting the future states of systems that are complex, stochastic and potentially chaotic is a major challenge. Model forecasting error (FE) is the usual measure of success; however model predictions provide no insights into the potential for improvement. In short, the realized predictability of a specific model is uninformative about whether the system is inherently predictable or whether the chosen model is a poor match for the system and our observations thereof. Ideally, model proficiency would be judged with respect to the systems intrinsic predictability - the highest achievable predictability given the degree to which system dynamics are the result of deterministic v. stochastic processes. Intrinsic predictability may be quantified with permutation entropy (PE), a model-free, information-theoretic measure of the complexity of a time series. By means of simulations we show that a correlation exists between estimated PE and FE and show how stochasticity, process error, and chaotic dynamics affect the relationship. This relationship is verified for a dataset of 461 empirical ecological time series. We show how deviations from the expected PE-FE relationship are related to covariates of data quality and the nonlinearity of ecological dynamics.\n\nThese results demonstrate a theoretically-grounded basis for a model-free evaluation of a systems intrinsic predictability. Identifying the gap between the intrinsic and realized predictability of time series will enable researchers to understand whether forecasting proficiency is limited by the quality and quantity of their data or the ability of the chosen forecasting model to explain the data. Intrinsic predictability also provides a model-free baseline of forecasting proficiency against which modeling efforts can be evaluated.\n\nGlossaryActive information: The amount of information that is available to forecasting models (redundant information minus lost information; Fig. 1).\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=45 SRC=\"FIGDIR/small/350017_fig1a.gif\" ALT=\"Figure 1A\">\nView larger version (9K):\norg.highwire.dtl.DTLVardef@934d9eorg.highwire.dtl.DTLVardef@ccdc10org.highwire.dtl.DTLVardef@1839ed9org.highwire.dtl.DTLVardef@31bd70_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1A.C_FLOATNO The total information content of an observation of a system at a given state in time, St, is depicted by filled circles with past states (St-1 and St-2) represented by shades of grey, i) lack of overlap between past and present states illustrating a case where no information is transmitted from past states (i.e. a purely stochastic system), with low redundancy and high Shannon entropy rate, ii) intermediate overlap indicating a case when some information is transferred from past to present (i.e. a deterministic system strongly driven by stochastic forcing), with intermediate redundancy and Shannon entropy rate, iii) large overlap indicating a case when the current state is mostly determined by the previous state (i.e. a highly deterministic system), with high redundancy and low Shannon entropy rate. Note that both the redundancy and Shannon entropy rate of a system are intrinsic properties of the system and will only change if the system itself changes.\n\nC_FIG Forecasting error (FE): A measure of the discrepancy between a models forecasts and the observed dynamics of a system. Common measures of forecast error are root mean squared error and mean absolute error.\n\nEntropy: Measures the average amount of information in the outcome of a stochastic process.\n\nInformation: Any entity that provides answers and resolves uncertainty about a process. When information is calculated using logarithms to the base two (i.e. information in bits), it is the minimum number of yes/no questions required, on average, to determine the identity of the symbol (Jost 2006). The information in an observation consists of information inherited from the past (redundant information), and of new information.\n\nIntrinsic predictability: the maximum achievable predictability of a system (Beckage et al. 2011).\n\nLost information: The part of the redundant information lost due to measurement or sampling error, or transformations of the data (Fig. 1).\n\nNew information, Shannon entropy rate: The Shannon entropy rate quantifies the average amount of information per observation in a time series that is unrelated to the past, i.e., the new information (Fig. 1).\n\nNonlinearity: When the deterministic processes governing system dynamics depend on the state of the system.\n\nPermutation entropy (PE): permutation entropy is a measure of the complexity of a time series (Bandt & Pompe, 2002) that is negatively correlated with a systems predictability (Garland et al. 2015). Permutation entropy quantifies the combined new and lost information. PE is scaled to range between a minimum of 0 and a maximum of 1.\n\nRealized predictability: the achieved predictability of a system from a given forecasting model.\n\nRedundant information: The information inherited from the past, and thus the maximum amount of information available for use in forecasting (Fig. 1).\n\nSymbols, words, permutations: symbols are simply the smallest unit in a formal language such as the letters in the English alphabet i.e., {\"A\", \"B\",..., \"Z\"}. In information theory the alphabet is more abstract, such as elements in the set {\"up\", \"down\"} or {\"1\", \"2\", \"3\"}. Words, of length m refer to concatenations of the symbols (e.g., up-down-down) in a set. Permutations are the possible orderings of symbols in a set. In this manuscript, the words are the permutations that arise from the numerical ordering of m data points in a time series.\n\nWeighted permutation entropy (WPE): a modification of permutation entropy (Fadlallah et al., 2013) that distinguishes between small-scale, noise-driven variation and large-scale, system-driven variation by considering the magnitudes of changes in addition to the rank-order patterns of PE.

ecology

The Genomic Formation of South and Central Asia

The genetic formation of Central and South Asian populations has been unclear because of an absence of ancient DNA. To address this gap, we generated genome-wide data from 362 ancient individuals, including the first from eastern Iran, Turan (Uzbekistan, Turkmenistan, and Tajikistan), Bronze Age Kazakhstan, and South Asia. Our data reveal a complex set of genetic sources that ultimately combined to form the ancestry of South Asians today. We document a southward spread of genetic ancestry from the Eurasian Steppe, correlating with the archaeologically known expansion of pastoralist sites from the Steppe to Turan in the Middle Bronze Age (2300-1500 BCE). These Steppe communities mixed genetically with peoples of the Bactria Margiana Archaeological Complex (BMAC) whom they encountered in Turan (primarily descendants of earlier agriculturalists of Iran), but there is no evidence that the main BMAC population contributed genetically to later South Asians. Instead, Steppe communities integrated farther south throughout the 2nd millennium BCE, and we show that they mixed with a more southern population that we document at multiple sites as outlier individuals exhibiting a distinctive mixture of ancestry related to Iranian agriculturalists and South Asian hunter-gathers. We call this group Indus Periphery because they were found at sites in cultural contact with the Indus Valley Civilization (IVC) and along its northern fringe, and also because they were genetically similar to post-IVC groups in the Swat Valley of Pakistan. By co-analyzing ancient DNA and genomic data from diverse present-day South Asians, we show that Indus Periphery-related people are the single most important source of ancestry in South Asia--consistent with the idea that the Indus Periphery individuals are providing us with the first direct look at the ancestry of peoples of the IVC--and we develop a model for the formation of present-day South Asians in terms of the temporally and geographically proximate sources of Indus Periphery-related, Steppe, and local South Asian hunter-gatherer-related ancestry. Our results show how ancestry from the Steppe genetically linked Europe and South Asia in the Bronze Age, and identifies the populations that almost certainly were responsible for spreading Indo-European languages across much of Eurasia.\n\nOne Sentence SummaryGenome wide ancient DNA from 357 individuals from Central and South Asia sheds new light on the spread of Indo-European languages and parallels between the genetic history of two sub-continents, Europe and South Asia.

genomics

Ancient genomes document multiple waves of migration in Southeast Asian prehistory

Southeast Asia is home to rich human genetic and linguistic diversity, but the details of past population movements in the region are not well known. Here, we report genome-wide ancient DNA data from thirteen Southeast Asian individuals spanning from the Neolithic period through the Iron Age (4100-1700 years ago). Early agriculturalists from Man Bac in Vietnam possessed a mixture of East Asian (southern Chinese farmer) and deeply diverged eastern Eurasian (hunter-gatherer) ancestry characteristic of Austroasiatic speakers, with similar ancestry as far south as Indonesia providing evidence for an expansive initial spread of Austroasiatic languages. In a striking parallel with Europe, later sites from across the region show closer connections to present-day majority groups, reflecting a second major influx of migrants by the time of the Bronze Age.

genetics

What drives interaction strengths in complex food webs? A test with feeding rates of a generalist stream predator

Describing the mechanisms that drive variation in species interaction strengths is central to understanding, predicting, and managing community dynamics. Multiple factors have been linked to trophic interaction strength variation, including species densities, species traits, and abiotic factors. Yet most empirical tests of the relative roles of multiple mechanisms that drive variation have been limited to simplified experiments that may diverge from the dynamics of natural food webs. Here, we used a field-based observational approach to quantify the roles of prey density, predator density, predator-prey body-mass ratios, prey identity, and abiotic factors in driving variation in feeding rates of reticulate sculpin (Cottus perplexus). We combined data on over 6,000 predator-prey observations with prey identification time functions to estimate 289 prey-specific feeding rates at nine stream sites in Oregon. Feeding rates on 57 prey types showed an approximately log-normal distribution, with few strong and many weak interactions. Model selection indicated that prey density, followed by prey identity, were the two most important predictors of prey-specific sculpin feeding rates. Feeding rates showed a positive, accelerating relationship with prey density that was inconsistent with predator saturation predicted by current functional response models. Feeding rates also exhibited four orders-of-magnitude in variation across prey taxonomic orders, with the lowest feeding rates observed on prey with significant anti-predator defenses. Body-mass ratios were the third most important predictor variable, showing a hump-shaped relationship with the highest feeding rates at intermediate ratios. Sculpin density was negatively correlated with feeding rates, consistent with the presence of intraspecific predator interference. Our results highlight how multiple co-occurring drivers shape trophic interactions in nature and underscore ways in which simplified experiments or reliance on scaling laws alone may lead to biased inferences about the structure and dynamics of species-rich food webs.

ecology

The mitotic spindle is chiral due to torques generated by motor proteins

Mitosis relies on forces generated in the spindle, a micro-machine composed of microtubules and associated proteins1,2. Forces are required for the congression of chromosomes to the metaphase plate and separation of chromatids in anaphase3-6. However, torques may also exist in the spindle, yet they have not been investigated. Here we show that the spindle is chiral. Chirality is evident from the finding that microtubule bundles follow a left-handed helical path, which cannot be explained by forces but rather by torques acting in the bundles. STED super-resolution microscopy, as well as confocal microscopy, of human spindles shows that the bundles have complex curved shapes. The average helicity of the bundles with respect to the spindle axis is 1.2{degrees}/m. Inactivation of kinesin-5 (Eg5/Kif11) abolished the chirality of the spindle, suggesting that this motor generates the helical shape of microtubule bundles. To explain the observed shapes, we introduce a theoretical model for the balance of forces and torques acting in the spindle, and show that torque is required to generate the helical shapes. We conclude that torques generated by motor proteins, in addition to forces, exist in the spindle and determine its architecture.

cell biology

The Genomic History Of Southeastern Europe

Farming was first introduced to southeastern Europe in the mid-7th millennium BCE - brought by migrants from Anatolia who settled in the region before spreading throughout Europe. To clarify the dynamics of the interaction between the first farmers and indigenous hunter-gatherers where they first met, we analyze genome-wide ancient DNA data from 223 individuals who lived in southeastern Europe and surrounding regions between 12,000 and 500 BCE. We document previously uncharacterized genetic structure, showing a West-East cline of ancestry in hunter-gatherers, and show that some Aegean farmers had ancestry from a different lineage than the northwestern Anatolian lineage that formed the overwhelming ancestry of other European farmers. We show that the first farmers of northern and western Europe passed through southeastern Europe with limited admixture with local hunter-gatherers, but that some groups mixed extensively, with relatively sex-balanced admixture compared to the male-biased hunter-gatherer admixture that prevailed later in the North and West. Southeastern Europe continued to be a nexus between East and West after farming arrived, with intermittent genetic contact from the Steppe up to 2,000 years before the migration that replaced much of northern Europes population.

genetics

Exact probabilities for the indeterminacy of complex networks as perceived through press perturbations

We consider the goal of predicting how complex networks respond to chronic (press) perturbations when characterizations of their network topology and interaction strengths are associated with uncertainty. Our primary result is the derivation of exact formulas for the expected number and probability of qualitatively incorrect predictions about a systems responses under uncertainties drawn form arbitrary distributions of error. These formulas obviate the current use of simulations, algorithms, and qualitative modeling techniques. Additional indices provide new tools for identifying which links in a network are most qualitatively and quantitatively sensitive to error, and for determining the volume of errors within which predictions will remain qualitatively determinate (i.e. sign insensitive). Together with recent advances in the empirical characterization of uncertainty in ecological networks, these tools bridge a way towards probabilistic predictions of network dynamics.

ecology

Quantifying predator dependence in the functional response of generalist predators.

A longstanding debate concerns whether functional responses are best described by prey-dependent versus ratio-dependent models. Theory suggests that ratio dependence can explain many food web patterns left unexplained by simple prey-dependent models. However, for logistical reasons, ratio dependence and predator dependence more generally have seen infrequent empirical evaluation and then only so in specialist predators, which are rare in nature. Here we develop an approach to simultaneously estimate the prey-specific attack rates and predator-specific interference rates of predators interacting with arbitrary numbers of prey and predator species. We apply the approach to field surveys and two field experiments involving two intertidal whelks and their full suite of potential prey. Our study provides strong evidence for the presence of weak predator dependence that is closer to being prey dependent than ratio dependent over manipulated and natural ranges of species abundances. It also indicates how, for generalist predators, even the qualitative nature of predator dependence can be prey-specific.\n\nAuthor contributionsCW contributed to method development, KC and IS performed the caging experiment, and MN conceived of the study, carried out the fieldwork and analyses, and wrote the manuscript.

ecology