Search bioRxiv⌕ Search

Biology subjects

Raskin, L. Y.

Publications and source records attributed to Raskin, L. Y..

6 recordsLinked to original sources

Estimating Divergence Times and Diversification Rates with the Unresolved Fossilized Birth-Death Process

Tip-dating using Fossilized Birth-Death (FBD) models offers a favorable alternative to conventional node-dating analyses of the evolutionary timescale, bypassing various conceptual and technical difficulties. However, despite their increasing popularity in phylogenetic studies, their use has remained rather limited in taxonomic scope because of the difficulty in collecting morphological data, which is often required for tip-dating analyses, and computational intractability when a large number of fossils are present. Recent studies, however, suggest that FBD models can also be used in a principled way as priors for molecular clock analyses, even in the absence of morphological data, relying only on temporal information and given clade assignments. Heath et al. (2014) introduced an "unresolved" FBD model in which uncertainty in fossil placements is analytically integrated out when morphological data are unavailable, providing a potential solution to this problem. Here, we show that the topological enumeration scheme used by Heath et al. (2014) to marginalize over unresolved fossil attachments overcounts configurations in which multiple fossil taxa can be resolved as a clade, and we provide a correction to this factor, yielding a valid marginalized unresolved likelihood. We apply our implementation to empirical datasets of Cetacea, Emydidae, and Palaeodictyopterida, and show that our method converges and yields reasonable estimates within a reasonable amount of time, even when hundreds of fossils are present.

evolutionary biology↗

DICAROS: Diffeomorphic Ancestral Shape Reconstruction on Phylogenies

Reconstructing ancestral morphologies on a phylogenetic tree is a central task in evolutionary morphometrics. Established reconstruction methods, including multivariate Brownian-motion approaches, rely on linear assumptions and do not directly model the correlations between landmarks within a shape, which can oversimplify the reconstructed morphology. The DICAROS method (Diffeomorphic Independent Contrasts for Ancestral Reconstruction of Shapes; Severinsen et al., 2026) instead fuses sibling shapes along branches with large-deformation diffeomorphic (LDDMM) landmark dynamics that model these correlations, so that ancestors remain on the shape manifold. DICAROS was shown to outperform ordinary least-squares, Brownian-motion, and penalized-likelihood reconstruction, particularly on non-symmetric trees. The dicaros package repackages that pipeline as a documented, pip-installable tool that runs on arbitrary landmark datasets from a single command. It handles 2D and 3D landmarks, Newick and NEXUS trees, a choice of Euclidean or Frechet species means, optional anchor-based alignment, and tips backed by a single specimen, and it returns the reconstructed shapes for all nodes together with the tree relabelled at its internal nodes. We demonstrate dicaros on two new datasets: a 2D leaf dataset (217 species) and a 3D guenon skull dataset (22 species).

evolutionary biology↗

A hierarchical Bayesian framework accommodates intraspecific and interspecific variation in multivariate traits

Phylogenetic comparative methods are a critical tool in biology, providing the framework to test evolutionary hypotheses of phenotypic diversification. Accommodating intraspecific variation in multivariate analyses is critical for accurate evolutionary inference, but current methods that incorporate intraspecific variation either 1) assume that traits evolve independently or 2) that all taxa share the same intraspecific covariance structure. Violations of these assumptions can produce biased estimates of evolutionary parameters. Here, we introduce a hierarchical Bayesian framework for multivariate traits that jointly estimates taxon-specific intraspecific covariance structures alongside the underlying evolutionary process. This framework propagates uncertainty from sample size discrepancies and missing data, enabling the incorporation of highly variable morphological traits into phylogenetic analyses. Analysis of simulated data confirms that the model and implementation are well calibrated under the assumed generative model, including challenging datasets with more traits than individuals and substantial missing observations. Applied to perikymata spacing across the great ape clade, including modern humans and Neandertals, the framework recovers intraspecific covariance structures that differ among taxa and yields evolutionary rate estimates markedly more uniform across the tooth crown than those obtained when taxon means are fixed. Our method, which is applicable to other multivariate traits, provide a flexible, tractable approach to joint estimation of intraspecific variation and evolutionary process in multivariate traits.

evolutionary biology↗

STEM-LM: Spatio-Temporal Ecological Modeling via Masked Language Model for Joint Species Distribution

Joint species distribution models (JSDMs) are central to biodiversity forecasting and conservation decision-making. As ecological datasets grow in size, dimensionality, and spatio-temporal resolution, there is a need for flexible yet scalable JSDMs tailored to large-scale species observation data. Recent advances in masked language modeling for text and genomics suggest a natural alternative: by treating each species presence or absence as a token, and a sites species assemblage together with its spatio-temporal and ecological covariates as a sentence, we can learn joint co-occurrence structure by reconstructing masked species from their neighboring sites. We propose STEM-LM1, a Transformer-based JSDM that frames joint species distribution modeling as masked language modeling. By varying the masking rate during training, a single trained model supports both purely spatiotemporal/ecological prediction and conditioning on arbitrary subsets of observed species for joint co-occurrence inference at a given site. On a North American butterfly and a global plant distribution dataset, STEM-LM performs better or on par with other statistical and deep-learning based methods in terms of discriminative ranking, while producing substantially better rank-calibrated occurrence probabilities. Utilizing partial species observations at the same site greatly enhances prediction performance.

ecology↗

Principal Components Analysis fails to recover phylogenetic structure in hominins

ObjectivesPaleoanthropologists often utilize geometric morphometrics and principal components analysis (PCA) to interpret shape variation within the hominin fossil record. It is common practice to interpret proximity in principal components (PC) space among taxa as indicative of not just morphological, but also phylogenetic affinity. This interpretation, however, has not been directly evaluated for hominins. Materials and MethodsFirst, we inferred the posterior distribution of hominin phylogenetic trees and subsampled trees from this distribution. On these phylogenies, we simulated 2D and 3D geometric morphometric datasets and traditional morphological datasets, containing traits analogous to measurements of size or length, with varying numbers of landmarks or traits and evolutionary rates. On each dataset, we conducted a PCA and used neighbor-joining to infer evolutionary relationships from the PC scores of each taxon. We measure the difference between the PCA tree and sampled tree with subtree pruning and regrafting distance and Robinson-Foulds distance. ResultsPCA trees inferred from traditional morphometric data were identical to the sampled tree in 0.11% of datasets when we only considered PC axes 1 and 2, and in 2.9% of datasets when we considered all axes. No PCA tree inferred from any of the 2,400,000 shape datasets was identical to the sampled tree, regardless of the number of axes. DiscussionPhylogenetic interpretations of the hominin fossil record based on proximity in PC space are inherently flawed and likely to be erroneous. Arguments in the hominin systematics literature based on PCA should therefore be reevaluated using phylogenetically-informed alternatives.

evolutionary biology↗

The effects of trait redundancy and information content on hominin phylogenetic inference

Paleoanthropological phylogenetic inference is based on characters assumed to be phylogenetically informative and independent. Yet, our understanding of whether these criteria are met in published character supermatrices is limited. We assess the phylogenetic information content (PHIC) of 107 discrete craniodental traits from a widely-used hominin character matrix. We compare test topologies -- inferred by permuting single traits or removing single traits, anatomical units (AUs), and operational taxonomic units -- to the baseline topology, inferred from the unmodified matrix. In this dataset, only 31 traits have some degree of PHIC: 23 uniquely informative traits -- sufficient, as a set, to closely approach the baseline topology -- and eight redundant traits. No single AU, nor the combination of mandibular and dentition AUs, contains sufficient PHIC to approach the baseline topology, and only the maxilla contains more PHIC than expected. Therefore, phylogenetic placements of fossil hominins represented by isolated AUs should be regarded as putative until better-preserved specimens and more informative traits can be incorporated. Given the ubiquity of discrete morphological data in paleontology and that most of the history of life on Earth was only recorded through fossils, our methods should be broadly applicable to phylogenetic inference involving other paleontological clades.

evolutionary biology↗