Search bioRxiv⌕ Search

Biology subjects

Sackett, P. W.

Publications and source records attributed to Sackett, P. W..

8 recordsLinked to original sources

CHARON: Estimating the drift time and the number of individuals in environmental DNA with diploid individuals

Environmental DNA (eDNA) offers a promising avenue for reconstructing the genetic diversity and demographic history of ancient populations. However, the analysis of eDNA from humans or forensic DNA, presents significant challenges, including low coverage, DNA degradation, and the uncertainty of the number of individuals contributing to a sample. This study introduces CHARON, a novel statistical method for jointly estimating the number of individuals and drift times between a human eDNA or forensic sample and a given population with known allele frequencies. We validate our method through simulations and synthetic empirical data and show that we can reliably estimate the number of individuals up to 8 at a coverage between 2X and 4X. Our method can also pinpoint the most likely population of origin for eDNA or forensic samples. This work provides a tool for the application of human eDNA in evolutionary and forensic studies and an implementation is available here: https://github.com/Jan-van-Waaij/Charon

evolutionary biology↗

AdDeam: A Fast and Scalable Tool for Estimating and Clustering Reference-Level Damage Profiles

MotivationDNA damage patterns, such as increased frequencies of C[->]T and G[->]A substitutions at fragment ends, are widely used in ancient DNA studies to assess authenticity and detect contamination. In metagenomic studies, fragments can be mapped against multiple references or de novo assembled contigs to identify those likely to be ancient. Generating and comparing damage profiles, however, can be both tedious and time-consuming. Although tools exist for estimating damage in single reference genomes and metagenomic datasets, none efficiently cluster damage patterns. ResultsTo address this methodological gap, we developed AdDeam, a tool that combines rapid damage estimation with clustering for streamlined analyses and easy identification of potential contaminants or outliers. Our tool takes aligned aDNA fragments from various samples or contigs as input, computes damage patterns, clusters them and outputs representative damage profiles per cluster, a probability of each sample of pertaining to a cluster as well as a PCA of the damage patterns for each sample for fast visualisation. We evaluated AdDeam on both simulated and empirical datasets. AdDeam effectively distinguishes different damage levels, such as UDG-treated samples, sample-specific damages from specimens of different time periods, and can also distinguish between contigs containing modern or ancient fragments, providing a clear framework for aDNA authentication and facilitating large-scale analyses. Availability and ImplementationAdDeam is publicly available at https://github.com/LouisPwr/AdDeam and can also be installed via Bioconda. It is implemented in Python and C++. All analysis scripts and datasets are available at https://github.com/LouisPwr/AdDeamAnalysis and on Zenodo under: 10.5281/zenodo.15052427.

bioinformatics↗

SAFARI: Pangenome Alignment of Ancient DNA Using Purine/Pyrimidine Encodings

Aligning DNA sequences retrieved from fossils or other paleontological artifacts, referred to as ancient DNA, is particularly challenging due to the short sequence length and chemical damage which creates a specific pattern of substitution (C[->]T and G[->]A) in addition to the heightened divergence between the sample and the reference genome thus exacerbating reference bias. This bias can be mitigated by aligning to pangenome graphs to incorporate documented organismic variation, but this approach still suffers from substitution patterns due to chemical damage. We introduce a novel methodology introducing the RYmer index, a variant of the commonly-used minimizer index which represents purines (A,G) and pyrimidines (C,T) as R and Y respectively. This creates an indexing scheme robust to the aforementioned chemical damage. We implemented SAFARI, an ancient DNA damage-aware version of the pangenome aligner vg giraffe which uses RYmers to rescue alignments containing deaminated seeds. We show that our approach produces more correct alignments from ancient DNA sequences than current approaches while maintaining a tolerable rate of spurious alignments. In addition, we demonstrate that our algorithm improves the estimate of the rate of ancient DNA damage, especially for highly damaged samples. Crucially, we show that this improved alignment can directly translate into better insights gained from the data by showcasing its integration with a number of extant pangenome tools.

bioinformatics↗

soibean: High-resolution Taxonomic Identification of Ancient Environmental DNA Using Mitochondrial Pangenome Graphs

Ancient environmental DNA (aeDNA) is becoming a powerful tool to gain insights about past ecosystems. However, several methodological challenges remain, particularly for classifying the DNA to species level and conducting phylogenetic placement. Current methods, primarily tailored for modern datasets, fail to capture several idiosyncrasies of aeDNA, including species mixtures from closely related species and ancestral divergence. We introduce soibean, a novel tool that utilises pangenomic graphs for identifying species from ancient environmental mitochondrial reads. It outperforms existing methods in accurately identifying species from multiple sources within a sample, enhancing phylogenetic analysis for aeDNA. soibean employs a damage-aware likelihood model for precise identification at low-coverage with high damage rate, demonstrating effectiveness through simulated data tests and empirical validation. Notably, our method uncovered new empirical results in published datasets, including using porpoise whales as food in a Mesolithic community in Sweden, demonstrating its potential to reveal previously unrecognised findings in aeDNA studies.

bioinformatics↗

euka: Robust detection of eukaryotic taxa from modern and ancient environmental DNA using pangenomic reference graphs.

1. Ancient environmental DNA (eDNA) is a crucial source of in-formation for past environmental reconstruction. However, the com-putational analysis of ancient eDNA involves not only the inherited challenges of ancient DNA (aDNA) but also the typical difficulties of eDNA samples, such as taxonomic identification and abundance esti-mation of identified taxonomic groups. Current methods for ancient eDNA fall into those that only perform mapping followed by taxo-nomic identification and those that purport to do abundance estima-tion. The former leaves abundance estimates to users, while methods for the latter are not designed for large metagenomic datasets and are often imprecise and challenging to use. 2. Here, we introduce euka, a tool designed for rapid and accurate characterisation of ancient eDNA samples. We use a taxonomy-based pangenome graph of reference genomes for robustly assigning DNA sequences and use a maximum-likelihood framework for abundance estimation. At the present time, our database is restricted to mito-chondrial genomes of tetrapods and arthropods but can be expanded in future versions. 3. We find euka to outperform current taxonomic profiling tools as well as their abundance estimates. Crucially, we show that regardless of the filtering threshold set by existing methods, euka demonstrates higher accuracy. Furthermore, our approach is robust to sparse data, which is idiosyncratic of ancient eDNA, detecting a taxon with an average of fifty reads aligning. We also show that euka is consistent with competing tools on empirical samples and about ten times faster than current quantification tools. 4. eukas features are fine-tuned to deal with the challenges of ancient eDNA, making it a simple-to-use, all-in-one tool. It is available on GitHub: https://github.com/grenaud/vgan. euka enables re-searchers to quickly assess and characterise their sample, thus allowing it to be used as a routine screening tool for ancient eDNA.

bioinformatics↗

MAVISp: Multi-layered Assessment of VarIants by Structure for proteins

The role of genomic variants in disease has expanded significantly with the advent of advanced sequencing techniques. The rapid increase in identified genomic variants has led to many variants being classified as Variants of Uncertain Significance or as having conflicting evidence, posing challenges for their interpretation and characterization. Additionally, current methods for predicting pathogenic variants often lack insights into the underlying molecular mechanisms. Here, we introduce MAVISp (Multi-layered Assessment of VarIants by Structure for proteins), a modular structural framework for variant effects, accompanied by a web server (https://services.healthtech.dtu.dk/services/MAVISp-1.0/) to enhance data accessibility, consultation, and reusability. MAVISp currently provides data over 1000 proteins, encompassing more than eight million variants. A team of biocurators regularly analyzes and updates protein entries using standardized workflows, incorporating free energy calculations or biomolecular simulations. We illustrate the utility of MAVISp through selected case studies. The framework facilitates the analysis of variant effects at the protein level and has the potential to advance the understanding and application of mutational data in disease research.

bioinformatics↗

HaploCart: Human mtDNA Haplogroup Classification Using a Pangenomic Reference Graph

Current mitochondrial DNA (mtDNA) haplogroup classification tools map reads to a single reference genome and perform inference based on the detected mutations to this reference. This approach biases haplogroup assignments towards the reference and prohibits accurate calculations of the uncertainty in assignment. We present HaploCart, an mtDNA haplogroup classifier which uses VGs pangenomic reference graph framework together with principles of Bayesian inference. We demonstrate that our approach significantly outperforms available tools by being more robust to lower coverage or incomplete consensus sequences and producing phylogenetically-aware confidence scores that are unbiased towards any haplogroup. HaploCart is available both as a command-line tool and through a user-friendly web interface. The program written in C++ accepts as input consensus FASTA, FASTQ, or GAM files, and outputs a text file with the haplogroup assignments along with confidence estimates. Our work considerably reduces the amount of data required to obtain a confident mitochondrial haplogroup assignment. HaploCart is available as a command-line tool at https://github.com/grenaud/vgan and as a web server at https://services.healthtech.dtu.dk/service. php?HaploCart.

bioinformatics↗

RosettaDDGPrediction for high-throughput mutational scans: from stability to binding

Reliable prediction of free energy changes upon amino acidic substitutions ({Delta}{Delta}Gs) is crucial to investigate their impact on protein stability and protein-protein interaction. Moreover, advances in experimental mutational scans allow high-throughput studies thanks to sophisticated multiplex techniques. On the other hand, genomics initiatives provide a large amount of data on disease-related variants that can benefit from analyses with structure-based methods. Therefore, the computational field should keep the same pace and provide new tools for fast and accurate high-throughput calculations of {Delta}{Delta}Gs. In this context, the Rosetta modeling suite implements effective approaches to predict the change in the folding free energy in a protein monomer upon amino acid substitutions and calculate the changes in binding free energy in protein complexes. Their application can be challenging to users without extensive experience with Rosetta. Furthermore, Rosetta protocols for {Delta}{Delta}G prediction are designed considering one variant at a time, making the setup of high-throughput screenings cumbersome. For these reasons, we devised RosettaDDGPrediction, a customizable Python wrapper designed to run free energy calculations on a set of amino acid substitutions using Rosetta protocols with little intervention from the user. RosettaDDGPrediction assists with checking whether the runs are completed successfully aggregates raw data for multiple variants, and generates publication-ready graphics. We showed the potential of the tool in selected case studies, including variants of unknown significance found in children who developed cancer, proteins with known experimental unfolding {Delta}{Delta}Gs values, interactions between target proteins and a disordered functional motif, and phospho-mimetic variants. RosettaDDGPrediction is available, free of charge and under GNU General Public License v3.0, at https://github.com/ELELAB/RosettaDDGPrediction.

bioinformatics↗