Search bioRxiv⌕ Search

Biology subjects

Daigle, A. T.

Publications and source records attributed to Daigle, A. T..

6 recordsLinked to original sources

The B-value calculator: expected diversity under background selection

Background selection (BGS), the indirect effect of purifying selection at linked and unlinked sites, is a key evolutionary process shaping genomic patterns of variation. Calculating the expected diversity at neutral sites experiencing BGS relative to that under strict neutrality (referred to as B-value or simply B) is important for developing null models when performing population genomic inference, in particular, demographic inference and detection of selective sweeps. We extend and integrate previous theory to estimate B-values analytically, assuming no selective interference, with novel expressions to account for gene conversion between proximal sites. Here, we present the B-value calculator, Bvalcalc, an easy-to-use command-line interface written in Python for efficient analytical calculation of expected genome-wide B at single base-pair resolution. Bvalcalc has several modules for calculating diversity as a function of distance from a single selected element or considering the multiplicative effects of all selected elements across the genome, accounting for recombination maps, gene conversion, self-fertilization, single population size changes, and unlinked effects from other chromosomes. We validated the effectiveness of Bvalcalc with comparisons against simulated results, and generated B-maps for the model species Homo sapiens, Drosophila melanogaster and Arabidopsis thaliana as a proof of concept using public data. Bvalcalc is available with documentation at johrilab.github.io/Bvalcalc/.

evolutionary biology↗

Towards an evolutionary baseline model of Plasmodium falciparum for population-genomic inference

Malaria has caused over 15.7 million deaths in the 21st century and was responsible for [~]600 thousand deaths globally in 2023 alone. Although many effective antimalarial drugs have been developed and widely adopted to reduce the occurrence and severity of the disease, recurrent resistance to the frontline treatment has been of major concern. Multiple drug resistance alleles at intermediate and high allele frequency have been identified in specific Asian and African populations of P. falciparum, the deadliest malaria parasite. With the improvement in throughput of sequencing technologies and global efforts such as the MalariaGEN project to build genomic surveillance, we now have access to tens of thousands of genomes of P. falciparum from across the world. With this data, it is becoming increasingly possible to employ powerful population genetics approaches to understand the selective pressures and demographic history of the parasite. While several empirically motivated outlier-based approaches have been employed to identify targets of drug resistance, there is a lack of a framework that jointly accounts for the multiple concurrent processes occurring in natural populations of P. falciparum. We argue that a baseline evolutionary model that accounts for simultaneously acting evolutionary processes is needed to understand patterns of genomic variation in P. falciparum populations. Here, we identify key components essential for building such a baseline model for the malaria-causing pathogen. The development of an appropriate null model will be important to test evolutionary hypotheses using genomic datasets, will provide a path forward to improve the accuracy of inference of evolutionary parameters, and will help identify new gene candidates involved in drug resistance.

evolutionary biology↗

The contribution of transposable element insertions to genetic diversity in Aedes aegypti populations

Aedes aegypti is a vector of multiple tropical diseases. The main strategy to control transmission is insecticide-based population control. However, mosquito populations rapidly evolve resistance, possibly enabled by their high levels of genetic diversity. Genome-wide surveys of diversity in Ae. aegypti have focused on single nucleotide polymorphisms (SNPs), and although structural variants such as transposable element (TE) insertions have been implicated in insecticide resistance (IR) in Drosophila, these have not been thoroughly characterized in Aedes. Here, we evaluated the TE content in 122 Ae. aegypti genomes from six countries across Africa, North, and South America. We found that TEs contribute substantially to genetic diversity and reflect population structure broadly consistent with that seen in SNPs. Although most TEs insertions are rare, some were observed at higher frequencies, suggesting that a small subset of these may be beneficial. For example, we identified numerous TEs with large frequency differences across populations, consistent with the possibility that these are in haplotypes underlying local adaptation. Specifically, we found three TEs near genes that may be involved in metabolic insecticide resistance: CYP6P12, GSTD11 and GSTZ1. In Colombian samples, we also identified a TE insertion that is in negative linkage disequilibrium with several insecticide resistance mutations that form an intermediate-frequency haplotype in the VGSC gene region. These results suggest the possibility that, just as TEs have been implicated in adaptation in other animals such as Drosophila, they may play an important role in the evolution of resistance to control efforts in Aedes and other pests.

evolutionary biology↗

Meiotic double strand DNA breaks and spontaneous mutation in Drosophila melanogaster

The exchange of genetic material during meiosis requires the formation and repair of DNA double-strand breaks (DSBs), which may not be repaired with perfect fidelity. If meiotic exchange is mutagenic, this would add to the costs of sexual reproduction and affect patterns of genome evolution, but much of the evidence for this is indirect. In the fruit fly Drosophila melanogaster, it is possible to completely suppress endogenous DSBs while retaining normal fertility. We took advantage of this system to generate fly strains with and without a mutant allele of mei-P22, a gene that is essential for meiotic DSB formation, on a common genetic background. This allowed us to investigate the relationship between DSBs and genome-wide mutation patterns, using a mutation accumulation design to allow un-selected spontaneous mutations to be observed. Following 30 generations of mutation accumulation, we identified over 1800 mutations by whole-genome sequencing. The presence of meiotic DSBs had little effect on the rate and spectrum of point mutations. We found that mutations were more likely to occur in areas of the genome with higher rates of crossover recombination, regardless of whether meiotic DSBs were occurring. We also found that the rate of transposable element insertions across multiple TE families was substantially elevated in the group lacking meiotic DSBs, suggesting that host suppression of mobile genetic elements is closely associated with meiotic recombination mechanisms.

evolutionary biology↗

Leveraging long-read assemblies and machine learning to enhance short-read transposable element detection and genotyping

Transposable elements (TEs) are parasitic genomic elements that are ubiquitous across the tree of life and play a crucial role in genome evolution. Advances in long-read sequencing have allowed highly accurate TE detection, though at a higher cost than short-read sequencing. Recent studies using long reads have shown that existing short-read TE detection methods perform inadequately when applied to real data. In this study, we use a machine learning approach (called TEforest) to discover and genotype TE insertions and deletions with short-read data by using TEs detected from long-read genome assemblies as training data. Our method first uses a highly sensitive algorithm to discover potential TE insertion or deletion sites in the genome, extracting relevant features from short-read alignments. To discriminate between true and false TE insertions, we train a random forest model with a labeled ground-truth dataset for which we have calculated the same set of short-read features. We conduct a comprehensive benchmark of TEforest and traditional TE detection methods using real data, finding that TEforest identifies more true positives and fewer false positives across datasets with different read lengths and coverages, while also accurately inferring genotypes and the precise breakpoints of insertions. By learning short-read signatures of TEs previously only discoverable using long reads, our approach bridges the gap between large-scale population genetic studies and the accuracy of long-read assemblies. This work provides a user-friendly tool to study the prevalence and phenotypic effects of TE insertions across the genome.

genomics↗

Bergerac Strains of C. elegans Revisited: Expansion of Tc1 elements Impose a Significant Genomic and Fitness Cost

The DNA transposon Tc1 was the first transposable element (TE) to be characterized in Caenorhabditis elegans and to date, remains the best-studied TE in Caenorhabditis worms. While Tc1 copy-number is regulated at approximately 30 copies in the laboratory N2/Bristol and the vast majority of C. elegans strains, the Bergerac strain and its derivatives have experienced a marked Tc1 proliferation. Given the historical importance of the Bergerac strain in the development of the C. elegans model, we implemented a modern genomic analysis of three Bergerac strains (CB4851, RW6999, and RW7000) in conjunction with multiple phenotypic assays to better elucidate the (i) genomic distribution of Tc1, and (ii) phenotypic consequences of TE deregulation for the host organism. The median estimates of Tc1 copy-number in the Bergerac strains ranged from 451 to 748, which is both (i) greater than previously estimated, and (ii) likely to be an underestimate of the actual copy-numbers since coverage-based estimates and ddPCR results both suggest higher Tc1 numbers. All three Bergerac strains had significantly reduced trait means compared to the N2 control for each of four fitness-related traits, with specific traits displaying significant differences between Bergerac strains. Tc1 proliferation was genome-wide, specific to Tc1, and particularly high on chromosomes V and X. There were fewer Tc1 insertions in highly expressed chromatin environments than expected by chance. Furthermore, Tc1 integration motifs were also less frequent in exon than non-coding sequences. The source of the proliferation of Tc1 in the Bergerac strains is specific to Tc1 and independent of other TEs. The Bergerac strains contain none of the alleles that have previously been found to derepress TE activity in C. elegans. However, the Bergerac strains had several Tc1 insertions near or within highly germline-transcribed genes which could account for the recent germline proliferation.

evolutionary biology↗