Search bioRxiv⌕ Search

Biology subjects

Vellnow, N.

Publications and source records attributed to Vellnow, N..

2 recordsLinked to original sources

SwarmGenomics: A Unified Pipeline for Individual-Based Whole-Genome Analyses

Advances in sequencing technologies have made whole-genome data widely accessible, enabling research in population genetics, evolutionary biology, and conservation. However, analyzing whole-genome sequencing (WGS) data remains challenging, often requiring multiple specialized tools and substantial bioinformatics expertise. We present SwarmGenomics, a modular, user-friendly command-line pipeline for reference-based genome assembly and individual-based genetic analyses. The pipeline integrates seven modules: heterozygosity estimation, runs of homozygosity detection, Pairwise Sequentially Markovian Coalescent (PSMC) analysis, unmapped reads classification, repeat analysis, mitochondrial genome assembly, and nuclear mitochondrial DNA segment (NUMT) identification. Each module can be run independently or as part of a complete workflow. We demonstrate the pipelines utility with a case study on the giant panda (Ailuropoda melanoleuca), revealing insights into genetic diversity, inbreeding history, historical population size changes, transposable element activity, and microbial contamination. SwarmGenomics lowers the entry barrier for genomic analysis of diploid, non-model species, serving both as a research and teaching tool. The pipeline and documentation are available at https://github.com/AureKylmanen/Swarmgenomics.

genomics↗

A comprehensive representation of selection at loci with multiple alleles that allows complex forms of genotypic fitness

Genetic diversity is central to evolutionary change, with both natural selection and random genetic drift depending on variation within a population. An individual in a diploid population carries two alleles per locus, yet the population as a whole can harbour many alleles, giving rise to a rich spectrum of homozygous and heterozygous genotypes. Such multiallelic variation is common at biologically and medically important loci such as the major histocompatibility complex, the ABO blood group system, and genes underlying monogenic diseases. However, much of population genetic theory and data analysis has focussed on biallelic loci. Here, we introduce a matrix representation of the genotypic selection acting at a multiallelic locus. This exploits the common mathematical structure underlying selection and drift, and separates the effects of genetic diversity and fitness. The representation accommodates diverse selection regimes, including additive, multiplicative, frequency-dependent, and temporally varying selection, as well as heterozygote advantage. We show how, under specific assumptions, genotype-specific fitness-effects can be estimated from allele frequency trajectories over microevolutionary timescales. Applying this estimation procedure to time-series data from experimental yeast evolution illustrates how multiallelic fitness interactions, including heterozygote advantage, may be characterised from haplotype frequency data. More broadly, this work provides a practical foundation for analysing evolutionary dynamics at multiallelic loci in experimental and natural populations.

evolutionary biology↗