Search bioRxiv⌕ Search

Biology subjects

Shumaker, K. A.

Publications and source records attributed to Shumaker, K. A..

4 recordsLinked to original sources

FIDDL: depth-matched negative controls distinguish genuine interspecific introgression from competitive-mapping artifact

Interspecific introgression is routinely detected by competitively mapping reads to a concatenated multi-species reference and calling regions where a non-focal species recruits coverage. Using strains that cannot contain the ancestry being detected, we show this design generates substantial false-positive signal through two mechanisms with opposite phylogenetic-distance signatures. Standard nuclear assemblies omit the mitochondrion and 2-micron plasmid, leaving high-copy cytoplasmic reads without a legitimate target; completing the reference preferentially removes signal from the most divergent donor. Genuine cross-species sequence conservation inflates the most closely related donor. Masking chromosome ends removes its subtelomeric part but plateaus at a non-zero floor, and the interior residual traces to conserved single-copy genes where a short read carries under one base of discriminating information. The floor grows with sequencing depth (1.19% of callable positions at 50x, 2.02% at 147x, 3.85% at 393x in a pure strain), is not mitigated by long reads, and appears at sub-diploid dosage - three properties widely read as evidence of authenticity. Because the discriminating information is below single-read resolution, no read-level filter separates artifact from introgression; we show three that fail. What works is locus-level: a consensus-phylogenetic test (29/29 specificity on confirmed artifact) and an allele-fraction donor-match test, complementary and validated in both directions on independent published introgression. We package the comparative controls as FIDDL (False Introgression Detection via Depth-matched controls and Loci-recurrence), an open-source tool, withdraw two of our own analysis-ready calls, and show re-analysis of published wild isolates reduces low-confidence introgression by [~]53% while leaving high-confidence signal intact.

bioinformatics↗

Genome-scale characterization of wild yeasts reveals cryptic diversity and population structure across three genera

Environmental surveys of wild yeasts typically rely on ribosomal barcodes, which cannot resolve cryptic species, interspecific gene flow, mixed cultures, or population structure. To determine what genome-scale characterization adds, we sequenced a representative panel of wild yeasts spanning the genera Saccharomyces, Schizosaccharomyces, and Lachancea using Oxford Nanopore long-read whole-genome sequencing and placed each isolate within published reference datasets. Whole-genome analyses revealed biologically important features that barcoding alone could not detect. A shagbark-hickory isolate resolved as a genuine two-species co-culture. An oak-bark isolate proved to be Schizosaccharomyces versatilis, a recently reinstated species represented by very few known strains, and its analysis demonstrated that standard assembly-quality benchmarks can be misleading for deep-branching taxa. Three Lachancea thermotolerans isolates formed a distinct, previously unsampled population within the wild tree-associated lineage, extending its known geographic range. In contrast, an apparent signal of Saccharomyces eubayanus introgression in two beer-associated S. cerevisiae isolates disappeared after analysis with matched negative controls and de novo assemblies, showing that it reflected mapping artifacts rather than genuine ancestry. Together, these results demonstrate that inexpensive long-read whole-genome sequencing transforms wild-yeast bioprospecting from species identification into a genome-scale framework for resolving cryptic diversity, population structure, and mixed cultures while providing stronger support - and stronger limits - for evolutionary inference. SIGNIFICANCEMost surveys of wild yeasts identify isolates using short DNA barcodes, which are well suited for naming species but often miss the evolutionary relationships and hidden diversity within them. By applying inexpensive whole-genome sequencing to a diverse collection of environmental yeasts, we uncovered previously undetected mixed cultures, a rare recently recognized species, and a distinct wild population, while also showing that an apparent case of interspecies gene exchange was instead a technical artifact. These results demonstrate that genome-scale analysis can both reveal biological diversity that simpler methods overlook and provide the evidence needed to avoid misleading evolutionary conclusions, making it a powerful new approach for studying natural microbial populations.

ecology↗

Declaration of Fermentation: Community-Embedded Wild Yeast Bioprospecting as a Model for Place-Based CURE Design

Course-based undergraduate research experiences (CUREs) are widely recognized as a high-impact practice in biology education, yet most existing CURE frameworks treat the research organism as an interchangeable teaching prop rather than a genuine scientific contribution. We argue that place-based, community-embedded CUREs - in which students isolate, characterize, and publicly deploy a locally meaningful wild organism - constitute a qualitatively distinct model warranting broader adoption. As proof of concept, we present the Declaration of Fermentation project at Indiana University Bloomington: graduate researchers isolated a wild Saccharomyces cerevisiae strain from the bark of a campus landmark tree, confirmed its wild provenance by whole-genome sequencing and phylogenomics, and partnered with local craft breweries to produce a colonial-era inspired ale released publicly for the 250th anniversary of the Declaration of Independence. Volunteer sensory panels at two independent public tasting events (combined n = 33-34 per attribute) confirmed a fruity-funky profile consistent with wild-strain fermentation, with no significant differences between events (Mann-Whitney U, Benjamini- Hochberg-corrected p > 0.05 for all 11 attributes). We describe three design principles - genomically confirmed strain identity, mandatory community partnership, and place-based historical narrative - that distinguish this model from prior wild yeast brewing CUREs, discuss how these principles generalize to other institutions and fermentation vehicles, and identify next steps for formal learning assessment. Complete implementation protocols are provided as supplemental Appendices 1-6, and the bioinformatics pipeline is freely available at https://doi.org/10.5281/zenodo.20679384.

microbiology↗

A screen for synthetic genetic interactions with the Saccharomyces cerevisiae hrq1ΔN allele

The Saccharomyces cerevisiae Hrq1 helicase is a functional homolog of the disease-linked human RECQL4 enzyme and has been used as a model to study RecQ4 helicase subfamily biology. Although the motor cores of Hrq1 and RECQL4 are quite similar, these proteins display distinct N-terminal domains (NTDs) of unknown function. Do these domains facilitate species-specific activities by the two helicases, or do they serve common roles despite their differences in sequence and predicted structure? We probed these questions here by analyzing an NTD-truncated isoform of Hrq1 (Hrq1{Delta}N) both in vitro and in vivo. We found that the Hrq1 NTD houses a cryptic DNA binding site and that the hrq1{Delta}N allele is distinct from both hrq1{Delta} and the catalytically inactive hrq1-K318A mutant. Using synthetic genetic array analysis of hrq1{Delta}N crossed to the yeast S. cerevisiae single-gene deletion and temperature-sensitive allele collections, we also identified hundreds of synthetic genetic interactions. As with similar analyses of hrq1{Delta} and hrq1-K318A, our results suggest roles for Hrq1 and its NTD in multiple physiological pathways that underpin genome integrity. Together, these data are guiding our ongoing efforts to understand the roles of Hrq1 and RECQL4 in genome maintenance, which will help to explain why RECQL4 mutations cause disease. ARTICLE SUMMARYThis work should interest researchers in the genome integrity and yeast disease model fields. It continues ongoing efforts to develop the Saccharomyces cerevisiae Hrq1 helicase as a model to understand the disease-linked human RECQL4 helicase. Hrq1 and RECQL4 share similar helicase cores but have divergent N-terminal domains (NTDs) of unknown function. The results demonstrate that the Hrq1 NTD contains a DNA binding site, and an HRQ1 allele lacking its NTD (hrq1{Delta}N) interacts with hundreds of mutants in defined allele collections. Thus, the hrq1{Delta}N allele is distinct from the previously characterized hrq1{Delta} and hrq1-K318A (catalytically inactive) mutants.

molecular biology↗