Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.08.05.606565

Inferring the composition of a mixed culture of natural microbial isolates by deep sequencing

Abstract

Next generation sequencing has unlocked a wealth of genotype information for microbial populations, but phenotyping remains a bottleneck for exploiting this information, particularly for pathogens that are difficult to manipulate. Here, we establish a method for high-throughput phenotyping of mixed cultures, in which the pattern of naturally occurring single-nucleotide polymorphisms in each isolate is used as intrinsic barcodes which can be read out by sequencing. We demonstrate that our method can correctly deconvolute strain proportions in simulated mixed-strain pools. As an experimental test of our method, we perform whole genome sequencing of 66 natural isolates of the thermally dimorphic pathogenic fungus Coccidioides posadasii and infer the strain compositions for large mixed pools of these strains after competition at 37{degrees}C and room temperature. We validate the results of these selection experiments by recapitulating the temperature-specific enrichment results in smaller pools. Additionally, we demonstrate that strain fitness estimated by our method can be used as a quantitative trait for genome-wide association studies. We anticipate that our method will be broadly applicable to natural populations of microbes and allow high-throughput phenotyping to match the rate of genomic data acquisition. Author summaryThe diversity of the gene pool in natural populations encodes a wealth of information about its molecular biology. This is an especially valuable resource for non-model organisms, from humans to many microbial pathogens, lacking traditional genetic approaches. An effective method for reading out this population genetic information is a genome wide association study (GWAS) which searches for genotypes correlated with a phenotype of interest. With the advent of cheap genotyping, high throughput phenotyping is the primary bottleneck for GWAS, particularly for microbes that are difficult to manipulate. Here, we take advantage of the fact that the naturally occurring genetic variation within each individual strain can be used as an intrinsic barcode, which can be used to read out relative abundance of each strain as a quantitative phenotype from a mixed culture. Coccidioides posadasii, the causative agent of Valley Fever, is a fungal pathogen that must be manipulated under biosafety level 3 conditions, precluding many high-throughput phenotyping approaches. We apply our method to pooled competitions of C. posadasii strains at environmental and host temperatures. We identify robustly growing and temperature-sensitive strains, confirm these inferences in validation pooled growth experiments, and successfully demonstrate their use in GWAS.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Voorhies, M. S., Joehnk, B., Uehling, J., Walcott, K., Dubin, C., Mead, H., Homer, C., Galgiani, J., Barker, B., Brem, R., Sil, A.. 2024-08-05. Inferring the composition of a mixed culture of natural microbial isolates by deep sequencing. https://doi.org/10.1101/2024.08.05.606565

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Generation of a transgenic cephalopod

Coleoid cephalopods (cuttlefish, octopus, and squid) are marine mollusks with elaborate nervous systems that support a diverse repertoire of complex behaviors. These include the neural control of the color, pattern, and texture of the skin, facilitating both adaptive camouflage and innate patterning that may reflect internal state. The development of transgenic cephalopods expressing fluorescent proteins, optogenetic actuators, and reporters of neural activity would contribute a new and important technology to cephalopod biology. The generation of transgenic cephalopods, however, has remained a major challenge. Here, we report the development of stable transgenic dwarf cuttlefish (Ascarosepion bandense) expressing ubiquitous nuclear-localized mScarlet, a red fluorescent protein. We evaluated multiple strategies for transgenesis, and established cuttlefish lines using both CRISPR and the transposons Sleeping Beauty and Minos. The stable expression of transgenes enabled live imaging of cell dynamics during embryonic development. The Minos transposon emerged as the most efficient transgenesis strategy and is adaptable to promoters and transgenes of choice. These strategies now enable the generation of diverse genetic tools for mechanistic studies of cephalopod biology.

genetics↗

Large language model-based bibliometric evaluation of population descriptors in human genetics

As the use of population descriptors such as race, ethnicity, and ancestry have become increasingly common in modern genetics research, there have been growing calls to critically examine their use. Most notably, in 2023, the National Academies of Science, Engineering, and Medicine (NASEM) published a report titled Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field, which included eight specific and actionable recommendations for researchers to implement the ethical and accurate use of population descriptors in genetic research. Here, we use the 2023 NASEM report as a benchmark to analyze the use of population descriptors in genome-wide association studies (GWAS). We develop a general toolkit for large language model-based bibliometrics, operationalize the report's recommendations into an evaluation framework, and apply this framework to evaluate all 4,007 papers from the GWAS Catalog published between 2007 and 2025 with full text available on PubMedCentral. We find significant improvements in adherence to NASEM report recommendations over time. However, most improvements predate the publication of the NASEM report itself, suggesting the report functioned primarily as a synthesis of existing best practices rather than a catalyst for change. We conclude by highlighting opportunities for growth in the field of human genetics.

genetics↗

Mitigating biases of rescaling in forward-in-time population genetic simulations

Forward-in-time population genetic simulations are widely used in evolutionary analyses, but simulating large populations and long genomic regions remains computationally demanding. To reduce this cost, parameter rescaling is widely employed, in which the original evolutionary process is approximated by one with a smaller population size and fewer generations. Recently, several studies using the SLiM simulator have raised concerns about the accuracy of this rescaling approach. In this study, we show that many of the biases reported in these studies can be mitigated by using a different simulation algorithm. These results reveal that the accuracy of parameter rescaling depends on how well the simulation algorithm preserves diffusion-limit properties under rescaling.

genetics↗