Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.02.27.640502

SyFi: generating and using sequence fingerprints to distinguish SynCom isolates

Abstract

The plant root microbiome is a complex community shaped by interactions among bacteria, the plant host, and the environment. Synthetic community (SynCom) experiments help disentangle these interactions by inoculating host plants with a representative set of culturable microbial isolates from the natural root microbiome. Studying these simplified communities provides valuable insights into microbiome assembly and function. However, as SynComs become increasingly complex to better represent natural communities, bioinformatics challenges arise. Specifically, accurately identifying, and quantifying SynCom members based on, for example, 16S rRNA amplicon sequencing becomes more difficult due to the high similarity of the target amplicon, limiting downstream interpretations. Here, we present SynCom Fingerprinting (SyFi), a bioinformatics workflow designed to improve the resolution and accuracy of SynCom member identification. SyFi consists of three modules: the first module constructs a genomic fingerprint for each SynCom member based on its genome sequence, accounting for both copy number and sequence variation in the target gene. The second module then extracts a specific region from this genomic fingerprint to create a secondary fingerprint focused on the target amplicon. The third module uses these fingerprints as a reference to perform pseudoalignment-based quantification of SynCom member abundance from amplicon sequencing reads. We demonstrate that SyFi outperforms standard amplicon analysis by leveraging natural intragenomic variation, enabling more precise differentiation of closely related SynCom members. As a result, SyFi enhances the reliability of microbiome experiments using complex SynComs, which more accurately reflect natural communities. This improved resolution is essential for advancing our understanding of the root microbiome and its impact on plant health and productivity in agricultural and ecological settings. SyFi is available at https://github.com/adriangeerre/SyFi. Impact statementSyFi represents a significant advancement in microbiome research by enhancing the accuracy and resolution of synthetic community (SynCom) member identification. By leveraging natural intragenomic variation, SyFi improves the differentiation of closely related microbial strains, addressing a key challenge in amplicon-based sequencing analysis. This increased precision allows researchers to more reliably track microbial dynamics in complex SynCom experiments, leading to deeper insights into microbiome assembly, function, and host-microbe interactions. As a result, SyFi strengthens the interpretability of microbiome studies, ultimately contributing to a better understanding of plant health and productivity in both agricultural and ecological contexts. Data SummaryThe data reported in this article have been deposited in the National Center for Biotechnology Information Short Read Archive BioProject database. SyFi fingerprint generation was run on a collection of 737 human gut-derived bacterial genomes from Forster et al. (2019) (Genomic read data deposited in the ENA under project numbers ERP105624 and ERP012217) and 447 Arabidopsis-derived bacterial genomes (Selten et al., 2024b) (NCBI Project numbers PRJNA1138681, PRJNA1139421 (Genomes), and PRJNA1131834 (Genomic reads)). A list of the closed genomes used for SyFi validation can be found at https://github.com/adriangeerre/SyFi. Subsequently, SyFi was validated on a complex SynCom dataset by pseudoaligning 16S rRNA V3-V4 and V5-V7 amplicon reads (PRJNA1191388) to SyFi-generated fingerprints and comparing this to shotgun metagenomics-sequenced dataset of the same samples in Selten et al. (2024a) (PRJNA1131994). This complex SynCom dataset included the inoculation of the 447 bacterial isolates on Arabidopsis, Barley, and Lotus roots.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Selten, G., Gomez Repolles, A., Lamouche, F., Radutoiu, S., de Jonge, R.. 2025-03-02. SyFi: generating and using sequence fingerprints to distinguish SynCom isolates. https://doi.org/10.1101/2025.02.27.640502

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Sequence and epigenetic characterization of chromosome 21 centromeres in a family with recurrent Trisomy 21

Trisomy 21 (T21) is the most common genetic cause of intellectual disability, yet the molecular mechanisms underlying maternal meiosis I errors--responsible for ~70% of free T21 cases--remain poorly understood. In this preliminary study, we used long-read sequencing and genome assembly to investigate the DNA sequence and epigenetic features of chromosome 21 (chr21) centromeres in a family with recurrent free T21 due to maternal meiosis I errors. The mother, who had two affected and three unaffected children, showed no mosaicism or structural rearrangements. One of her two chr21 centromeres lacked a pronounced centromere dip region (CDR), displaying instead a diffuse hypomethylation pattern (dCDR) with much higher methylated CpG levels (55%) compared to its homologue (36%). This dCDR was transmitted to an unaffected child and the affected proband analyzed, suggesting it was present in one of the maternal chr21 since she was at least 32 years of age. Chr21 dCDRs were not observed in seven young mothers with children with T21 or previously described in the literature in 108 population haplotypes. We hypothesize that dCDRs may weaken kinetochore function, increasing nondisjunction risk, and propose two models linking such epigenetic variation to maternal age-related T21 risk. These findings highlight the value of complete centromere characterization in families with children with T21 and suggest centromere methylation status of chr21 as a potential T21 risk factor for future investigation.

genomics↗

Single-Cell Analytics for Dose Response (SCADR) discriminates PTEN missense variants by lipid and protein phosphatase dysfunction

The proliferation of sequencing efforts has revealed a vast and expanding catalog of single nucleotide gene variants, many associated to, but with unclear roles in disease. Fully charactering variant impacts and linking specific protein dysfunctions to disease are challenging due to the multi-functional nature of many proteins and varying degree of variant effects on these functions. Lagging are sensitive approaches to empirically assess the impact of missense variant-induced single amino acid changes on a wide range of protein functions. To address these issues, we have developed an open-source computational analysis tool called SCADR (Single-Cell Analytics for Dose Response) for simultaneously measuring and comparing impacts of exogenously-expressed variants on multiple signaling pathways using multiplex phospho-antibody spectral flow cytometry in human cell lines. SCADR retains and correlates single-cell measures of signal protein activity states along with expression levels of exogenously-expressed variants, providing rich characterization of multiple protein functions, signaling protein interactions, and enhanced discrimination of variant impacts on different signaling pathways, highlighting each variants unique dysfunction profile. Here, we apply SCADR for analyses of the impact of 6 variants of the tumor-suppressor protein PTEN (P38H, C124S, G129E, Y138L, D268E, 4A) expressed in HEK293 cells on the phosphorylation states of the canonical and noncanonical downstream signaling proteins Akt, S6, CREB, ERK, and p38 detected with fluorophore-conjugated phospho-antibodies, along with an antibody detecting an N-terminal HA tag on PTEN variants allowing measures of dose-response effects of each variants expression on signaling cascades. Results identify variant-specific impacts on downstream signaling cascades.

genomics↗

Microsecond molecular dynamics of SOD1 variants suggest a structural basis for divergent ALS clinical outcomes

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease characterised by progressive motor neuron degeneration. Mutations in the SOD1 gene represent the second most common genetic cause of ALS (ALS), and distinct SOD1 missense variants present with markedly different clinical profiles. A4V leads to an aggressive form of the disease (median survival [~]1y), H46R confers a mild, slowly progressive course and I113T exhibits an intermediate phenotype. The molecular basis by which these mutations produce divergent clinical outcomes remains poorly understood. We performed extensive classical molecular dynamics simulations of wild-type SOD1 and the three ALS-associated variants in the apo monomeric state to attempt to investigate the mechanisms behind such phenotypic differences. Structural stability, global compactness, and conformational flexibility, as well as analysis of collective motions between residues and estimation of free energy, were assessed. The H46R, A4V, and I113T variants exhibited distinct dynamic behaviours, highlighting differences in structural stability, local flexibility, and intramolecular interactions. These findings suggest that specific structural regions may contribute differently to protein dysfunction and could represent key elements for understanding the relationship between molecular dynamic properties and the differing clinical severity associated with these variants. Most strikingly, H46R exhibited exceptional structural stability across every analytical level, the lowest global deviation, most attenuated local flexibility, strongest internal dynamic coordination, and the deepest, most confined free energy basins of any system examined. This convergent multi-layered evidence of structural restraint provides a compelling mechanistic basis for the mild and slowly progressive clinical course of H46R ALS, suggesting that enhanced conformational rigidity, rather than bulk destabilisation, is the defining biophysical feature of this variant, and that its pathogenic mechanism operates through a route fundamentally decoupled from the aggregation-driven toxicity that characterises the more aggressive SOD1-ALS mutations.

genomics↗