Search bioRxivSearch

EXPLORE THE ARCHIVE

Yang, H.

Publications and source records attributed to Yang, H..

2 recordsLinked to original sources

Half-match recombination drives bridge RNA-guided excision and off-target insertion

IS110-family bridge recombinases are a recently identified class of compact, RNA-guided editors in which a bridge RNA (bRNA) directs the recombination of a donor DNA into a target site. In the current model, the bRNA engages fully complementary donor and target sequences within a single synaptic complex to drive double-stranded recombination, implying that the transposon is cut from its donor site rather than copied, yet neither the strandedness of the excised intermediate nor the requirement for full complementarity has been tested directly. Here we reconstituted IS621 recombination in a cell-free transcription-translation system, building representative arrangements of the excision and insertion reactions and characterizing the outcomes. We find that IS621 predominantly excises a single strand, releasing a single-stranded circle and leaving the donor site intact, consistent with copy-and-paste transposition. By introducing mismatches into the bRNA target sequences, we further find that excision proceeds independently of target-site complementarity, relying strictly on donor-arm recognition; we term this "half-match" recombination, because a substrate matching only half of the bRNA is sufficient. We also find half-match activity during insertion, both in vitro and in a published genome-editing experiment, where it accounts for approximately half of non-target insertion reads. Half-match recombination provides both a mechanistic explanation for off-target insertion and a framework for the rational design of high-fidelity bridge recombinases.

molecular biology

TomatoPGFM: A graph-conditioned foundation model for tomato pangenomes

Most genomic foundation models are pretrained on independent linear assemblies and therefore do not explicitly represent population-level segment sharing or local graph connectivity. We developed TomatoPGFM, a graph-conditioned model pretrained on 54.65 Gb of sequence from 66 tomato (Solanum spp.) accessions. Sequence tokens were conditioned on pangenome node attributes and local adjacency, and the model was optimised using masked language modelling and graph-feature reconstruction. To evaluate model responses to graph-conditioned input, we compared aligned, shuffled and disabled graph inputs in 25,000 windows from the training panel. Sequence-aligned graph input produced lower masked language modelling loss than graph-off at all five curriculum stages in both training-panel strata, while the shuffled perturbation generally yielded intermediate losses. We then assessed sequence-only transfer in Solanum sitiens LA1974 and S. lycopersicum MicroTom, neither of which was used for graph construction or pretraining. Frozen-probe AUROC values for gene-versus-intergenic and coding-sequence-versus-intergenic classification ranged from 0.8489 to 0.9593. TomatoPGFM produced higher AUROC point estimates than DNABERT-2 in all four comparisons. Enabling the zero-feature GraphAdapter pathway with adjacency messaging disabled changed throughput by less than 1% at 512-2,048 positions under the tested configuration. Together, these results show that TomatoPGFM responds consistently to sequence-aligned pangenome context in training-panel sequences and provides informative sequence representations for genic-region classification in accessions excluded from graph construction and pretraining.

bioinformatics