bioRxiv · 10.64898/2025.12.05.692400
KuPID: Kmer-based Upstream Preprocessing of Long Reads forIsoform Discovery
Abstract
Eukaryotic genes can encode multiple protein isoforms based on alternative splicing of their transcribed regions. Most modern novel isoform discovery methods function by identifying and assembling exon splice junctions from an RNAseq sample. However, splice junctions can only be accurately annotated with time-intensive dynamic programming alignment. This manuscript introduces KuPID, a method for preprocessing long RNAseq reads with the goal of better identifying novel isoform transcripts. KuPID utilizes kmer sketching as a pre-filter to quickly pseudo-align reads to known reference isoforms. Full alignment need only then be applied to reads that are most relevant to isoform discovery. Not only does KuPID speed up the discovery pipeline, it also increases downstream accuracy by filtering out extraneous reads. KuPID preprocessing simultaneously increases the f1 accuracy of isoform discovery pipelines by up to 16.7 points while decreasing the runtime by a factor of 2-3x. An optional mode permits a KuPID sample to be paired with both isoform discovery and transcript quantification. Code availability: https://github.com/mboro2000/KuPID.git
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Borowiak, M., Yu, Y. W.. 2025-12-09. KuPID: Kmer-based Upstream Preprocessing of Long Reads forIsoform Discovery. https://doi.org/10.64898/2025.12.05.692400
Cite the original work for its findings. Save a collection to share your selection of sources.