Search bioRxiv⌕ Search

Biology subjects

Paulsen, R.

Publications and source records attributed to Paulsen, R..

2 recordsLinked to original sources

Evidence Aggregator: AI reasoning applied to rare disease diagnostics

Variant assessment of rare disease diagnostics depends on using domain knowledge in the time- consuming process of retrieving, reviewing, and synthesizing clinical and technical information. To address these challenges, we developed the Evidence Aggregator (EvAgg), an open-source, generative-AI-based tool designed for rare disease diagnosis that systematically extracts relevant information from the scientific literature for any human gene. EvAgg provides a thorough and current summary of observed genetic variants and their associated clinical features, enabling rapid synthesis of evidence concerning gene-disease relationships. We constructed an expert-curated dataset and evaluated EvAggs performance. EvAgg achieves 92% recall in identifying relevant papers, 96% recall in detecting instances of genetic variation within those papers, and [~]80% accuracy in extracting individual case and variant-level content (e.g. zygosity, inheritance, variant type, and phenotype). Further, EvAgg complemented the process of manual literature review by identifying substantial additional relevant information. When tested with analysts in rare disease case analysis, EvAgg reduced review time by 34% (p-value < 0.002) and increased the number of papers, variants, and cases evaluated per unit time. These savings have the potential to reduce diagnostic latency and increase solve rates for challenging rare disease cases.

genomics↗

Blended Length Genome Sequencing (blend-seq): Combining Short Reads with Low-Coverage Long Reads to Maximize Variant Discovery

We introduce blend-seq, a workflow for combining data from traditional short-read sequencing pipelines with low-coverage long reads, to improve variant discovery for single samples without the full cost of high-coverage long reads. We demonstrate that with only 4x long-read coverage augmenting 30x short reads, we can improve SNP discovery across the genome, exceeding performance beyond even high-coverage short reads (60x). For genotype-agnostic discovery of structural variants, we see a threefold improvement in recall while maintaining precision by using the low-coverage long reads on their own, and show how we can improve genotyping accuracy by adding in the short-read data. In addition, we demonstrate how the long reads can better phase these variants, incorporating long-context information in the genome to substantially outperform phasing with short reads alone. Our experiments highlight the complementary nature of short- and long-read technologies: the former contributing higher depth for genotyping and the latter better resolution of larger events or those in difficult regions.

genomics↗