Search bioRxiv⌕ Search

Biology subjects

DeFelice, M.

Publications and source records attributed to DeFelice, M..

4 recordsLinked to original sources

Long-Read Sequencing of the MUC1 VNTR: Genomic Variation, Mutational Landscape, and Its Impact on ADTKD Diagnosis and Progression

BackgroundADTKD-MUC1 is caused by frameshift mutations in MUC1 gene that produce a frameshifted protein (MUC1fs) toxic to kidney cells. The genes variable number of tandem repeats (VNTR), with high GC content, makes it largely inaccessible to standard sequencing. As a result, both the reference sequence and natural variation in this region remain poorly defined, complicating mutation detection and data interpretation. Standard methods also fail to pinpoint the exact VNTR unit affected, limiting insight into mutation mechanisms and genotype-phenotype correlations. MethodsWe employed Single Molecule, Real-Time (SMRT) sequencing and characterized the genomic sequence of MUC1 in 300 individuals including 279 individuals from 143 families suspected of having ADTKD-MUC1. We compared these results to those obtained using the CLIA-approved mass spectrometry-based probe extension (PE) assay, which specifically detect the most prevalent 59dupC mutation. We correlated the structural features of the MUC1 VNTR with the rate of kidney function decline in affected individuals. ResultsWe identified MUC1 consensus sequences for 205 unique VNTR alleles, with 9 distinct types of frameshift mutations present on 52 distinct mutated VNTR alleles. MUC1 frameshift mutations were identified in 71 of 143 families (50%) with suspected ADTKD, comprising 135 genetically affected individuals (48%). The SMRT assay exhibited complete concordance and revealed that the PE assay is capable of detecting frameshift mutations in approximately 85% of affected families. The constellation of VNTR structures supports a genotype-progression model, in which fast progressors exhibit a significantly lower number of repeat units on the wild-type allele and a higher number of repeats on the mutation-bearing allele, including an increased number of frameshifted repeat units. ConclusionsSMRT sequencing outperforms current diagnostic methods for ADTKD-MUC1 and reveals the prognostic value of VNTR structures. Although their contribution to disease progression is modest ([~]6% variance explained), it remains biologically and clinically meaningful. Key Points 3O_LISingle Molecule, Real-Time (SMRT) sequencing of MUC1 outperforms existing genetic and immunohistochemical methods for ADTKD-MUC1 diagnosis. C_LIO_LICLIA-approved mass spectrometry-based genotyping assay, complemented by SMRT sequencing, can detect nearly all ADTKD-MUC1 patients. C_LIO_LIDisease progression correlates with VNTR pattern: fast progressors have fewer WT repeats and more mutant/frameshifted repeats C_LI

genetics↗

A blended genome and exome sequencing method captures genetic variation in an unbiased, high-quality, and cost-effective manner

We deployed the Blended Genome Exome (BGE), a DNA library blending approach that generates low pass whole genome (1-4x mean depth) and deep whole exome (30-40x mean depth) data in a single sequencing run. This technology is cost-effective, empowers most genomic discoveries possible with deep whole genome sequencing, and provides an unbiased method to capture the diversity of common SNP variation across the globe. To evaluate this new technology at scale, we applied BGE to sequence >53,000 samples from the Populations Underrepresented in Mental Illness Associations Studies (PUMAS) Project, which included participants across African, African American, and Latin American populations. We evaluated the accuracy of BGE imputed genotypes against raw genotype calls from the Illumina Global Screening Array. All PUMAS cohorts had R2 concordance [&ge;]95% among SNPs with MAF[&ge;]1%, and never fell below [&ge;]90% R2 for SNPs with MAF<1%. Furthermore, concordance rates among local ancestries within two recently admixed cohorts were consistent among SNPs with MAF[&ge;]1%, with only minor deviations in SNPs with MAF<1%. We also benchmarked the discovery capacity of BGE to access protein-coding copy number variants (CNVs) against deep whole genome data, finding that deletions and duplications spanning at least 3 exons had a positive predicted value of [~]90%. Our results demonstrate BGE scalability and efficacy in capturing SNPs, indels, and CNVs in the human genome at 28% of the cost of deep whole-genome sequencing. BGE is poised to enhance access to genomic testing and empower genomic discoveries, particularly in underrepresented populations.

genomics↗

Sharing Data from the Human Tumor Atlas Network through Standards, Infrastructure, and Community Engagement

The Data Coordinating Center (DCC) of the Human Tumor Atlas Network (HTAN) has played a crucial role in enabling the broad sharing and effective utilization of HTAN data within the scienti[fi]c community. Data from the [fi]rst phase of HTAN are now available publicly. We describe the diverse datasets and modalities shared, multiple access routes to HTAN assay data and metadata, data standards, technical infrastructure and governance approaches, as well as our approach to sustained community engagement. HTAN data can be accessed via the HTAN Portal, explored in visualization tools--including CellxGene, Minerva, and cBioPortal--and analyzed in the cloud through the NCI Cancer Research Data Commons nodes. We have developed a streamlined infrastructure to ingest and disseminate data by leveraging the Synapse platform. Taken together, the HTAN DCCs approach demonstrates a successful model for coordinating, standardizing, and disseminating complex cancer research data via multiple resources in the cancer data ecosystem, offering valuable insights for similar consortia, and researchers looking to leverage HTAN data.

cancer biology↗

Blended Genome Exome (BGE) as a Cost Efficient Alternative to Deep Whole Genomes or Arrays

Genomic scientists have long been promised cheaper DNA sequencing, but deep whole genomes are still costly, especially when considered for large cohorts in population-level studies. More affordable options include microarrays + imputation, whole exome sequencing (WES), or low-pass whole genome sequencing (WGS) + imputation. WES + array + imputation has recently been shown to yield 99% of association signals detected by WGS. However, a method free from ascertainment biases of arrays or the need for merging different data types that still benefits from deeper exome coverage to enhance novel coding variant detection does not exist. We developed a new, combined, "Blended Genome Exome" (BGE) in which a whole genome library is generated, an aliquot of that genome is amplified by PCR, the exome regions are selected and enriched, and the genome and exome libraries are combined back into a single tube for sequencing (33% exome, 67% genome). This creates a single CRAM with a low-coverage whole genome (2-3x) combined with a higher coverage exome (30-40x). This BGE can be used for imputing common variants throughout the genome as well as for calling rare coding variants. We tested this new method and observed >99% r2 concordance between imputed BGE data and existing 30x WGS data for exome and genome variants. BGE can serve as a useful and cost-efficient alternative sequencing product for genomic researchers, requiring ten-fold less sequencing compared to 30x WGS without the need for complicated harmonization of array and sequencing data.

genomics↗