Search bioRxivSearch

Biology subjects

Ying Hu

Publications and source records attributed to Ying Hu.

2 recordsLinked to original sources

Unbiased K-mer Analysis Reveals Changes in Copy Number of Highly Repetitive Sequences During Maize Domestication and Improvement

The major component of complex genomes is repetitive elements, which remain recalcitrant to characterization. Using maize as a model system, we analyzed whole genome shotgun (WGS) sequences for the two maize inbred lines B73 and Mo17 using k-mer analysis to quantify the differences between the two genomes. Significant differences were identified in highly repetitive sequences, including centromere repeats, 45S ribosomal DNA (rDNA), knob, and telomere repeats. Previously unknown genotype specific 45S rDNA sequences were discovered. The B73-specific 45S rDNA is not only located on the nucleolus organizer region (NOR) on chromosome 6 but also dispersed on all the chromosomes in B73, indicating the relatively recent spread of 45S rDNA from the NOR. The B73 and Mo17 polymorphic k-mers were used to examine allele-specific expression of 45S rDNA. Although Mo17 contains higher copy number than B73, equivalent levels of overall 45S rDNA expression indicates that dosage compensation operates for the 45S rDNA in the hybrids. Using WGS sequences of B73xMo17 double haploids (DHs), genomic locations showing differential repetitive contents were genetically mapped. Analysis of WGS sequences of HapMap2 lines, including maize wild progenitor teosintes, landraces, and improved lines, decreases and increases in abundance of additional sets of k-mers associated with centromere repeats, 45S rDNA, knob, and retrotransposon sequences were found between teosinte and maize lines, revealing global evolutionary trends of genomic repeats during maize domestication and improvement.

Genomics

Improving genetic diagnosis in Mendelian disease with transcriptome sequencing

Exome and whole-genome sequencing are becoming increasingly routine approaches in Mendelian disease diagnosis. Despite their success, the current diagnostic rate for genomic analyses across a variety of rare diseases is approximately 25-50%. Here, we explore the utility of transcriptome sequencing (RNA-seq) as a complementary diagnostic tool in a cohort of 50 patients with genetically undiagnosed rare muscle disorders. We describe an integrated approach to analyze patient muscle RNA-seq, leveraging an analysis framework focused on the detection of transcript-level changes that are unique to the patient compared to over 180 control skeletal muscle samples. We demonstrate the power of RNA-seq to validate candidate splice-disrupting mutations and to identify splice-altering variants in both exonic and deep intronic regions, yielding an overall diagnosis rate of 35%. We also report the discovery of a highly recurrent de novo intronic mutation in COL6A1 that results in a dominantly acting splice-gain event, disrupting the critical glycine repeat motif of the triple helical domain. We identify this pathogenic variant in a total of 27 genetically unsolved patients in an external collagen VI-like dystrophy cohort, thus explaining approximately 25% of patients clinically suggestive of collagen VI dystrophy in whom prior genetic analysis is negative. Overall, this study represents a large systematic application of transcriptome sequencing to rare disease diagnosis and highlights its utility for the detection and interpretation of variants missed by current standard diagnostic approaches.\n\nOne Sentence SummaryTranscriptome sequencing improves the diagnostic rate for Mendelian disease in patients for whom genetic analysis has not returned a diagnosis.

Genomics