Search bioRxivSearch

Biology subjects

Cogan, J.

Publications and source records attributed to Cogan, J..

2 recordsLinked to original sources

Programmatic Detection of Diploid-Triploid Mixoploidy via Whole Genome Sequencing

PurposeMixoploidy is a type of mosaicism where an organism is a mixture of cells with different numbers of chromosomes. There are a broad range of phenotypes associated with mixoploidy that vary greatly depending on the fraction of cells that are non-diploid, their chromosome number, their distribution, and presumably the specific variation present in the patient. Clinical detection of mixoploidy is important for diagnosis.\n\nMethodsWe developed a method to detect mixoploidy from clinical whole genome sequencing (WGS) data through the identification of excess of variant calls centered on unusual B-allele frequencies. Our method isolates the signal from these variants using trio calls and then solves a basic linear equation to estimate levels of diploid-triploid mixoploidy within the sample.\n\nResultsWe show that our method reflects the results from a cytogenetic test. We provide examples detailing how our method has been used to identify diploid-triploid mixoploid individuals from within the NIH Undiagnosed Diseases Network. We present confirmatory findings obtained by clinical cytogenetic testing and show that our method can be used to identify the diploid-triploid ratio in these cases.\n\nConclusionWGS data from patients with rare diseases can be used to identify mixoploid individuals. Individuals with certain characteristics as discussed should be tested for mixoploidy as part of standard clinical pipeline procedures. Scripts that perform this calculation are publicly available at https://github.com/HudsonAlpha/mixoviz.

genomics

Comprehensive Analysis of Constraint on the Spatial Distribution of Missense Variants in Human Protein Structures

The spatial distribution of genetic variation within proteins is shaped by evolutionary constraint and thus can provide insights into the functional importance of protein regions and the potential pathogenicity of protein alterations. Here, we comprehensively evaluate the 3D spatial patterns of constraint on human germline and somatic variation in 4,568 solved protein structures. Different classes of coding variants have significantly different spatial distributions. Neutral missense variants exhibit a range of 3D constraint patterns, with a general trend of spatial dispersion driven by constraint on core residues. In contrast, germline and somatic disease-causing variants are significantly more likely to be clustered in protein structure space. We demonstrate that this difference in the spatial distributions of disease-associated and benign germline variants provides a signature for accurately classifying variants of unknown significance (VUS) that is complementary to current approaches for VUS classification. We further illustrate the clinical utility of our approach by classifying new mutations identified from patients with familial idiopathic pneumonia (FIP) that segregate with disease.

genetics