Search bioRxivSearch

Biology subjects

Myers, R. M.

Publications and source records attributed to Myers, R. M..

9 recordsLinked to original sources

De novo mutations in the GTP/GDP-binding region of RALA, a RAS-like small GTPase, cause intellectual disability and developmental delay

Mutations that alter signaling of RAS/MAPK-family proteins give rise to a group of Mendelian diseases known as RASopathies, but the matrix of genotype-phenotype relationships is still incomplete, in part because there are many RAS-related proteins, and in part because the phenotypic consequences may be variable and/or pleiotropic. Here, we describe a cohort of ten cases, drawn from six clinical sites and over 16,000 sequenced probands, with de novo protein-altering variation in RALA, a RAS-like small GTPase. All probands present with speech and motor delays, and most have intellectual disability, low weight, short stature, and facial dysmorphism. The observed rate of de novo RALA variants in affected probands is significantly higher (p=4.93 x 10-11) than expected from the estimated mutation rate. Further, all de novo variants described here affect conserved residues within the GTP/GDP-binding region of RALA; in fact, six alleles arose at only two codons, Val25 and Lys128. We directly assayed GTP hydrolysis and RALA effector-protein binding, and all but one tested variant significantly reduced both activities. The one exception, S157A, reduced GTP hydrolysis but significantly increased RALA-effector binding, an observation similar to that seen for oncogenic RAS variants. These results show the power of data sharing for the interpretation and analysis of rare variation, expand the spectrum of molecular causes of developmental disability to include RALA, and provide additional insight into the pathogenesis of human disease caused by mutations in small GTPases.\n\nAuthor SummaryWhile many causes of developmental disabilities have been identified, a large number of affected children cannot be diagnosed despite extensive medical testing. Previously unknown genetic factors are likely to be the culprits in many of these cases. Using DNA sequencing, and by sharing information among many doctors and researchers, we have identified a set of individuals with developmental problems who all have changes to the same gene, RALA. The affected individuals all have similar symptoms, including intellectual disability, speech delay (or no speech), and problems with motor skills like walking. In nearly all of these cases (10 of 11), the genetic change found in the child was not inherited from either parent. The locations and biological properties of these changes suggest that they are likely to disrupt the normal functions of RALA and cause significant health problems. We also performed experiments to show that the genetic changes found in these individuals alter two key functions of RALA. Together, we have provided evidence that genetic changes in RALA can cause DD/ID. These results will allow doctors and researchers to identify additional children with the same condition, providing a clinical diagnosis to these families and leading to new research opportunities.

genetics

Regional collapsing of rare variation implicates specific genic regions in ALS

Large-scale sequencing efforts in amyotrophic lateral sclerosis (ALS) have implicated novel genes using gene-based collapsing methods. However, pathogenic mutations may be concentrated in specific genic regions. To address this, we developed two collapsing strategies, one focuses rare variation collapsing on homology-based protein domains as the unit for collapsing and another gene-level approach that, unlike standard methods, leverages existing evidence of purifying selection against missense variation on said domains. The application of these two collapsing methods to 3,093 ALS cases and 8,186 controls of European ancestry, and also 3,239 cases and 11,808 controls of diversified populations, pinpoints risk regions of ALS genes including SOD1, NEK1, TARDBP and FUS. While not clearly implicating novel ALS genes, the new analyses not only pinpoint risk regions in known genes but also highlight candidate genes as well.

genetics

CRISPR/Cas9-targeted removal of unwanted sequences from small-RNA sequencing libraries

In small RNA (smRNAs) sequencing studies, highly abundant molecules such as adapter dimer products and tissue-specific microRNAs (miRNAs) inhibit accurate quantification of lowly expressed species. We previously developed a method to selectively deplete highly abundant miRNAs. However, this method does not deplete adapter dimer ligation products that, unless removed by gelseparation, comprise most of the library. Here, we have adapted and modified recently described methods for CRISPR/Cas9-based Depletion of Abundant Species by Hybridization (\"DASH\") to smRNA-seq, which we have termed miRNA and Adapter Dimer - DASH (MAD-DASH). In MAD-DASH, Cas9 is complexed with sgRNAs targeting adapter dimer ligation products, alongside highly expressed tissue-specific smRNAs, for cleavage in vitro. This process dramatically reduces (>90%) adapter dimer and targeted smRNA sequences, is multiplexable, shows minimal off-target effects, improves the quantification of lowly expressed miRNAs from human plasma and tissue derived RNA, and obviates the need for gel-separation, greatly increasing sample throughput. Additionally, the method is fully customizable to other smRNA-seq preparation methods. Like depletion of ribosomal RNA for mRNA-seq and mitochondrial DNA for ATAC-seq, our method allows for greater proportional read-depth of nontargeted sequences.

genomics

powerTCR: a model-based approach to comparative analysis of the clone size distribution of the T cell receptor repertoire

Sequencing of the T cell receptor repertoire is a powerful tool for deeper study of immune response, but the unique structure of this type of data makes its meaningful quantification challenging. We introduce a new method, the Gamma-GPD spliced threshold model, to address this difficulty. This biologically interpretable model captures the distribution of the TCR repertoire, demonstrates stability across varying sequencing depths, and permits comparative analysis across any number of sampled individuals. We apply our method to several datasets and obtain insights regarding the differentiating features in the T cell receptor repertoire among sampled individuals across conditions. We have implemented our method in the open-source R package powerTCR.\n\nAuthor summaryA more detailed understanding of the immune response can unlock critical information concerning diagnosis and treatment of disease. Here, in particular, we study T cells through T cell receptor sequencing, as T cells play a vital role in immune response. One important feature of T cell receptor sequencing data is the frequencies of each receptor in a given sample. These frequencies harbor global information about the landscape of the immune response. We introduce a flexible method that extracts this information by modeling the distribution of these frequencies, and show that it can be used to quantify differences in samples from individuals of different biological conditions.

immunology

Genomic sequencing identifies secondary findings in a cohort of parent study participants

PURPOSEClinically relevant secondary variants were identified in parents enrolled with a child with developmental delay and intellectual disability.\n\nMETHODSExome/genome sequencing and analysis of 789 unaffected parents was performed.\n\nRESULTSPathogenic/likely pathogenic variants were identified in 21 genes within 25 individuals (3.2%), with 11 (1.4%) participants harboring variation in a gene defined as clinically actionable by the ACMG. Of the 25 individuals, five carried a variant consistent with a previous clinical diagnosis, thirteen were not previously diagnosed but had symptoms or family history with probable association with the detected variant, and seven reported no symptoms or family history of disease. A limited carrier screen was performed yielding 15 variants in 48 (6.1%) parents. Parents were also analyzed as mate-pairs to identify cases in which both parents were carriers for the same recessive disease; this led to one finding in ATP7B. Four participants had two findings (one carrier and one non-carrier variant). In total, 71 of the 789 enrolled parents (9.0%) received secondary findings.\n\nCONCLUSIONWe provide an overview of the rates and types of clinically relevant secondary findings, which may be useful in the design, and implementation of research and clinical sequencing efforts to identify such findings.

genomics

A genome-wide interactome of DNA-associated proteins in the human liver

Large-scale efforts like the Encyclopedia of DNA Elements (ENCODE) Project have made tremendous progress in cataloging the genomic binding patterns of DNA-associated proteins (DAPs), such as transcription factors (TFs). However most chromatin immunoprecipitation-sequencing (ChIP-seq) analyses have focused on a few immortalized cell lines whose activities and physiology deviate in important ways from endogenous cells and tissues. Consequently, binding data from primary human tissue are essential to improving our understanding of in vivo gene regulation. Here we analyze ChIP-seq data for 20 DAPs assayed in two healthy human liver tissue samples, identifying more than 450,000 binding sites. We integrated binding data with transcriptome and phased whole genome data to investigate allelic DAP interactions and the impact of heterozygous sequence variation on the expression of neighboring genes. We find our tissue-based dataset demonstrates binding patterns more consistent with liver biology than cell lines, and describe uses of these data to better prioritize impactful non-coding variation. Collectively, our rich dataset offers novel insights into genome function in healthy liver tissue and provides a valuable research resource for assessing disease-related disruptions.

genomics

Extremely rare variants reveal patterns of germline mutation rate heterogeneity in humans

A detailed understanding of the genome-wide variability of single-nucleotide germline mutation rates is essential to studying human genome evolution. Here we use [~]36 million singleton variants from 3,560 whole-genome sequences to infer fine-scale patterns of mutation rate heterogeneity. Mutability is jointly affected by adjacent nucleotide context and diverse genomic features of the surrounding region, including histone modifications, replication timing, and recombination rate, sometimes suggesting specific mutagenic mechanisms. Remarkably, GC content, DNase hypersensitivity, CpG islands, and H3K36 trimethylation are associated with both increased and decreased mutation rates depending on nucleotide context. We validate these estimated effects in an independent dataset of [~]46,000 de novo mutations, and confirm our estimates are more accurate than previously published estimates based on ancestrally older variants without considering genomic features. Our results thus provide the most refined portrait to date of the factors contributing to genome-wide variability of the human germline mutation rate.

genomics

INFERENCE OF CELL TYPE COMPOSITION FROM HUMAN BRAIN TRANSCRIPTOMIC DATASETS ILLUMINATES THE EFFECTS OF AGE, MANNER OF DEATH, DISSECTION, AND PSYCHIATRIC DIAGNOSIS

Psychiatric illness is unlikely to arise from pathology occurring uniformly across all cell types in affected brain regions. Despite this, transcriptomic analyses of the human brain have typically been conducted using macro-dissected tissue due to the difficulty of performing single-cell type analyses with donated post-mortem brains. To address this issue statistically, we compiled a database of several thousand transcripts that were specifically-enriched in one of 10 primary cortical cell types in previous publications. Using this database, we predicted the relative cell type composition for 833 human cortical samples using microarray or RNA-Seq data from the Pritzker Consortium (GSE92538) or publicly-available databases (GSE53987, GSE21935, GSE21138, CommonMind Consortium). These predictions were generated by averaging normalized expression levels across transcripts specific to each cell type using our R-package BrainInABlender (validated and publicly-released: https://github.com/hagenaue/BrainInABlender). Using this method, we found that the principal components of variation in the datasets strongly correlated with the neuron to glia ratio of the samples.\n\nThis variability was not simply due to dissection - the relative balance of brain cell types appeared to be influenced by a variety of demographic, pre- and post-mortem variables. Prolonged hypoxia around the time of death predicted increased astrocytic and endothelial gene expression, illustrating vascular upregulation. Aging was associated with decreased neuronal gene expression. Red blood cell gene expression was reduced in individuals who died following systemic blood loss. Subjects with Major Depressive Disorder had decreased astrocytic gene expression, mirroring previous morphometric observations. Subjects with Schizophrenia had reduced red blood cell gene expression, resembling the hypofrontality detected in fMRI experiments. Finally, in datasets containing samples with especially variable cell content, we found that controlling for predicted sample cell content while evaluating differential expression improved the detection of previously-identified psychiatric effects. We conclude that accounting for cell type can greatly improve the interpretability of transcriptomic data.

bioinformatics

Genomic diagnosis for children with intellectual disability and/or developmental delay

BackgroundDevelopmental disabilities have diverse genetic causes that must be identified to facilitate precise diagnoses. We describe genomic data from 371 affected individuals, 309 of which were sequenced as proband-parent trios.\n\nMethodsWhole exome sequences (WES) were generated for 365 individuals (127 affected) and whole genome sequences (WGS) were generated for 612 individuals (244 affected).\n\nResultsPathogenic or likely pathogenic variants were found in 100 individuals (27%), with variants of uncertain significance in an additional 42 (11.3%). We found that a family history of neurological disease, especially the presence of an affected 1st degree relative, reduces the pathogenic/likely pathogenic variant identification rate, reflecting both the disease relevance and ease of interpretation of de novo variants. We also found that improvements to genetic knowledge facilitated interpretation changes in many cases. Through systematic reanalyses we have thus far reclassified 15 variants, with 11.3% of families who initially were found to harbor a VUS, and 4.7% of families with a negative result, eventually found to harbor a pathogenic or likely pathogenic variant. To further such progress, the data described here are being shared through ClinVar, GeneMatcher, and dbGAP.\n\nConclusionOur data strongly support the value of large-scale sequencing, especially WGS within proband-parent trios, as both an effective first-choice diagnostic tool and means to advance clinical and research progress related to pediatric neurological disease.

genomics