Search bioRxivSearch

Biology subjects

Li, J. Z.

Publications and source records attributed to Li, J. Z..

7 recordsLinked to original sources

Coverage-based detection of copy number alterations in mixed samples using DNA sequencing data: a theoretical framework for evaluating statistical power

1DNA sequencing can discover not only single-base variants but also copy-number alterations (CNAs). In shotgun sequencing, regions of CNAs show step-wise changes in read depth when compared to adjacent \"normal\" regions, allowing their detection by parametric statistical tests that compare the mean coverage in suspected regions against that of a baseline distribution. Traditionally, the power of such a test depends on (1) the integer number of copy number change, (2) the overall sequencing depth, (3) the length of the CNA region, (4) the read length and (5) the variation of coverage along the genome, which depends on many experimental factors, including whether the chosen platform is whole-genome, whole-exome, or targeted-panel sequencing. In cases involving inadvertent sample mixing or genuine somatic mosaicism, power also depends on the mixing ratio. However, the analysis of statistical power that considers the interplay of all these factors has not been systematically developed. Here we present a general analytical framework and a series of simulations that explore situations from the simplest to the increasingly multifactorial. Specifically, we expand the expression of power to include not just the known factors but also one or both of two complications: (1) the dispersion of read depth around the mean beyond the independent sampling-by-sequencing assumption, and (2) the reduced fraction of the CNA-bearing sample (\"purity\") as seen in studies of intratumor heterogeneity or in clinical monitoring of minimal residual disease. We describe the analytical formula and their simplifications in special cases, and share the extendable scripts for others to perform customized power analysis using study-specific parameters. As study designs vary and technologies continue to evolve, the input data and the noise characteristics will change depending on the practical situation. We present two use cases commonly encountered in cancer research: ultra-shallow whole-genome sequencing for detecting large, chromosome-scale events, and targeted ultra-deep sequencing for surveillance of known CNAs in rare tumor clones in the task of sensitive detection of cancer relapse or metastasis. We also present an online calculator at https://shiny.med.umich.edu/apps/hanyou/CNV_Detection_Power_Calculator/.

bioinformatics

Linkage disequilibrium connects genetic records of relatives typed with disjoint genomic marker sets

In familial searching in forensic genetics, a query DNA profile is tested against a database to determine whether it represents a relative of a database entrant. We examine the potential for using linkage disequilibrium to identify pairs of profiles as belonging to relatives when the query and database rely on nonoverlapping genetic markers. Considering data on individuals genotyped with both microsatellites used in forensic applications and genome-wide SNPs, we find that ~30-32% of parent-offspring pairs and ~35-36% of sib pairs can be identified from the SNPs of one member of the pair and the microsatellites of the other. The method suggests the possibility of performing familial searches of microsatellite databases using query SNP profiles, or vice versa. It also reveals that privacy concerns arising from computations across multiple databases that share no genetic markers in common entail risks not only for database entrants, but for their close relatives as well.

genetics

Extended regions of suspected mis-assembly in the rat reference genome

We performed whole-genome sequencing for eight inbred rat strains commonly used in genetic mapping studies, and they are the founders of the NIH heterogeneous stock (HS) outbred colony. We provide their sequences and variant calls to the rat genomics community. When analyzing the variant calls we identified regions with unusually high heterozygosity. We show that these regions are consistent across the eight inbred strains, including the BN strain, which was the basis of the rat reference genome. These regions show significantly higher read depth than other regions in the genome. The evidence suggests that these regions may contain segmental duplications that are incorrectly overlaid in the reference genome. We provide masks for these suspected regions of mis-assembly as a resource for the community to flag potentially false interpretations of mapping results or functional data.

genomics

Refinement of highly flexible protein structures using simulation-guided spectroscopy

Highly flexible proteins present a special challenge for structure determination because they are multi-structured yet not disordered, and the resulting conformational ensembles are essential for understanding function. Determining such ensembles is difficult because many measurements that capture multiple conformational populations provide sparse data. A powerful opportunity exists to leverage molecular simulations for spectroscopic experiment selection. We have developed an information-theoretic approach to guide experiments by identifying which measurements best refine the underlying conformational ensemble. We have tested this approach on three flexible bacterial proteins. For proteins where a clear mechanistic hypothesis drives label selection, our approach systematically identifies labels that would test this hypothesis. Furthermore, when available data do not yield an obvious mechanistically-guided label selection strategy, our approach guides label selection and produces conformational refinement that significantly outperforms standard structure-guided approaches. Our information-theoretic approach to label selection thus offers a particular advantage when refining challenging, underdetermined protein conformational ensembles.

bioinformatics

Extremely rare variants reveal patterns of germline mutation rate heterogeneity in humans

A detailed understanding of the genome-wide variability of single-nucleotide germline mutation rates is essential to studying human genome evolution. Here we use [~]36 million singleton variants from 3,560 whole-genome sequences to infer fine-scale patterns of mutation rate heterogeneity. Mutability is jointly affected by adjacent nucleotide context and diverse genomic features of the surrounding region, including histone modifications, replication timing, and recombination rate, sometimes suggesting specific mutagenic mechanisms. Remarkably, GC content, DNase hypersensitivity, CpG islands, and H3K36 trimethylation are associated with both increased and decreased mutation rates depending on nucleotide context. We validate these estimated effects in an independent dataset of [~]46,000 de novo mutations, and confirm our estimates are more accurate than previously published estimates based on ancestrally older variants without considering genomic features. Our results thus provide the most refined portrait to date of the factors contributing to genome-wide variability of the human germline mutation rate.

genomics

INFERENCE OF CELL TYPE COMPOSITION FROM HUMAN BRAIN TRANSCRIPTOMIC DATASETS ILLUMINATES THE EFFECTS OF AGE, MANNER OF DEATH, DISSECTION, AND PSYCHIATRIC DIAGNOSIS

Psychiatric illness is unlikely to arise from pathology occurring uniformly across all cell types in affected brain regions. Despite this, transcriptomic analyses of the human brain have typically been conducted using macro-dissected tissue due to the difficulty of performing single-cell type analyses with donated post-mortem brains. To address this issue statistically, we compiled a database of several thousand transcripts that were specifically-enriched in one of 10 primary cortical cell types in previous publications. Using this database, we predicted the relative cell type composition for 833 human cortical samples using microarray or RNA-Seq data from the Pritzker Consortium (GSE92538) or publicly-available databases (GSE53987, GSE21935, GSE21138, CommonMind Consortium). These predictions were generated by averaging normalized expression levels across transcripts specific to each cell type using our R-package BrainInABlender (validated and publicly-released: https://github.com/hagenaue/BrainInABlender). Using this method, we found that the principal components of variation in the datasets strongly correlated with the neuron to glia ratio of the samples.\n\nThis variability was not simply due to dissection - the relative balance of brain cell types appeared to be influenced by a variety of demographic, pre- and post-mortem variables. Prolonged hypoxia around the time of death predicted increased astrocytic and endothelial gene expression, illustrating vascular upregulation. Aging was associated with decreased neuronal gene expression. Red blood cell gene expression was reduced in individuals who died following systemic blood loss. Subjects with Major Depressive Disorder had decreased astrocytic gene expression, mirroring previous morphometric observations. Subjects with Schizophrenia had reduced red blood cell gene expression, resembling the hypofrontality detected in fMRI experiments. Finally, in datasets containing samples with especially variable cell content, we found that controlling for predicted sample cell content while evaluating differential expression improved the detection of previously-identified psychiatric effects. We conclude that accounting for cell type can greatly improve the interpretability of transcriptomic data.

bioinformatics

A sensitized mutagenesis screen in Factor V Leiden mice identifies novel thrombosis suppressor loci

Factor V Leiden (F5L) is a common genetic risk factor for venous thromboembolism in humans. We conducted a sensitized ENU mutagenesis screen for dominant thrombosuppressor genes based on perinatal lethal thrombosis in mice homozygous for F5L (F5L/L) and haploinsufficient for tissue factor pathway inhibitor (Tfpi+/-). F8 deficiency enhanced survival of F5L/L Tfpi+/- mice, demonstrating that F5L/L Tfpi+/- lethality is genetically suppressible. ENU-mutagenized F5L/L males and F5L/+ Tfpi+/- females were crossed to generate 6,729 progeny, with 98 F5L/L Tfpi+/- offspring surviving until weaning. Sixteen lines exhibited transmission of a putative thrombosuppressor to subsequent generations, with these lines referred to as MF5L (Modifier of Factor 5 Leiden) 1-16. Linkage analysis in MF5L6 identified a chromosome 3 locus containing the tissue factor gene (F3). Though no ENU-induced F3 mutation was identified, haploinsufficiency for F3 (F3+/-) suppressed F5L/L Tfpi+/- lethality. Whole exome sequencing in MF5L12 identified an Actr2 gene point mutation (p.R258G) as the sole candidate. Inheritance of this variant is associated with suppression of F5L/L Tfpi+/- lethality (p=1.7x10-6), suggesting that Actr2p.R258G is thrombosuppressive. CRISPR/Cas9 experiments to generate an independent Actr2 knockin/knockout demonstrated that Actr2 haploinsufficiency is lethal, supporting a hypomorphic or gain of function mechanism of action for Actr2p.R258G. Our findings identify F8 and the Tfpi/F3 axis as key regulators in determining thrombosis balance in the setting of F5L and also suggest a novel role for Actr2 in this process.\n\nSignificance StatementVenous thromboembolism (VTE) is a common disease characterized by the formation of inappropriate blood clots. Inheritance of specific genetic variants, such as the Factor V Leiden polymorphism, increases VTE susceptibility. However, only ~10% of people inheriting Factor V Leiden develop VTE, suggesting the involvement of other genes that are currently unknown. By inducing random genetic mutations into mice with a genetic predisposition to VTE, we identified two genomic regions that reduce VTE susceptibility. The first includes the gene for blood coagulation Factor 3 and its role was confirmed by analyzing mice with an independent mutation in this gene. The second contains a mutation in the Actr2 gene. These findings identify critical genes for the regulation of blood clotting risk.

genetics