Search bioRxivSearch

Biology subjects

Robinson, M.

Publications and source records attributed to Robinson, M..

11 recordsLinked to original sources

Bayesian reassessment of the epigenetic architecture of complex traits

1Epigenetic DNA modification is partly under genetic control, and occurs in response to a wide range of environmental exposures. Linking epigenetic marks to clinical outcomes may provide greater insight into underlying molecular processes of disease, assist in the identification of therapeutic targets, and improve risk prediction. Here, we present a statistical approach, based on Bayesian inference, that estimates associations between disease risk and all measured epigenetic probes jointly, automatically controlling for both data structure (including cell-count effects, relatedness, and experimental batch effects) and correlations among probes. We benchmark our approach in simulation study, finding improved estimation of probe associations across a wide range of scenarios over existing approaches. Our method estimates the total proportion of disease risk captured by epigenetic probe variation, and when we applied it to measures of body mass index (BMI) and cigarette consumption behaviour in 5,101 individuals, we find that 66.7% (95% CI 60.0-72.8) of the variation in BMI and 67.7% (95% CI 58.4-76.9) of the variation in cigarette consumption can be captured by methylation array data from whole blood, independent of the variation explained by single nucleotide polymorphism markers. We find novel associations, with smoking behaviour associated with a methylation probe at the MNDA gene with >95% posterior inclusion probability, which is a myeloid cell nuclear differentiation antigen gene previously implicated as a biomarker for inflammation and non-Hodgkin lymphoma risk. We conduct unique genome-wide enrichment analyses, identifying blood cholesterol, lipid transport and sterol metabolism pathways for BMI, and response to xenobiotic stimulus and negative regulation of RNA polymerase II promoter transcription for smoking, all with >95% posterior inclusion probability of having methylation probes with associations >1.5 times larger than the average. Finally, we improve phenotypic prediction in two independent cohorts by 28.7% and 10.2% for BMI and smoking respectively over a LASSO model. These results imply that probe measures may capture large amounts of variance because they are likely a consequence of the phenotype rather than a cause. As a result, 'omics' data may enable accurate characterization of disease progression and identification of individuals who are on a path to disease. Our approach facilitates better understanding of the underlying epigenetic architecture of complex common disease and is applicable to any kind of genomics data.

genomics

Scale-invariant geometric data analysis (SIGDA) provides robust, detailed visualizations of human ancestry specific to individuals and populations

Scale invariance is a common property of physical laws and a key concept in perspective drawing, which aims to provide a meaningful two-dimensional representation of a more complex, three-dimensional scene. Here we describe Scale Invariant Geometric Data Analysis (SIGDA), a new, general exploratory data analysis (EDA) method based on normalization of data to scale invariance. We discuss similarities and differences between SIGDA and two widely-used EDA methods, Correspondence Analysis (CA) and Principal Components Analysis (PCA). We then illustrate SIGDAs ability to analyze and visualize population structure relationships within the data that inspired its development: genetic marker data, in which context PCA is considered a standard method. We show that SIGDA provides significant advantages over PCA of the same data, including: (a) robust detection and separation of a larger number of population axes, leading to (b) better separation of annotated populations; (c) separation of an independent allele frequency axis interpretable as a proxy for allele age, (d) visualization of marker flow between populations (population history), and (d) robust detection and visualization of relationships between closely-related individuals and among family groups. Although this illustration focuses on a specific task, SIGDA is a general-purpose EDA method and derives its advantages from its novel approach to fundamental issues in data analysis, rather than clever sampling or other task-specific methodology.\n\nOne Sentence SummaryWe illustrate the advantages of Scale Invariant Geometric Data Analysis (SIGDA), a new exploratory data analysis method similar to PCA, by applying SIGDA to derive detailed, robust visualizations of the complex history of human population structure from a large sample of single nucleotide variants.

bioinformatics

Virus-inclusive single cell RNA sequencing reveals molecular signature predictive of progression to severe dengue infection

Dengue virus (DENV) infection can result in severe complications. Yet, the understanding of the molecular correlates of severity is limited, partly due to difficulties in defining the peripheral blood mononuclear cells (PBMCs) that are associated with DENV in vivo. Additionally, there are currently no biomarkers predictive of progression to severe dengue (SD). Bulk transcriptomics data are difficult to interpret because blood consists of multiple cell types that may react differently to infection. Here we applied virus-inclusive single cell RNA-seq approach (viscRNA-Seq) to profile transcriptomes of thousands of single PBMCs derived early in the course of disease from six dengue patients and four healthy controls, and to characterize distinct DENV-associated leukocytes. Multiple genes, particularly interferon response genes, were upregulated in a cell-specific manner prior to progression to SD. Expression of MX2 in naive B cells and CD163 in CD14+ CD16+ monocytes was predictive of SD. The majority of DENV-associated cells in the blood of two patients who progressed to SD were naive IgM B cells expressing the CD69 and CXCR4 receptors and antiviral genes, followed by monocytes. Bystander uninfected B cells also demonstrated immune activation, and plasmablasts from two patients exhibited antibody lineages with convergently hypermutated heavy chain sequences. Lastly, assembly of the DENV genome revealed diversity at unexpected genomic sites. This study presents a multi-faceted molecular elucidation of natural dengue infection in humans and proposes biomarkers for prediction of SD, with implications for profiling any tissue and viral infection, and for the development of a dengue prognostic assay.\n\nSignificanceA fraction of the 400 million people infected with dengue annually progresses to severe dengue (SD). Yet, there are currently no biomarkers to effectively predict disease progression. We profiled the landscape of host transcripts and viral RNA in thousands of single blood cells from dengue patients prior to progressing to SD. We discovered cell-type specific immune activation and candidate predictive biomarkers. We also revealed preferential virus association with specific cell populations, particularly naive B cells and monocytes. We then explored immune activation of bystander cells, clonality and somatic evolution of adaptive immune repertoires, and viral genomics. This multi-faceted approach could advance understanding of pathogenesis of any viral infection, map an atlas of infected cells and promote the development of prognostics.

immunology

Fast and simple comparison of semi-structured data, with emphasis on electronic health records

We present a locality-sensitive hashing strategy for summarizing semi-structured data (e.g., in JSON or XML formats) into data fingerprints: highly compressed representations which cannot recreate details in the data, yet simplify and greatly accelerate the comparison and clustering of semi-structured data by preserving similarity relationships. Computation on data fingerprints is fast: in one example involving complex simulated medical records, the average time to encode one record was 0.53 seconds, and the average pairwise comparison time was 3.75 microseconds. Both processes are trivially parallelizable.\n\nApplications include detection of duplicates, clustering and classification of semi-structured data, which support larger goals including summarizing large and complex data sets, quality assessment, and data mining. We illustrate use cases with three analyses of electronic health records (EHRs): (1) pairwise comparison of patient records, (2) analysis of cohort structure, and (3) evaluation of methods for generating simulated patient data.

bioinformatics

BDQC: a general-purpose analytics tool for domain-blind validation of Big Data

Translational biomedical research is generating exponentially more data: thousands of whole-genome sequences (WGS) are now available; brain data are doubling every two years. Analyses of Big Data, including imaging, genomic, phenotypic, and clinical data, present qualitatively new challenges as well as opportunities. Among the challenges is a proliferation in ways analyses can fail, due largely to the increasing length and complexity of processing pipelines. Anomalies in input data, runtime resource exhaustion or node failure in a distributed computation can all cause pipeline hiccups that are not necessarily obvious in the output. Flaws that can taint results may persist undetected in complex pipelines, a danger amplified by the fact that research is often concurrent with the development of the software on which it depends. On the positive side, the huge sample sizes increase statistical power, which in turn can shed new insight and motivate innovative analytic approaches. We have developed a framework for Big Data Quality Control (BDQC) including an extensible set of heuristic and statistical analyses that identify deviations in data without regard to its meaning (domain-blind analyses). BDQC takes advantage of large sample sizes to classify the samples, estimate distributions and identify outliers. Such outliers may be symptoms of technology failure (e.g., truncated output of one step of a pipeline for a single genome) or may reveal unsuspected \" signal\" in the data (e.g., evidence of aneuploidy in a genome). We have applied the framework to validate real-world WGS analysis pipelines. BDQC successfully identified data outliers representing various failure classes, including genome analyses missing a whole chromosome or part thereof, hidden among thousands of intermediary output files. These failures could then be resolved by reanalyzing the affected samples. BDQC both identified hidden flaws as well as yielded new insights into the data. BDQC is designed to complement quality software development practices. There are multiple benefits from the application of BDQC at all pipeline stages. By verifying input correctness, it can help avoid expensive computations on flawed data. Analysis of intermediary and final results facilitates recovery from aberrant termination of processes. All these computationally inexpensive verifications reduce cryptic analytical artifacts that could seriously preclude clinical-grade genome interpretation. BDQC is available at https://github.com/ini-bdds/bdqc.

bioinformatics

Genotype fingerprints enable fast and private comparison of genetic testing results for research and direct-to-consumer applications

As genetic testing expands out of the research laboratory into medical practice as well as the direct-to-consumer market, the efficiency with which the resulting genotype data can be compared between individuals is of increasing importance.\n\nWe present a method for summarizing personal genotypes, yielding genotype fingerprints that can be derived from any single nucleotide polymorphism (SNP)-based assay and readily compared to estimate relatedness. The resulting fingerprints remain comparable as chip designs evolve to higher marker densities. We demonstrate that they support applications including distinguishing genotypes of closely related individuals by relationship type, distinguishing closely related individuals from individuals from the same background population, identification of individuals in known background populations, and de novo identification of subpopulations within a large cohort in a high-throughput manner.\n\nAn important feature of genotype fingerprints is that, while fingerprints do not preserve anonymity, they summarize individual marker data in a way that prevents phenotype prediction. Genotype fingerprints are therefore well-suited to public sharing for ancestry determination purposes, without revealing personal health risk status.

bioinformatics

Novel metrics for quantifying bacterial genome composition skews

BackgroundBacterial genomes have characteristic compositional skews, which are differences in nucleotide frequency between the leading and lagging DNA strands across a segment of a genome. It is thought that these strand asymmetries arise as a result of mutational biases and selective constraints, particularly for energy efficiency. Analysis of compositional skews in a diverse set of bacteria provides a comparative context in which mutational and selective environmental constraints can be studied. These analyses typically require finished and well-annotated genomic sequences.\n\nResultsWe present three novel metrics for examining genome composition skews; all three metrics can be computed for unfinished or partially-annotated genomes. The first two metrics, (dot-skew and cross-skew) depend on sequence and gene annotation of a single genome, while the third metric (residual skew) highlights unusual genomes by subtracting a GC content-based model of a library of genome sequences. We applied these metrics to all 7738 available bacterial genomes, including partial drafts, and identified outlier species. A number of these outliers (i.e., Borrelia, Ehrlichia, Kinetoplastibacterium, and Phytoplasma) display similar skew patterns despite only distant phylogenetic relationship. While unrelated, some of the outlier bacterial species share lifestyle characteristics, in particular intracellularity and biosynthetic dependence on their hosts.\n\nConclusionsOur novel metrics appear to reflect the effects of biosynthetic constraints and adaptations to life within one or more hosts on genome composition. We provide results for each analyzed genome, software and interactive visualizations at http://db.systemsbiology.net/gestalt/skew_metrics.

genomics

Causal associations between risk factors and common diseases inferred from GWAS summary data

Health risk factors such as body mass index (BMI), serum cholesterol and blood pressure are associated with many common diseases. It often remains unclear whether the risk factors are cause or consequence of disease, or whether the associations are the result of confounding. Genetic methods are useful to infer causality because genetic variants are present from birth and therefore unlikely to be confounded with environmental factors. We develop and apply a method (GSMR) that performs a multi-SNP Mendelian Randomization analysis using summary-level data from large genome-wide association studies (sample sizes of up to 405,072) to test the causal associations of BMI, waist-to-hip ratio, serum cholesterols, blood pressures, height and years of schooling (EduYears) with a range of common diseases. We identify a number of causal associations including a protective effect of LDL-cholesterol against type-2 diabetes (T2D) that might explain the side effects of statins on T2D, a protective effect of EduYears against Alzheimers disease, and bidirectional associations with opposite effects (e.g. higher BMI increases the risk of T2D but the effect T2D of BMI is negative). HDL-cholesterol has a significant risk effect on age-related macular degeneration, and the effect size remains significant accounting for the other risk factors. Our study develops powerful tools to integrate summary data from large studies to infer causality, and provides important candidates to be prioritized for further studies in medical research and for drug discovery.

genetics

Ultrafast comparison of personal genomes

We present an ultra-fast method for comparing personal genomes. We transform the standard genome representation (lists of variants relative to a reference) into genome fingerprints that can be readily compared across sequencing technologies and reference versions. Because of their reduced size, computation on the genome fingerprints is fast and requires little memory. This enables scaling up a variety of important genome analyses, including quantifying relatedness, recognizing duplicative sequenced genomes in a set, population reconstruction, and many others. The original genome representation cannot be reconstructed from its fingerprint; the method thus has significant implications for privacy-preserving genome analytics.

bioinformatics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics

Highly efficient DNA-free gene disruption in the agricultural pest Ceratitis capitata by CRISPR-Cas9 RNPs

The Mediterranean fruitfly Ceratitis capitata (medfly) is an invasive agricultural pest of high economical impact and has become an emerging model for developing new genetic control strategies as alternative to insecticides. Here, we report the successful adaptation of CRISPR-Cas9-based gene disruption in the medfly by injecting in vitro pre-assembled, solubilized Cas9 ribonucleoprotein complexes (RNPs) loaded with gene-specific sgRNAs into early embryos. When targeting the eye pigmentation gene white eye (we), we observed a high rate of somatic mosaicism in surviving G0 adults. Germline transmission of mutated we alleles by G0 animals was on average above 70%, with individual cases achieving a transmission rate of nearly 100%. We further recovered large deletions in the we gene when two sites were simultaneously targeted by two sgRNAs. CRISPR-Cas9 targeting of the Ceratitis ortholog of the Drosophila segmentation paired gene (Ccprd) caused segmental malformations in late embryos and in hatched larvae. Mutant phenotypes correlate with repair by non-homologous end joining (NHEJ) lesions in the two targeted genes. This simple and highly effective Cas9 RNP-based gene editing to introduce mutations in Ceratitis capitata will significantly advance the design and development of new effective strategies for pest control management.

genetics