Search bioRxivSearch

Biology subjects

Graffelman, J.

Publications and source records attributed to Graffelman, J..

3 recordsLinked to original sources

Multidimensional Scaling and Relatedness Research

Multidimensional scaling is a well-known multivariate technique, that is often used in genetics for studying population substructure. In this paper we show that multidimensional scaling of marker data is of relevance for relatedness research. Relatedness is usually investigated by estimating and plotting identity-by-state and identity-by-descent allele-sharing statistics. We show that outlying individuals in a map obtained by multidimensional scaling of genetic variables do not necessarily stem from a different human population, but can be the consequence of relatedness. We propose a method for classifying pairs of individuals into the standard relationship categories that combines genetic bootstrapping, multidimensional scaling and discriminant analysis. We validate our method with simulation studies. Given the variant filtering procedures, our method classifies relationships up to and including the fourth degree with high accuracy (96-97%), using only identity by state. The usefulness of the method is illustrated with data from the 1,000 genomes and the GCAT projects.

genetics

Multi-allelic Exact tests for Hardy-Weinberg equilibrium that account for gender

Statistical tests for Hardy-Weinberg equilibrium are important elementary tools in genetic data analysis. X-chromosomal variants have long been tested by applying autosomal test procedures to females only, and gender is usually not considered when testing autosomal variants for equilibrium. Recently, we proposed specific X-chromosomal exact test procedures for bi-allelic variants that include the hemizygous males, as well as autosomal tests that consider gender. In this paper we present the extension of the previous work for variants with multiple alleles. A full enumeration algorithm is used for the exact calculations of tri-allelic variants. For variants with many alternate alleles we use a permutation test. Some empirical examples with data from the 1000 genomes project are discussed.

genetics

Compositional Canonical Correlation Analysis

The study of the relationships between two compositions by means of canonical correlation analysis is addressed A coimnositional version of canonical correlation analysis is developed. and called CODA-CCO. We consider two approaches, using the centred log-ratio transformation and the calculation of all possible pairwise log-ratios within sets. The relationships between both approaches are pointed out, and their merits are discussed. The related covariance matrices are structurally singular, and this is efficiently dealt with by using generalized inverses. We develop compositional canonical biplots and detail their properties. The canonical biplots are shown to be powerful tools for discovering the most salient relationships between two compositions. Some guidelines for compositional canonical biplots construction are discussed. A geological data set with X-ray fluorescence spectrometry measurements on major oxides and trace elements is used to illustrate the proposed method. The relationships between an analysis based on centred log-ratios and on isometric log-ratios are also shown.

bioinformatics