Search bioRxivSearch

Biology subjects

Leopold Parts

Publications and source records attributed to Leopold Parts.

5 recordsLinked to original sources

Accurate classification of protein subcellular localization from high throughput microscopy images using deep learning

High throughput microscopy of many single cells generates high-dimensional data that are far from straightforward to analyze. One important problem is automatically detecting the cellular compartment where a fluorescently tagged protein resides, a task relatively simple for an experienced human, but difficult to automate on a computer. Here, we train an 11-layer neural network on data from mapping thousands of yeast proteins, achieving per cell localization classification accuracy of 91%, and per protein accuracy of 99% on held out images. We confirm that low-level network features correspond to basic image characteristics, while deeper layers separate localization classes. Using this network as a feature calculator, we train standard classifiers that assign proteins to previously unseen compartments after observing only a small number of training examples. Our results are the most accurate subcellular localization classifications to date, and demonstrate the usefulness of deep learning for high throughput microscopy.

Bioinformatics

Powerful decomposition of complex traits in a diploid model using Phased Outbred Lines

Explaining trait differences between individuals is a core but challenging aim of life sciences. Here, we introduce a powerful framework for complete decomposition of trait variation into its underlying genetic causes in diploid model organisms. We intercross two natural genomes over many sexual generations, sequence and systematically pair the recombinant gametes into a large array of diploid hybrids with fully assembled and phased genomes, termed Phased Outbred Lines (POLs). We demonstrate the capacity of the framework by partitioning fitness traits of 7310 yeast POLs across many environments, achieving near complete trait heritability (mean H2 = 91%) and precisely estimating additive (74%), dominance (8%), second (9%) and third (1.8%) order epistasis components. We found nonadditive quantitative trait loci (QTLs) to outnumber (3:1) but to be weaker than additive loci; dominant contributions to heterosis to outnumber overdominant (3:1); and pleiotropy to be the rule rather than the exception. The POL approach presented here offers the most complete decomposition of diploid traits to date and can be adapted to most model organisms.

Genetics

Predicting quantitative traits from genome and phenome with near perfect accuracy

In spite of decades of linkage and association studies and its potential impact on human health1, reliable prediction of an individual's risk for heritable disease remains difficult2-4. Large numbers of mapped loci do not explain substantial fractions of the heritable variation, leaving an open question of whether accurate complex trait predictions can be achieved in practice5,6. Here, we use a full genome sequenced population of 7396 yeast strains of varying relatedness, and predict growth traits from family information, effects of segregating genetic variants, and growth measurements in other environments with an average coefficient of determination R2 of 0.91. This accuracy exceeds narrow-sense heritability, approaches limits imposed by measurement repeatability, and is higher than achieved with a single replicate assay in the lab. We find that both relatedness and variant-based predictions are greatly aided by availability of closer relatives, while information from a large number of more distant relatives does not improve predictive performance when close relatives can be used. Our results prove that very accurate prediction of heritable traits is possible, and recommend prioritizing collection of deeper family-based data over large reference cohorts.

Genetics

Pathway based factor analysis of gene expression data produces highly heritable phenotypes that associate with age

Statistical factor analysis methods have previously been used to remove noise components from high dimensional data prior to genetic association mapping, and in a guided fashion to summarise biologically relevant sources of variation. Here we show how the derived factors summarising pathway expression can be used to analyse the relationships between expression, heritability and ageing. We used skin gene expression data from 647 twins from the MuTHER Consortium and applied factor analysis to concisely summarise patterns of gene expression, both to remove broad confounding influences and to produce concise pathway-level phenotypes. We derived 930 \"pathway phe-notypes\" which summarised patterns of variation across 186 KEGG pathways (five phenotypes per pathway). We identified 69 significant associations of age with phenotype from 57 distinct KEGG pathways at a stringent Bon-ferroni threshold (P < 5.38 x 10-5). These phenotypes are more heritable (h2 = 0.32) than gene expression levels. On average, expression levels of 16% of genes within these pathways are associated with age. Several significant pathways relate to metabolising sugars and fatty acids, others with insulin signalling. We have demonstrated that factor analysis methods combined with biological knowledge can produce more reliable phenotypes with less stochastic noise than the individual gene expression levels, which increases our power to discover biologically relevant associations. These phenotypes could also be applied to discover associations with other environmental factors.

Genomics

Molecular phenotypes that are causal to complex traits can have low heritability and are expected to have small influence.

Work on genetic makeup of complex traits has led to some unexpected findings. Molecular trait heritability estimates have consistently been lower than those of common diseases, even though it is intuitively expected that the genotype signal weakens as it becomes more dissociated from DNA. Further, results from very large studies have not been sufficient to explain most of the heritable signal, and suggest hundreds if not thousands of responsible alleles. Here, I demonstrate how trait heritability depends crucially on the definition of the phenotype, and is influenced by the variability of the assay, measurement strategy, and the quantification approach used. For a phenotype downstream of many molecular traits, it is possible that its heritability is larger than for any of its upstream determinants. I also rearticulate via models and data that if a phenotype has many dependencies, a large number of small effect alleles are expected. However, even if these alleles do drive highly heritable causal intermediates that can be modulated, it does not imply that large changes in phenotype can be obtained.

Genetics