Search bioRxivSearch

Biology subjects

Lofgren, S.

Publications and source records attributed to Lofgren, S..

3 recordsLinked to original sources

Interpretation of biological experiments changes with evolution of Gene Ontology and its annotations

Gene Ontology (GO) enrichment analysis is ubiquitously used for interpreting high throughput molecular data and generating hypotheses about underlying biological phenomena of experiments. However, the two building blocks of this analysis -- the ontology and the annotations -- evolve rapidly. We used gene signatures derived from 104 disease analyses to systematically evaluate how enrichment analysis results were affected by evolution of the GO over a decade. We found low consistency between enrichment analyses results obtained with early and more recent GO versions. Furthermore, there continues to be strong annotation bias in the GO annotations where 58% of the annotations are for 16% of the human genes. Our analysis suggests that GO evolution may have affected the interpretation and possibly reproducibility of experiments over time. Hence, researchers must exercise caution when interpreting GO enrichment analyses and should reexamine previous analyses with the most recent GO version.

bioinformatics

Integrated molecular and clinical analysis for understanding human disease relationships

Existing knowledge of human disease relationships is incomplete. To establish a comprehensive understanding of disease, we integrated transcriptome profiles of 41,000 human samples with clinical profiles of 2 million patients, across 89 diseases. Based on transcriptome data, autoimmune diseases clustered with their specific infectious triggers, and brain disorders clustered by disease class. Clinical profiles clustered diseases according to the similarity of their initial manifestation and later complications, identifying disease relationships absent in prior co-occurrence analyses. Our integrated analysis of transcriptome and clinical profiles identified overlooked, therapeutically actionable disease relationships, such as between myositis and interstitial cystitis. Our improved understanding of disease relationships will identify disease mechanisms, offer novel therapeutic targets, and create synergistic research opportunities.

bioinformatics

Leveraging heterogeneity across multiple data sets increases accuracy of cell-mixture deconvolution and reduces biological and technical biases

In silico quantification of cell proportions from mixed-cell transcriptomics data (deconvolution) requires a reference expression matrix, called basis matrix. We hypothesized that matrices created using only healthy samples from a single microarray platform would introduce biological and technical biases in deconvolution. We show presence of such biases in two existing matrices, IRIS and LM22, irrespective of the deconvolution method used. Here, we present immunoStates, a basis matrix built using 6160 samples with different disease states across 42 microarray platforms. We found that immunoStates significantly reduced biological and technical biases. We further show that cellular proportion estimates using immunoStates are consistently more correlated with measured proportions than IRIS and LM22, across all methods. Importantly, we found that different methods have virtually no effect once the basis matrix is chosen. Our results demonstrate the need and importance of incorporating biological and technical heterogeneity in a basis matrix for achieving consistently high accuracy.

bioinformatics