Search bioRxivSearch

Biology subjects

Brett K Beaulieu-Jones

Publications and source records attributed to Brett K Beaulieu-Jones.

2 recordsLinked to original sources

Reproducible Computational Workflows with Continuous Analysis

Reproducing experiments is vital to science. Being able to replicate, validate and extend previous work also speeds new research projects. Reproducing computational biology experiments, which are scripted, should be straightforward. But reproducing such work remains challenging and time consuming. In the ideal world we would be able to quickly and easily rewind to the precise computing environment where results were generated. We would then be able to reproduce the original analysis or perform new analyses. We introduce a process termed \"continuous analysis\" which provides inherent reproducibility to computational research at a minimal cost to the researcher. Continuous analysis combines Docker, a container service similar to virtual machines, with continuous integration, a popular software development technique, to automatically re-run computational analysis whenever relevant changes are made to the source code. This allows results to be reproduced quickly, accurately and without needing to contact the original authors. Continuous analysis also provides an audit trail for analyses that use data with sharing restrictions. This allows reviewers, editors, and readers to verify reproducibility without manually downloading and rerunning any code. Example configurations are available at our online repository (https://github.com/greenelab/continuous_analysis).

Bioinformatics

Semi-Supervised Learning of the Electronic Health Record for Phenotype Stratification

Patient interactions with health care providers result in entries to electronic health records (EHRs). EHRs were built for clinical and billing purposes but contain many data points about an individual. Mining these records provides opportunities to extract electronic phenotypes, which can be paired with genetic data to identify genes underlying common human diseases. This task remains challenging: high quality phenotyping is costly and requires physician review; many fields in the records are sparsely filled; and our definitions of diseases are continuing to improve over time. Here we develop and evaluate a semi-supervised learning method for EHR phenotype extraction using denoising autoencoders for phenotype stratification. By combining denoising autoencoders with random forests we find classification improvements across multiple simulation models and improved survival prediction in ALS clinical trial data. This is particularly evident in cases where only a small number of patients have high quality phenotypes, a common scenario in EHR-based research. Denoising autoencoders perform dimensionality reduction enabling visualization and clustering for the discovery of new subtypes of disease. This method represents a promising approach to clarify disease subtypes and improve genotype-phenotype association studies that leverage EHRs.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC=\"FIGDIR/small/039800_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (31K):\norg.highwire.dtl.DTLVardef@567074org.highwire.dtl.DTLVardef@f0e45corg.highwire.dtl.DTLVardef@1207659org.highwire.dtl.DTLVardef@39f8b8_HPS_FORMAT_FIGEXP M_FIG C_FIG HIGHLIGHTSO_LIDenoising autoencoders (DAs) can model electronic health records.\nC_LIO_LISemi-supervised learning with DAs improves ALS patient survival predictions.\nC_LIO_LIDAs improve patient cluster visualization through dimensionality reduction.\nC_LI

Bioinformatics