Search bioRxivSearch

Biology subjects

Leek, J.

Publications and source records attributed to Leek, J..

3 recordsLinked to original sources

crsra: A package for Cleaning and Analyzing Coursera Research Export Data

Due to the fundamental differences between traditional education and Massive Open Online Courses (MOOCs) and the ever-increasing popularity of MOOCs more research is needed to under- stand current and future trends in education. Although research in the field has rapidly grown in recent years, one of the main challenges facing researchers remains to be the complexity and messiness of the data. Therefore, it is imperative to provide tools that pave the way for more research on the new subject of MOOCs. This paper introduces a package called crsra based on the statistical software R to help clean and analyze massive loads of data provided by Coursera. The advantages of the package are as follows: a) faster loading and organizing data for analysis, b) an efficient method for combining data from multiple courses and even across institutions, and c) provision of a set of functions for analyzing student behaviors.

scientific communication and education

Improving the value of public RNA-seq expression data by phenotype prediction

BackgroundPublicly available genomic data are a valuable resource for studying normal human variation and disease, but these data are often not well labeled or annotated. The lack of phenotype information for public genomic data severely limits their utility for addressing targeted biological questions.\n\nResultsWe develop an in silico phenotyping approach for predicting critical missing annotation directly from genomic measurements using, well-annotated genomic and phenotypic data produced by consortia like TCGA and GTEx as training data. We apply in silico phenotyping to a set of 70,000 RNA-seq samples we recently processed on a common pipeline as part of the recount2 project (https://jhubiostatistics.shinyapps.io/recount/). We use gene expression data to build and evaluate predictors for both biological phenotypes (sex, tissue, sample source) and experimental conditions (sequencing strategy). We demonstrate how these predictions can be used to study cross-sample properties of public genomic data, select genomic projects with specific characteristics, and perform downstream analyses using predicted phenotypes. The methods to perform phenotype prediction are available in the phenopredict R package (https://github.com/leekgroup/phenopredict) and the predictions for recount2 are available from the recount R package (https://bioconductor.org/packages/release/bioc/html/recount.html)\n\nConclusionHaving leveraging massive public data sets to generate a well-phenotyped set of expression data for more than 70,000 human samples, expression data is available for use on a scale that was not previously feasible.

bioinformatics

A statistical definition for reproducibility and replicability

Everyone agrees that reproducibility and replicability are fundamental characteristics of scientific studies. These topics are attracting increasing attention, scrutiny, and debate both in the popular press and the scientific literature. But there are no formal statistical definitions for these concepts, which leads to confusion since the same words are used for different concepts by different people in different fields. We provide formal and informal definitions of scientific studies, reproducibility, and replicability that can be used to clarify discussions around these concepts in the scientific and popular press.

Scientific Communication and Education