Search bioRxivSearch

Biology subjects

Martin, L. S.

Publications and source records attributed to Martin, L. S..

2 recordsLinked to original sources

UMI-Reducer: Collapsing duplicate sequencing reads via Unique Molecular Identifiers

Short Structured AbstractO_ST_ABSSummaryC_ST_ABSEvery sequencing library contains duplicate reads. While many duplicates arise during polymerase chain reaction (PCR), some duplicates derive from multiple identical fragments of mRNA present in the original lysate (termed \"biological duplicates\"). Unique Molecular Identifiers (UMIs) are random oligonucleotide sequences that allow differentiation between technical and biological duplicates. Here we report the development of UMI-Reducer, a new computational tool for processing and differentiating PCR duplicates from biological duplicates. UMI-Reducer uses UMIs and the mapping position of the read to identify and collapse reads that are technical duplicates. Remaining true biological reads are further used for bias-free estimate of mRNA abundance in the original lysate. This strategy is of particular use for libraries made from low amounts of starting material, which typically require additional cycles of PCR and therefore are most prone to PCR duplicate bias.\n\nAvailability and ImplementationThe UMI-Reducer is an open source Python software and is freely available for non-commercial use (GPL-3.0) at https://sergheimangul.wordpress.com/umi-reducer/. Documentation and tutorials are available at https://github.com/smangul1/UMI-Reducer/wiki/.\n\nContactsmangul@ucla.edu, SVanDriesche@mednet.ucla.edu\n\nSupplementary informationFlowchart of Library Construction

bioinformatics

Review: Population Structure in Genetic Studies: Confounding Factors and Mixed Models

A genome-wide association study (GWAS) seeks to identify genetic variants that contribute to the development and progression of a specific disease. Over the past 10 years, new approaches using mixed models have emerged to mitigate the deleterious effects of population structure and relatedness in association studies. However, developing GWAS techniques to effectively test for association while correcting for population structure is a computational and statistical challenge. Using laboratory mouse strains as an example, our review characterizes the problem of population structure in association studies and describes how it can cause false positive associations. We then motivate mixed models in the context of unmodeled factors.

bioinformatics