Search bioRxivSearch

Biology subjects

Tan, G.

Publications and source records attributed to Tan, G..

3 recordsLinked to original sources

Long fragments achieve lower base quality in Illumina paired-end sequencing

Illuminas technology provides high quality reads of DNA fragments with error rates below 1/1000 per base. Runs typically generate a millions of reads where the vast majority of the reads has also an average error rate below 1/1000. However, some paired-end sequencing data show the presence of a subpopulation of reads where the second read has lower average qualities. We show that the fragment length is a major driver of increased error rates in the R2 reads. Fragments above 500 nt tend to yield lower base qualities and higher error rates than shorter fragments. We demonstrate the fragment length dependency of the R2 read qualities using publicly available Illumina data generated by various library protocols, in different labs and using different sequencer models. Our finding extends the understanding of the Illumina read quality and has implications on error models for Illumina reads. It also sheds a light on the importance of the fragmentation during library preparation and the resulting fragment length distribution.

bioinformatics

Expanding the Atlas of Functional Missense Variation for Human Genes

Although we now routinely sequence human genomes, we can confidently identify only a fraction of the sequence variants that have a functional impact. Here we developed a deep mutational scanning framework that produces exhaustive maps for human missense variants by combining random codon-mutagenesis and multiplexed functional variation assays with computational imputation and refinement. We applied this framework to four proteins corresponding to six human genes: UBE2I (encoding SUMO E2 conjugase), SUMO1 (small ubiquitin-like modifier), TPK1 (thiamin pyrophosphokinase), and CALM1/2/3 (three genes encoding the protein calmodulin). The resulting maps recapitulate known protein features, and confidently identify pathogenic variation. Assays potentially amenable to deep mutational scanning are already available for 57% of human disease genes, suggesting that DMS could ultimately map functional variation for all human disease genes.

molecular biology

Tracking Subclonal Mutation Frequencies Throughout Lymphomagenesis Identifies Cancer Drivers in Mouse Models of Lymphoma.

Determining whether recurrent but rare cancer mutations are bona fide driver mutations remains a bottleneck in cancer research. Here we present the most comprehensive analysis of retrovirus driven lymphomagenesis produced to date, sequencing 700,000 mutations from >500 malignancies collected at time points throughout tumor development. This enabled identification of positively selected events, and the first demonstration of negative selection of mutations that may be deleterious to tumor development indicating novel avenues for therapy. Customized sequencing and bioinformatics methodologies were developed to quantify subclonal mutations in both premalignant and malignant tissue, greatly expanding the statistical power for identifying driver mutations and yielding a high-resolution, genome wide map of the selective forces surrounding cancer gene loci. Screening two BCL2 transgenic models confirms known drivers of human B-cell non-Hodgkin lymphoma, and implicates novel candidates including modifiers of immunosurveillance such as co-stimulatory molecules and MHC loci. Correlating mutations with genotypic and phenotypic features also gives robust identification of known cancer genes independently of local variance in mutation density. An online resource http://mulv.lms.mrc.ac.uk allows customized queries of the entire dataset.

cancer biology