Search bioRxivSearch

Biology subjects

Kai Wang

Publications and source records attributed to Kai Wang.

4 recordsLinked to original sources

Assessing the measurement transfer function of single-cell RNA sequencing

Recently, measurement of RNA at single cell resolution has yielded surprising insights. Methods for single-cell RNA sequencing (scRNA-seq) have received considerable attention, but the broad reliability of single cell methods and the factors governing their performance are still poorly known. Here, we conducted a large-scale control experiment to assess the transfer function of three scRNA-seq methods and factors modulating the function. All three methods detected greater than 70% of the expected number of genes and had a 50% probability of detecting genes with abundance greater than 2 to 4 molecules. Despite the small number of molecules, sequencing depth significantly affected gene detection. While biases in detection and quantification were qualitatively similar across methods, the degree of bias differed, consistent with differences in molecular protocol. Measurement reliability increased with expression level for all methods and we conservatively estimate the measurement transfer functions to be linear above ~5-10 molecules. Based on these extensive control studies, we propose that RNA-seq of single cells has come of age, yielding quantitative biological information.

Genomics

iCAGES: integrated CAncer GEnome Score for comprehensively prioritizing cancer driver genes in personal genomes

All cancers arise as a result of the acquisition of somatic mutations that drive the disease progression. A number of computational tools have been developed to identify driver genes for a specific cancer from a group of cancer samples. However, it remains a challenge to identify driver mutations/genes for an individual patient and design drug therapies. We developed iCAGES, a novel statistical framework to rapidly analyze patient-specific cancer genomic data, prioritize personalized cancer driver events and predict personalized therapies. iCAGES includes three consecutive layers: the first layer integrates contributions from coding, non-coding and structural variations to infer driver variants. For coding mutations, we developed a radial support vector machine using manually curated mutations to predict their driver potential. The second layer identifies driver genes, by using information from the first layer and integrating prior biological knowledge on gene-gene and gene-phenotype networks. The third layer prioritizes personalized drug treatment, by classifying potential driver genes into different categories and querying drug-gene databases. Compared to currently available tools, iCAGES achieves better performance by correctly classifying point coding driver mutations (AUC=0.97, 95% CI: 0.97-0.97, significantly better than the second best tool with P=0.01) and genes (AUC=0.93, 95% CI: 0.93-0.94, significantly better than MutSigCV with P<1x10-15). We also illustrated two examples where iCAGES correctly nominated two targeted drugs for two advanced cancer patients with exceptional response, based on their somatic mutation profiles. iCAGES leverages personal genomic information and prior biological knowledge, effectively identifies cancer driver genes and predicts treatment strategies. iCAGES is available at http://icages.usc.edu.

Genomics

A variant in TAF1 is associated with a new syndrome with severe intellectual disability and characteristic dysmorphic features

We describe the discovery of a new genetic syndrome, RykDax syndrome, driven by a whole genome sequencing (WGS) study of one family from Utah with two affected male brothers, presenting with severe intellectual disability (ID), a characteristic intergluteal crease, and very distinctive facial features including a broad, upturned nose, sagging cheeks, downward sloping palpebral fissures, prominent periorbital ridges, deep-set eyes, relative hypertelorism, thin upper lip, a high-arched palate, prominent ears with thickened helices, and a pointed chin. This Caucasian family was recruited from Utah, USA. Illumina-based WGS was performed on 10 members of this family, with additional Complete Genomics-based WGS performed on the nuclear portion of the family (mother, father and the two affected males). Using WGS datasets from 10 members of this family, we can increase the reliability of the biological inferences with an integrative bioinformatic pipeline. In combination with insights from clinical evaluations and medical diagnostic analyses, these DNA sequencing data were used in the study of three plausible genetic disease models that might uncover genetic contribution to the syndrome. We found a 2 to 5-fold difference in the number of variants detected as being relevant for various disease models when using different sets of sequencing data and analysis pipelines. We de-rived greater accuracy when more pipelines were used in conjunction with data encompassing a larger portion of the family, with the number of putative de-novo mutations being reduced by 80%, due to false negative calls in the parents. The boys carry a maternally inherited mis-sense variant in a X-chromosomal gene TAF1, which we consider as disease relevant. TAF1 is the largest subunit of the general transcription factor IID (TFIID) multi-protein complex, and our results implicate mutations in TAF1 as playing a critical role in the development of this new intellectual disability syndrome.

Genetics

Ligation-anchored PCR unveils immune repertoire of TCR-beta from whole blood

BackgroundAs one of the genetic mechanisms for adaptive immunity, V(D)J recombination generates an enormous repertoire of T-cell receptors (TCRs). With the development of high-throughput sequencing techniques, systematic exploration of V(D)J recombination becomes possible. Multiplex PCR method has been previously developed to assay immune repertoire, however the usage of primer pools has inherent bias in target amplification. In our study, we developed a ligation-anchored PCR method to unbiasedly amplify the repertoire.\n\nResultsBy utilizing a universal primer paired with a single primer targeting the conserved constant region, we amplified TCR-beta (TRB) variable regions from total RNA extracted from blood. Next-generation sequencing libraries were then prepared for Illumina HiSeq 2500 sequencer, which provided 151 bp read length to cover the entire V(D)J recombination region. We evaluated this approach on blood samples from patients with malignant and benign meningiomas. Mapping of sequencing data showed 64% to 91% of mapped TCRV-containing reads belong to TRB subtype. An increased usage of TRBV29-1 was observed in malignant meningiomas. Also distinct signatures were identified from CDR3 sequence logos, with predominant subset as 42 nt for benign and 45 nt for malignant samples, respectively.\n\nConclusionsIn summary, we report an integrative approach to monitor immune repertoire in a systematic manner.

Genomics