Search bioRxivSearch

Biology subjects

Rao, J.

Publications and source records attributed to Rao, J..

3 recordsLinked to original sources

An integrated Asian human SNV and indel benchmark combining multiple sequencing methods

Precision medicine of human requires an accurate and complete reference variant benchmark for different populations. A human standard cell line of NA12878 provides a good reference for part of the human populations, but it is still lack of a fine reference standard sample and variant benchmark for the Asians. Here, we constructed a stabilized cell line of a Chinese Han volunteer. We received about 4.16T clean data of the sample using eight sequencing strategies in different laboratories, including two BGI regular NGS platforms, three Illumina regular NGS platforms, two linked-read libraries, and PacBio CCS model. The sequencing depth and reference coverage of eight sequencing strategies have reached the saturation. We detected small variants of SNPs and Indels using the eight data sets and obtained eight variant sets by performing a series of strictly quality control. Finally, we got 3.35M SNPs and 349K indels supported by all of sequencing data, which could be considered as a high confidence standard small variant sets for the studies. Besides, we also detected 5,913 high quality SNPs located in the high homologous regions supported by both linked-reads and CCS data benefited by their long-range information, while these regions are recalcitrant to regular NGS data due to the limited mappability and read length. We compared the later SNPs against the public databases and 969 sites of them were novel SNPs, indicating these SNPs provide a vital complement for the variant database. Moreover, we also phased more than 99% heterozygous SNPs also supported by linked-reads and CCS data. This work provided an integrated Asians SNV and indel benchmark for the further basic studies and precision medicine.

genomics

Accurate Prediction of Genome-wide RNA Secondary Structure Profile Based On Extreme Gradient Boosting

MotivationMany studies have shown that RNA secondary structure plays a vital role in fundamental cellular processes, such as protein synthesis, mRNA processing, mRNA assembly, ribosome function and eukaryotic spliceosomes. Identification of RNA secondary structure is a key step to understand the common mechanisms underlying the translation process. Recently, a few experimental methods were developed to measure genome-wide RNA secondary structure profile through high-throughput sequencing techniques, and have been successfully applied to genomes including yeast and human. However, these high-throughput methods usually have low precision and are hard to cover all nucleotides on the RNA due to limited sequencing coverage.\n\nResultsIn this study, we developed a new method for the prediction of genome-wide RNA secondary structure profile (TH-GRASP) from RNA sequence based on eXtreme Gradient Boosting (XGBoost). The method achieves an prediction with areas under the receiver operating characteristic curve (AUC) values greater than 0.9 on three different datasets, and AUC of 0.892 by an independent test on the recently released Zika virus RNA dataset. These AUCs represent a consistent increase of >6% than the recently developed method CROSS trained by a shallow neural network. A further analysis on the 1000-Genome Project data showed that our predicted unpaired probability at mutations sites are highly correlated with the minor allele frequencies (MAF) of synonymous, non-synonymous mutations, and mutations in 3 and 5UTR with Pearson Correlation Coefficients all above 0.8. These PCCs are consistently higher than those generated by RNAplfold method. Moreover, an investigation over all human mRNA indicated a periodic distribution of the predicted unpaired probability on codons, and a decrease of paired probability in the boundary with 5 and 3 untranslated regions. These results highlighted TH-GRASP is effective to remove experimental noises and to have ability to make predictions on nucleotides with low or no coverage by fitting high-throughput genomic data for RNA secondary structure profiles, and also suggested that building model on high throughput experimental data might be a future direction to substitute analytical methods.\n\nAvailabilityThe TH-GRASP is available for academic use at https://github.com/sysu-yanglab/TH-GRASP.\n\nSupplementary informationSupplementary data are available online.

bioinformatics

Leptin deficient rats develop nonalcoholic steatohepatitis with unique disease progression

Nonalcoholic steatohepatitis (NASH) is an aggressive liver disease threatening public health, however its natural history is poorly understood. Unlike ob/ob mice, Lep{triangleup}I14/{triangleup}I14 rats develop unique NASH phenotype with steatosis, lymphocyte infiltration and ballooning after postnatal week 16. Using Lep{triangleup}I14/{triangleup}I14 rats as NASH model, we studied the natural history of NASH progression by performing an integrated analysis of hepatic transcriptome from postnatal week 4 to 48. Leptin deficiency results in a robust increase in expression of genes encoding 9 rate-limiting enzymes in lipid metabolism such as ACC and FASN. However, genes in positive regulation of inflammatory response are highly expressed at week 16 and then remain the steady elevated expression till week 48. The high expression of cytokines and chemokines including CCL2, TNF, IL6 and IL1{beta} is correlated with the phosphorylation of several key molecules in pathways such as JNK and NF-{kappa}B. Meanwhile, we observed cell infiltration of MPO+ neutrophils, CD8+ T cells, CD68+ hepatic macrophages and CCR2+ inflammatory monocyte-derived macrophages, together with macrophage polarization from M2 to M1. Importantly, Lep{triangleup}I14/{triangleup}I14 rats share more homologous genes with NASH patients than previously established mouse models and crab eating monkeys with spontaneous hepatic steatosis. Transcriptomic analysis showed that many drug targets in clinical trials can be evaluated in Lep{triangleup}I14/{triangleup}I14 rats.\n\nConclusionWe characterize Lep{triangleup}I14/{triangleup}I14 rats as a unique NASH model by performing a long-term (i.e., 4 to 48 postnatal weeks) integrated transcriptomic analysis. This work reveal the temporal dynamics of hepatic gene expression in lipid metabolism and inflammation, and shed light on understanding the natural history of NASH in human beings.

pathology