Search bioRxivSearch

Biology subjects

Xuan, Y.

Publications and source records attributed to Xuan, Y..

3 recordsLinked to original sources

Full-length genome sequence of segmented RNA virus from ticks was obtained using small RNA sequencing data

In 2014, A novel tick-borne virus of the genus Flavivirus was first reported from the Mogiana region in Brazil. This virus was named the Mogiana tick virus (MGTV). Later, MGTV was also named as Jingmen tick virus (JMTV), Kindia tick virus (KDTV), Guangxi tick virus (GXTV) etc. In the present study, we used small RNA sequencing (sRNA-seq) to detect viruses in ticks and detected MGTV in Amblyomma testudinarium ticks, which had been captured in Yunnan province of China in the year of 2016. The full-length genome sequence of a new MGTV strain Yunnan2016 (GenBank: MT080097, MT080098, MT080099 and MT080100) was obtained and recommended to be included into the NCBI RefSeq database for the future studies of MGTV. Our phylogenetic analyses showed that viruses named MGTV, JMTV, KDTV and GXTV are monophyletic: the MGTV group (lineage) of viruses. We show, for the first time, that 5' and 3' sRNAs can be used to obtain full-length sequences of the 5 and 3 ends of, but not limited to genome sequences of RNA viruses. And we proved the feasibility of using the sRNA-seq based method for the detection of viruses in a sample containing miniscule RNA.

bioinformatics

Standardization and Harmonization of Distributed Multi-National Proteotype Analysis supporting Precision Medicine Studies

Cancer has no borders: Generation and analysis of molecular data across multiple centers worldwide is necessary to gain statistically significant clinical insights for the benefit of patients. Here we conceived and standardized a proteotype data generation and analysis workflow enabling distributed data generation and evaluated the quantitative data generated across laboratories of the international Cancer Moonshot consortium. Using harmonized mass spectrometry (MS) instrument platforms and standardized data acquisition procedures, we demonstrated robust, sensitive, and reproducible data generation across eleven sites in nine countries on seven consecutive days in a 24/7 operation mode. The data presented from the high-resolution MS1-based quantitative data-independent acquisition (HRMS1-DIA) workflow shows that coordinated proteotype data acquisition is feasible from clinical specimens using such standardized strategies. This work paves the way for the distributed multi-omic digitization of large clinical specimen cohorts across multiple sites as a prerequisite for turning molecular precision medicine into reality.

molecular biology

DPHL: A pan-human protein mass spectrometry library for robust biomarker discovery using Data-Independent Acquisition and Parallel Reaction Monitoring

To answer the increasing need for detecting and validating protein biomarkers in clinical specimens, proteomic techniques are required that support the fast, reproducible and quantitative analysis of large clinical sample cohorts. Targeted mass spectrometry techniques, specifically SRM, PRM and the massively parallel SWATH/DIA technique have emerged as a powerful method for biomarker research. For optimal performance, they require prior knowledge about the fragment ion spectra of targeted peptides. In this report, we describe a mass spectrometric (MS) pipeline and spectral resource to support data-independent acquisition (DIA) and parallel reaction monitoring (PRM) based biomarker studies. To build the spectral resource we integrated common open-source MS computational tools to assemble an open source computational workflow based on Docker. It was then applied to generate a comprehensive DIA pan-human library (DPHL) from 1,096 data dependent acquisition (DDA) MS raw files, and it comprises 242,476 unique peptide sequences from 14,782 protein groups and 10,943 SwissProt-annotated proteins expressed in 16 types of cancer samples. In particular, tissue specimens from patients with prostate cancer, cervical cancer, colorectal cancer, hepatocellular carcinoma, gastric cancer, lung adenocarcinoma, squamous cell lung carcinoma, diseased thyroid, glioblastoma multiforme, sarcoma and diffuse large B-cell lymphoma (DLBCL), as well as plasma samples from a range of hematologic malignancies were collected from multiple clinics in China, the Netherlands and Singapore and included in the resource. This extensive spectral resource was then applied to a prostate cancer cohort of 17 patients, consisting of 8 patients with prostate cancer (PCa) and 9 with benign prostate hyperplasia (BPH), respectively. Data analysis of DIA data from these samples identified differential expressions of FASN, TPP1 and SPON2 in prostate tumors. Thereafter, PRM validation was applied to a larger PCa cohort of 57 patients and the differential expressions of FASN, TPP1 and SPON2 in prostate tumors were validated. As a second application, the DPHL spectral resource was applied to a patient cohort consisting of samples from 19 DLBCL patients and 18 healthy individuals. Differential expressions of CRP, CD44 and SAA1 between DLBCL cases and healthy controls were detected by DIA-MS and confirmed by PRM. These data demonstrate that the DPHL supported that DIA-PRM MS pipeline enables robust protein biomarker discoveries.

bioinformatics