Search bioRxivSearch

Biology subjects

Qiu, W.

Publications and source records attributed to Qiu, W..

12 recordsLinked to original sources

Novel Data Transformations for RNA-seq Data Analysis

We propose eight data transformations for RNA-seq data analysis aiming to make the transformed sample mean to be representative of the distribution center since it is not always possible to transform count data to satisfy the normality assumption. Simulation studies showed that limma based on transformed data by using the rv transformation (denoted as limma+rv) performed best compared with limma based on transformed data by using other transformation methods in term of high accuracy and low FNR, while keeping FDR at the nominal level. For large sample size, limma based on transformed data by using the 8 proposed transformation methods had similar performance to limma based on transformed data by using existing transformation methods for equal library size scenarios. Otherwise, limma based on transformed data by using the rv, lv, rv2, or lv2 transformation, or by using the existing voom transformation performed better than limma based on data from other transformation methods. Real data analysis results showed that limma+ l2 performed best, while limma+ rv also had good performance.

bioinformatics

The Orphan Kinesin PAKRP2 Achieves Processive Motility Via Noncanonical Stepping

PAKRP2 is an orphan kinesin in Arabidopsis thaliana that is thought to transport vesicles along phragmoplast microtubules for cell plate formation. Here, using single-molecule fluorescence microscopy, we show that PAKRP2 exhibits processive plus-end-directed motility on single microtubules as individual homodimers despite having an exceptionally long (32 residues) neck linker. Furthermore, using high-resolution nanoparticle tracking to visualize motor stepping dynamics, we find that PAKRP2 achieves processivity via a noncanonical stepping mechanism that includes small step sizes and frequent lateral steps to adjacent protofilaments. We propose that the small steps sizes are due to a transient intermediate step that involves a prolonged diffusional search of the tethered head due to its long neck linker. Despite this different stepping behavior, ATP is tightly coupled to each 8-nm step. Collectively, this study reveals PAKRP2 as the first orphan kinesin to demonstrate processive motility and broadens our understanding of the diverse kinesin stepping mechanisms.

biophysics

Noninvasive prenatal test of methylmalonic academia cblC type through targeted sequencing of cell-free DNA in maternal plasma

Methylmalonic acidemia (MMA) cblC type is the most frequent inborn error of intracellular cobalamin metabolism which is caused by mutations of MMACHC gene. Non-invasive test of MMA for pregnant women facilitates safe and timely prenatal diagnosis of the disease. In our study, we aimed to design and validate a haplotype-based noninvasive prenatal test (NIPT) method for cblC type of MMA. Targeted capture sequencing using customized hybridization was performed utilizing gDNA (genomic DNA) of trios including parents and an affected proband to determine parental haplotypes associated with the mutant and wild allele. The fetal haplotype was inferred later based on the high depth sequencing data of maternal plasma as well as haplotype linkage analysis. The fetal genotypes deduced by NIPT were further validated by amniocentesis. Haplotype-based NIPT was successfully performed in 21 families. The results of NIPT of 21 families were all consistent with invasive prenatal diagnosis, which was interpreted in a blinded fashion. Three fetuses were identified as compound heterozygosity of MMACHC, 9 fetuses were carriers of MMACHC variant, and 9 fetuses were normal. These results indicated that the haplotype-based NIPT for MMA through small target capture region sequencing is technically accurate and feasible.

genetics

A draft reference genome sequence for Scutellaria baicalensis Georgi

Scutellaria baicalensis Georgi is an important medicinal plant used worldwide. Information about the genome of this species is important for scientists studying the metabolic pathways that synthesise the bioactive compounds in this plant. Here, we report a draft reference genome sequence for S. baicalensis obtained by a combination of Illumina and PacBio sequencing, which was assembled using 10 X Genomics and Hi-C technologies. We assembled 386.63 Mb of the 408.14 Mb genome, amounting to about 94.73% of the total genome size, and the sequences were anchored onto 9 pseudochromosomes with a super-N50 of 33.2 Mb. The reference genome sequence of S. baicalensis offers an important foundation for understanding the biosynthetic pathways for bioactive compounds in this medicinal plant and for its improvement through molecular breeding.

plant biology

New Statistical Methods for Constructing Robust Differential Correlation Networks

The interplay among microRNAs (miRNAs) plays an important role in the developments of complex human diseases. Co-expression networks can characterize the interactions among miRNAs. Differential correlation network is a powerful tool to investigate the differences of co-expression networks between cases and controls. To construct a differential correlation network, the Fishers Z-transformation test is usually used. However, the Fishers Z-transformation test requires the normality assumption, the violation of which would result in inflated Type I error rate. Several bootstrapping-based improvements for Fishers Z test have been proposed. However, these methods are too computationally intensive to be used to construct differential correlation networks for high-throughput genomic data. In this article, we proposed six novel robust equal-correlation tests that are computationally efficient. The systematic simulation studies and a real microRNA data analysis showed that one of the six proposed tests (ST5) overall performed better than other methods.

bioinformatics

Introduce a New Approach to Detect Genes Associated to Oral Squamous Cell Carcinoma

Oral squamous cell carcinoma (OSCC) represents the most frequent of all oral neoplasms in the world. Genetics plays an important role in the etiopathogenesis of OSCC. However, the investigation of the molecular mechanism of OSCC is still incomplete. In this article, we introduced a new approach to detect OSCC-associated genes, in which we not only compare mean difference, but also variance difference between cases and controls. Based on two OSCC datasets from Gene Expression Omnibus, we identified 456 differentially variable (DV) gene probes, in addition to 2,375 differentially expressed (DE) gene probes. There are 2,193 DE-only probes, 274 DV-only probes, and 182 DE-and-DV probes. DAVID functional analysis showed that genes corresponding to DE-only, DV-only, and DE-and-DV probes were enriched in different KEGG pathways, indicating they play different roles in OSCC. This new approach can be used to investigate the genetic risk factors for other complex human diseases.

genetics

The Association between Alcohol Consumption and Telomere Length: A Meta-Analysis Focusing on Observational Studies

BackgroundBoth telomere length and alcohol consumption play important roles in carcinogenesis and biological age. Many efforts have been made to investigate the association between alcohol consumption and telomere length. However, no consensus has been reached yet.\n\nMethodsIn this article, we performed a meta-analysis to integrate the investigation results in the literature about the association between alcohol consumption and telomere length. After searching articles published between 2000 and 2016, 21 articles (including 27 analyses, total sample size 35,891) met our eligibility criteria.\n\nResultsWe found a significant association between alcohol consumption and telomere length (Fishers combined p-value = 3.52E-8 and Liptaks weighted p-value = 8.24E-3). We also found that the significance of the association between alcohol consumption and telomere length varies with study type (cohort, case-control, or cross-sectional) and study population (Europe, Asia, American, or Australia).\n\nConclusionsCombined evidence showed that alcohol consumption is associated with telomere length. The consistent quantifications of alcohol consumption and telomere length would benefit the future aggregation of the evidence from different studies.

cancer biology

Identification and quantification of Lyme pathogen strains by deep sequencing of outer surface protein C (ospC) amplicons

Mixed infection of a single tick or host by Lyme disease spirochetes is common and a unique challenge for diagnosis, treatment, and surveillance of Lyme disease. Here we describe a novel protocol for differentiating Lyme strains based on deep sequencing of the hypervariable outer-surface protein C locus (ospC). Improving upon the traditional DNA-DNA hybridization method, the next-generation sequencing-based protocol is high-throughput, quantitative, and able to detect new pathogen strains. We applied the method to over one hundred infected Ixodes scapularis ticks collected from New York State, USA in 2015 and 2016. Analysis of strain distributions within individual ticks suggests an overabundance of multiple infections by five or more strains, inhibitory interactions among co-infecting strains, and presence of a new strain closely related to Borreliella bissettiae. A supporting bioinformatics pipeline has been developed. With the newly designed pair of universal ospC primers targeting intergenic sequences conserved among all known Lyme pathogens, the protocol could be used for culture-free identification and quantification of Lyme pathogens in wildlife and clinical specimens across the globe.

microbiology

Detecting Differential Variable microRNAs via Model-Based Clustering

Identifying genomic probes (e.g., DNA methylation marks) is becoming a new approach to detect novel genomic risk factors for complex human diseases. The F test is the standard equal-variance test in Statistics. For high-throughput genomic data, the probe-wise F test has been successfully used to detect biologically relevant DNA methylation marks that have different variances between two groups of subjects (e.g., cases vs. controls). In addition to DNA methylation, microRNA is another mechanism of epigenetics. However, to the best of our knowledge, no studies have identified differentially variable (DV) microRNAs. In this article, we proposed a novel model-based clustering to improve the power of the probe-wise F test to detect DV microRNAs. We imposed special structures on covariance matrices for each cluster of microRNAs based on the prior information about the relationship between variance in cases and variance in controls and about the independence among cases and controls. To the best of our knowledge, the proposed method is the first clustering algorithm that aims to detect DV genomic probes. Simulation studies showed that the proposed method outperformed the probe-wise F test and had certain robustness to the violation of the normality assumption. Based on two real datasets about human hepatocellular carcinoma (HCC), we identified 7 DV-only microRNAs (hsa-miR-1826, hsa-miR-191, hsa-miR-194-star, hsa-miR-222, hsa-miR-502-3p, hsa-miR-93, and hsa-miR-99b) using the proposed method, one (hsa-miR-1826) of which has not yet been reported to relate to HCC in the literature.

bioinformatics

Phylogeny Recapitulates Learning: Self-Optimization of Genetic Code

Learning algorithms have been proposed as a non-selective mechanism capable of creating complex adaptive systems in life. Evolutionary learning however has not been demonstrated to be a plausible cause for the origin of a specific molecular system. Here we show that genetic codes as optimal as the Standard Genetic Code (SGC) emerge readily by following a molecular analog of the Hebbs rule (\"neurons fire together, wire together\"). Specifically, error-minimizing genetic codes are obtained by maximizing the number of physio-chemically similar amino acids assigned to evolutionarily similar codons. Formulating genetic code as a Traveling Salesman Problem (TSP) with amino acids as \"cities\" and codons as \"tour positions\" and implemented with a Hopfield neural network, the unsupervised learning algorithm efficiently finds an abundance of genetic codes that are more error-minimizing than SGC. Drawing evidence from molecular phylogenies of contemporary tRNAs and aminoacyl-tRNA synthetases, we show that co-diversification between gene sequences and gene functions, which cumulatively captures functional differences with sequence differences and creates a genomic \"memory\" of the living environment, provides the biological basis for the Hebbian learning algorithm. Like the Hebbs rule, the locally acting phylogenetic learning rule, which may simply be stated as increasing phylogenetic divergence for increasing functional difference, could lead to complex and robust life systems. Natural selection, while essential for maintaining gene function, is not necessary to act at system levels. For molecular systems that are self-organizing through phylogenetic learning, the TSP model and its Hopfield network solution offer a promising framework for simulating emerging behavior, forecasting evolutionary trajectories, and designing optimal synthetic systems.

evolutionary biology

The preprophase band-associated kinesin-14 OsKCH2 is a processive minus-end-directed microtubule motor

In animals and fungi, cytoplasmic dynein is a processive motor that plays dominant roles in various intracellular processes. In contrast, land plants lack cytoplasmic dynein but contain many minus-end-directed kinesin-14s. No plant kinesin-14 is known to produce processive motility as a homodimer. OsKCH2 is a plant-specific kinesin-14 with an N-terminal actin-binding domain and a central motor domain flanked by two predicted coiled-coils (CC1 and CC2). Here, we show that OsKCH2 specifically decorates preprophase band microtubules in vivo and transports actin filaments along microtubules in vitro. Importantly, OsKCH2 exhibits processive minus-end-directed motility on single microtubules as individual homodimers. We find that CC1 but not CC2 forms the coiled-coil for OsKCH2 dimerization. Instead, CC2 functions to enable OsKCH2 processivity by enhancing its binding to microtubules. Collectively, these results show that land plants have evolved unconventional kinesin-14 homodimers with inherent minus-end-directed processivity that may function to compensate for the loss of cytoplasmic dynein.

biophysics

A SUMO-Ubiquitin Relay Recruits Proteasomes to Chromosome Axes to Regulate Meiotic Recombination

Meiosis produces haploid gametes through a succession of chromosomal events including pairing, synapsis and recombination. Mechanisms that orchestrate these events remain poorly understood. We found that the SUMO-modification and ubiquitin-proteasomes systems regulate the major events of meiotic prophase in mouse. Interdependent localization of SUMO, ubiquitin and proteasomes along chromosome axes was mediated largely by RNF212 and HEI10, two E3 ligases that are also essential for crossover recombination. RNF212-dependent SUMO conjugation effected a checkpoint-like process that stalls recombination by rendering the turnover of a subset of recombination factors dependent on HEI10-mediated ubiquitylation. We propose that SUMO conjugation establishes a precondition for designating crossover sites via selective protein stabilization. Thus, meiotic chromosome axes are hubs for regulated proteolysis via SUMO-dependent control of the ubiquitin-proteasome system.\n\nOne Sentence SummaryChromosomal events of meiotic prophase in mouse are regulated by proteasome-dependent protein degradation.

genetics