Search bioRxivSearch

Biology subjects

Gao, X.

Publications and source records attributed to Gao, X..

At least 19 recordsLinked to original sources

Proteomic Profile of TGF-β1 treated Lung Fibroblasts identifies Novel Markers of Activated Fibroblasts in the Silica Exposed Rat Lung

We performed liquid chromatography-tandem mass spectrometry (LC-MS/MS) on control and TGF-{beta}1-exposed rat lung fibroblasts to identify proteins differentially expressed between cell populations. A total of 1648 proteins were found to be differentially expressed in response to TGF-{beta}1 treatment and 196 proteins were expressed at [≥] 1.2 fold relative to control. Guided by these results, we next determined whether similar changes in protein expression were detectable in the rat lung after chronic exposure to silica dust. Of the five proteins selected for further analysis, we found that levels of all proteins were markedly increased in the silica-exposed rat lung, including the proteins for the very low density lipoprotein receptor (VLDLR) and the transmembrane (type I) heparin sulfate proteoglycan called syndecan 2 (SDC2). Because VLDLR and SDC2 have not, to our knowledge, been previously linked to the pathobiology of silicosis, we next examined whether knockdown of either gene altered responses to TGF-{beta}1 in MRC-5 lung fibroblasts. Interestingly, we found knockdown of either VLDLR or SDC2 dramatically reduced collagen production to TGF-{beta}1, suggesting that both proteins might play a novel role in myofibroblast biology and pathogenesis of silica-induced pulmonary fibrosis. In summary, our findings suggest that performing LC-MS/MS on TGF-{beta}1 stimulated lung fibroblasts can uncover novel molecular targets of activated myofibroblasts in silica-exposed lung.\n\nHighlightsWe identified 196 proteins differentially expressed between control and TGF-{beta}1 treated fibroblastsby LC-MS/MS.\n\nSeveral proteins identified by LC-MS/MS were also found to be differentially expressed in whole lung tissues and isolated fibroblasts after chronic exposure to silica dust, including the very low density lipoprotein receptor (VLDLR) and the transmembrane type I heparan sulfate proteoglycan called syndecan 2\n\nKnockdown of SDC2 or VLDLR markedly inhibited collagen production in MRC-5 fibroblasts, suggesting a novel pathogenic role for these proteins in myofibroblast biology.

pharmacology and toxicology

Carbon starvation-induced lipoprotein Slp directs the synthesis of catalase and expression of OxyR regulator to protect against hydrogen peroxide stress in Escherichia coli

Escherichia coli can induce a group of stress-response proteins, including carbon starvation-induced lipoprotein (Slp), which is an outer membrane lipoprotein expressed in response to stressful environments. In this paper, slp null mutant\n\nE. coli were constructed by insertion of the group II intron, and then the growth sensitivity of the slp mutant strain was measured under 0.6% (vol/vol) hydrogen peroxide. The changes in resistance to hydrogen peroxide stress were investigated by detecting antioxidant activity and gene expression in the slp mutant strain. The results showed that deletion of the slp gene increased the sensitivity of E. coli under 0.6% (vol/vol) hydrogen peroxide oxidative stress. Analysis of the unique mapping rates from the transcriptome libraries revealed that four of thirteen remarkably up/down-regulated genes in E. coli were involved in antioxidant enzymes after mutation of the slp gene. Mutation of the slp gene caused a significant increase in catalase activity, which contributed to an increase in glutathione peroxidase activity. The katG gene was activated by the OxyR regulator, which was activated directly by 0.6% (vol/vol) hydrogen peroxide, and HPI encoded by katG was induced against oxidative stress. Therefore, the carbon starvation-induced lipoprotein Slp regulates the expression of antioxidant enzymes and the transcriptional activator OxyR in response to the hydrogen peroxide environment, ensuring that cells are protected from hydrogen peroxide oxidative stress at the level of enzyme activity and gene expression.

molecular biology

RBM-5 modulates U2AF large subunit-dependent alternative splicing in C. elegans

A key step in pre-mRNA splicing is the recognition of 3 splicing sites by the U2AF large and small subunits, a process regulated by numerous trans-acting splicing factors. How these trans-acting factors interact with U2AF in vivo is unclear. From a screen for suppressors of the temperature-sensitive (ts) lethality of the C. elegans U2AF large subunit gene uaf-1(n4588) mutants, we identified mutations in the RNA binding motif gene rbm-5, a homolog of the tumor suppressor RBM5. rbm-5 mutations can suppress uaf-1(n4588) ts-lethality by loss of function and neuronal expression of rbm-5 was sufficient to rescue the suppression. Transcriptome analyses indicate that uaf-1(n4588) affected the expression of numerous genes and rbm-5 mutations can partially reverse the abnormal gene expression to levels similar to that of wild type. Though rbm-5 mutations did not obviously affect alternative splicing per se, they can suppress or enhance, in a gene-specific manner, the altered splicing of genes in uaf-1(n4588) mutants. Specifically, the recognition of a weak 3 splice site was more susceptible to the effect of rbm-5. Our findings provide novel in vivo evidence that RBM-5 can modulate UAF-1-dependent RNA splicing and suggest that RBM5 might interact with U2AF large subunit to affect tumor formation.\n\nAuthor summaryRNA splicing is a critical regulatory step for eukaryotic gene expression and has been involved in the pathogenesis of multiple diseases. How RNA splicing factors interact in vivo to affect the splicing and expression of genes is unclear. In studying the temperature-sensitive lethal phenotypes of a mutation affecting the splicing factor U2AF large subunit gene uaf-1 in the nematode Caenorhabditis elegans, we isolated suppressive mutations in the rbm-5 gene, a homolog of the human tumor suppressor gene RBM5. rbm-5 is broadly expressed in neurons to enhance the lethality of the uaf-1 mutants. We found that the uaf-1 mutation causes aberrant expression of genes in numerous biological pathways, a large portion of which can be corrected by rbm-5 mutations. The abnormal splicing of multiple genes caused by the uaf-1 mutation is either corrected or enhanced by rbm-5 mutations in a gene-specific manner. We propose that RBM-5 interacts with UAF-1 to affect RNA splicing and the tumor suppressor function of RBM5 might involve U2AF-dependent RNA splicing.

genetics

Genome-wide characterization, evolutionary analysis of WRKY genes in Cucurbitaceae species and assessment of its roles in resisting to powdery mildew disease

The WRKY proteins constitute a large family of transcription factors that have been known to play a wide range of regulatory roles in multiple biological processes. Over the past few years, many reports have focused on analysis of evolution and biological function of WRKY genes at the whole genome level in different plant species. However, little information is known about WRKY genes in melon (Cucumis melo L.). In the present study, a total of 56 putative WRKY genes were identified in melon, which were randomly distributed on their respective chromosomes. A multiple sequence alignment and phylogenetic analysis using melon, cucumber and watermelon predicted WRKY domains indicated that melon WRKY proteins could be classified into three main groups (I-III). Our analysis indicated that no recent duplication events of WRKY genes were detected in melon, and strong purifying selection was observed among the 85 orthologous pairs of Cucurbitaceae species. Expression profiles of CmWRKY derived from RNA-seq data and quantitative RT-PCR (qRT-PCR) analyses showed distinct expression patterns in various tissues, and the expression of 16 CmWRKY were altered following powdery mildew infection in melon. Besides, we also found that a total of 24 WRKY genes were co-expressed with 11 VQ family genes in melon. Our comparative genomic analysis provides a foundation for future functional dissection and understanding the evolution of WRKY genes in cucurbitaceae species, and will promote powdery mildew resistance study in melon.

plant biology

Validation of Prostate Cancer Risk Variants by CRISPR/Cas9 Mediated Genome Editing

GWAS have identified numerous SNPs associated with prostate cancer risk. One such SNP is rs10993994. It is located in the MSMB promoter, associates with MSMB encoded {beta}-microseminoprotein prostate secretion levels, and is associated with mRNA expression changes in MSMB and the adjacent gene NCOA4. In addition, our previous work showed a second SNP, rs7098889, is in LD with rs10993994 and associated with MSMB expression independent of rs10993994. Here, we generate a series of clones with single alleles removed by double guide RNA (gRNA) mediated CRISPR/Cas9 deletions, through which we demonstrate that each of these SNPs independently and greatly alters MSMB expression in an allele-specific manner. We further show that these SNPs have no substantial effect on the expression of NCOA4. These data demonstrate that a single SNP can have a large effect on gene expression and illustrate the importance of functional validation to deconvolute observed correlations. The method we have developed is generally applicable to test any SNP for which a relevant heterozygous cell line is available.\n\nAuthor summaryIn pursuing the underlying biological mechanism of prostate cancer pathogenesis, scientists utilized the existence of common single nucleotide polymorphisms (SNPs) in human genome as genetic markers to perform large scale genome wide association studies (GWAS) and have so far identified more than a hundred prostate cancer risk variants. Such variants provide an unbiased and systematic new venue to study the disease mechanism, and the next big challenge is to translate these genetic associations to the causal role of altered gene function in oncogenesis. The majority of these variants are waiting to be studied and lots of them may act in oncogenesis through gene expression regulation. To prove the concept, we took rs10993994 and its linked rs7098889 as an example and engineered single cell clones by allelic-specific CRISPR/Cas9 deletion to separate the effect of each allele. We observed that a single nucleotide difference would lead to surprisingly high level of MSMB gene expression change in a gene specific and tissue specific manner. Our study strongly supports the notion that differential level of gene expression caused by risk variants and their associated genetic locus play a major role in oncogenesis and also highlights the importance of studying the function of MSMB encoded {beta}-MSP in prostate cancer pathogenesis.

genetics

ClusterMap: Compare analysis across multiple Single Cell RNA-Seq profiling

Single cell RNA-Seq facilitates the characterization of cell type heterogeneity and developmental processes. Further study of single cell profiles across different conditions enables the understanding of biological processes and underlying mechanisms at the sub-population level. However, developing proper methodology to compare multiple scRNA-Seq datasets remains challenging. We have developed ClusterMap, a systematic method and workflow to facilitate the comparison of scRNA profiles across distinct biological contexts. Using hierarchical clustering of the marker genes of each sub-group, ClusterMap matches the sub-types of cells across different samples and provides \"similarity\" as a metric to quantify the quality of the match. We introduce a purity tree cut method designed specifically for this matching problem. We use Circos plot and regrouping method to visualize the results concisely. Furthermore, we propose a new metric \"separability\" to summarize sub-population changes among all sample pairs. In three case studies, we demonstrate that ClusterMap has the ability to offer us further insight into the different molecular mechanisms of cellular sub-populations across different conditions. ClusterMap is implemented in R and available at https://github.com/xgaoo/ClusterMap.

bioinformatics

SupportNet: a novel incremental learning framework through deep learning and support data

MotivationIn most biological data sets, the amount of data is regularly growing and the number of classes is continuously increasing. To deal with the new data from the new classes, one approach is to train a classification model, e.g., a deep learning model, from scratch based on both old and new data. This approach is highly computationally costly and the extracted features are likely very different from the ones extracted by the model trained on the old data alone, which leads to poor model robustness. Another approach is to fine tune the trained model from the old data on the new data. However, this approach often does not have the ability to learn new knowledge without forgetting the previously learned knowledge, which is known as the catastrophic forgetting problem. To our knowledge, this problem has not been studied in the field of bioinformatics despite its existence in many bioinformatic problems.\n\nResultsHere we propose a novel method, SupportNet, to solve the catastrophic forgetting problem efficiently and effectively. SupportNet combines the strength of deep learning and support vector machine (SVM), where SVM is used to identify the support data from the old data, which are fed to the deep learning model together with the new data for further training so that the model can review the essential information of the old data when learning the new information. Two powerful consolidation regularizers are applied to ensure the robustness of the learned model. Comprehensive experiments on various tasks, including enzyme function prediction, subcellular structure classification and breast tumor classification, show that SupportNet drastically outperforms the state-of-the-art incremental learning methods and reaches similar performance as the deep learning model trained from scratch on both old and new data.\n\nAvailabilityOur program is accessible at: https://github.com/lykaust15/SupportNet.

bioinformatics

Transcription Factor Regulation of RNA polymerase’s Torsional Capacity

During transcription, RNA polymerase (RNAP) supercoils DNA as it forsward-translocates. Accumulation of this torsional stress in DNA can become a roadblock for an elongating RNAP and thus should be subject to regulation during transcription. Here, we investigate whether, and how, a transcription factor may regulate the torque generation capacity of RNAP and torque-induced RNAP stalling. Using a real-time assay based on an angular optical trap, we found that under a resisting torque, RNAP was highly prone to extensive backtracking. However, the presence of GreB, a transcription factor that facilitates the cleavage of the 3 end of the extruded RNA transcript, greatly suppressed backtracking and remarkably increased the torque that RNAP was able to generate by 65%, from 11.2 to 18.5 pN{middle dot}nm. Analysis of the real-time trajectories of RNAP position at a stall revealed the kinetic parameters of backtracking and GreB rescue. These results demonstrate that backtracking is the primary mechanism that limits transcription against DNA supercoiling and the transcription factor GreB effectively enhances the torsional capacity of RNAP. These findings broaden the potential impact of transcription factors on RNAP functionality.

biophysics

Inference of significant microbial interactions from longitudinal metagenomics sequencing data

Data of next-generation sequencing (NGS) and their analysis have been facilitating advances in our understanding of microbial ecosystems such as human gut microbiota. However, inference of microbial interactions occurring within an ecosystem is still a challenge mainly due to metagenomics sequencing (e.g., 16S rDNA sequences) providing relative abundance of microbes instead of absolute cell count. In order to describe the population dynamics in microbial communities and estimate the involved microbial interactions, we introduce a procedure by integrating generalized Lotka-Volterra equations, forward stepwise regression and bootstrap aggregation. First, we successfully identify experimentally confirmed microbial interactions with relative abundance data of a cheese microbial community. Then, we apply the procedure to time-series of 16S rDNA sequences of gut microbiomes of children who were progressing to Type 1 diabetes (T1D progressors), and compare their gut microbial interactions to a healthy control group. Our results suggest that the number of inferred microbial interactions increased over time during the first three years of life. More microbial interactions are found in the gut flora of healthy children than the T1D progressors. The inhibitory effects from Actinobacteria and Bacilli to Bacteroidia, from Bacteroidia to Clostridia, and the benifit effect from Clostridia to Bacteroidia are shared between healthy children and T1D progressors. An inhibition of Clostridia by Gammaproteobacteria is found in healthy children that maintains through their first three years of life. This suppression appears in T1D progressors during the first year of life, which transforms to a commensalism relationship at the age of three years old. Gammaproteobacteria is found exerting an inhibition on Bacteroidia in the T1D progressors, which is not identified in the healthy controls.

ecology

Probing compression versus stretch activated recruitment of cortical actin and apical junction proteins using mechanical stimulations of suspended doublets.

We report an experimental approach to study the mechanosensitivity of cellcell contact upon mechanical stimulation in suspended cell-doublets. The doublet is placed astride an hourglass aperture, and a hydrodynamic force is selectively exerted on only one of the cells. The geometry of the device concentrates the mechanical shear over the junction area. Together with mechanical shear, the system also allows confocal quantitative live imaging of the recruitment of junction proteins (e.g. E-cadherin, ZO-1, Occludin and actin). We observed the time sequence over which proteins were recruited to the stretched region of the contact. The compressed side of the contact showed no response. We demonstrated how this mechanism polarizes the stress-induced recruitment of junctional components within one single junction. Finally, we demonstrated that stabilizing the actin cortex dynamics abolishes the mechanosensitive response of the junction. Our experimental design provides an original approach to study the role of mechanical force at a cell-cell contact with unprecedented control over stress application and quantitative optical analysis.

biophysics

Whole-exome sequencing identified rare variants associated with body length and girth in cattle

Body measurements can be used in determining body size to monitor the cattle growth and examine the response to selection. Despite efforts putting into the identification of common genetic variants, the mechanism understanding of the rare variation in complex traits about body size and growth remains limited. Here, we firstly performed GWAS study for body measurement traits in Simmental cattle, however there were no SNPs exceeding significant level associated with body measurements. To further investigate the mechanism of growth traits in beef cattle, we conducted whole exome analysis of 20 cattle with phenotypic differences on body girth and length, representing the first systematic exploration of rare variants on body measurements in cattle. By carrying out a three-phase process of the variant calling and filtering, a sum of 1158, 1151, 1267, and 1303 rare variants were identified in four phenotypic groups of two growth traits, higher/ lower body girth (BG_H and BG_L) and higher/lower body length (BL_H and BL_L) respectively. The subsequent functional enrichment analysis revealed that these rare variants distributed in 886 genes associated with collagen formation and organelle organization, indicating the importance of collagen formation and organelle organization for body size growth in cattle. The integrative network construction distinguished 62 and 66 genes with different co-expression patterns associated with higher and lower phenotypic groups of body measurements respectively, and the two sub-networks were distinct. Gene ontology and pathway annotation further showed that all shared genes in phenotypic differences participate in many biological processes related to the growth and development of the organism. Together, these findings provide a deep insight into rare genetic variants of growth traits in cattle and this will have a promising application in animal breeding.

genetics

Proteome-level assessment of origin, prevalence and function of Leucine-Aspartic Acid (LD) motifs

Short Linear Motifs (SLiMs) contribute to almost every cellular function by connecting appropriate protein partners. Accurate prediction of SLiMs is difficult due to their shortness and sequence degeneracy. Leucine-aspartic acid (LD) motifs are SLiMs that link paxillin family proteins to factors controlling (cancer) cell adhesion, motility and survival. The existence and importance of LD motifs beyond the paxillin family is poorly understood. To enable a proteome-wide assessment of these motifs, we developed an active-learning based framework that iteratively integrates computational predictions with experimental validation. Our analysis of the human proteome identified a dozen proteins that contain LD motifs, all being involved in cell adhesion and migration, and revealed a new type of inverse LD motif consensus. Our evolutionary analysis suggested that LD motif signalling originated in the common unicellular ancestor of opisthokonts and amoebozoa by co-opting nuclear export sequences. Inter-species comparison revealed a conserved LD signalling core, and reveals the emergence of species-specific adaptive connections, while maintaining a strong functional focus of the LD motif interactome. Collectively, our data elucidate the mechanisms underlying the origin and adaptation of an ancestral SLiM.

bioinformatics

Differential Expression of Coding and Long Noncoding RNAs in Keratoconus-affected Corneas

PURPOSEKeratoconus (KC) is the most common corneal ectasia. We aimed to determine the differential expression of coding and long noncoding RNAs (lncRNAs) in human corneas affected with KC.\n\nMETHODS200ng total RNA from the corneas of 10 KC patients and 8 non-KC normal controls was used to prepare sequencing libraries with the SMARTer Stranded RNA-Seq kit after ribosomal RNA depletion. Paired-end 50bp sequences were generated using Illumina HiSeq 2500 Sequencer. Differential analysis was done using TopHat and Cufflinks with a gene file from Ensembl and a lncRNA file from NONCODE. Pathway analysis was performed using WebGestalt. Using the expression level of differentially expressed coding and noncoding RNAs in each sample, we correlated their expression levels in KC and controls separately and identified significantly different correlations in KC against controls followed by visualization using Cytoscape.\n\nRESULTSUsing |fold change| [≥] 2 and a false discovery rate [≤] 0.05, we identified 436 coding RNAs and 584 lncRNAs with differential expression in the KC-affected corneas. Pathway analysis indicated the enrichment of genes involved in extracellular matrix, protein binding, glycosaminoglycan binding, and cell migration. Our correlation analysis identified 296 pairs of significant KC-specific correlations containing 117 coding genes enriched in functions related with cell migration/motility, extracellular space, cytokine response, and cell adhesion, suggesting the potential functions of these correlated lncRNAs, especially those with multiple pairs of correlations.\n\nCONCLUSIONSOur RNA-Seq based differential expression and correlation analyses have identified many potential KC contributing coding and noncoding RNAs.

genomics

Xist Intron 1 Repression by TALE Transcriptional Factor Improves Somatic Cell Reprogamming in Mice

Xist is the master regulator of X chromosome inactivation (XCI). In order to further understand the Xist locus in reprogramming of somatic cells to induced pluripotent stem cells (iPSCs) and in somatic cell nuclear transfer (SCNT), we tested transcription-factor-like effectors (TALE)-based designer transcriptional factors (dTFs), which were specific to numerous regions at the Xist locus. We report that the selected dTF repressor 6 (R6) binding the intron 1 of Xist, which did not affect Xist expression in mouse embryonic fibroblasts (MEFs), substantially improved the iPSC generation and the SCNT preimplantation embryo development. Conversely, the dTF activator targeting the same genomic region of R6 decreased iPSC formation, and blocked SCNT-embryo development. These results thus uncover the critical requirement for the Xist locus in epigenetic resetting, which is not directly related to Xist transcription. This may provide a unique route to improving the reprogramming.

developmental biology

Methionine metabolism influences the genomic architecture of H3K4me3 with the link to gene expression encoded in peak width

Nutrition and metabolism are known to influence chromatin biology and epigenetics by modifying the levels of post-translational modifications on histones, yet how changes in nutrient availability influence specific aspects of genomic architecture and connect to gene expression is unknown. To investigate this question we considered, as a model, the metabolically-driven dynamics of H3K4me3, a histone methylation mark that is known to encode information about active transcription, cell identity, and tumor suppression. We analyzed the genome-wide changes in H3K4me3 and gene expression in response to alterations in methionine availability under conditions that are known to affect the global levels of histone methylation in both normal rodent physiology and in human cancer cells. Surprisingly, we found that the location of H3K4me3 peaks at specific genomic loci was largely preserved under conditions of methionine restriction. However, upon examining different geometrical features of peak shape, it was found that the response of H3K4me3 peak width encoded almost all aspects of H3K4me3 biology including changes in expression levels, and the presence of cell identity and cancer associated genes. These findings reveal simple yet new and profound principles for how nutrient availability modulates specific aspects of chromatin dynamics to mediate key biological features.

genetics

Spatial gradient in activity within the insula reflects dissociable neural mechanisms underlying context-dependent advantageous and disadvantageous inequity aversion

Humans are capable of integrating social contextual information into decision-making processes to adjust their attitudes towards inequity. This context-dependency emerges both when individual is better off (i.e. advantageous inequity) and worse off (i.e. disadvantageous inequity) than others. It is not clear however, whether the context-dependent processing of advantageous and disadvantageous inequity rely on dissociable or shared neural mechanisms. Here, by combining an interpersonal interactive game that gave rise to interpersonal guilt and different versions of the dictator games that enabled us to characterize individual weights on aversion to advantageous and disadvantageous inequity, we investigated the neural mechanisms underlying the two forms of inequity aversion in the interpersonal guilt context. In each round, participants played a dot-estimation task with an anonymous co-player. The co-players received pain stimulation with 50% probability when anyone responded incorrectly. At the end of each round, participants completed a dictator game, which determined payoffs of him/herself and the co-player. Both computational model-based and model-free analyses demonstrated that when inflicting pain upon co-players (i.e., the guilt context), participants cared more about advantageous inequity and became less sensitive to disadvantageous inequity, compared with other social contexts. The contextual effects on two forms of inequity aversion are uncorrelated with each other at the behavioral level. Neuroimaging results revealed that the context-dependent representation of inequity aversion exhibited a spatial gradient in activity within the insula, with anterior parts predominantly involved in the aversion to advantageous inequity and posterior parts predominantly involved in the aversion to disadvantageous inequity. The dissociable mechanisms underlying the two forms of inequity aversion are further supported by the involvement of right dorsolateral prefrontal cortex and dorsomedial prefrontal cortex in advantageous inequity processing, and the involvement of right amygdala and dorsal anterior cingulate cortex in disadvantageous inequity processing. These results extended our understanding of decision-making processes involving inequity and the social functions of inequity aversion.

neuroscience

An accurate and rapid continuous wavelet dynamic time warping algorithm for unbalanced global mapping in nanopore sequencing

Long-reads, point-of-care, and PCR-free are the promises brought by nanopore sequencing. Among various steps in nanopore data analysis, the global mapping between the raw electrical current signal sequence and the expected signal sequence from the pore model serves as the key building block to base calling, reads mapping, variant identification, and methylation detection. However, the ultra-long reads of nanopore sequencing and an order of magnitude difference in the sampling speeds of the two sequences make the classical dynamic time warping (DTW) and its variants infeasible to solve the problem. Here, we propose a novel multi-level DTW algorithm, cwDTW, based on continuous wavelet transforms with different scales of the two signal sequences. Our algorithm starts from low-resolution wavelet transforms of the two sequences, such that the transformed sequences are short and have similar sampling rates. Then the peaks and nadirs of the transformed sequences are extracted to form feature sequences with similar lengths, which can be easily mapped by the original DTW. Our algorithm then recursively projects the warping path from a lower-resolution level to a higher-resolution one by building a context-dependent boundary and enabling a constrained search for the warping path in the latter. Comprehensive experiments on two real nanopore datasets on human and on Pandoraea pnomenusa, as well as two benchmark datasets from previous studies, demonstrate the efficiency and effectiveness of the proposed algorithm. In particular, cwDTW can almost always generate warping paths that are very close to the original DTW, which are remarkably more accurate than the state-of-the-art methods including Fast-DTW and PrunedDTW. Meanwhile, on the real nanopore datasets, cwDTW is about 440 times faster than FastDTW and 3000 times faster than the original DTW. Our program is available at https://github.com/realbigws/cwDTW.

bioinformatics

DeepSimulator: a deep simulator for Nanopore sequencing

MotivationOxford Nanopore sequencing is a rapidly developed sequencing technology in recent years. To keep pace with the explosion of the downstream data analytical tools, a versatile Nanopore sequencing simulator is needed to complement the experimental data as well as to benchmark those newly developed tools. However, all the currently available simulators are based on simple statistics of the produced reads, which have difficulty in capturing the complex nature of the Nanopore sequencing procedure, the main task of which is the generation of raw electrical current signals.\n\nResultsHere we propose a deep learning based simulator, DeepSimulator, to mimic the entire pipeline of Nanopore sequencing. Starting from a given reference genome or assembled contigs, we simulate the electrical current signals by a context-dependent deep learning model, followed by a base-calling procedure to yield simulated reads. This workflow mimics the sequencing procedure more naturally. The thorough experiments performed across four species show that the signals generated by our context-dependent model are more similar to the experimentally obtained signals than the ones generated by the official context-independent pore model. In terms of the simulated reads, we provide a parameter interface to users so that they can obtain the reads with different accuracies ranging from 83% to 97%. The reads generated by the default parameter have almost the same properties as the real data. Two case studies demonstrate the application of DeepSimulator to benefit the development of tools in de novo assembly and in low coverage SNP detection.\n\nAvailabilityThe software can be accessed freely at: https://github.com/lykaust15/deep_simulator.

bioinformatics