Search bioRxivSearch

Biology subjects

Akalin, A.

Publications and source records attributed to Akalin, A..

6 recordsLinked to original sources

Long non-coding RNAs defining major subtypes of B cell precursor acute lymphoblastic leukemia

Recent studies implicated that long non-coding RNAs (lncRNAs) may play a role in the progression and development of acute lymphoblastic leukemia, however, this role is not yet clear. In order to unravel the role of lncRNAs associated with B-cell precursor Acute Lymphoblastic Leukemia (BCP-ALL) subtypes, we performed transcriptome sequencing and DNA methylation array across 82 BCP-ALL samples from three molecular subtypes (DUX4, Ph-like, and Near Haploid or High Hyperdiploidy). Unsupervised clustering of BCP-ALL samples on the basis of their lncRNAs on transcriptome and DNA methylation profiles revealed robust clusters separating three molecular subtypes. Using extensive computational analysis, we developed a comprehensive catalog of 1235 aberrantly dysregulated BCP-ALL subtype-specific lncRNAs with altered expression and methylation patterns from three subtypes of BCP-ALL. By analyzing the co-expression of subtype-specific lncRNAs and protein-coding genes, we inferred key molecular processes in BCP-ALL subtypes. A strong correlation was identified between the DUX4 specific lncRNAs and activation of TGF-{beta} and Hippo signaling pathways. Similarly, Ph-like specific lncRNAs were correlated with genes involved in activation of PI3K-AKT, mTOR, and JAK-STAT signaling pathways. Interestingly, the relapse-specific differentially expressed lncRNAs correlated with the activation of metabolic and signaling pathways. Finally, we showed a set of epigenetically altered lncRNAs facilitating the expression of tumor genes located at their cis location. Overall, our study provides a comprehensive set of novel subtype and relapse-specific lncRNAs in BCP-ALL. Our findings suggest a wide range of molecular pathways are associated with lncRNAs in BCP-ALL subtypes and provide a foundation for functional investigations that could lead to new therapeutic approaches.\n\nAuthor SummaryAcute lymphoblastic leukemia is a heterogeneous blood cancer, with multiple molecular subtypes, and with high relapse rate. We are far from the complete understanding of the rationale behind these subtypes and high relapse rate. Long non-coding (lncRNAs) has emerged as a novel class of RNA due to its diverse mechanism in cancer development and progression. LncRNAs does not code for proteins and represent around 70% of human transcripts. Recently, there are a number of studies used lncRNAs expression profile in the classification of various cancers subtypes and displayed their correlation with genomic, epigenetic, pathological and clinical features in diverse cancers. Therefore, lncRNAs can account for heterogeneity and has independent prognostic value in various cancer subtypes. However, lncRNAs defining the molecular subtypes of BCP-ALL are not portrayed yet. Here, we describe a set of relapse and subtype-specific lncRNAs from three major BCP-ALL subtypes and define their potential functions and epigenetic regulation. Our data uncover the diverse mechanism of action of lncRNAs in BCP-ALL subtypes defining how lncRNAs are involved in the pathogenesis of disease and the relevance in the stratification of BCP-ALL subtypes.

cancer biology

Reproducible genomics analysis pipelines with GNU Guix

In bioinformatics, as well as other computationally-intensive research fields, there is a need for workflows that can reliably produce consistent output, independent of the software environment or configuration settings of the machine on which they are executed. Indeed, this is essential for controlled comparison between different observations or for the wider dissemination of workflows. Providing this type of reproducibility, however, is often complicated by the need to accommodate the myriad dependencies included in a larger body of software, each of which generally come in various versions. Moreover, in many fields (bioinformatics being a prime example), these versions are subject to continual change due to rapidly evolving technologies, further complicating problems related to reproducibility. Here, we propose a principled approach for building analysis pipelines and managing their dependencies. As a case study to demonstrate the utility of our approach, we present a set of highly reproducible pipelines for the analysis of RNA-seq, ChIP-seq, Bisulfite-seq, and single-cell RNA-seq. All pipelines process raw experimental data, and generate reports containing publication-ready plots and figures, with interactive report elements and standard observables. Users may install these highly reproducible packages and apply them to their own datasets without any special computational expertise beyond the use of the command line. We hope such a toolkit will provide immediate benefit to laboratory workers wishing to process their own data sets or bioinformaticians seeking to automate all, or parts of, their analyses. In the long term, we hope our approach to reproducibility will serve as a blueprint for reproducible workflows in other areas. Our pipelines, along with their corresponding documentation and sample reports, are available at http://bioinformatics.mdc-berlin.de/pigx

bioinformatics

netSmooth: Network-smoothing based imputation for single cell RNA-seq

Single cell RNA-seq (scRNA-seq) experiments suffer from a range of characteristic technical biases, such as dropouts (zero or near zero counts) and high variance. Current analysis methods rely on imputing missing values by various means of local averaging or regression, often amplifying biases inherent in the data. We present netSmooth, a network-diffusion based method that uses priors for the covariance structure of gene expression profiles on scRNA-seq experiments in order to smooth expression values. We demonstrate that netSmooth improves clustering results of scRNA-seq experiments from distinct cell populations, time-course experiments, and cancer genomics. We provide an R package for our method, available at: https://github.com/BIMSBbioinfo/netSmooth.

bioinformatics

FACT sets a barrier for cell fate reprogramming in C. elegans and Human

O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC=\"FIGDIR/small/185116_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (59K):\norg.highwire.dtl.DTLVardef@15ae7aaorg.highwire.dtl.DTLVardef@11f5c08org.highwire.dtl.DTLVardef@1d32442org.highwire.dtl.DTLVardef@f1808f_HPS_FORMAT_FIGEXP M_FIG C_FIG The chromatin regulator FACT (Facilitates Chromatin Transcription) is essential for ensuring stable gene expression by promoting transcription. In a genetic screen using C. elegans we identified that FACT maintains cell identities and acts as a barrier for transcription factor-mediated cell fate reprogramming. Strikingly, FACTs role as a reprogramming barrier is conserved in humans as we show that FACT depletion enhances reprogramming of fibroblasts into stem cells and neurons. Such activity of FACT is unexpected since known reprogramming barriers typically repress gene expression by silencing chromatin. In contrast, FACT is a positive regulator of gene expression suggesting an unprecedented link of cell fate maintenance with counteracting alternative cell identities. This notion is supported by ATAC-seq analysis showing that FACT depletion results in decreased but also increased chromatin accessibility for transcription factors. Our findings identify FACT as a cellular reprogramming barrier in C. elegans and humans, revealing an evolutionarily conserved mechanism for cell fate protection.

developmental biology

Mutations In Disordered Regions Cause Disease By Creating Endocytosis Motifs

Mutations in intrinsically disordered regions (IDRs) of proteins can cause a wide spectrum of diseases. Since IDRs lack a fixed three-dimensional structure, the mechanism by which such mutations cause disease is often unknown. Here, we employ a proteomic screen to investigate the impact of mutations in IDRs on protein-protein interactions. We find that mutations in disordered cytosolic regions of three transmembrane proteins (GLUT1, ITPR1 and CACNA1H) lead to an increased binding of clathrins. In all three cases, the mutation creates a dileucine motif known to mediate clathrin-dependent trafficking. Follow-up experiments on GLUT1 (SLC2A1), a glucose transporter involved in GLUT1 deficiency syndrome, revealed that the mutated protein mislocalizes to intracellular compartments. A systematic analysis of other known disease-causing variants revealed a significant and specific overrepresentation of gained dileucine motifs in cytosolic tails of transmembrane proteins. Dileucine motif gains thus appear to be a recurrent cause of disease.

biochemistry

Strategies for analyzing bisulfite sequencing data

DNA methylation is one of the main epigenetic modifications in the eukaryotic genome; it has been shown to play a role in cell-type specific regulation of gene expression, and therefore cell-type identity. Bisulfite sequencing is the gold-standard for measuring methylation over the genomes of interest. Here, we review several techniques used for the analysis of high-throughput bisulfite sequencing. We introduce specialized short-read alignment techniques as well as pre/post-alignment quality check methods to ensure data quality. Furthermore, we discuss subsequent analysis steps after alignment. We introduce various differential methylation methods and compare their performance using simulated and real bisulfite sequencing datasets. We also discuss the methods used to segment methylomes in order to pinpoint regulatory regions. We introduce annotation methods that can be used for further classification of regions returned by segmentation and differential methylation methods. Finally, we review software packages that implement strategies to efficiently deal with large bisulfite sequencing datasets locally and we discuss online analysis workflows that do not require any prior programming skills. The analysis strategies described in this review will guide researchers at any level to the best practices of bisulfite sequencing analysis.

bioinformatics