Search bioRxivSearch

Biology subjects

Conesa, A.

Publications and source records attributed to Conesa, A..

7 recordsLinked to original sources

Building gene regulatory networks from single-cell ATAC-seq and RNA-seq using Linked Self-Organizing Maps

Rapid advances in single-cell assays have outpaced methods for analysis of those data types. Different single-cell assays show extensive variation in sensitivity and signal to noise levels. In particular, scATAC-seq generates extremely sparse and noisy datasets. Existing methods developed to analyze this data require cells amenable to pseudo-time analysis or require datasets with drastically different cell-types. We describe a novel approach using self-organizing maps (SOM) to link scATAC-seq and scRNA-seq data that overcomes these challenges and can generate draft regulatory networks. Our SOMatic package generates chromatin and gene expression SOMs separately and combines them using a linking function. We applied SOMatic on a mouse pre-B cell differentiation time-course using controlled Ikaros over-expression to recover gene ontology enrichments, identify motifs in genomic regions showing similar single-cell profiles, and generate a gene regulatory network that both recovers known interactions and predicts new Ikaros targets during the differentiation process. The ability of linked SOMs to detect emergent properties from multiple types of highly-dimensional genomic data with very different signal properties opens new avenues for integrative analysis of single-cells.

genomics

MOSim: Multi-Omics Simulation in R

As multi-omics sequencing technologies advance, the need for simulation tools capable of generating realistic and diverse (bulk and single-cell) multi-omics datasets for method testing and benchmarking becomes increasingly important. We present MOSim, an R package that simulates both bulk (via mosim function) and single-cell (via sc_mosim function) multi-omics data. The mosim function generates bulk transcriptomics data (RNA-seq) and additional regulatory omics layers (ATAC-seq, miRNA-seq, ChIP-seq, Methyl-seq and Transcription Factors), while sc_mosim simulates single-cell transcriptomics data (scRNA-seq) with scATAC-seq and Transcription Factors as regulatory layers. The tool supports various experimental designs, including simulation of gene co-expression patterns, biological replicates, and differential expression between conditions. MOSim enables users to generate quantification matrices for each simulated omics data type, capturing the heterogeneity and complexity of bulk and single-cell multi-omics datasets. Furthermore, MOSim provides differentially abundant features within each omics layer and elucidates the active regulatory relationships between regulatory omics and gene expression data at both bulk and single-cell levels. By leveraging MOSim, researchers will be able to generate realistic and customizable bulk and single-cell multi-omics datasets to benchmark and validate analytical methods specifically designed for the integrative analysis of diverse regulatory omics data. Key PointsO_LIMOSim is capable of generating synthetic datasets for a broad spectrum of omics types, supporting bulk RNA-seq, ChIP-seq, ATAC-seq, miRNA-seq, Methyl-seq, and transcription factor data, as well as single-cell omics, including scRNA-seq, scATAC-seq, and transcription factors. C_LIO_LIMOSim enables the robust simulation of complex, many-to-many regulatory relationships across molecular layers, faithfully capturing intricate regulatory patterns. C_LIO_LIOffering extensive options for customization, MOSims flexible experimental design and parameterization empowers users to simulate count matrices and multilayer regulatory networks, tailoring simulations to diverse experimental scenarios and omics types. C_LI

bioinformatics

Vascular brassinosteroid receptors confer drought resistance without penalizing plant growth

Drought represents a major threat to food security. Mechanistic data describing plant responses to drought have been studied extensively and genes conferring drought resistance have been introduced into crop plants. However, plants with enhanced drought resistance usually display lower growth, highlighting the need for strategies to uncouple drought resistance from growth. Here, we show that overexpression of BRL3, a vascular-enriched member of the brassinosteroid receptor family, can confer drought stress tolerance in Arabidopsis. Whereas loss-of-function mutations in the ubiquitously expressed BRI1 receptor leads to drought resistance at the expense of growth, overexpression of BRL3 receptor confers drought tolerance without penalizing overall growth. Systematic analyses reveal that upon drought stress, increased BRL3 triggers the accumulation of osmoprotectant metabolites including proline and sugars. Transcriptomic analysis suggests that this results from differential expression of genes in the vascular tissues. Altogether, this data suggests that manipulating BRL3 expression could be used to engineer drought tolerant crops.

plant biology

PaintOmics 3: a web resource for the pathway analysis and visualization of multi-omics data

The increasing availability of multi-omic platforms poses new challenges to data analysis. Joint visualization of multi-omics data is instrumental to understand interconnections across molecular layers and to fully leverage the biology discovery power offered by the multi-omics approach.\n\nWe present here PaintOmics 3, a web-based resource for the integrated visualization of multiple omic data types onto KEGG pathway diagrams. PaintOmics 3 combines server-end capabilities for data analysis with the potential of modern web resources for data visualization, providing researchers with a powerful framework for interactive exploration of their multi-omics information.\n\nUnlike other visualization tools, PaintOmics 3 covers a complete pathway analysis workflow, including automatic feature name/identifier conversion, multi-layered feature matching, pathway enrichment, network analysis, interactive heatmaps, trend charts, etc. It accepts a wide variety of omic types, including transcriptomics, proteomics and metabolomics, as well as region-based approaches such as ATAC-seq or ChIP-seq data. The tool is freely available at http://bioinfo.cipf.es/paintomics/.

bioinformatics

A benchmarking of workflows for detecting differential splicing and differential expression at isoform level in human RNA-seq studies

Over the last few years, RNA-seq has been used to study alterations in alternative splicing related to several diseases. Bioinformatics workflows used to perform these studies can be divided into two groups, those finding changes in the absolute isoform expression and those studying differential splicing. Many computational methods for transcriptomics analysis have been developed, evaluated and compared; however, there are not enough reports of systematic and objective assessment of processing pipelines as a whole. Moreover, comparative studies have been performed considering separately the changes in absolute or relative isoform expression levels. Consequently, no consensus exists about the best practices and appropriate workflows to analyse alternative and differential splicing. To assist the adequate pipeline choice, we present here a benchmarking of nine commonly used workflows to detect differential isoform expression and splicing. We evaluated the workflows performance over three different experimental scenarios where changes in absolute and relative isoform expression occurred simultaneously. In addition, the effect of the number of isoforms per gene, and the magnitude of the expression change over pipeline performances were also evaluated. Our results suggest that workflow performance is influenced by the number of replicates per condition and the conditions heterogeneity. In general, workflows based on DESeq, DEXSeq, Limma and NOISeq performed well over a wide range of transcriptomics experiments. In particular, we suggest the use of workflows based on Limma when high precision is required, and DESeq2 and DEXseq pipelines to prioritize sensitivity. When several replicates per condition are available, NOISeq and Limma pipelines are indicated.

bioinformatics

Identification and visualization of differential isoform expression in RNA-seq time series

As sequencing technologies improve their capacity to detect distinct transcripts of the same gene and to address complex experimental designs such as longitudinal studies, there is a need to develop statistical methods for the analysis of isoform expression changes in time series data. Iso-maSigPro is a new functionality of the R package maSigPro for transcriptomics time series data analysis. Iso-maSigPro identifies genes with a differential isoform usage across time. The package also includes new clustering and visualization functions that allow grouping of genes with similar expression patterns at the isoform level, as well as those genes with a shift in major expressed isoform. The package is freely available under the LGPL license from the Bioconductor web site (http://bioconductor.org).

bioinformatics

SQANTI: extensive characterization of long read transcript sequences for quality control in full-length transcriptome identification and quantification

High-throughput sequencing of full-length transcripts using long reads has paved the way for the discovery of thousands of novel transcripts, even in very well annotated organisms as mice and humans. Nonetheless, there is a need for studies and tools that characterize these novel isoforms. Here we present SQANTI, an automated pipeline for the classification of long-read transcripts that computes 47 descriptors that can be used to assess the quality of the data and of the preprocessing pipelines. We applied SQANTI to a neuronal mouse transcriptome using PacBio long reads and illustrate how the tool is effective in readily describing the composition of and characterizing the full-length transcriptome. We perform extensive evaluation of ToFU PacBio transcripts by PCR to reveal that an important number of the novel transcripts are technical artifacts of the sequencing approach, and that SQANTI quality descriptors can be used to engineer a filtering strategy to remove them. Most novel transcripts in this curated transcriptome are novel combinations of existing splice sites, result more frequently in novel ORFs than novel UTRs and are enriched in both general metabolic and neural specific functions. We show that these new transcripts have a major impact in the correct quantification of transcript levels by state-of-the-art short-read based quantification algorithms. By comparing our iso-transcriptome with public proteomics databases we find that alternative isoforms are elusive to proteogenomics detection and are variable in protein changes with respect to the principal isoform of their genes. SQANTI allows the user to maximize the analytical outcome of long read technologies by providing the tools to deliver quality-evaluated and curated full-length transcriptomes. SQANTI is available at https://bitbucket.org/ConesaLab/sqanti.

bioinformatics