Search bioRxivSearch

Biology subjects

Klein, K.

Publications and source records attributed to Klein, K..

4 recordsLinked to original sources

Visualisation and analysis of RNA-Seq assembly graphs

RNA-sequencing (RNA-Seq) is a powerful transcriptome profiling technology enabling transcript discovery and quantification. RNA-Seq data are large, and most commonly used as a source of genelevel quantification measurements, whilst the underlying assemblies of reads, if inspected, are usually viewed as sequence reads mapped on to a reference genome. Whilst sufficient for many needs, when the underlying transcript assemblies are complex, this visualisation approach can be limiting; errors in assembly can be difficult to spot and interpretation of splicing events is challenging.\n\nHere we report on the development of a graph-based visualisation method as a complementary approach to understanding transcript diversity and read assembly from short-read RNA-Seq data. Following the mapping of reads to the reference genome, read-to-read comparison is performed on all reads mapping to a given gene, producing a matrix of weighted similarity scores between reads. This is used to produce an RNA assembly graph where nodes represent reads derived from a cDNA and edges similarity scores between reads, above a defined threshold. Visualisation of resulting graphs is performed using Graphia Professional. This tool can render the often large and complex graph topologies that result from DNA/RNA sequence assembly in 3D space and supports info rmatio no verlay on to nodes, e.g. transcript models. We have also implemented an analysis pipeline for the creation of RNA assembly graphs with both a command-line and web-based interface that allows users to create and visualise these data. Here we demonstrate the utility of this approach on RNA-Seq data, including the unusual structure of these graphs and how they can be used to identify issues in assembly, repetitive sequences within transcripts and splice variants. We believe this approach has the potential to significantly improve our understanding of transcript complexity.

bioinformatics

Genome Scale Epigenetic Profiling Reveals Five Distinct Subtypes of Colorectal Cancer

BACKGROUNDColorectal cancer is an epigenetically heterogeneous disease, however the extent and spectrum of the CpG Island Methylator Phenotype (CIMP) is not clear.\n\nRESULTSAn unselected cohort of 216 colorectal cancers clustered into five clinically and molecularly distinct subgroups using Illumina 450K DNA methylation arrays. CIMP-High cancers were most frequent in the proximal colons of female patients. These dichotomised into CIMP-Hl and CIMP-H2 based on methylation profile which was supported by over representation of BRAF (74%, P<0.0001) or KRAS (55%, P<0.0001) mutation, respectively. Congruent with increasing methylation, there was a stepwise increase in patient age from 62 years in the CI MP-Negative subgroup to 75 years in the CIMP-Hl subgroup (P<0.0001). There was a striking association between PRC2-marked loci and those subjected to significant gene body methylation in CIMP-type cancers (P<1.6xl078). We identified oncogenes susceptible to gene body methylation and Wnt pathway antagonists resistant to gene body methylation. CIMP cluster specific mutations were observed for genes involved in chromatin remodelling, such as in the SWI/SNF and NuRD complexes, suggesting synthetic lethality.\n\nCONCLUSIONThere are five clinically and molecularly distinct subgroups of colorectal cancer based on genome wide epigenetic profiling. These analyses highlighted an unidentified role for gene body methylation in progression of serrated neoplasia. Subgroup-specific mutation of distinct epigenetic regulator genes revealed potentially druggable vulnerabilities for these cancers, which may provide novel precision medicine approaches.

genomics

Multi-omics approach identifies novel pathogen-derived prognostic biomarkers in patients with Pseudomonas aeruginosa bloodstream infection

Pseudomonas aeruginosa is a human pathogen that causes health-care associated blood stream infections (BSI). Although P. aeruginosa BSI are associated with high mortality rates, the clinical relevance of pathogen-derived prognostic biomarker to identify patients at risk for unfavorable outcome remains largely unexplored. We found novel pathogen-derived prognostic biomarker candidates by applying a multi-omics approach on a multicenter sepsis patient cohort. Multi-level Cox regression was used to investigate the relation between patient characteristics and pathogen features (2298 accessory genes, 1078 core protein levels, 107 parsimony-informative variations in reported virulence factors) with 30-day mortality. Our analysis revealed that presence of the helP gene encoding a putative DEAD-box helicase was independently associated with a fatal outcome (hazard ratio 2.01, p = 0.05). helP is located within a region related to the pathogenicity island PAPI-1 in close proximity to a pil gene cluster, which has been associated with horizontal gene transfer. Besides helP, elevated protein levels of the bacterial flagellum protein FliL (hazard ratio 3.44, p < 0.001) and of a bacterioferritin-like protein (hazard ratio 1.74, p = 0.003) increased the risk of death, while high protein levels of a putative aminotransferase were associated with an improved outcome (hazard ratio 0.12, p < 0.001). The prognostic potential of biomarker candidates and clinical factors was confirmed with different machine learning approaches using training and hold-out datasets. The helP genotype appeared the most attractive biomarker for clinical risk stratification due to its relevant predictive power and ease of detection.

microbiology

Genome-Wide Identification of Early-Firing Human Replication Origins by Optical Replication Mapping

The timing of DNA replication is largely regulated by the location and timing of replication origin firing. Therefore, much effort has been invested in identifying and analyzing human replication origins. However, the heterogeneous nature of eukaryotic replication kinetics and the low efficiency of individual origins in metazoans has made mapping the location and timing of replication initiation in human cells difficult. We have mapped early-firing origins in HeLa cells using Optical Replication Mapping, a high-throughput single-molecule approach based on Bionano Genomics genomic mapping technology. The single-molecule nature and 290-fold coverage of our dataset allowed us to identify origins that fire with as little as 1% efficiency. We find sites of human replication initiation in early S phase are not confined to well-defined efficient replication origins, but are instead distributed across broad initiation zones consisting of many inefficient origins. These early-firing initiation zones co-localize with initiation zones inferred from Okazaki-fragment-mapping analysis and are enriched in ORC1 binding sites. Although most early-firing origins fire in early-replication regions of the genome, a significant number fire in late-replicating regions, suggesting that the major difference between origins in early and late replicating regions is their probability of firing in early S-phase, as opposed to qualitative differences in their firing-time distributions. This observation is consistent with stochastic models of origin timing regulation, which explain the regulation of replication timing in yeast.

genomics