Search bioRxivSearch

Biology subjects

Watson, M.

Publications and source records attributed to Watson, M..

9 recordsLinked to original sources

Open prediction of polysaccharide utilisation loci (PUL) in 5414 public Bacteroidetes genomes using PULpy

Polysaccharide utilisation loci (PUL) are regions within bacterial genomes that encode all the necessary machinery for the cleavage of particular carbohydrates. For the Bacteroidetes phylum, prediction of PUL from genomic data alone involves the identification of carbohydrate-active enzymes (CAZymes) co-localised with susCD gene pairs. Here we present the open prediction of PUL in 5414 public Bacteroidetes genomes, and an open-source pipeline to reproduce or extend the results.

microbiology

Visualisation and analysis of RNA-Seq assembly graphs

RNA-sequencing (RNA-Seq) is a powerful transcriptome profiling technology enabling transcript discovery and quantification. RNA-Seq data are large, and most commonly used as a source of genelevel quantification measurements, whilst the underlying assemblies of reads, if inspected, are usually viewed as sequence reads mapped on to a reference genome. Whilst sufficient for many needs, when the underlying transcript assemblies are complex, this visualisation approach can be limiting; errors in assembly can be difficult to spot and interpretation of splicing events is challenging.\n\nHere we report on the development of a graph-based visualisation method as a complementary approach to understanding transcript diversity and read assembly from short-read RNA-Seq data. Following the mapping of reads to the reference genome, read-to-read comparison is performed on all reads mapping to a given gene, producing a matrix of weighted similarity scores between reads. This is used to produce an RNA assembly graph where nodes represent reads derived from a cDNA and edges similarity scores between reads, above a defined threshold. Visualisation of resulting graphs is performed using Graphia Professional. This tool can render the often large and complex graph topologies that result from DNA/RNA sequence assembly in 3D space and supports info rmatio no verlay on to nodes, e.g. transcript models. We have also implemented an analysis pipeline for the creation of RNA assembly graphs with both a command-line and web-based interface that allows users to create and visualise these data. Here we demonstrate the utility of this approach on RNA-Seq data, including the unusual structure of these graphs and how they can be used to identify issues in assembly, repetitive sequences within transcripts and splice variants. We believe this approach has the potential to significantly improve our understanding of transcript complexity.

bioinformatics

A conserved role of the insulin-like signaling pathway in uric acid pathologies revealed in Drosophila melanogaster

Elevated uric acid (UA) is a key factor for disorders, including gout or kidney stones and result from abrogated expression of Urate Oxidase (Uro) and diet. To understand the genetic pathways influencing UA metabolism we established a Drosophila melanogaster model with elevated UA using Uro knockdown. Reduced Uro expression resulted in the accumulation of UA concretions and diet-dependent shortening of lifespan. Inhibition of insulin-like signaling (ILS) pathway genes reduced UA and concretion load. In humans, SNPs in the ILS genes AKT2 and FOXO3 were associated with UA levels or gout, supporting a conserved role for ILS in modulating UA metabolism. Downstream of the ILS pathway UA pathogenicity was mediated partly by NADPH Oxidase, whose inhibition attenuated the reduced lifespan and concretion accumulation. Thus, genes in the ILS pathway represent potential therapeutic targets for treating UA associated pathologies, including gout and kidney stones.\n\nHighlightsO_LIIn Drosophila high uric acid (UA) levels shorten lifespan and cause UA aggregation\nC_LIO_LIConserved in flies and humans, the ILS pathway associates with UA pathologies\nC_LIO_LIFoxO dampens concretion formation by reducing UA levels and ROS formation\nC_LIO_LIInhibition of NOX alleviates the lifespan attenuation and UA aggregation\nC_LI

pathology

Meta-analysis of 1,200 transcriptomic profiles identifies a prognostic model for pancreatic ductal adenocarcinoma

BackgroundWith a dismal 8% median 5-year overall survival (OS), pancreatic ductal adenocarcinoma (PDAC) is highly lethal. Only 10-20% of patients are eligible for surgery, and over 50% of these will die within a year of surgery. Identify molecular predictors of early death would enable the selection of PDAC patients at high risk.\n\nMethodsWe developed the Pancreatic Cancer Overall Survival Predictor (PCOSP), a prognostic model built from a unique set of 89 PDAC tumors where gene expression was profiled using both microarray and sequencing platforms. We used a meta-analysis framework based on the binary gene pair method to create gene expression barcodes robust to biases arising from heterogeneous profiling platforms and batch effects. Leveraging the largest compendium of PDAC transcriptomic datasets to date, we show that PCOSP is a robust single-sample predictor of early death ([≤]1 yr) after surgery in a subset of 823 samples with available transcriptomics and survival data.\n\nResultsThe PCOSP model was strongly and significantly prognostic with a meta-estimate of the area under the receiver operating curve (AUROC) of 0.70 (P=1.9e-18) and hazard ratio (HR) of 1.95(1.6-2.3) (P=2.6e-16) for binary and survival predictions, respectively. The prognostic value of PCOSP was independent of clinicopathological parameters and molecular subtypes. Over-representation analysis of the PCOSP 2619 gene-pairs (1070 unique genes) unveiled pathways associated with Hedgehog signalling, epithelial mesenchymal transition (EMT) and extracellular matrix (ECM) signalling.\n\nConclusionPCOSP could improve treatment decision by identifying patients who will not benefit from standard surgery/chemotherapy and may benefit from alternate approaches.\n\nAbbreviations

genomics

MAGpy: a reproducible pipeline for the downstream analysis of metagenome-assembled genomes (MAGs)

Recent advances in bioinformatics have enabled the rapid assembly of genomes from metagenomes (MAGs), and there is a need for reproducible pipelines that can annotate and characterise thousands of genomes simultaneously. Here we present MAGpy, a Snakemake pipeline that takes FASTA input and compares MAGs to several public databases, checks quality, assigns a taxonomy and draws a phylogenetic tree.

bioinformatics

Assembly of hundreds of microbial genomes from the cow rumen reveals novel microbial species encoding enzymes with roles in carbohydrate metabolism

The cow rumen is a specialised organ adapted for the efficient breakdown of plant material into energy and nutrients, and it is the rumen microbiome that encodes the enzymes responsible. Many of these enzymes are of huge industrial interest. Despite this, rumen microbes are under-represented in the public databases. Here we present 220 high quality bacterial and archaeal genomes assembled directly from 768 gigabases of rumen metagenomic sequence data. Comparative analysis with current publicly available genomes reveals that the majority of these represent previously unsequenced strains and species of bacteria and archaea. The genomes contain over 13,000 proteins predicted to be involved in carbohydrate metabolism, over 90% of which do not have a good match in the public databases. Inclusion of the 220 genomes presented here improves metagenomic read classification by 2-3-fold, both in our data and in other publicly available rumen datasets. This release improves the coverage of rumen microbes in the public databases, and represents a hugely valuable resource for biomass-degrading enzyme discovery and studies of the rumen microbiome

genomics

A High Resolution Atlas Of Gene Expression In The Domestic Sheep (Ovis aries)

Sheep are a key source of meat, milk and fibre for the global livestock sector, and an important biomedical model. Global analysis of gene expression across multiple tissues has aided genome annotation and supported functional annotation of mammalian genes. We present a large-scale RNA-Seq dataset representing all the major organ systems from adult sheep and from several juvenile, neonatal and prenatal developmental time points. The Ovis aries reference genome (Oar v3.1) includes 27,504 genes (20,921 protein coding), of which 25,350 (19,921 protein coding) had detectable expression in at least one tissue in the sheep gene expression atlas dataset. Network-based cluster analysis of this dataset grouped genes according to their expression pattern. The principle of guilt by association was used to infer the function of uncharacterised genes from their co-expression with genes of known function. We describe the overall transcriptional signatures present in the sheep gene expression atlas and assign those signatures, where possible, to specific cell populations or pathways. The findings are related to innate immunity by focusing on clusters with an immune signature, and to the advantages of cross-breeding by examining the patterns of genes exhibiting the greatest expression differences between purebred and crossbred animals. This high-resolution gene expression atlas for sheep is, to our knowledge, the largest transcriptomic dataset from any livestock species to date. It provides a resource to improve the annotation of the current reference genome for sheep, presenting a model transcriptome for ruminants and insight into gene, cell and tissue function at multiple developmental stages.\n\nAuthor SummarySheep are ruminant mammals kept as livestock for the production of meat, milk and wool in agricultural industries across the globe. Genetic and genomic information can be used to improve production traits such as disease resiliance. The sheep genome is however missing important information relating to gene function and many genes, which may be important for productivity, have no informative gene name. This can be remedied using RNA-Sequencing to generate a global expression profile of all protein-coding genes, across multiple organ systems and developmental stages. Clustering genes based on their expression profile across tissues and cells allows us to assign function to those genes. If for example a gene with no informative gene name is expressed in macrophages and is found within a cluster of known macrophage related genes it is likely to be involved in macrophage function and play a role in innate immunity. This information improves the quality of the reference genome and provides insight into biological processes underlying the complex traits that influence the productivity of sheep and other livestock species.

genomics

poRe GUIs for parallel and real-time processing of MinION sequence data

MotivationOxford Nanopores MinION device has matured rapidly and is now capable of producing over one million reads and several gigabases of sequence data per run. The nature of the MinION output requires new tools that are easy to use by scientists with a range of computational skills and which enable quick and simple QC and data extraction from MinION runs.\n\nResultsWe have developed two GUIs for the R package poRe that allow parallel and real-time processing of MinION datasets. Both GUIs are capable of extracting sequence- and meta- data from large MinION datasets via a friendly point-and-click interface using commodity hardware.\n\nAvailabilityThe GUIs are packaged within poRe which is available on SourceForge: https://source-forge.net/projects/rpore/files/. Documentation is available on GitHub: https://github.com/mw55309/poRe_docs

bioinformatics