Search bioRxiv⌕ Search

Biology subjects

Kabiljo, R.

Publications and source records attributed to Kabiljo, R..

4 recordsLinked to original sources

Molecular dynamics analysis of Superoxide Dismutase 1 mutations suggests decoupling between mechanisms underlying ALS onset and progression

Mutations in the superoxide dismutase 1 (SOD1) gene are the second most common known cause of ALS. SOD1 variants express high phenotypic variability and over 200 have been reported in people with ALS. Investigating how different SOD1 variants affect the protein dynamics might help in understanding their pathogenic mechanism and explaining their heterogeneous clinical presentation. It was previously proposed that variants can be broadly classified in two groups, wild-type like (WTL) and metal binding region (MBR) variants, based on their structural location and biophysical properties. MBR variants are associated with a loss of SOD1 enzymatic activity. In this study we used molecular dynamics and large clinical datasets to characterise the differences in the structural and dynamic behaviour of WTL and MBR variants with respect to the wild-type SOD1, and how such differences influence the ALS clinical phenotype. Our study identified marked structural differences, some of which are observed in both variant groups, while others are group specific. Moreover, applying graph theory to a network representation of the proteins, we identified differences in the intramolecular contacts of the two classes of variants. Finally, collecting clinical data of approximately 500 SOD1 ALS patients carrying variants from both classes, we showed that the survival time of patients carrying an MBR variant is generally longer (~6 years median difference, p < 0.001) with respect to patients with a WTL variant. In conclusion, our study highlights key differences in the dynamic behaviour of the WTL and MBR SOD1 variants, and wild-type SOD1 at an atomic and molecular level. We identified interesting structural features that could be further investigated to explain the associated phenotypic variability. Our results support the hypothesis of a decoupling between mechanisms of onset and progression of SOD1 ALS, and an involvement of loss-of-function of SOD1 with the disease progression.

neuroscience↗

DNAscan2: a versatile, scalable, and user-friendly analysis pipeline for next-generation sequencing data

The current widespread adoption of next-generation sequencing (NGS) in all branches of basic and clinical genetics fields means that users with highly variable informatics skills, computing facilities and application purposes need to process, analyse, and interpret NGS data. In this landscape, versatility, scalability, and user-friendliness are key characteristics for an NGS analysis tool. We developed DNAscan2, a highly flexible, end-to-end pipeline for the analysis of NGS data, which (i) can be used for the detection of multiple variant types, including SNVs, small indels, transposable elements, short tandem repeats and other large structural variants; (ii) covers all steps of the analysis, from quality control of raw data to the generation of html reports for the interpretation and prioritisation of results; (iii) is highly adaptable and scalable as it can be deployed and run via either a graphic user interface for non-bioinformaticians, a command line tool for personal computer usage, or as a Snakemake workflow that facilitates parallel multi-sample execution for high-performance computing environments; (iv) is computationally efficient by minimising RAM and CPU time requirements. Availability and ImplementationDNAscan2 is implemented in Python3 and is available to download as a command-line tool and graphical-user interface at https://github.com/KHP-Informatics/DNAscanv2 or a Snakemake workflow at https://github.com/KHP-Informatics/DNAscanv2_snakemake.

bioinformatics↗

RetroSnake: a Modular End-to-End Pipeline for Detection of Human Endogenous Retrovirus (HERV) Transposable Elements in Next Generation Sequencing (NGS) Data

Human Endogenous Retroviruses (HERVs) integrated into the genome of vertebrates as a result of ancient exogenous infections and currently comprise [~]8% of our genome. The majority of these elements have accumulated mutations rendering them inactive. The most recently acquired members, HERV-K have potential to produce viral particles and have been linked to a wide range of diseases including cancer and neurodegeneration. Although a range of tools for HERV discovery exist, most of them lack wet-lab validation of their results and are not end-to-end as they do not cover all steps of the analysis. These factors greatly limit their use. Here we describe RetroSnake, an end-to-end, modular, computationally efficient and customisable pipeline for the discovery of HERVs in short-read NGS data. RetroSnake presents important advantages with respect to other available tools. For instance, it is the only pipeline based on an extensively wet-lab validated protocol, and it is the most complete transposable elements detection pipeline, producing annotated insertions presented as an interactive html file, easy enough to use by life scientists without substantial computational training. Availability and implementationThe Pipeline and an extensive documentation are available at https://github.com/KHP-Informatics/RetroSnake Contactalfredo.iacoangeli@kcl.ac.uk

genomics↗

An assessment of bioinformatics tools for the detection of human endogenous retroviral insertions in short-read genome sequencing data

There is a growing interest in the study of human endogenous retroviruses (HERVs) given the substantial body of evidence that implicates them in many human diseases. Although their genomic characterization presents numerous technical challenges, next-generation sequencing (NGS) has shown potential to detect HERV insertions and their polymorphisms in humans, and a number of computational tools to detect them in short-read NGS data exist. In order to design optimal analysis pipelines, an independent evaluation of the currently available tools is required. We evaluated the performance of a set of such tools using a variety of experimental designs and types of NGS datasets. These included 50 human short read whole-genome sequencing samples, matching long and short read NGS data, and simulated short-read NGS data. Our results highlight the performance variability of the tools across the datasets and suggest that different tools might be suitable for different study designs. Using multiple tools and a consensus approach is advisable if computationally feasible and wet-lab validation via PCR is advisable where biological samples are available.

bioinformatics↗