Search bioRxivSearch

Biology subjects

Lio, P.

Publications and source records attributed to Lio, P..

5 recordsLinked to original sources

A parameter-efficient deep learning approach to predict conversion from mild cognitive impairment to Alzheimer’s disease within three years

Some forms of mild cognitive impairment (MCI) can be the clinical precursor of severe dementia like Alzheimers disease (AD), while other types of MCI tend to remain stable over-time and do not progress to AD pathology. To choose an effective and personalized treatment for AD, we need to identify which MCI patients are at risk of developing AD and which are not.\n\nHere, we present a novel deep learning architecture, based on dual learning and an ad hoc layer for 3D separable convolutions, which aims at identifying those people with MCI who have a high likelihood of developing AD. Our deep learning procedures combine structural magnetic resonance imaging (MRI), demographic, neuropsychological, and APOe4 genotyping data as input measures. The most novel characteristics of our machine learning model compared to previous ones are as follows: 1) multi-tasking, in the sense that our deep learning model jointly learns to simultaneously predict both MCI to AD conversion, and AD vs healthy classification which facilitates the relevant feature extraction for prognostication; 2) the neural network classifier employs relatively few parameters compared to other deep learning architectures (we use ~550,000 network parameters, orders of magnitude lower than other network designs) without compromising network complexity and hence significantly limits data-overfitting; 3) both structural MRI images and warp field characteristics, which quantify the amount of volumetric change compared to the common template, were used as separate input streams to extract as much information as possible from the MRI data. All the analyses were performed on a subset of the Alzheimers Disease Neuroimaging Initiative (ADNI) database, for a total of n=785 participants (192 AD, 409 MCI, and184 healthy controls (HC)).\n\nWe found that the most predictive combination of inputs included the structural MRI images and the demographic, neuropsychological, and APOe4 data, while the warp field metric added little predictive value. We achieved an area under the ROC curve (AUC) of 0.925 with a 10-fold cross-validated accuracy of 86%, a sensitivity of 87.5% and specificity of 85% in classifying MCI patients who developed AD in three years time from those individuals showing stable MCI over the same time-period. To the best of our knowledge, this is the highest performance reported on a test set achieved in the literature using similar data. The same network provided an AUC of 1 and 100% accuracy, sensitivity and specificity when classifying NC from AD. We also demonstrated that our classification framework was robust to different co-registration templates and possibly irrelevant features / image sections.\n\nOur approach is flexible and can in principle integrate other imaging modalities, such as PET, and a more diverse group of clinical data. The convolutional framework is potentially applicable to any 3D image dataset and gives the flexibility to design a computer-aided diagnosis system targeting the prediction of any medical condition utilizing multi-modal imaging and tabular clinical data.

bioinformatics

GEESE: Metabolically driven latent space learning for gene expression data

Gene expression microarrays provide a characterisation of the transcriptional activity of a particular biological sample. Their high dimensionality hampers the process of pattern recognition and extraction. Several approaches have been proposed for gleaning information about the hidden structure of the data. Among these approaches, deep generative models provide a powerful way for approximating the manifold on which the data reside.\n\nHere we develop GEESE, a deep learning based framework that provides novel insight into the manifold learning for gene expression data, employing a metabolic model to constrain the learned representation. We evaluated the proposed framework, showing its ability to capture biologically relevant features, and encoding that features in a much simpler latent space. We showed how using a metabolic model to drive the autoencoder learning process helps in achieving better generalisation to unseen data. GEESE provides a novel perspective on the problem of unsupervised learning for biological data.\n\nAvailabilitySource code of GEESE is available at https://bitbucket.org/mbarsacchi/geese/.

systems biology

Network-Based Biomarkers Enable Cross-Disease Biomarker Discovery

Biomarkers lie at the heart of precision medicine, biodiversity monitoring, agricultural pathogen detection, amongst others. Surprisingly, while rapid genomic profiling is becoming ubiquitous, the development of biomarkers almost always involves the application of bespoke techniques that cannot be directly applied to other datasets. There is an urgent need for a systematic methodology to create biologically-interpretable molecular models that robustly predict key phenotypes. We therefore created SIMMS: an algorithm that fragments pathways into functional modules and uses these to predict phenotypes. We applied SIMMS to multiple data-types across four diseases, and in each it reproducibly identified subtypes, made superior predictions to the best bespoke approaches, and identified known and novel signaling nodes. To demonstrate its ability on a new dataset, we measured 33 genes/nodes of the PIK3CA pathway in 1,734 FFPE breast tumours and created a four-subnetwork prediction model. This model significantly out-performed existing clinically-used molecular tests in an independent 1,742-patient validation cohort. SIMMS is generic and can work with any molecular data or biological network, and is freely available at: https://cran.r-project.org/web/packages/SIMMS.

bioinformatics

Paratope Prediction using Convolutional and Recurrent Neural Networks

Antibodies play an essential role in the immune system of vertebrates and are vital tools in research and diagnostics. While hypervariable regions of antibodies, which are responsible for binding, can be readily identified from their amino acid sequence, it remains challenging to accurately pinpoint which amino acids will be in contact with the antigen (the paratope). In this work, we present a sequence-based probabilistic machine learning algorithm for paratope prediction, named Parapred. Parapred uses a deep-learning architecture to leverage features from both local residue neighbourhoods and across the entire sequence. The method outperforms the current state-of-the-art methodology, and only requires a stretch of amino acid sequence corresponding to a hypervariable region as an input, without any information about the antigen. We further show that our predictions can be used to improve both speed and accuracy of a rigid docking algorithm. The Parapred method is freely available at https://github.com/eliberis/parapred for download.

bioinformatics

Single-cell RNA-Sequencing uncovers transcriptional states and fate decisionsin haematopoiesis

The success of marker-based approaches for dissecting haematopoiesis in mouse and human is reliant on the presence of well-defined cell-surface markers specific for diverse progenitor populations. An inherent problem with this approach is that the presence of specific cell surface markers does not directly reflect the transcriptional state of a cell. Here we used a marker-free approach to computationally reconstruct the blood lineage tree in zebrafish and order cells along their differentiation trajectory, based on their global transcriptional differences. Within the population of transcriptionally similar stem and progenitor cells our analysis revealed considerable cell-to-cell differences in their probability to transition to another, committed state. Once fate decision was executed, the suppression of transcription of ribosomal genes and up-regulation of lineage specific factors coordinately controlled lineage differentiation. Evolutionary analysis further demonstrated that this haematopoietic program was highly conserved between zebrafish and higher vertebrates.

cell biology