Search bioRxivSearch

Biology subjects

Laurent Mesnard

Publications and source records attributed to Laurent Mesnard.

2 recordsLinked to original sources

Adaptive Somatic Mutations Calls with Deep Learning and Semi-Simulated Data

A number of approaches have been developed to call somatic variation in high-throughput sequencing data. Here, we present an adaptive approach to calling somatic variations. Our approach trains a deep feed-forward neural network with semi-simulated data. Semi-simulated datasets are constructed by planting somatic mutations in real datasets where no mutations are expected. Using semi-simulated data makes it possible to train the models with millions of training examples, a usual requirement for successfully training deep learning models. We initially focus on calling variations in RNA-Seq data. We derive semi-simulated datasets from real RNA-Seq data, which offer a good representation of the data the models will be applied to. We test the models on independent semi-simulated data as well as pure simulations. On independent semi-simulated data, models achieve an AUC of 0.973. When tested on semi-simulated exome DNA datasets, we find that the models trained on RNA-Seq data remain predictive (sens 0.4 & spec 0.9 at cutoff of P > = 0.9), albeit with lower overall performance (AUC=0.737). Interestingly, while the models generalize across assay, training on RNA-Seq data lowers the confidence for a group of mutations. Haloplex exome specific training was also performed, demonstrating that the approach can produce probabilistic models tuned for specific assays and protocols. We found that the method adapts to the characteristics of experimental protocol. We further illustrate these points by training a model for a trio somatic experimental design when germline DNA of both parents is available in addition to data about the individual. These models are distributed with Goby (http://goby.campagnelab.org).

Bioinformatics

Exome Sequencing and Prediction of Long-Term Kidney Allograft Function

AbstractCurrent strategies to improve graft outcome following kidney transplantation consider information at the HLA loci. Here, we used exome sequencing of DNA from ABO compatible kidney graft recipients and their living donors to determine recipient and donor mismatches at the amino acid level over entire exomes. We estimated the number of amino acid mismatches in transmembrane proteins, more likely to be seen as foreign by the recipients immune system, and designated this tally as the allogenomics mismatch score (AMS). The AMS can be measured prior to transplantation with DNA for potential donor and recipient pairs. We examined the degree of relationship between the AMS and post-transplantation kidney allograft function by linear regression. In a discovery cohort, we found a significant inverse correlation between the AMS and kidney graft function at 36 months post-transplantation (n=10 recipient/donor pairs; 20 exomes) (r2>=0.57, P<0.05). The predictive ability of the AMS persists when the score is restricted to regions outside of the HLA loci. This relationship was validated using an independent cohort of 24 recipient donor pairs (n=48 exomes) (r2>=0.39, P<0.005). In an additional cohort of living and mostly intra-familial recipient/donor pairs (n=19, 38 exomes), we validated the association after controlling for donor age at time of transplantation. Finally, a model that controls for donor age, HLA mismatches and time post-transplantation yields a consistent AMS effect across these three independent cohorts (P<0.05). Taken together, these results show that the AMS is a strong predictor of long-term graft function in kidney transplant recipients.\n\nOne Sentence SummaryPrediction of long-term kidney graft function with exome sequencing

Genomics