Search bioRxiv⌕ Search

Biology subjects

Ashford, J.

Publications and source records attributed to Ashford, J..

2 recordsLinked to original sources

Estimated limits of organism-specific training for epitope prediction

BackgroundThe identification of linear B-cell epitopes remains an important task in the development of vaccines, therapeutic antibodies and several diagnostic tests. Machine learning predictors are trained to flag potential epitope candidates for experimental validation and currently, most predictors are trained as generalist models using large, heterogeneous data sets. Recently, organism-specific training has been shown to improve prediction performance for data-rich organisms. Unfortunately, for most organisms, large volumes of validated epitope data are not yet available. This article investigates the limits of organism-specific training for epitope prediction. It explores the validity of organism-specific training for data-poor organisms by examining how the size of the training data set affects prediction performance. It also compares the performance of organism-specific training under simulated data-poor conditions to that of models trained using traditional large heterogeneous and hybrid data sets. ResultsThis work shows how models trained on small organism-specific data sets can outperform similar models trained on (potentially much larger) heterogeneous and mixed data sets. The results reported indicate that as few as 20 labelled peptides from a given pathogen can be sufficient to generate models that outperform widely-used predictors from the literature, which are trained on heterogeneous data. Models trained using more than about 100 to 150 organism-specific peptides perform consistently better than most generalist models across a wide variety of performance measures, and in some cases can even approach the performance of organism-specific models trained on considerably larger data sets. ConclusionsOrganism-specific training improves linear B-cell epitope prediction performance even in situations when only small training sets are available, which opens new possibilities for the development of bespoke, high-performance predictive models when studying data-poor organisms such as emerging or neglected pathogens.

bioinformatics↗

Multi-objective prioritisation of candidate epitopes for diagnostic test development

BackgroundThe development of peptide-based diagnostic tests requires the identification of epitopes that are at the same time highly immunogenic and, ideally, unique to the pathogen of interest, to minimise the chances of cross-reactivity. Existing computational pipelines for the prediction of linear B-cell epitopes tend to focus exclusively on the first objective, leaving considerations of cross-reactivity to later stages of test development. ResultsWe present a multi-objective approach to the prioritisation of candidate epitopes for experimental validation, in the context of diagnostic test development. The dual objectives of uniqueness (measured as dissimilarity from known epitope sequences from other pathogens) and predicted immunogenicity (measured as the probability score returned by the prediction model) are considered simultaneously. Validation was performed using data from three distinct pathogens (namely the nematode Onchocerca volvulus, the Epstein-Barr Virus and the Hepatitis C Virus), with predictions derived using an organism-specific prediction approach. The multi-objective rankings returned sets of non-dominated solutions as potential targets for the development of diagnostic tests with lower probability of false positives due to cross-reactivity. ConclusionsThe application of the proposed approach to three test pathogens led to the identification of 20 new potential epitopes, with both high probability and a high degree of exclusivity to the target organisms. The results indicate the potential of the proposed approach to provide enhanced filtering and ranking of potential candidates, highlighting potential cross-reactivities and including this information into the test development process right from the target identification and prioritisation step.

bioinformatics↗