Search bioRxiv⌕ Search

Biology subjects

Makau, D. N.

Publications and source records attributed to Makau, D. N..

2 recordsLinked to original sources

Machine learning approaches for estimating cross-neutralization potential among FMD serotype O viruses

In this study, we aimed to develop an algorithm that uses sequence data to estimate cross-neutralization between serotype O foot-and-mouth disease viruses (FMDV) based on r1 values, while identifying key genomic sites associated with high or low r1 values. The ability to estimate cross-neutralization potential among co-circulating FMDVs in silico is significant for vaccine developers, animal health agencies making herd immunization decisions, and disease preparedness. Using published data on virus neutralization titer (VNT) assays and associated VP1 sequences from GenBank, we applied machine learning algorithms (BORUTA and random forest) to predict potential cross-reaction between serum/vaccine-virus pairs for 73 distinct serotype O FMDV strains. Model optimization involved tenfold cross-validation and sub-sampling to address data imbalance and improve performance. Model predictors included amino acid distances, site-wise amino acid polymorphisms, and differences in potential N-glycosylation sites. The dataset comprised 108 observations (serum-virus pairs) from 73 distinct viruses with r1 values. Observations were dichotomized using a 0.3 threshold, yielding putative non-cross-neutralizing (< 0.3 r1 values) and cross-neutralizing groups ([&ge;] 0.3 r1 values). The best model had a training accuracy, sensitivity, and specificity of 0.96 (95% CI: 0.88-0.99), 0.93, and 0.96, respectively, and an accuracy of 0.94 (95% CI: 0.71-1.00), sensitivity of 1.00, and specificity of 0.93, positive, and negative predictive values of 0.60 and 1.00, respectively, on one testing dataset and an accuracy, AUC, sensitivity, specificity, and predictive values all approaching 1.00 on a second testing dataset. Additionally, amino acid positions 48, 100, 135, 150, and 151 in the VP1 region alongside amino acid distance were found to be important predictors of cross-neutralization. Our study highlights the value of genetic/genomic data for informing immunization strategies in disease management and understanding potential immune-mediated competition amongst related endemic strains of serotype O FMDVs in the field. We also showcase leveraging routinely generated sequence data and applying a parsimonious machine learning model to expedite decision-making in selection of vaccine candidates and application of vaccines for controlling FMD, particularly serotype O. A similar approach can be applied to other serotypes.

immunology↗

Phylogenetic-based methods for fine-scale classification of PRRSV-2 ORF5 sequences: a comparison of their robustness and reproducibility

Disease management and epidemiological investigations of porcine reproductive and respiratory syndrome virus-type 2 (PRRSV-2) often rely on grouping together highly related sequences. In the USA, the last five years have seen a major paradigm shift within the swine industry when classifying PRRSV-2, beginning to move away from RFLP (restriction fragment length polymorphisms)-typing and adopting the use of phylogenetic lineage-based classification. However, lineages and sub-lineages are large and genetically diverse, and the rapid mutation rate of PRRSV coupled with the global prevalence of the disease has made it challenging to identify new and emerging variants. Thus, within the lineage system, a dynamic fine-scale classification scheme is needed to provide better resolution on the relatedness of PRRSV-2 viruses to inform disease management and monitoring efforts and facilitate research and communication surrounding circulating PRRSV viruses. Here, we compare potential fine-scale systems for classifying PRRSV-2 variants (i.e., genetic clusters of closely related ORF5 sequences at finer scales than sub-lineage) using a database of 28,730 sequences from 2010 to 2021, representing >55% of the U.S. pig population. In total, we compared 140 approaches that differed in their tree-building method, criteria, and thresholds for defining variants within phylogenetic trees using TreeCluster. Three approaches produced epidemiologically meaningful variants (i.e., [&ge;]5 sequences per cluster), and resulted in reproducible and robust outputs even when the input data or input phylogenies were changed. In the three best performing approaches, the average genetic distance amongst sequences belonging to the same variant was 2.1 - 2.5%, and the genetic divergence between variants was 2.5-2.7%. Machine learning classification algorithms were also trained to assign new sequences to an existing variant with >95% accuracy, which shows that newly generated sequences could be assigned without repeating the phylogenetic and clustering analyses. Finally, we identified 73 sequence-clusters (dated <1 year apart with close phylogenetic relatedness) associated with circulation events on single farms. The percent of farm sequence-clusters with an ID change was 6.5-8.7% for our best approaches. In contrast, [~]43% of farm sequence-clusters had variation in their RFLP-type, further demonstrating how our proposed fine-scale classification system addresses shortcomings of RFLP-typing. Through identifying robust and reproducible classification approaches for PRRSV-2, this work lays the foundation for a fine-scale system that would more reliably group related field viruses and provide better improved clarity for decision-making surrounding disease management.

bioinformatics↗