Search bioRxivSearch

Biology subjects

Washburne, A.

Publications and source records attributed to Washburne, A..

2 recordsLinked to original sources

Phylogenetic factorization of mammalian viruses complements trait-based analyses and guides surveillance efforts

Predicting which novel microorganisms may spill over from animals to humans has become a major priority in infectious disease biology. However, there are few tools to help assess the zoonotic potential of the enormous number of potential pathogens, the majority of which are undiscovered or unclassified and may be unlikely to infect or cause disease in humans. We adapt a new biological machine learning technique - phylofactorization - to partition viruses into clades based on their non-human host range and whether or not there exist evidence they have infected humans. Our cladistic analyses identify clades of viruses with common within-clade patterns - unusually high or low propensity for spillover. Phylofactorization by spillover yields many clades of viruses containing few to no representatives that have spilled over to humans, including the families Papillomaviridae and Herpesviridae, and the genus Parvovirus. Removal of these non-zoonotic clades from previous trait-based analyses changed the relative significance of traits determining spillover due to strong associations of traits with non-zoonotic clades. Phylofactorization by host breadth yielded clades with unusually high host breadth, including the family Togaviridae. We identify putative life-history traits differentiating clades host breadth and propensities for zoonosis, and discuss how these results can prioritize sequencing-based surveillance of emerging infectious diseases.

epidemiology

Phylofactorization - theory and challenges

Data from biological communities are composed of species connected by the phylogeny. A greedy algorithm phylofactorization - was developed to construct an isometric log-ratio transform whose balances correspond to edges along which traits arose, controlling for previously made inferences.\n\nIn this paper, the general theory of phylofactorization is presented as a graph-partitioning algorithm. A special case-regression phylofactorization-chooses coordinates based on sequential maximization of objective functions from regression on \"contrast\" variables such as an isometric log-ratio transform. The connections between regression phylofactorization and other methods is discussed, including matrix factorization, hierarchical regression, factor analysis and latent variable models. Open challenges in the statistical analysis of phylofactorization are presented, including criteria for choosing the number of factors and approximating null-distributions of commonly used test statistics and objective functions. As a graph-partitioning algorithm, cross-validation of phylo factorization across datasets requires graph-topological considerations, such as how to deal with novel nodes and edges and whether or not to control for partition order. Overcoming these challenges can accelerate our analysis of phylogenetically-structured data and allow annotations of edges in an online tree of life.

evolutionary biology