Search bioRxivSearch

Biology subjects

Jacob Machado, D.

Publications and source records attributed to Jacob Machado, D..

2 recordsLinked to original sources

Evidence of Absence Treated as Absence of Evidence: The Effects of Variation in the Number and Distribution of Gaps Treated as Missing Data on the Results of Standard Maximum Likelihood Analysis

We evaluated the effects of variation in the number and distribution of gaps (i.e., no base; coded as IUPAC \".\" or \"-\") treated as missing data (i.e., any base, coded as \"?\" or IUPAC \"N\") in standard maximum likelihood (ML) analysis. We obtained alignments with variable numbers and arrangements of gaps by aligning seven diverse empirical datasets under different gap opening costs using MAFFT. We selected the optimal substitution model for each alignment using the corrected Akaike Information Criterion (AICc) in jModelTest2 and searched for the optimal trees for each alignment using default search parameters and the selected models in GARLI. We also employed a Monte Carlo approach to randomly insert gaps (treated as missing data) into an empirical dataset to understand more precisely the effects of their variable numbers and distributions. To compare alignments quantitatively, we used several measures to quantify the number and distribution of gaps in all alignments (e.g., alignment length, total number of gaps, total number of characters containing gaps, number of gap openings). We then used these variables to derive four indices (ranging from 0 to 1) that summarize the distribution of gaps both within and among terminals, including an index that takes into account their optimization on the tree. Our most important observation is that ML scores correlate negatively with gap opening costs, and the amount of missing data. These variables also cause unpredictable effects on tree topologies. We discuss the implications of our results for the traditional and tree-alignment approaches in ML.

evolutionary biology

Enhanced genome annotation strategy provides novel insights on the phylogeny of Flaviviridae

The ongoing and severe public health threat of viruses of the family Flaviviridae, including dengue, hepatitis C, West Nile, yellow fever, and zika, demand a greater understanding of how these viruses evolve, emerge and spread in order to respond. Central to this understanding is an updated phylogeny of the entire family. Unfortunately, most cladograms of Flaviviridae focus on specific lineages, ignore outgroups, and rely on midpoint rooting, hampering their ability to test ingroup monophyly and estimate ingroup relationships. This problem is partly due to the lack of fully annotated genomes of Flaviviridae, which has genera with slightly different gene content, hindering genome analysis without partitioning. To tackle these problems, we developed an annotation pipeline for Flaviviridae that uses a combination of ab initio and homology-based strategies. The pipeline recovered 100% of the genes in reference genomes and annotated over 97% of the expected genes in the remaining non curated sequences. We further demonstrate that the combined analysis of genomes of all genera of Flaviviridae (Flavivirus, Hepacivirus, Pegivirus, and Pestivirus), as made possible by our annotation strategy, enhances the phylogenetic analyses of these viruses for all optimality criteria that we tested (parsimony, maximum likelihood, and posterior probability). The final tree sheds light on the phylogenetic relationship of viruses that are divergent from most Flaviviridae and should be reclassified, especially the soybean cyst nematode virus 5 (SbCNV-5) and the Tamana bat virus. We also corroborate the close phylogenetic relationship of dengue and zika viruses with an unprecedented degree of support.

evolutionary biology