Search bioRxivSearch

Biology subjects

Soltis, D. E.

Publications and source records attributed to Soltis, D. E..

4 recordsLinked to original sources

For comparing phylogenetic diversity among communities, go ahead and use synthesis phylogenies

Should we build our own phylogenetic trees based on gene sequence data, or can we simply use available synthesis phylogenies? This is a fundamental question that any study involving a phylogenetic framework must face at the beginning of the project. Building a phylogeny from gene sequence data (purpose-built phylogeny) requires more effort, expertise, and cost than subsetting an already available phylogeny (synthesis-based phylogeny). However, we still lack a comparison of how these two approaches to building phylogenetic trees influence common community phylogenetic analyses such as comparing community phylogenetic diversity and estimating trait phylogenetic signal. Here, we generated three purpose-built phylogenies and their corresponding synthesis-based trees (two from Phylomatic and one from the Open Tree of Life [OTL]). We simulated 1,000 communities and 12,000 continuous traits along each purpose-built phylogeny. We then compared the effects of different trees on estimates of phylogenetic diversity (alpha and beta) and phylogenetic signal (Pagels {lambda} and Blombergs K). Synthesis-based phylogenies generally yielded higher estimates of phylogenetic diversity when compared to purpose-built phylogenies. However, resulting measures of phylogenetic diversity from both types of phylogenies were highly correlated (Spearmans{rho} > 0.8 in most cases). Mean pairwise distance (both alpha and beta) is the index that is most robust to the differences in tree construction that we tested. Measures of phylogenetic diversity based on the OTL showed the highest correlation with measures based on the purpose-built phylogenies. Trait phylogenetic signal estimated with synthesis-based phylogenies, especially from the OTL, were also highly correlated with estimates of Blombergs K or close to Pagels {lambda} from purpose-built phylogenies when traits were simulated under Brownian Motion. For commonly employed community phylogenetic analyses, our results justify taking advantage of recently developed and continuously improving synthesis trees, especially the Open Tree of Life.

ecology

A Universal Probe Set for Targeted Sequencing of 353 Nuclear Genes from Any Flowering Plant Designed Using k-medoids Clustering

Sequencing of target-enriched libraries is an efficient and cost-effective method for obtaining DNA sequence data from hundreds of nuclear loci for phylogeny reconstruction. Much of the cost associated with developing targeted sequencing approaches is preliminary data needed for identifying orthologous loci for probe design. In plants, identifying orthologous loci has proven difficult due to a large number of whole-genome duplication events, especially in the angiosperms (flowering plants). We used multiple sequence alignments from over 600 angiosperms for 353 putatively single-copy protein-coding genes to design a set of targeted sequencing probes for phylogenetic studies of any angiosperm lineage. To maximize the phylogenetic potential of the probes while minimizing the cost of production, we introduce a k-medoids clustering approach to identify the minimum number of sequences necessary to represent each coding sequence in the final probe set. Using this method, five to 15 representative sequences were selected per orthologous locus, representing the sequence diversity of angiosperms more efficiently than if probes were designed using available sequenced genomes alone. To test our approximately 80,000 probes, we hybridized libraries from 42 species spanning all higher-order lineages of angiosperms, with a focus on taxa not present in the sequence alignments used to design the probes. Out of a possible 353 coding sequences, we recovered an average of 283 per species and at least 100 in all species. Differences among taxa in sequence recovery could not be explained by relatedness to the representative taxa selected for probe design, suggesting that there is no phylogenetic bias in the probe set. Our probe set, which targeted 260 kbp of coding sequence, achieved a median recovery of 137 kbp per taxon in coding regions, a maximum recovery of 250 kbp, and an additional median of 212 kbp per taxon in flanking non-coding regions across all species. These results suggest that the Angiosperms353 probe set described here is effective for any group of flowering plants and would be useful for phylogenetic studies from the species level to higher-order lineages, including all angiosperms.

evolutionary biology

Divergent gene expression levels between diploid and autotetraploid Tolmiea (Saxifragaceae) relative to the total transcriptome, the cell, and biomass

O_LIStudies of gene expression and polyploidy are typically restricted to characterizing differences in transcript concentration. Integrating multiple methods of transcript analysis, we document a difference in transcriptome size, and make multiple comparisons of transcript abundance in diploid and autotetraploid Tolmiea.\nC_LIO_LIWe use RNA spike-in standards to identify and correct for differences in transcriptome size, and compare levels of gene expression across multiple scales: per transcriptome, per cell, and per biomass.\nC_LIO_LIIn total, ~17% of all loci were identified as differentially expressed (DEGs) between the diploid and autopolyploid species. A shift in total transcriptome size resulted in only ~58% of the total DEGs being identified as differentially expressed following a per transcriptome normalization. When transcript abundance was normalized per cell, ~82% of the total DEGs were recovered. The discrepancy between per-transcriptome and per-cell recovery of DEGs occurs because per-transcriptome normalizations are concentration-based and therefore blind to differences in transcriptome size.\nC_LIO_LIWhile each normalization enables valid comparisons at biologically relevant scales, a holistic comparison of multiple normalizations provides additional explanatory power not available from any single approach. Notably, autotetraploid loci tend to conserve diploid-like transcript abundance per biomass through increased gene expression per cell, and these loci are enriched for photosynthesis-related functions.\nC_LI

genomics

Geographic range dynamics drove ancient hybridization in a lineage of angiosperms

Factors explaining global distribution patterns have been central to biology since the 19th century, yet failure to combine dispersal-based biogeography with shifts in habitat suitability remains a present-day setback in understanding geographic distributions present and past, and time-extended trajectories of lineages. The lack of methods in a suitable integrative framework stands as a conspicuous shortcoming for reconstructing these dynamics. Here we showcase novel methods to overcome these methodological gaps, broadening the prospects for phyloclimatic modeling. We focus on a clade in the angiosperm genus Heuchera endemic to southern California that experienced ancient introgression from circumboreally distributed species of Mitella, testing hypotheses regarding biotic contact in the past between ancestral species lacking a fossil record. We obtain strong support for a past contact zone in northwestern North America, resolving this paradox of hybridization between ancestors of taxa currently separated by [~]1300 km.

evolutionary biology