Search bioRxivSearch

bioRxiv · 10.1101/053843

Comparing the Statistical Fate of Paralogous and Orthologous Sequences

Abstract

Since several decades, sequence alignment is a widely used tool in bioinformatics. For instance, finding homologous sequences with known function in large databases is used to get insight into the function of non-annotated genomic regions. Very efficient tools, like BLAST have been developed to identify and rank possible homologous sequences. To estimate the significance of the homology, the ranking of alignment scores takes a background model for random sequences into account. Using this model one can estimate the probability to find two exactly matching subsequences by chance in two unrelated sequences. The corresponding probability for two homologous sequences is much higher allowing to identify them. Here we focus on the distribution of lengths of exact sequence matches in protein coding regions pairs of evolutionary distant genomes. We show that this distribution exhibits a power-law tail with exponent = --5. Developing a simple model of sequence evolution by substitutions and segmental duplications, we show analytically that paralogous and orthologous gene pairs contribute differently to this distribution. Our model explains the differences observed in the comparison of coding and non-coding parts of genomes, thus providing with a better understanding of statistical properties of genomic sequences and their evolution.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Florian Massip, Michael Sheinman, Sophie Schbath, Peter Arndt. 2016-05-17. Comparing the Statistical Fate of Paralogous and Orthologous Sequences. https://doi.org/10.1101/053843

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Using the Price equation to analyze multi-level selection on the reproductive policing mechanism of bacterial plasmids

The replication control system of non-conjugative bacterial plasmids constitutes a simple and elegant example of a reproductive policing mechanism that moderates competition in the intra-cellular replication pool and establishes a mutually beneficial partnership among plasmids within a bacterial host and between plasmids and their hosts. The emergence of these partnerships is a product of the conflict between the evolutionary interests of hosts, who seek to maximize their growth rates within the population, and plasmids, who seek to maximize their growth rates within hosts. We employ a multi-scale computational model describing the growth, division and death of hosts, as well as the independent replication of plasmids within hosts, in order to investigate the implications of this conflict for the evolution of the plasmid replication parameters. We apply the multi-level form of the Price equation in order to quantify and elucidate the various selective pressures that drive the evolution of plasmid replication control. Our analysis shows how the evolution of the constituent components of the plasmid replication control system are shaped by selection acting at the level of hosts and the level of plasmids. In addition, we calculate finer-grained selective pressures that are attributed to atomic plasmid-related events (such as intra-cellular replication and plasmid loss due to host death) and demonstrate their special role at the early stages of the evolution of policing. Our approach constitutes a novel application of the Price equation for discerning and discussing the synergies between the levels of selection given the availability of a mechanistic model for the generation of the systems dynamics. We show how the Price equation, particularly in its multi-level form, can provide significant insight by quantifying the relative importance of the various selective forces that shape the evolution of policing in bacterial plasmids.

Evolutionary Biology

Sexual selection drives floral scent diversification in carnivorous pitcher plants (NA Sarraceniaceae)

Plant volatiles play vital roles in signaling with their insect associates. Empirical studies show that both pollinators and herbivores exert strong selective pressures on plant phenotypes. While studies often evoke the assumption that volatiles from floral and vegetative tissues are distinct due to strong pollinator-mediated selection operating on the flowers or selection from herbivores acting on the leaves, explicit tests of these assumptions are often lacking. In this study, we examined the evolution of floral and vegetative volatiles in the North American (NA) pitcher plants (Sarraceniaceae). In these taxa, insects are attracted for both pollination and prey capture, providing an ideal opportunity to understand the evolution of scent compounds across different plant organs. We collected a comprehensive dataset of floral and vegetative volatiles from across the NA Sarraceniaceae. We used multivariate analysis methods to examine whether volatile profiles are distinct between plant tissues, and investigated rates of scent evolution in these unique taxa. Our major findings revealed that (i) flowers and pitchers produced highly distinct scent profiles, consistent with the hypothesis that volatiles alleviate trade-offs due to incidental pollinator consumption; (ii) across species, floral scent separated into distinct regions of scent space, while pitchers showed little evidence of clustering - this may be due to convergence on a generalist strategy for insect capture; and (iii) rates of scent evolution depended on tissue type, suggesting that pollinators and herbivores differentially influence the evolution of chemical traits. We emphasize the need for additional functional studies to further distinguish between volatile functions.

Evolutionary Biology

Contrasting patterns of genome-level diversity across distinct co-occurring bacterial populations

To understand the forces driving differentiation and diversification in wild bacterial populations, we must be able to delineate and track ecologically relevant units through space and time. Mapping metagenomic sequences to reference genomes derived from the same environment can reveal genetic heterogeneity within populations, and in some cases, be used to identify boundaries between genetically similar, but ecologically distinct, populations. Here we examine population-level heterogeneity within abundant and ubiquitous freshwater bacterial groups such as the acI Actinobacteria and LD12 Alphaproteobacteria (the freshwater sister clade to the marine SAR11) using 33 single cell genomes and a 5-year metagenomic time series. The single cell genomes grouped into 15 monophyletic clusters (termed \"tribes\") that share at least 97.9% 16S rRNA identity. Distinct populations were identified within most tribes based on the patterns of metagenomic read recruitments to single-cell genomes representing these tribes. Genetically distinct populations within tribes of the acI actinobacterial lineage living in the same lake had different seasonal abundance patterns, suggesting these populations were also ecologically distinct. In contrast, sympatric LD12 populations were less genetically differentiated. This suggests that within one lake, some freshwater lineages harbor genetically discrete (but still closely related) and ecologically distinct populations, while other lineages are composed of less differentiated populations with overlapping niches. Our results point at an interplay of evolutionary and ecological forces acting on these communities that can be observed in real time.

Evolutionary Biology