Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “evolutionary biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Evolution of the D. melanogaster chromatin landscape and its associated proteins

In the nucleus of eukaryotic cells, genomic DNA associates with numerous protein complexes and RNAs, forming the chromatin landscape. Through a genome-wide study of chromatin-associated proteins in Drosophila cells, five major chromatin types were identified as a refinement of the traditional binary division into hetero- and euchromatin. These five types are defined by distinct but overlapping combinations of proteins and differ in biological and biochemical properties, including transcriptional activity, replication timing and histone modifications. In this work, we assess the evolutionary relationships of chromatin-associated proteins and present an integrated view of the evolution and conservation of the fruit fly D. melanogaster chromatin landscape. We combine homology prediction across a wide range of species with gene age inference methods to determine the origin of each chromatin-associated protein. This provides insight into the emergence of the different chromatin types. Our results indicate that the two euchromatic types, YELLOW and RED, were one single activating type that split early in eukaryotic history. Next, we provide evidence that GREEN-associated proteins are involved in a centromere drive and expanded in a lineage-specific way in D. melanogaster. Our results on BLUE chromatin support the hypothesis that the emergence of Polycomb Group proteins is linked to eukaryotic multicellularity. In light of these results, we discuss how the regulatory complexification of chromatin links to the origins of eukaryotic multicellularity.

evolutionary biology

Environmental disturbances compromise synthetic symbiosis

By virtue of complex interactions, the behaviour of mutualistic systems is difficult to study and nearly impossible to predict. We have developed a theoretical model of a modifiable experimental yeast system that is amenable to exploring self-organised cooperation while considering the production and use of specific metabolites. Leveraging the simplicity of an artificial yeast system, a simple model of mutualism, we develop and test the assumptions and stability of this theoretical model. We examine how one-off, recurring and permanent changes to an ecological niche affect a cooperative interaction and identify an ecological \"Goldilocks zone\" in which the mutualism can survive. Moreover, we explore how a factor like the cost of mutualism - the cellular burden of cooperating - influences the stability of mutualism and how environmental changes shape this stability. Our results highlight the fragility of mutualisms and suggest the use of synthetic biology to stave off an ecological collapse.

evolutionary biology

Model adequacy and the macroevolution of angiosperm functional traits

Making meaningful inferences from phylogenetic comparative data requires a meaningful model of trait evolution. It is thus important to determine whether the model is appropriate for the data and the question being addressed. One way to assess this is to ask whether the model provides a good statistical explanation for the variation in the data. To date, researchers have focused primarily on the explanatory power of a model relative to alternative models. Methods have been developed to assess the adequacy, or absolute explanatory power, of phylogenetic trait models but these have been restricted to specific models or questions. Here we present a general statistical framework for assessing the adequacy of phylogenetic trait models. We use our approach to evaluate the statistical performance of commonly used trait models on 337 comparative datasets covering three key Angiosperm functional traits. In general, the models we tested often provided poor statistical explanations for the evolution of these traits. This was true for many different groups and at many different scales. Whether such statistical inadequacy will qualitatively alter inferences draw from comparative datasets will depend on the context. Regardless, assessing model adequacy can provide interesting biological insights -- how and why a model fails to describe variation in a dataset gives us clues about what evolutionary processes may have driven trait evolution across time.

Evolutionary Biology

The importance and adaptive value of life history evolution for metapopulation dynamics

O_LIThe spatial configuration and size of patches influence metapopulation dynamics by altering colonisation-extinction dynamics and local density-dependency. This spatial forcing as determined by the metapopulation typology then imposes strong selection pressures on life history traits, which will in turn feedback on the ecological metapopulation dynamics. Given the relevance of metapopulation persistence for biological conservation, and the potential rescuing role of evolution, a firm understanding of the relevance of these eco-evolutionary processes is essential.\nC_LIO_LIWe here follow a systems modelling approach to quantify the importance of spatial forcing and experimentally observed life history evolution for metapopulation demography as quantified by (meta)population size and variability. We therefore developed an individual based model matching an earlier experimental evolution with spider mites to perform virtual translocation and invasion experiments that would have been otherwise impossible to conduct.\nC_LIO_LIWe show that (1) metapopulation demography is more affected by spatial forcing than by life history evolution, but that life history evolution contributes substantially to changes in local and especially metapopulation-level population sizes, (2) extinction rates are minimised by evolution in classical metapopulations, and (3) evolution is optimising individual performance in metapopulations when considering the importance of more cryptic stress resistance evolution.\nC_LIO_LIEcological systems modelling opens up a promising avenue to quantify the importance of eco-evolutionary feedbacks for larger-scale population dynamics. Metapopulation sizes are especially impacted by evolution but its variability is mainly determined by the spatial forcing.\nC_LIO_LIEco-evolutionary dynamics can increase the persistence of classical metapopulations. The maintenance of evolutionary dynamics in spatially structured populations is thus not only essential in the face of environmental change; it also generates feedbacks that impact metapopulation persistence.\nC_LI\n\nData-archiveMetadata population dynamics in artificial metapopulations data are available from the Dryad Digital Repository: http://dx.doi.org/10.5061/dryad.18r5f (De Roissart, Wang & Bonte 2015). Modeling code is available on: https://github.ugent.be/pages/dbonte/eco_evo-metapop/ (ODD protocol in supplementary material)

evolutionary biology

The mysterious orphans of Mycoplasmataceae

BackgroundThe length of a protein sequence is largely determined by its function, i.e. each functional group is associated with an optimal size. However, comparative genomics revealed that proteins length may be affected by additional factors. In 2002 it was shown that in bacterium Escherichia coli and the archaeon Archaeoglobus fulgidus, protein sequences with no homologs are, on average, shorter than those with homologs [1]. Most experts now agree that the length distributions are distinctly different between protein sequences with and without homologs in bacterial and archaeal genomes. In this study, we examine this postulate by a comprehensive analysis of all annotated prokaryotic genomes and focusing on certain exceptions.\n\nResultsWe compared lengths distributions of \"having homologs proteins\" (HHPs) and \"non-having homologs proteins\" (orphans or ORFans) in all currently annotated completely sequenced prokaryotic genomes. As expected, the HHPs and ORFans have strikingly different length distributions in almost all genomes. As previously established, the HHPs, indeed, are, on average, longer than the ORFans, and the length distributions for the ORFans have a relatively narrow peak, in contrast to the HHPs, whose lengths spread over a wider range of values. However, about thirty genomes do not obey these rules. Practically all genomes of Mycoplasma and Ureaplasma have atypical ORFans distributions, with the mean lengths of ORFan larger than the mean lengths of HHPs. These genera constitute over 80% of atypical genomes.\n\nConclusionsWe confirmed on a ubiquitous set of genomes the previous observation that HHPs and ORFans have different gene length distributions. We also showed that Mycoplasmataceae genomes have very distinctive distributions of ORFans lengths. We offer several possible biological explanations of this phenomenon.

Evolutionary Biology

Effects of multiple sources of genetic drift on pathogen variation within hosts

Changes in pathogen genetic variation within hosts alter the severity and spread of infectious diseases, with important implications for clinical disease and public health. Genetic drift may play a strong role in shaping pathogen variation, but analyses of drift in pathogens have oversimplified pathogen population dynamics, either by considering dynamics only at a single scale (within hosts, between hosts), or by making drastic simplifying assumptions (host immune systems can be ignored, transmission bottlenecks are complete). Moreover, previous studies used genetic data to infer the strength of genetic drift, whereas we test whether the genetic drift imposed by pathogen population processes can be used to explain genetic data. We first constructed and parameterized a mathematical model of gypsy moth baculovirus dynamics that allows genetic drift to act within and between hosts. We then quantified the genome-wide diversity of baculovirus populations within each of 143 field-collected gypsy moth larvae using Illumina sequencing. Finally, we determined whether the genetic drift imposed by host-pathogen population dynamics in our model explains the levels of pathogen diversity in our data. We found that when the model allows drift to act at multiple scales, including within hosts, between hosts, and between years, it can accurately reproduce the data, but when the effects of drift are simplified by neglecting transmission bottlenecks and stochastic variation in virus replication within hosts, the model fails. A de novo mutation model and a purifying selection model similarly fail to explain the data. Our results show that genetic drift can play a strong role in determining pathogen variation, and that mathematical models that account for pathogen population growth at multiple scales of biological organization can be used to explain this variation.

evolutionary biology

Rapid evolution of the human mutation spectrum

DNA is a remarkably precise medium for copying and storing biological information. This high fidelity results from the action of hundreds of genes involved in replication, proofreading, and damage repair. Evolutionary theory suggests that in such a system, selection has limited ability to remove genetic variants that change mutation rates by small amounts or in specific sequence contexts. Consistent with this, using SNV variation as a proxy for mutational input, we report here that mutational spectra differ substantially among species, human continental groups and even some closely-related populations. Close examination of one signal, an increased TCC[->]TTC mutation rate in Europeans, indicates a burst of mutations from about 15,000 to 2,000 years ago, perhaps due to the appearance, drift, and ultimate elimination of a genetic modifier of mutation rate. Our results suggest that mutation rates can evolve markedly over short evolutionary timescales and suggest the possibility of mapping mutational modifiers.

evolutionary biology

Spatial selection and local adaptation jointly shape life-history evolution during range expansion

In the context of climate change and species invasions, range shifts increasingly gain attention because the rates at which they occur in the Anthropocene induce fast shifts in biological assemblages. During such range shifts, species experience multiple selection pressures. Especially for poleward expansions, a straightforward interpretation of the observed evolutionary dynamics is hampered because of the joint action of evolutionary processes related to spatial selection and to adaptation towards local climatic conditions. To disentangle the effects of these two processes, we integrated stochastic modeling and empirical approaches, using the spider mite Tetranychus urticae as a model species. We demonstrate considerable latitudinal quantitative genetic divergence in life-history traits in T. urticae, that was shaped by both spatial selection and local adaptation. The former mainly affected dispersal behavior, while development was mainly shaped by adaptation to the local climate. Divergence in life-history traits in species shifting their range poleward can consequently be jointly determined by fast local adaptation to the environmental gradient and contemporary evolutionary dynamics resulting from spatial selection. The integration of modeling with common garden experiments provides a powerful tool to study the contribution of these two evolutionary processes on life-history evolution during range expansion.

Evolutionary Biology

Long read sequencing reveals poxvirus evolution through rapid homogenization of gene arrays

Large DNA viruses rapidly evolve to defeat host defenses. Poxvirus adaptation can involve combinations of recombination-driven gene copy number variation and beneficial single nucleotide variants (SNVs) at the same locus, yet how these distinct mechanisms of genetic diversification might simultaneously facilitate adaptation to immune blocks is unknown. We performed experimental evolution with a vaccinia virus population harboring a SNV in a gene actively undergoing copy number amplification. Comparisons of virus genomes using the Oxford Nanopore Technologies sequencing platform allowed us to phase SNVs within large gene copy arrays for the first time, and uncovered a mechanism of adaptive SNV homogenization reminiscent of gene conversion, which is actively driven by selection. Our work reveals a new mechanism for the fluid gain of beneficial mutations in genetic regions undergoing active recombination in viruses, and illustrates the value of long read sequencing technologies for investigating complex genome dynamics in diverse biological systems.

evolutionary biology

HIV-1 Protease Evolvability is Affected by Synonymous Nucleotide Recoding

One unexplored aspect of HIV-1 genetic architecture is how codon choice influences population diversity and evolvability. Here we compared the development of HIV-1 resistance to protease inhibitors (PIs) between wild-type (WT) virus and a synthetic virus (MAX) carrying a codon-pair re-engineered protease sequence including 38 (13%) synonymous mutations. WT and MAX viruses showed indistinguishable replication in MT-4 cells or PBMCs. Both viruses were subjected to serial passages in MT-4 cells with selective pressure from the PIs atazanavir (ATV) and darunavir (DRV). After 32 successive passages, both the WT and MAX viruses developed phenotypic resistance to PIs (IC50 14.6 {+/-} 5.3 and 21.2 {+/-} 9 nM for ATV, and 5. 9 {+/-} 1.0 and 9.3 {+/-} 1.9 for DRV, respectively). Ultra-deep sequence clonal analysis revealed that both viruses harbored previously described resistance mutations to ATV and DRV. However, the WT and MAX virus proteases showed different resistance variant repertoires, with the G16E and V77I substitutions observed only in WT, and the L33F, S37P, G48L, Q58E/K, and L89I substitutions detected only in MAX. Remarkably, G48L and L89I are rarely found in vivo in PI-treated patients. The MAX virus showed significantly higher nucleotide and amino acid diversity of the propagated viruses with and without PIs (P < 0.0001), suggesting higher selective pressure for change in this recoded virus. Our results indicate that HIV-1 protease position in sequence space delineates the evolution of its mutant spectra. Nevertheless, the investigated synonymously recoded variant showed mutational robustness and evolvability similar to the WT virus.\n\nIMPORTANCELarge-scale synonymous recoding of virus genomes is a new tool for exploring various aspects of virus biology. Synonymous virus genome recoding can be used to investigate how a viruss position in sequence space defines its mutant spectrum, evolutionary trajectory, and pathogenesis. In this study, we evaluated how synonymous recoding of the human immunodeficiency virus type 1 (HIV-1) protease impacts the development of protease inhibitor (PI) resistance. HIV-1 protease is a main target of current antiretroviral therapies. Our present results demonstrate that the wild-type (WT) virus and the virus with the recoded protease exhibited different patterns of resistance mutations after PI treatment. Nevertheless, the developed PI resistance phenotype was indistinguishable between the recoded virus and the WT virus, suggesting that the synonymously recoded protease HIV-1 and the WT protease virus were equally robust and evolvable.

evolutionary biology

Trait evolution with jumps: illusionary normality

Phylogenetic comparative methods for real-valued traits usually make use of stochastic process whose trajectories are continuous. This is despite biological intuition that evolution is rather punctuated than gradual. On the other hand, there has been a number of recent proposals of evolutionary models with jump components. However, as we are only beginning to understand the behaviour of branching Ornstein-Uhlenbeck (OU) processes the asymptotics of branching OU processes with jumps is an even greater unknown. In this work we build up on a previous study concerning OU with jumps evolution on a pure birth tree. We introduce an extinction component and explore via simulations, its effects on the weak convergence of such a process. We furthermore, also use this work to illustrate the simulation and graphic generation possibilities of the mvSLOUCH package.

evolutionary biology

Integrative analysis of large scale transcriptome data draws a comprehensive landscape of Phaeodactylum tricornutum functional genome and evolutionary origin of diatoms

Diatoms are one of the most successful and ecologically important groups of eukaryotic phytoplankton in the modern ocean. Deciphering their genomes is a key step towards better understanding of their biological innovations, evolutionary origins, and ecological underpinnings. Here, we have used 90 RNA-Seq datasets from different growth conditions combined with published expressed sequence tags and protein sequences from multiple taxa to explore the genome of the model diatom Phaeodactylum tricornutum, and introduce 1,489 novel genes. The new annotation additionally permitted the discovery for the first time of extensive alternative splicing (AS) in diatoms, including intron retention and exon skipping which increases the diversity of transcripts to regulate gene expression in response to nutrient limitations. In addition, we have used up-to-date reference sequence libraries to dissect the taxonomic origins of diatom genomes. We show that the P. tricornutum genome is replete in lineage-specific genes, with up to 47% of the gene models present only possessing orthologues in other stramenopile groups. Finally, we have performed a comprehensive de novo annotation of repetitive elements showing novel classes of TEs such as SINE, MITE, LINE and TRIM/LARD. This work provides a solid foundation for future studies of diatom gene function, evolution and ecology.

bioinformatics

GADMA: Genetic Algorithm for Automatic Inferring Joint Demographic History of Multiple Populations from Allele Frequency Spectrum

The demographic history of any population is imprinted in the genomes of the individuals that make up the population. One of the most popular and convenient representations of genetic information is the allele frequency spectrum or AFS, the distribution of allele frequencies in populations. The joint allele frequency spectrum is commonly used to reconstruct the demographic history of multiple populations and several methods based on diffusion approximation (e.g.,{partial} a{partial}i) and ordinary differential equations (e.g., moments) have been developed and applied for demographic inference. These methods provide an opportunity to simulate AFS under a variety of researcher-specified demographic models and to estimate the best model and associated parameters using likelihood-based local optimizations. However, there are no known algorithms to perform global searches of demographic models with a given AFS. Here, we introduce a new method that implements a global search using a genetic algorithm for the automatic and unsupervised inference of demographic history from joint allele frequency spectrum data. Our method is implemented in the software GADMA (Genetic Algorithm for Demographic Analysis, https://github.com/ctlab/GADMA). We demonstrate the performance of GADMA by applying it to sequence data from humans and non-model organisms and show that it is able to automatically infer a demographic model close to or even better than the one that was previously obtained manually. Moreover, GADMA is able to infer demographic models at different local optima close to the global one, making it is possible to detect more biology corrected model during further research.

evolutionary biology

Molecular Analysis of Parasites in the Choreocolacaceae (Rhodophyta) Reveals a Reduced Harveyella mirabilis Plastid Genome and Supports the Transfer of Genera to the Rhodomelaceae (Rhodophyta)

Parasitism is a life strategy that has repeatedly evolved within the Florideophyceae. Historically, the terms adelphoparasite and alloparasite have been used to distinguish parasites based on the relative phylogenetic relationship of host and parasite. However, analyses using molecular phylogenetics indicate that nearly all red algal parasites infect within their taxonomic family, and a range of relationships exist between host and parasite. To date, all investigated adelphoparasites have lost their plastid, and instead, incorporate a host derived plastid when packaging spores. In contrast, a highly reduced plastid lacking photosynthesis genes was sequenced from the alloparasite Choreocolax polysiphoniae. Here we present the complete Harveyella mirabilis plastid genome, which has also lost genes involved in photosynthesis, and a partial plastid genome from Leachiella pacifica. The H. mirabilis plastid shares more synteny with free-living red algal plastids than that of C. polysiphoniae. Phylogenetic analysis demonstrates that C. polysiphoniae, H. mirabilis, and L. pacifica form a robustly supported clade of parasites, which retain their own plastid genomes, within the Rhodomelaceae. We therefore transfer all three genera from the exclusively parasitic family, Choreocolacaceae, to the Rhodomelaceae. Additionally, we recommend applying the terms archaeplastic parasites (formerly alloparasites), and neoplastic parasites (formerly adelphoparasites) to distinguish red algal parasites using a biological framework rather than taxonomic affiliation with their hosts.

evolutionary biology

Differences between the de novo proteome and its non-functional precursor can result from neutral constraints on its birth process, not necessarily from natural selection alone

Proteins are among the most important constituents of biological systems. Because all proteins ultimately evolved from previously non-coding DNA, the properties of these non-coding sequences and how they shape the birth of novel proteins are also expected to influence the organization of biological networks. When trying to explain and predict the properties of novel proteins, it is of particular importance to distinguish the contributions of natural selection and other evolutionary forces. Studies in the field typically use non-coding DNA and GC-content-based random-sequence models to generate random expectations for the properties of novel functional proteins. Deviations from these expectations have been interpreted as the result of natural selection. However, interpreting such deviations requires a yet-unattained understanding of the raw material of de novo gene birth and its relation to novel functional proteins. We mathematically show how the importance of the \"junk\" polypeptides that make up this raw material goes beyond their average properties and their filtering by natural selection. We find that the mean of any property among novel functional proteins also depends on its variance among junk polypeptides and its correlation with their rate of evolutionary turnover. In order to exemplify the use of our general theoretical results, we combine them with a simple model that predicts the means and variances of the properties of junk polypeptides from the genomic GC content alone. Under this model, we predict the effect of GC content on the mean length and mean intrinsic disorder of novel functional proteins as a function of evolutionary parameters. We use these predictions to formulate new evolutionary interpretations of published data on the length and intrinsic disorder of novel functional proteins. This work provides a theoretical framework that can serve as a guide for the prediction and interpretation of past and future results in the study of novel proteins and their properties under various evolutionary models. Our results provide the foundation for a better understanding of the properties of cellular networks through the evolutionary origin of their components.

evolutionary biology

The evolution and maintenance of extraordinary allelic diversity

AbstractThe majority of highly polymorphic genes are related to immune functions and with over 100 alleles within a population, genes of the major histocompatibility complex (MHC) are the most polymorphic loci in vertebrates. How such extraordinary polymorphism arose and is maintained is controversial. One possibility is heterozygote advantage (HA), which can in principle maintain any number of alleles, but biologically explicit models based on this mechanism have so far failed to reliably predict the coexistence of significantly more than ten alleles. We here present an eco-evolutionary model showing that evolution can result in the emergence and maintenance of more than 100 alleles under HA if the following two assumptions are fulfilled: first, pathogens are lethal in the absence of an appropriate immune defence; second, the effect of pathogens depends on host condition, with hosts in poorer condition being affected more strongly. Thus, our results show that HA can be a more potent force in explaining the extraordinary polymorphism found at MHC loci than currently recognized.

evolutionary biology

Destructive Disinfection Of Infected Brood Prevents Systemic Disease Spread In Ant Colonies

Social insects protect their colonies from infectious disease through collective defences that result in social immunity. In ants, workers first try to prevent infection of colony members. Here, we show that if this fails and a pathogen establishes an infection, ants employ an efficient multicomponent behaviour - \"destructive disinfection\" - to prevent further spread of disease through the colony. Ants specifically target infected pupae during the pathogens non-contagious incubation period, relying on chemical sickness cues emitted by pupae. They then remove the pupal cocoon, perforate its cuticle and administer antimicrobial poison, which enters the body and prevents pathogen replication from the inside out. Like the immune system of a body that specifically targets and eliminates infected cells, this social immunity measure sacrifices infected brood to stop the pathogen completing its lifecycle, thus protecting the rest of the colony. Hence, the same principles of disease defence apply at different levels of biological organisation.

evolutionary biology

Genome wide association analysis identifies genetic variants associated with reproductive variation across domestic dog breeds and uncovers links to domestication

The diversity of eutherian reproductive strategies has led to variation in many traits, such as number of offspring, age of reproductive maturity, and gestation length. While reproductive trait variation has been extensively investigated and is well established in mammals, the genetic loci contributing to this variation remain largely unknown. The domestic dog, Canis lupus familiaris is a powerful model for studies of the genetics of inherited disease due to its unique history of domestication. To gain insight into the genetic basis of reproductive traits across domestic dog breeds, we collected phenotypic data for four traits - cesarean section rate (n = 97 breeds), litter size (n = 60), stillbirth rate (n = 57), and gestation length (n = 23) - from primary literature and breeders handbooks. By matching our phenotypic data to genomic data from the Cornell Veterinary Biobank, we performed genome wide association analyses for these four reproductive traits, using body mass and kinship among breeds as co-variates. We identified 14 genome-wide significant associations between these traits and genetic loci, including variants near CACNA2D3 with gestation length, MSRB3 with litter size, SMOC2 with cesarean section rate, MITF with litter size and still birth rate, KRT71 with cesarean section rate, litter size, and stillbirth rate, and HTR2C with stillbirth rate. Some of these loci, such as CACNA2D3 and MSRB3, have been previously implicated in human reproductive pathologies. Many of the variants that we identified have been previously associated with domestication-related traits, including brachycephaly (SMOC2), coat color (MITF), coat curl (KRT71), and tameness (HTR2C). These results raise the hypothesis that the artificial selection that gave rise to dog breeds also shaped the observed variation in their reproductive traits. Overall, our work establishes the domestic dog as a system for studying the genetics of reproductive biology and disease.

evolutionary biology