Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Evolutionary Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Phylofactorization - theory and challenges

Data from biological communities are composed of species connected by the phylogeny. A greedy algorithm phylofactorization - was developed to construct an isometric log-ratio transform whose balances correspond to edges along which traits arose, controlling for previously made inferences.\n\nIn this paper, the general theory of phylofactorization is presented as a graph-partitioning algorithm. A special case-regression phylofactorization-chooses coordinates based on sequential maximization of objective functions from regression on \"contrast\" variables such as an isometric log-ratio transform. The connections between regression phylofactorization and other methods is discussed, including matrix factorization, hierarchical regression, factor analysis and latent variable models. Open challenges in the statistical analysis of phylofactorization are presented, including criteria for choosing the number of factors and approximating null-distributions of commonly used test statistics and objective functions. As a graph-partitioning algorithm, cross-validation of phylo factorization across datasets requires graph-topological considerations, such as how to deal with novel nodes and edges and whether or not to control for partition order. Overcoming these challenges can accelerate our analysis of phylogenetically-structured data and allow annotations of edges in an online tree of life.

evolutionary biology

Evolution of shell flattening and the loss of coiling in top shells (Gastropoda: Trochidae: Fossarininae) on wave-swept rock reefs

Flattening of coiled shells has occurred in numerous gastropod lineages, probably as an adaptation to life in narrow protected spaces, such as crevices or the undersides of rocks. While several genera in the top snail family (Trochidae) have flattened shells, two Fossarininae genera, Broderipia and Roya, are unique in having shells that are limpet-like and zygomorphic, lacking any trace of coiling. The sister genera of these two genera are Fossarina and Synaptocochlea, both of which have coiled shells and live in rock crevices or the vacant shells of sessile organisms. Although Broderipia has recently been identified as living symbiotically in the pits of sea urchins, the habitat and biology of Roya are poorly known. After an extensive search for rare Roya snails on rocky shores of the Japanese Archipelago, we found live Roya eximia snails on intertidal/subtidal rock surfaces exposed to strong waves. The Roya snails crawled swiftly over wave-swept rock surfaces at low tide, while they retreated into the vacant shells of barnacles at high tide, where they adhered firmly to the inner wall. A survey of the macrobenthic communities around the snail habitat showed that Roya snails inhabited only wave-swept rocks of exposed reefs, where the substrata was covered by encrusting red algae and barnacles. Despite the abnormal shell morphology, the radula was similar to other species in the subfamily, and the diet of Roya snails was mainly pennate diatoms. The limpet-like shell of Roya caused loss of coiling and contraction of the soft body, acquisition of a zygomorphic flat body, expansion of the foot sole and loss of the operculum. All of these changes improved tolerance of strong waves and the ability to cling to rock surfaces, and thus enabled a lifestyle split between wave-swept rock surfaces and refugia of vacant barnacle shells.

evolutionary biology

Widespread Historical Contingency in Influenza Viruses

In systems biology and genomics, epistasis characterizes the impact that a substitution at a particular location in a genome can have on a substitution at another location. This phenomenon is often implicated in the evolution of drug resistance or to explain why particular disease-causing mutations do not have the same outcome in all individuals. Hence, uncovering these mutations and their locations in a genome is a central question in biology. However, epistasis is notoriously difficult to uncover, especially in fast-evolving organisms. Here, we present a novel statistical approach that replies on a model developed in ecology and that we adapt to analyze genetic data in fast-evolving systems such as the influenza A virus. We validate the approach using a two-pronged strategy: extensive simulations demonstrate a low-to-moderate sensitivity with excellent specificity and precision, while analyses of experimentally-validated data recover known interactions, including in a eukaryotic system. We further evaluate the ability of our approach to detect correlated evolution during antigenic shifts or at the emergence of drug resistance. We show that in all cases, correlated evolution is prevalent in influenza A viruses, involving many pairs of sites linked together in chains, a hallmark of historical contingency. Strikingly, interacting sites are separated by large physical distances, which entails either long-range conformational changes or functional tradeoffs, for which we find support with the emergence of drug resistance. Our work paves a new way for the unbiased detection of epistasis in a wide range of organisms by performing whole-genome scans.

Evolutionary Biology

Divergent genome evolution caused by regional variation in DNA gain and loss between human and mouse

The forces driving the accumulation and removal of non-coding DNA and ultimately the evolution of genome size in complex organisms are intimately linked to genome structure and organisation. Our analysis provides a novel method for capturing the regional variation of lineage-specific DNA gain and loss events in their respective genomic contexts. To further understand this connection we used comparative genomics to identify genome-wide individual DNA gain and loss events in the human and mouse genomes. Focusing on the distribution of DNA gains and losses, relationships to important structural features and potential impact on biological processes, we found that in autosomes, DNA gains and losses both followed separate lineage-specific accumulation patterns. However, in both species chromosome X was particularly enriched for DNA gain, consistent with its high L1 retrotransposon content required for X inactivation. We found that DNA loss was associated with gene-rich open chromatin regions and DNA gain events with gene-poor closed chromatin regions. Additionally, we found that DNA loss events tended to be smaller than DNA gain events suggesting that they were more tolerated in open chromatin regions. GO term enrichment in human gain hotspots showed terms related to cell cycle/metabolism, human loss hotspots were enriched for terms related to gene silencing, and mouse gain hotspots were enriched for terms related to transcription regulation. Interestingly, mouse loss hotspots were strongly enriched for terms related to developmental processes, suggesting that DNA loss in mouse is associated with phenotypic changes in mouse morphology. This is consistent with a model in which DNA gain and loss results in turnover or \"churning\" of regulatory regions that are then subjected to selection, resulting in the differences we now observe, both genomic and phenotypic/morphological.

evolutionary biology

More than meets the eye: diversity and geographic patterns in sea cucumbers

Estimates for the number of species in the sea vary by orders of magnitude. Molecular taxonomy can greatly speed up screening for diversity and evaluating species boundaries, while gaining insights into the biology of the species. DNA barcoding with a region of cytochrome oxidase 1 (COI) is now widely used as a first pass for molecular evaluation of diversity, as it has good potential for identifying cryptic species and improving our understanding of marine biodiversity. We present the results of a large scale barcoding effort for holothuroids (sea cucumbers). We sequenced 3048 individuals from numerous localities spanning the diversity of habitats in which the group occurs, with a particular focus in the shallow tropics (Indo-Pacific and Caribbean) and the Antarctic region. The number of cryptic species is much higher than currently recognized. The vast majority of sister species have allopatric distributions, with species showing genetic differentiation between ocean basins, and some are even differentiated among archipelagos. However, many closely related and sympatric forms, that exhibit distinct color patterns and/or ecology, show little differentiation in, and cannotbe separatedby, COI sequence data. This pattern is much more common among echinoderms than among molluscs or arthropods, and suggests that echinoderms acquire reproductive isolation at a much faster pace than other marine phyla. Understanding the causes behind such patterns will refine our understanding of diversification and biodiversity in the sea.

Evolutionary Biology

Antagonistic pleiotropy is unexpectedly rare in new mutations

Pleiotropic effects of mutations may underlie diverse biological phenomena such as ageing and specialization. In particular, antagonistic pleiotropy (\"AP\": when a mutation has opposite fitness effects in different environments) generates tradeoffs, which may constrain adaptation. Models of adaptation typically assume that AP is common - especially among large-effect mutations - and that pleiotropic effect sizes are positively correlated. The rare empirical tests of these assumptions have largely focused on beneficial mutations observed under strong selection, whereas most mutations are actually deleterious or neutral, and are removed by selection. We quantified the incidence, nature and effect size of pleiotropy for carbon utilization across 80 single mutations in Escherichia coli that arose under mutation accumulation (i.e. weak selection). Although ~46% of the mutations were pleiotropic, only 11% showed AP, which is lower than expected given the distributions of fitness effects for each resource. In some environments, AP was more common in large-effect mutations (but not synergistic pleiotropy, SP); whereas pleiotropic effect sizes were positively correlated for SP (but not AP). Thus, AP is generally rare, is not consistently enriched in large-effect mutations, and often involves weakly deleterious antagonistic effects. Our unbiased quantification of mutational effects therefore suggests that antagonistic pleiotropy is unlikely to cause maladaptive tradeoffs.

evolutionary biology

Medaka population genome structure and demographic history unveiled via Genotyping-by-Sequencing

Medaka is a model organism in medicine, genetics, developmental biology and population genetics. Lab stocks composed of more than 100 local wild populations are available for research in these fields. Thus, medaka represents a potentially excellent bioresource for screening disease-risk- and adaptation-related genes in genome-wide association studies. Although the genetic population structure should be known before performing such an analysis, a comprehensive study on the genome-wide diversity of wild medaka populations has not been performed. Here, we performed genotyping-by-sequencing (GBS) for 81 and 12 medakas captured from a bioresource and the wild, respectively. Based on the GBS data, we evaluated the genetic population structure and estimated the demographic parameters using an approximate Bayesian computation (ABC) framework. The autosomal data confirmed that there were substantial differences between local populations and supported our previously proposed hypothesis on medaka dispersal based on mitochondrial genome (mtDNA) data. A new finding was that a local group that was thought to be a hybrid between the northern and the southern Japanese groups was actually a sister group of the northern Japanese group. Thus, this paper presents the first population-genomic study of medaka and reveals its population structure and history based on autosomal diversity.

evolutionary biology

The Effect of Horizontal Gene Transfer on the Dynamics of Antibiotic Drug Resistance in a Unicellular Population with a Dynamic Fitness Landscape, Repression and De-repression

Antibiotic drug resistance spreads through horizontal gene transfer (HGT) via bacterial conjugation in unicellular populations of bacteria. Consequently, the efficiency of antibiotics is limited and the expected \"grace period\" of novel antibiotics is typically quite short. One of the mechanisms that allow the accelerated adaptation of bacteria to antibiotics is bacterial conjugation. However, bacterial conjugation is regulated by several biological factors, with one of the most important ones being repression and derepression.\n\nIn recent work, we have studied the effects that repression and de-repression on the mutation-selection balance of an HGT-enabled bacterial population in a static environment. Two of our main findings were that conjugation has a deleterious effect on the mean fitness of the population and that repression is expected to allow a restoration of the fitness cost due to plasmid hosting.\n\nHere, we consider the effect that conjugation-mediated HGT has on the speed of adaptation in a dynamic environment and the effect that repression will have on the dynamics of antibiotic drug resistance. We find that, the effect of repression is dynamic in its possible outcome, that a conjugators to non-conjugators phase transition exists in a dynamic landscape as we have previously found for a static landscape and we quantify the time required for a unicellular population to adapt to a new antibiotic in a periodically changing fitness landscape. Our results also confirmed that HGT accelerates adaptation for a population of prokaryotes which agrees with current knowledge, that HGT rates increase when a population is put under stress.

Evolutionary Biology

Individual- versus group-optimality in the production of secreted bacterial compounds

How unicellular organisms optimize the production of compounds is a fundamental biological question. While it is typically thought that production is optimized at the individual-cell level, secreted compounds could also allow for optimization at the group level, leading to a division of labor where a subset of cells produces and shares the compound with everyone. Using mathematical modelling, we show that the evolution of such division of labor depends on the cost function of compound production. Specifically, for any trait with saturating benefits, linear costs promote the evolution of uniform production levels across cells. Conversely, production costs that diminish with higher output levels favor the evolution of specialization - especially when compound shareability is high. When experimentally testing these predictions with pyoverdine, a secreted iron-scavenging compound produced by Pseudomonas aeruginosa, we found linear costs and, consistent with our model, detected uniform pyoverdine production levels across cells. We conclude that for shared compounds with saturating benefits, the evolution of division of labor is facilitated by a diminishing cost function. More generally, we note that shifts in the level of selection from individuals to groups do not solely require cooperation, but critically depend on mechanistic factors, including the distribution of compound synthesis costs.

evolutionary biology

Deconstructing isolation-by-distance: the genomic consequences of limited dispersal

Geographically limited dispersal can shape genetic population structure and result in a correlation between genetic and geographic distance, commonly called isolation-bydistance. Despite the prevalence of isolation-by-distance in nature, to date few studies have empirically demonstrated the processes that generate this pattern, largely because few populations have direct measures of individual dispersal and pedigree information. Intensive, long-term demographic studies and exhaustive genomic surveys in the Florida Scrub-Jay (Aphelocoma coerulescens) provide an excellent opportunity to investigate the influence of dispersal on genetic structure. Here, we used a panel of genome-wide SNPs and extensive pedigree information to explore the role of limited dispersal in shaping patterns of isolation-by-distance in both sexes, and at an exceedingly fine spatial scale (within ~10 km). Isolation-by-distance patterns were stronger in male-male and male-female comparisons than in female-female comparisons, consistent with observed differences in dispersal propensity between the sexes. Using the pedigree, we demonstrated how various genealogical relationships contribute to fine-scale isolation-by-distance. Simulations using field-observed distributions of male and female natal dispersal distances showed good agreement with the distribution of geographic distances between breeding individuals of different pedigree relationship classes. Furthermore, we extended Malecots theory of isolation-by-distance by building coalescent simulations parameterized by the observed dispersal curve, population density, and immigration rate, and showed how incorporating these extensions allows us to accurately reconstruct observed sex-specific isolation-by-distance patterns in autosomal and Z-linked SNPs. Therefore, patterns of fine-scale isolation-by-distance in the Florida Scrub-Jay can be well understood as a result of limited dispersal over contemporary timescales.\n\nAuthor SummaryDispersal is a fundamental component of the life history of most organisms and therefore influences many biological processes. Dispersal is particularly important in creating genetic structure on the landscape. We often observe a pattern of decreased genetic relatedness between individuals as geographic distances increases, or isolation-by-distance. This pattern is particularly pronounced in organisms with extremely short dispersal distances. Despite the ubiquity of isolation-by-distance patterns in nature, there are few examples that explicitly demonstrate how limited dispersal influences spatial genetic structure. Here we investigate the processes that result in spatial genetic structure using the Florida Scrub-Jay, a bird with extremely limited dispersal behavior and extensive genome-wide data. We take advantage of the long-term monitoring of a contiguous population of Florida Scrub-Jays, which has resulted in a detailed pedigree and measurements of dispersal for hundreds of individuals. We show how limited dispersal results in close genealogical relatives living closer together geographically, which generates a strong pattern of isolation-by-distance at an extremely small spatial scale (<10 km) in just a few generations. Given the detailed dispersal, pedigree, and genomic data, we can achieve a fairly complete understanding of how dispersal shapes patterns of genetic diversity over short spatial scales.

evolutionary biology

Variation and evolution of the glutamine-rich repeat region of Drosophila Argonaute-2

RNA interference pathways mediate multiple biological processes through Argonaute-family proteins, which bind small RNAs as guides to silence complementary target nucleic acids. In insects and crustaceans, Argonaute-2 silences viral nucleic acids, and therefore acts as a primary effector of innate antiviral immunity. Although the function of the major of Argonaute-2 domains, which are conserved across most Argonaute-family proteins, are known, many invertebrate Argonaute-2 homologs contain a glutamine-rich repeat (GRR) region of unknown function at the N-terminus. Here we combine long-read amplicon sequencing of Drosophila Genetic Reference Panel (DGRP) lines with publicly available sequence data from many insect species to show that this region evolves extremely rapidly and is hypervariable within species. We identify distinct GRR haplotype groups in D. melanogaster, and suggest that one of these haplotype groups has recently risen to high frequency in North American populations. Finally, we use published data from genome-wide association studies of viral resistance in D. melanogaster to test whether GRR haplotypes are associated with survival after virus challenge. We find a marginally significant association with survival after challenge with Drosophila C Virus in the DGRP, but we were unable to replicate this finding using lines from the Drosophila Synthetic Population Resource panel.

Evolutionary Biology

MicroRNA-205 affects mouse granulosa cell apoptosis and estradiol synthesis by targeting CREB1

MicroRNAs-205 (miR-205), were reportedly to be involved in various physiological and pathological processes, but its biological function in follicular atresia remain unknown. In this study, we investigated the expression of miR-205 in mouse granulosa cells (mGCs), and explored its functions in primary mGCs using a serial of in vitro experiments. The result of qRT-PCR demonstrated that miR-205 expression was significantly increased in early atretic follicles (EAF), and progressively atretic follicles (PAF) compared to healthy follicles (HF). Our results also revealed that overexpression of miR-205 in mGCs significantly promoted apoptosis, caspas-3/9 activities, and inhibited estrogen E2 release, and cytochrome P450 family 19 subfamily A polypeptide 1 (CYP19A1, a key gene in E2 production) expression. Bioinformatics and luciferase reporter assays revealed that the gene of cyclic AMP response element (CRE)-binding protein 1 (CREB1) was a potential target of miR-205. qRT-PCR and western blot assays revealed that overexpression of miR-205 inhibited the expression of CREB1 in mGCs. Importantly, CREB1 upregulation partially rescued the effects of miR-205 on apoptosis, caspase-3/9 activities, E2 production and CYP19A1 expression in mGCs. Our results indicate that miR-205 may play an important role in ovarian follicular development and provide new insights into follicular atresia.

evolutionary biology

Short-term insurance versus long-term bet-hedging strategies as adaptations to variable environments

Understanding how organisms adapt to environmental variation is a key challenge of biology. Central to this are bet-hedging strategies that maximize geometric mean fitness across generations, either by being conservative or diversifying phenotypes. Theoretical models of bet-hedging and the multiplicative fitness effects of environmental variation across generations have traditionally assumed that environmental conditions are constant within lifetimes. However, behavioral ecology has revealed adaptive responses to additive fitness effects of environmental variation within lifetimes, either through insurance or risk-sensitive strategies. Here we explore whether the effects of adaptive insurance interact with the evolution of bet-hedging by varying the position and skew of fitness functions within and between lifetimes. When insurance causes the optimal phenotype to shift from the peak to down the less steeply decreasing side of the fitness function, then conservative bet-hedging does not generally evolve on top of this, even if diversifying bet-hedging can. Canalization to reduce phenotypic variation within a lifetime is almost always favored, except when the tails of the fitness function are steeply convex and produce a novel risk-sensitive increase in phenotypic variance akin to diversifying bet-hedging. Importantly, using skewed fitness functions, we provide the first example of how conservative and diversifying bet-hedging strategies might coexist.

evolutionary biology

Microbial composition of enigmatic bird parasites: Wolbachia and Spiroplasma are the most important bacterial associates of quill mites (Acari:Syringophilidae)

The microbiome is an integral component of many animal species, potentially affecting behaviour, physiology, and other biological properties. Despite this importance, bacterial communities remain vastly understudied in many groups of invertebrates, including mites. Quill mites (Acariformes: Syringophilidae) are a poorly known group of permanent bird ectoparasites that occupy quills of feathers and feed on bird subcutaneous tissue and fluids. Most species have strongly female biased sex ratios and it was hypothesized that this is caused by endosymbiotic bacteria. Their peculiar lifestyle further makes them potential vectors for bird diseases. Previously, Anaplasma phagocytophilum and a high diversity of Wolbachia strains were detected in quill mites via targeted PCR screens. Here, we use an unbiased 16S amplicon sequencing approach to determine other Bacteria that potentially impact quill mite biology.\n\nWe performed 16S V4 amplicon sequencing of 126 quill mite individuals from eleven species parasitizing twelve bird species (four families) of passeriform birds. In addition to Wolbachia, we found Spiroplasma as potential symbiont of quill mites. Interestingly, consistently high Spiroplasma titres were only found in individuals of two mite species associated with finches of the genus Cardfuelis, suggesting a history of horizontal transfers of Spiroplasma via the bird host. Furthermore, there was evidence for Spiroplasma negatively affecting Wolbachia titres. We found no evidence for the previously reported Anaplasma in quill mites, but detected the potential pathogens Brucella and Bartonella at low abundances. Other amplicon sequence variants (ASVs) could be assigned to a diverse number of bacterial taxa, including several that were previously isolated from bird skin. We observed a relatively uniform distribution of these ASVs across mite taxa and bird hosts, i.e, there was a lack of host-specificity for most detected ASVs. Further, many frequently found ASVs were assigned to taxa that show a very broad distribution with no strong prior evidence for symbiotic association with animals. We interpret these findings as evidence for a scarcity or lack of resident microbial associates (other than inherited symbionts) in quill mites, or for abundances of these taxa below our detection threshold.

evolutionary biology

Paleophenotype Reconstruction Of Carbon Fixation Proteins As A Window Into Historic Biological States

Two datasets, the geologic record and the genetic content of extant organisms, provide complementary insights into the history of how key molecular components have shaped or driven global environmental and macroevolutionary trends. Changes in global physicochemical modes over time are thought to be a consistent feature of this relationship between Earth and life, as life is thought to have been optimizing protein functions for the entirety of its [~]3.8 billion years of history on Earth. Organismal survival depends on how well critical genetic and metabolic components can adapt to their environments, reflecting an ability to optimize efficiently to changing conditions. The geologic record provides an array of biologically independent indicators of macroscale atmospheric and oceanic composition, but provides little in the way of the exact behavior of the molecular components that influenced the compositions of these reservoirs. By reconstructing sequences of proteins that might have been present in ancient organisms, we can identify a subset of possible sequences that may have been optimized to these ancient environmental conditions. How can extant life be used to reconstruct ancestral phenotypes? Configurations of ancient sequences can be inferred from the diversity of extant sequences, and then resurrected in the lab to ascertain their biochemical attributes. One way to augment sequence-based, single-gene methods to obtain a richer and more reliable picture of the deep past, is to resurrect inferred ancestral protein sequences in living organisms, where their phenotypes can be exposed in a complex molecular-systems context, and to then link consequences of those phenotypes to biosignatures that were preserved in the independent historical repository of the geological record. As a first-step beyond single molecule reconstruction to the study of functional molecular systems, we present here the ancestral sequence reconstruction of the beta-carbonic anhydrase protein. We assess how carbonic anhydrase proteins meet our selection criteria for reconstructing ancient biosignatures in the lab, which we term paleophenotype reconstruction.

evolutionary biology

Inferring individual-level processes from population-level patterns in cultural evolution

Our species is characterized by a great degree of cultural variation, both within and between populations. Understanding how group-level patterns of culture emerge from individual-level behaviour is a long-standing question in the biological and social sciences. We develop a simulation model capturing demographic and cultural dynamics relevant to human cultural evolution, focusing on the interface between population-level patterns and individual-level processes. The model tracks the distribution of variants of cultural traits across individuals in a population over time, conditioned on different pathways for the transmission of information between individuals. From these data we obtain theoretical expectations for a range of statistics commonly used to capture population-level characteristics (e.g. the degree of cultural diversity). Consistent with previous theoretical work, our results show that the patterns observed at the level of groups are rooted in the interplay between the transmission pathways and the age structure of the population. We also explore whether, and under what conditions, the different pathways can be distinguished based on their group-level signatures, in an effort to establish theoretical limits to inference. Our results show that the temporal dynamic of cultural change over time retains a stronger signature than the cultural composition of the population at a specific point in time. Overall, the results suggest a shift in focus from identifying the one individual-level process that likely produced the observed data to excluding those that likely did not. We conclude by discussing the implications for empirical studies of human cultural evolution.

evolutionary biology

High-density linkage map and QTLs for growth in snapper (Chrysophrys auratus)

Characterizing the genetic variation underlying phenotypic traits is a central objective in biological research. This research has been hampered in the past by the limited genomic resources available for most non-model species. However, recent advances in sequencing technology and related genotyping methods are rapidly changing this. Here we report the use of genome-wide SNP data from the ecologically and commercially important marine fish species Chrysophrys auratus (snapper) to 1) construct the first linkage map for this species, 2) scan for growth QTLs, and 3) search for candidate genes in the surrounding QTL regions. The newly constructed linkage map contained ~11K SNP markers and is the densest map to date in the fish family Sparidae. Comparisons with available genome scaffolds indicated that overall marker placement was strongly correlated between the scaffolds and linkage map (R = 0.7), but at fine scales (< 5 cM) there were some precision limitations. Of the 24 linkage groups, which reflect the 24 chromosomes of this species, three were found to contain QTLs with genome-wide significance for growth-related traits. A scan for 13 known candidate growth genes located the genes for growth hormone, parvalbumin, and myogenin within 13.2, 2.6, and 5.0 cM of these genome-wide significant QTLs, respectively. The linkage map and QTLs found in this study will advance the investigation of genome structure and selective breeding in snapper.

evolutionary biology

Natural selection defines the cellular complexity

Current biology is perplexed by the lack of a theoretical framework for understanding the organization principles of the molecular system within a cell. Here we first studied growth rate, one of the seemingly most complex cellular traits, using functional data of yeast single-gene deletion mutants. We observed nearly one thousand expression informative genes (EIGs) whose expression levels are linearly correlated to the trait within an unprecedentedly large functional space. A simple model considering six EIG-formed protein modules revealed a variety of novel mechanistic insights, and also explained [~]50% of the variance of cell growth rates measured by Bar-seq technique for over 400 yeast mutants (Pearsons R = 0.69), a performance comparable to the microarray-based (R = 0.77) or colony-size-based (R = 0.66) experimental approach. We then applied the same strategy to 501 morphological traits of the yeast and achieved successes in most fitness-coupled traits each with hundreds of trait-specific EIGs. Surprisingly, there is no any EIG found for most fitness-uncoupled traits, indicating that they are controlled by super-complex epistases that allow no simple expression-trait correlation. Thus, EIGs are recruited exclusively by natural selection, which builds a rather simple functional architecture for fitness-coupled traits, and the endless complexity of a cell lies primarily in its fitness-uncoupled features.

Evolutionary Biology