Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Mining unknown porcine protein isoforms by tissue-based map of proteome enhances the pig genome annotation

A lack of the complete pig proteome has left a gap in our knowledge of the pig genome and has restricted the feasibility of using pigs as a biomedical model. We developed the tissue-based proteome maps using 34 major normal pig tissues. A total of 7,319 unknown protein isoforms were identified and systematically characterized, including 3,703 novel protein isoforms, 669 protein isoforms from 460 genes symbolized beginning with LOC, and 2,947 protein isoforms without clear NCBI annotation in current pig reference genome. These newly identified protein isoforms were functionally annotated through profiling the pig transcriptome with high-throughput RNA sequencing (RNA-seq) of the same pig tissues, further improving the genome annotation of corresponding protein coding genes. Combining the well-annotated genes that having parallel expression pattern and subcellular witness, we predicted the tissue related subcellular components and potential function for these unknown proteins. Finally, we mined 3,656 orthologous genes for 49.95% of unknown protein isoforms across multiple species, referring to 65 KEGG pathways and 25 disease signaling pathways. These findings provided valuable insights and a rich resource for enhancing studies of pig genomics and biology as well as biomedical model application to human medicine.

genomics

New human chromosomal safe harbor sites for genome engineering with CRISPR/Cas9, TAL effector and homing endonucleases

Safe Harbor Sites (SHS) are genomic locations where new genes or genetic elements can be introduced without disrupting the expression or regulation of adjacent genes. We have identified 35 potential new human SHS in order to substantially expand SHS options beyond the three widely used canonical human SHS, AAVS1, CCR5 and hROSA26. All 35 potential new human SHS and the three canonical sites were assessed for SHS potential using 9 different criteria weighted to emphasize safety that were broader and more genomics-based than previous efforts to assess SHS potential. We then systematically compared and rank-ordered our 35 new sites and the widely used human AAVS1, hROSA26 and CCR5 sites, then experimentally validated a subset of the highly ranked new SHS together versus the canonical AAVS1 site. These characterizations included in vitro and in vivo cleavage-sensitivity tests; the assessment of population-level sequence variants that might confound SHS targeting or use for genome engineering; homology-dependent and -independent, SHS-targeted transgene integration in different human cell lines; and comparative transgene integration efficiencies at two new SHS versus the canonical AAVS1 site. Stable expression and function of new SHS-integrated transgenes were demonstrated for transgene-encoded fluorescent proteins, selection cassettes and Cas9 variants including a transcription transactivator protein that were shown to drive large deletions in a PAX3/FOXO1 fusion oncogene and induce expression of the MYF5 gene that is normally silent in human rhabdomyosarcoma cells. We also developed a SHS genome engineering toolkit to enable facile use of the most extensively characterized of our new human SHS located on chromosome 4p. We anticipate our newly identified human SHS, located on 16 chromosomes including both arms of the human X chromosome, will be useful in enabling a wide range of basic and more clinically-oriented human gene editing and engineering.

genomics

A comprehensive genomics solution for HIV surveillance and clinical monitoring in a global health setting

High-throughput viral genetic sequencing is needed to monitor the spread of drug resistance, direct optimal antiretroviral regimes, and to identify transmission dynamics in generalised HIV epidemics. Public health efforts to sequence HIV genomes at scale face three major technical challenges: (i) minimising assay cost and protocol complexity, (ii) maximising sensitivity, and (iii) recovering accurate and unbiased sequences of both the genome consensus and the within-host viral diversity. Here we present a novel, high-throughput, virus-enriched sequencing method and computational pipeline tailored specifically to HIV (veSEQ-HIV), which addresses all three technical challenges, and can be used directly on leftover blood drawn for routine CD4 testing. We demonstrate its performance on 1,620 plasma samples collected from consenting individuals attending 10 large urban clinics in Zambia, partners of HPTN 071 (PopART). We show that veSEQ-HIV consistently recovers complete HIV genomes from the majority of samples of different subtypes, and is also quantitative: the number of HIV reads per sample obtained by veSEQ-HIV estimates viral load without the need for additional testing. Both quantitativity and sensitivity were assessed on a subset of 126 samples with clinically measured viral loads, and with standardized quantification controls (VL 100 - 5,000,000 RNA copies/ml). Complete HIV genomes were recovered from 93% (85/91) of samples when viral load was over 1,000 copies per ml. The quantitative nature of the assay implies that variant frequencies estimated with veSEQ-HIV are representative of true variant frequencies in the sample. Detection of minority variants can be exploited for epidemiological analysis of transmission and drug resistance, and we show how the information contained in individual reads of a veSEQ-HIV sample can be used to detect linkage between multiple mutations associated with resistance to antiretroviral therapy. Less than 2% of reads obtained by veSEQ-HIV were identified as in silico contamination events using updates to the phyloscanner software (phyloscanner clean) that we show to be 95% sensitive and 99% specific at decontaminating NGS data. The cost of the assay -- approximately 45 USD per sample -- compares favourably with existing VL and HIV genotyping tests, and provides the additional value of viral load quantification and inference of drug resistance with a single test. veSEQ-HIV is well suited to large public health efforts and is being applied to all [~]9000 samples collected for the HPTN 071-2 (PopART Phylogenetics) study.

genomics

Genome-scale Capture C promoter interaction analysis implicates novel effector genes at GWAS loci for bone mineral density

ASBTRACTOsteoporosis is a devastating disease with an essential genetic component. Genome wide association studies (GWAS) have discovered genetic variants robustly associated with bone mineral density (BMD), however they only report genomic signals and not necessarily the precise localization of culprit effector genes. Therefore, we sought to carry out physical and direct variant to gene mapping in a relevant primary human cell type. We developed SPATIaL-seq (genome-Scale, Promoter-focused Analysis of chromaTIn Looping), a massively parallel, high resolution Capture-C based method to simultaneously characterize the genome-wide interactions of all human promoters. By intersecting our SPATIaL-seq and ATAC-seq data from human mesenchymal progenitor cell -derived osteoblasts, we observed consistent contacts between candidate causal variants and putative target gene promoters in open chromatin for ~30% of the 110 BMD loci investigated. Knockdown of two novel implicated genes, ING3 at CPED1-WNT16 and EPDR1 at STARD3NL, had pronounced inhibitory effects on osteoblastogenesis. Our approach therefore aids target discovery in osteoporosis and can be applied to other common genetic diseases.

genomics

An improved high-quality genome assembly and annotation of Qingke, Tibetan hulless barley

BackgroundThe Tibetan hulless barley (Hordeum vulgare L. var. nudum), also called \"Qingke\" in Chinese and \"Ne\" in Tibetan, is the staple food for Tibetans and an important livestock feed in the Tibetan Plateau. The Tibetan hulless barley in China has about 3500 years of cultivation history, mainly produced in Tibet, Qinghai, Sichuan, Yunnan and other areas. In addition, Tibetan hulless barley has rich nutritional value and outstanding health effects, including the beta glucan, dietary fiber, amylopectin, the contents of trace elements, which are higher than any other cereal crops.\n\nFindingsHere, we reported an improved high-quality assembly of Tibetan hulless barley genome with 4.0 Gb in size. We employed the falcon assembly package, scaffolding and error correction tools to finish improvement using PacBio long reads sequencing technology, with contig and scaffold N50 lengths of 1.563Mb and 4.006Mb, respectively, representing more continuous than the original Tibetan hulless barley genome nearly two orders of magnitude. We also re-annotated the new assembly, and reported 61,303 stringent confident putative protein-coding genes, of which 40,457 is HC genes. We have developed a new Tibetan hulless barley genome database (THBGD) to download and use friendly, as well as to better manage the information of the Tibetan hulless barley genetic resources.\n\nConclusionsThe availability of new Tibetan hulless barley genome and annotations will take the genetics of Tibetan hulless barley to a new level and will greatly simplify the breeders effort. It will also enrich the granary of the Tibetan people.

genomics

Process-specific somatic mutation distributions vary with three-dimensional genome structure

Somatic mutations arise during the life history of a cell. Mutations occurring in cancer driver genes may ultimately lead to the development of clinically detectable disease. Nascent cancer lineages continue to acquire somatic mutations throughout the neoplastic process and during cancer evolution (Martincorena and Campbell, 2015). Extrinsic and endogenous mutagenic factors contribute to the accumulation of these somatic mutations (Zhang and Pellman, 2015). Understanding the underlying factors generating somatic mutations is crucial for developing potential preventive, therapeutic and clinical decisions. Earlier studies have revealed that DNA replication timing (Stamatoyannopoulos et al., 2009) and chromatin modifications (Schuster-Bockler and Lehner, 2012) are associated with variations in mutational density. What is unclear from these early studies, however, is whether all extrinsic and exogenous factors that drive somatic mutational processes share a similar relationship with chromatin state and structure. In order to understand the interplay between spatial genome organization and specific individual mutational processes, we report here a study of 3000 tumor-normal pair whole genome datasets from more than 40 different human cancer types. Our analyses revealed that different mutational processes lead to distinct somatic mutation distributions between chromatin folding domains. APOBEC- or MSI-related mutations are enriched in transcriptionally-active domains while mutations occurring due to tobacco-smoke, ultraviolet (UV) light exposure or a signature of unknown aetiology (signature 17) enrich predominantly in transcriptionally-inactive domains. Active mutational processes dictate the mutation distributions in cancer genomes, and we show that mutational distributions shift during cancer evolution upon mutational processes switch. Moreover, a dramatic instance of extreme chromatin structure in humans, that of the unique folding pattern of the inactive X-chromosome leads to distinct somatic mutation distribution on X chromosome in females compared to males in various cancer types. Overall, the interplay between three-dimensional genome organization and active mutational processes has a substantial influence on the large-scale mutation rate variations observed in human cancer.

genomics

Modular dynamics of DNA co-methylation networks exposes the functional organization of colon cancer cells’ genome

Epigenomic plasticity is interconnected with chromatin structure and gene regulation. In tumor progression, orchestrated remodeling of genome organization accompanies the acquisition of malignant properties. DNA methylation, a key epigenetic mark extensively altered in cancer, is also linked to genome architecture and function. Based on this association, we postulate that the dissection of long-range co-methylation structure unveils cancer cells genome architecture remodeling. We applied network-modeling of DNA methylation co-variation in two colon cancer cohorts and found abundant and consistent transchromosomal structures in both normal and tumor tissue. Normal-tumor comparison indicated substantial remodeling of the epigenome covariation and revealed novel genomic compartments with a unique signature of DNA methylation rank inversion.

genomics

Leveraging evolutionary relationships to improve Anopheles genome assemblies

While new sequencing technologies have lowered financial barriers to whole genome sequencing, resulting assemblies are often fragmented and far from finished. Subsequent improvements towards chromosomal-level status can be achieved by both experimental and computational approaches. Requiring only annotated assemblies and gene orthology data, comparative genomics approaches that aim to capture evolutionary signals to predict scaffold neighbours (adjacencies) offer potentially substantive improvements without the costs associated with experimental scaffolding or re-sequencing. We leverage the combined detection power of three such gene synteny-based methods applied to 21 Anopheles mosquito assemblies with variable contiguity levels to produce consensus sets of scaffold adjacency predictions. Three complementary validations were performed on subsets of assemblies with additional supporting data: six with physical mapping data; 13 with paired-end RNA sequencing (RNAseq) data; and three with new assemblies based on re-scaffolding or incorporating Pacific Biosciences (PacBio) sequencing data. Improved assemblies were built by integrating the consensus adjacency predictions with supporting experimental data, resulting in 20 new reference assemblies with improved contiguities. Combined with physical mapping data for six anophelines, chromosomal positioning of scaffolds improved assembly anchoring by 47% for A. funestus and 38% A. stephensi. Reconciling an A. funestus PacBio assembly with synteny-based and RNAseq-based adjacencies and physical mapping data resulted in a new 81.5% chromosomally mapped reference assembly and cytogenetic photomap. While complementary experimental data are clearly key to achieving high-quality chromosomal-level assemblies, our assessments and validations of gene synteny-based computational methods highlight the utility of applying comparative genomics approaches to improve community genomic resources.

genomics

Comparison of single-cell whole-genome amplification strategies

Single-cell genomics is an alluring area that holds the potential to change the way we understand cell populations. Due to the small amount of DNA within a single cell, whole-genome amplification becomes a mandatory step in many single-cell applications. Unfortunately, single-cell whole-genome amplification (scWGA) strategies suffer from several technical biases that complicate the posterior interpretation of the data. Here we compared the performance of six different scWGA methods (GenomiPhi, REPLIg, TruePrime, Ampli1, MALBAC, and PicoPLEX) after amplifying and low-pass sequencing the complete genome of 230 healthy/tumoral human cells. Overall, REPLIg outperformed competing methods regarding DNA yield, amplicon size, amplification breadth, amplification uniformity -being the only method with a random amplification bias-, and false single-nucleotide variant calls. On the other hand, non-MDA methods, and in particular Ampli1, showed less allelic imbalance and ADO, more reliable copy-number profiles and less chimeric amplicons. While no single scWGA method showed optimal performance for every aspect, they clearly have distinct advantages. Our results provide a convenient guide for selecting a scWGA method depending on the question of interest while revealing relevant weaknesses that should be considered during the analysis and interpretation of single-cell sequencing data.

genomics

Genome-wide association study of suicide attempt in psychiatric disorders identifies association with major depression polygenic risk scores

ObjectiveOver 90% of suicide attempters have a psychiatric diagnosis, however twin and family studies suggest that the genetic etiology of suicide attempt (SA) is partially distinct from that of the psychiatric disorders themselves. Here, we present the largest genome-wide association study (GWAS) on suicide attempt using major depressive disorder (MDD), bipolar disorder (BIP) and schizophrenia (SCZ) cohorts from the Psychiatric Genomics Consortium.\n\nMethodSamples comprise 1622 suicide attempters and 8786 non-attempters with MDD, 3264 attempters and 5500 non-attempters with BIP and 1683 attempters and 2946 non-attempters with SCZ. SA GWAS were performed comparing attempters to non-attempters in each disorder followed by meta-analysis across disorders. Polygenic risk scoring investigated the genetic relationship between SA and the psychiatric disorders.\n\nResultsThree genome-wide significant loci for SA were found: one associated with SA in MDD, one in BIP, and one in the meta-analysis of SA in mood disorders. These associations were not replicated in independent mood disorder cohorts from the UK Biobank and iPSYCH. Polygenic risk scores for major depression were significantly associated with SA in MDD (P=0.0002), BIP (P=0.0006) and SCZ (P=0.0006).\n\nConclusionsThis study provides new information on genetic associations and the genetic etiology of SA across psychiatric disorders. The finding that polygenic risk scores for major depression predict suicide attempt across disorders provides a possible starting point for predictive modelling and preventative strategies. Further collaborative efforts to increase sample size hold potential to robustly identify genetic associations and gain biological insights into the etiology of suicide attempt.

genetics

Population genomics of parallel hybrid zones in the mimetic butterflies, H. melpomene and H. erato

Hybrid zones can be valuable tools for studying evolution and identifying genomic regions responsible for adaptive divergence and underlying phenotypic variation. Hybrid zones between subspecies of Heliconius butterflies can be very narrow and are maintained by strong selection acting on colour pattern. The co-mimetic species H. erato and H. melpomene have parallel hybrid zones where both species undergo a change from one colour pattern form to another. We use restriction associated DNA sequencing to obtain several thousand genome wide sequence markers and use these to analyse patterns of population divergence across two pairs of parallel hybrid zones in Peru and Ecuador. We compare two approaches for analysis of this type of data; alignment to a reference genome and de novo assembly, and find that alignment gives the best results for species both closely (H. melpomene) and distantly (H. erato, ~15% divergent) related to the reference sequence. Our results confirm that the colour pattern controlling loci account for the majority of divergent regions across the genome, but we also detect other divergent regions apparently unlinked to colour pattern differences. We also use association mapping to identify previously unmapped colour pattern loci, in particular the Ro locus. Finally, we identify within our sample a new cryptic population of H. timareta in Ecuador, which occurs at relatively low altitude and is mimetic with H. melpomene malleti.

Evolutionary Biology

Species Delimitation using Genome-Wide SNP Data

The multi-species coalescent has provided important progress for evolutionary inferences, including increasing the statistical rigor and objectivity of comparisons among competing species delimitation models. However, Bayesian species delimitation methods typically require brute force integration over gene trees via Markov chain Monte Carlo (MCMC), which introduces a large computation burden and precludes their application to genomic-scale data. Here we combine a recently introduced dynamic programming algorithm for estimating species trees that bypasses MCMC integration over gene trees with sophisticated methods for estimating marginal likelihoods, needed for Bayesian model selection, to provide a rigorous and computationally tractable technique for genome-wide species delimitation. We provide a critical yet simple correction that brings the likelihoods of different species trees, and more importantly their corresponding marginal likelihoods, to the same common denominator, which enables direct and accurate comparisons of competing species delimitation models using Bayes factors. We test this approach, which we call Bayes factor delimitation (*with genomic data; BFD*), using common species delimitation scenarios with computer simulations. Varying the numbers of loci and the number of samples suggest that the approach can distinguish the true model even with few loci and limited samples per species. Misspecification of the prior for population size{theta} has little impact on support for the true model. We apply the approach to West African forest geckos (Hemidactylus fasciatus complex) using genome-wide SNP data data. This new Bayesian method for species delimitation builds on a growing trend for objective species delimitation methods with explicit model assumptions that are easily tested.

Evolutionary Biology

An optimized CRISPR/Cas toolbox for efficient germline and somatic genome engineering in Drosophila

The type II CRISPR/Cas system has recently emerged as a powerful method to manipulate the genomes of various organisms. Here, we report a novel toolbox for high efficiency genome engineering of Drosophila melanogaster consisting of transgenic Cas9 lines and versatile guide RNA (gRNA) expression plasmids. Systematic evaluation reveals Cas9 lines with ubiquitous or germline restricted patterns of activity. We also demonstrate differential activity of the same gRNA expressed from different U6 snRNA promoters, with the previously untested U6:3 promoter giving the most potent effect. Choosing an appropriate combination of Cas9 and gRNA allows targeting of essential and non-essential genes with transmission rates ranging from 25% - 100%. We also provide evidence that our optimized CRISPR/Cas tools can be used for offset nicking-based mutagenesis and, in combination with oligonucleotide donors, to precisely edit the genome by homologous recombination with efficiencies that do not require the use of visible markers. Lastly, we demonstrate a novel application of CRISPR/Cas-mediated technology in revealing loss-of-function phenotypes in somatic cells following efficient biallelic targeting by Cas9 expressed in a ubiquitous or tissue-restricted manner. In summary, our CRISPR/Cas tools will facilitate the rapid evaluation of mutant phenotypes of specific genes and the precise modification of the genome with single nucleotide precision. Our results also pave the way for high throughput genetic screening with CRISPR/Cas.

Genetics

A genome-wide analysis of Cas9 binding specificity using ChIP-seq and targeted sequence capture

Clustered regularly interspaced short palindromic repeat (CRISPR) RNA-guided nucleases have gathered considerable excitement as a tool for genome engineering. However, questions remain about the specificity of their target site recognition. Most previous studies have examined predicted off-target binding sites that differ from the perfect target site by one to four mismatches, which represent only a subset of genomic regions. Here, we use ChIP-seq to examine genome-wide CRISPR binding specificity at gRNA-specific and gRNA-independent sites. For two guide RNAs targeting the murine Snurf gene promoter, we observed very high binding specificity at the intended target site while off-target binding was observed at 2- to 6-fold lower intensities. We also identified significant gRNA-independent off-target binding. Interestingly, we found that these regions are highly enriched in the PAM site, a sequence required for target site recognition by CRISPR. To determine the relationship between Cas9 binding and endonuclease activity, we used targeted sequence capture as a high-throughput approach to survey a large number of the potential off-target sites identified by ChIP-seq or computational prediction. A high frequency of indels was observed at both target sites and one off-target site, while no cleavage activity could be detected at other ChIP-bound regions. Our data is consistent with recent finding that most interactions between the CRISPR nuclease complex and genomic PAM sites are transient and do not lead to DNA cleavage. The interactions are stabilized by gRNAs with good matches to the target sequence adjacent to the PAM site, resulting in target cleavage activity.

Molecular Biology

Extraordinarily wide genomic impact of a selective sweep associated with the evolution of sex ratio distorter suppression

Symbionts that distort their hosts sex ratio by favouring the production and survival of females are common in arthropods. Their presence produces intense Fisherian selection to return the sex ratio to parity, typified by the rapid spread of host suppressor loci that restore male survival/development. In this study, we investigated the genomic impact of a selective event of this kind in the butterfly Hypolimnas bolina. Through linkage mapping we first identified a genomic region that was necessary for males to survive Wolbachia-induced killing. We then investigated the genomic impact of the rapid spread of suppression that converted the Samoan population of this butterfly from a 100:1 female-biased sex ratio in 2001, to a 1:1 sex ratio by 2006. Models of this process revealed the potential for a chromosome-wide selective sweep. To measure the impact directly, the pattern of genetic variation before and after the episode of selection was compared. Significant changes in allele frequencies were observed over a 25cM region surrounding the suppressor locus, alongside generation of linkage disequilibrium. The presence of novel allelic variants in 2006 suggests that the suppressor was introduced via immigration rather than through de novo mutation. In addition, further sampling in 2010 indicated that many of the introduced variants were lost or had reduced in frequency since 2006. We hypothesise that this loss may have resulted from a period of purifying selection - removing deleterious material that introgressed during the initial sweep. Our observations of the impact of suppression of sex ratio distorting activity reveal an extraordinarily wide genomic imprint, reflecting its status as one of the strongest selective forces in nature.

Evolutionary Biology

Benchmarking undedicated cloud computing providers for analysis of genomic datasets.

A major bottleneck in biological discovery is now emerging at the computational level. Cloud computing offers a dynamic means whereby small and medium-sized laboratories can rapidly adjust their computational capacity. We benchmarked two established cloud computing services, Amazon Web Services Elastic MapReduce (EMR) on Amazon EC2 instances and Google Compute Engine (GCE), using publicly available genomic datasets (E.coli CC102 strain and a Han Chinese male genome) and a standard bioinformatic pipeline on a Hadoop-based platform. Wall-clock time for complete assembly differed by 52.9% (95%CI: 27.5-78.2) for E.coli and 53.5% (95%CI: 34.4-72.6) for human genome, with GCE being more efficient than EMR. The cost of running this experiment on EMR and GCE differed significantly, with the costs on EMR being 257.3% (95%CI: 211.5-303.1) and 173.9% (95%CI: 134.6-213.1) more expensive for E.coli and human assemblies respectively. Thus, GCE was found to outperform EMR both in terms of cost and wall-clock time. Our findings confirm that cloud computing is an efficient and potentially cost-effective alternative for analysis of large genomic datasets. In addition to releasing our cost-effectiveness comparison, we present available ready-to-use scripts for establishing Hadoop instances with Ganglia monitoring on EC2 or GCE.

Bioinformatics

Rate and cost of adaptation in the Drosophila genome

Recent studies have consistently inferred high rates of adaptive molecular evolution between Drosophila species. At the same time, the Drosophila genome evolves under different rates of recombination, which results in partial genetic linkage between alleles at neighboring genomic loci. Here we analyze how linkage correlations affect adaptive evolution. We develop a new inference method for adaptation that takes into account the effect on an allele at a focal site caused by neighboring deleterious alleles (background selection) and by neighboring adaptive substitutions (hitchhiking). Using complete genome sequence data and fine-scale recombination maps, we infer a highly heterogeneous scenario of adaptation in Drosophila. In high-recombining regions, about 50% of all amino acid substitutions are adaptive, together with about 20% of all substitutions in proximal intergenic regions. In low-recombining regions, only a small fraction of the amino acid substitutions are adaptive, while hitchhiking accounts for the majority of these changes. Hitchhiking of deleterious alleles generates a substantial collateral cost of adaptation, leading to a fitness decline of about 30/2N per gene and per million years in the lowest-recombining regions. Our results show how recombination shapes rate and efficacy of the adaptive dynamics in eukaryotic genomes.

Evolutionary Biology

GC-content evolution in bacterial genomes: the biased gene conversion hypothesis expands.

The characterization of functional elements in genomes relies on the identification of the footprints of natural selection. In this quest, taking into account neutral evolutionary processes such as mutation and genetic drift is crucial because these forces can generate patterns that may obscure or mimic signatures of selection. In mammals, and probably in many eukaryotes, another such confounding factor called GC-Biased Gene Conversion (gBGC) has been documented. This mechanism generates patterns identical to what is expected under selection for higher GC-content, specifically in highly recombining genomic regions. Recent results have suggested that a mysterious selective force favouring higher GC-content exists in Bacteria but the possibility that it could be gBGC has been excluded. Here, we show that gBGC is probably at work in most if not all bacterial species. First we find a consistent positive relationship between the GC-content of a gene and evidence of intra-genic recombination throughout a broad spectrum of bacterial clades. Second, we show that the evolutionary force responsible for this pattern is acting independently from selection on codon usage, and could potentially interfere with selection in favor of optimal AU-ending codons. A comparison with data from human populations shows that the intensity of gBGC in Bacteria is comparable to what has been reported in mammals. We propose that gBGC is not restricted to sexual Eukaryotes but also widespread among Bacteria and could therefore be an ancestral feature of cellular organisms. We argue that if gBGC occurs in bacteria, it can account for previously unexplained observations, such as the apparent non-equilibrium of base substitution patterns and the heterogeneity of gene composition within bacterial genomes. Because gBGC produces patterns similar to positive selection, it is essential to take this process into account when studying the evolutionary forces at work in bacterial genomes.

Evolutionary Biology