Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Pelagiphages in the Podoviridiae family integrate into host genomes

The Pelagibacterales order (SAR11) in Alphaproteobacteria dominates marine surface bacterioplankton communities, where it plays a key role in carbon and nutrient cycling. SAR11 phages, known as pelagiphages, are among the most abundant phages in the ocean. Four pelagiphages that infect Pelagibacter HTCC1062 have been reported. Here we report 11 new pelagiphages in the Podoviridae family. Comparative genomic analysis revealed that they are all closely related to previously reported pelagiphages HTVC011P and HTVC019P, in the HTVC019Pvirus genus. HTVC019Pvirus pelagiphages share a core genome of 15 genes, with a pan-genome of 234 genes. Phylogenomic analysis clustered these pelagiphages into three subgroups. Integrases were identified in all but one pelagiphage genomes. Evidence of site-specific integration was obtained by high-throughput sequencing and sequencing PCR amplicons containing predicted integration sites, demonstrating the capacity of these pelagiphages to propagate by both lytic and lysogenic infection. HTVC019P, HTVC021P, HTVC022P, HTVC201P and HTVC121P integrate into tRNA-Cys genes. HTVC011P, HTVC025P, HTVC105P, HTVC109P, HTVC119P and HTVC200P target tRNA-Leu genes, while HTVC120P integrates into the tRNA-Arg. Evidence of pelagiphage integration was also retrieved from Global Ocean Survey (GOS) database, suggesting the occurrence of pelagiphage integration in situ. The capacity of HTVC019Pvirus pelagiphages to integrate into host genomes suggests they could impact SAR11 populations by a variety of mechanisms, including mortality, genetic transduction, and prophage-induced viral immunity. HTVC019Pvirus pelagiphages are a rare example of a lysogenic phage that can be implicated in ecological processes on broad scales, and thus have potential to become a useful model for investigating strategies of host infection and phage-dependent horizontal gene transfer.\n\nIMPORTANCEPelagiphages are ecologically important because of their extraordinarily high census numbers, which makes them potentially significant agents in the viral shunt, a concept that links viral predation to the recycling of dissolved organic matter released from lysing plankton cells. Lysogenic Pelagiphages, such as the HTVC019Pvirus pelagiphages we investigate here, are also important because of their potential to contribute to the hypothesized processes such as the \"Piggy-Back-the-Winner\" and \"King-of-the-Mountain\". The former explains nonlinearities in virus to host ratios by postulating increased lysogenization of successful host cells, while the latter postulates host-density dependent propagation of defensive alleles. Here we report multiple Pelagiphage isolates, and provided detailed evidence of their integration into SAR11 genomes. The development of this ecologically significant experimental system for studying phage-dependent processes is progress towards the validation of broad hypotheses about phage ecology with specific examples based on knowledge of mechanisms.

microbiology

Genotypic clustering does not imply recent tuberculosis transmission in a high prevalence setting: A genomic epidemiology study in Lima, Peru

BackgroundWhole genome sequencing (WGS) can elucidate Mycobacterium tuberculosis (Mtb) transmission patterns but more data is needed to guide its use in high-burden settings. In a household-based transmissibility study of 4,000 TB patients in Lima, Peru, we identified a large MIRU-VNTR Mtb cluster with a range of resistance phenotypes and studied host and bacterial factors contributing to its spread.\n\nMethodsWGS was performed on 61 of 148 isolates in the cluster. We compared transmission link inference using epidemiological or genomic data with and without the inclusion of controversial variants, and estimated the dates of emergence of the cluster and antimicrobial drug resistance acquisition events by generating a time-calibrated phylogeny. We validated our findings in genomic data from an outbreak of 325 TB cases in London. Using a larger set of 12,032 public Mtb genomes, we determined bacterial factors characterizing this cluster and under positive selection in other Mtb lineages.\n\nFindingsFour isolates were distantly related and the remaining 57 isolates diverged ca. 1968 (95% HPD: 1945-1985). Isoniazid resistance arose once, whereas rifampicin resistance emerged subsequently at least three times. Amplification of other drug resistance occurred as recently as within the last year of sampling. High quality PE/PPE variants and indels added information for transmission inference. We identified five cluster-defining SNPs, including esxV S23L to be potentially contributing to transmissibility.\n\nInterpretationClusters defined by MIRU-VNTR typing, could be circulating for decades in a high-burden setting. WGS allows for an improved understanding of transmission, as well as bacterial resistance and fitness factors.\n\nFundingThe study was funded by the National Institutes of Health (Peru Epi study U19-AI076217 and K01-ES026835 to MRF). The funding sources had no role in any aspect of the study, manuscript or decision to submit it for publication.\n\nResearch in contextO_ST_ABSEvidence before this studyC_ST_ABSUse of whole genome sequencing (WGS) to study tuberculosis (TB) transmission has proven to have higher resolution that traditional typing methods in low-burden settings. The implications of its use in high-burden settings are not well understood.\n\nAdded value of this studyUsing WGS, we found that TB clusters defined by traditional typing methods may be circulating for several decades. Genomic regions typically excluded from WGS analysis contain large amount of genetic variation that may affect interpretation of transmission events. We also identified five bacterial mutations that may contribute to transmission fitness.\n\nImplications of all the available evidenceAdded value of WGS for understanding TB transmission may be even higher in high-burden vs. low-burden settings. Methods integrating variants found in polymorphic sites and insertions and deletions are likely to have higher resolution. Several host and bacterial factors may be responsible for higher transmissibility that can be targets of intervention to interrupt TB transmission in communities.

microbiology

Processing of grassland soil C-N compounds into soluble and volatile molecules is depth stratified and mediated by genomically novel bacteria and archaea

Soil microbial activity drives the carbon and nitrogen cycles and is an important determinant of atmospheric trace gas turnover, yet most soils are dominated by organisms with unknown metabolic capacities. Even Acidobacteria, among the most abundant bacteria in soil, remain poorly characterized, and functions across groups such as Verrucomicrobia, Gemmatimonadetes, Chloroflexi, Rokubacteria are understudied. Here, we resolved sixty metagenomic, and twenty proteomic datasets from a grassland soil ecosystem and recovered 793 near-complete microbial genomes from 18 phyla, representing around one third of all organisms detected. Importantly, this enabled extensive genomics-based metabolic predictions for these understudied communities. Acidobacteria from multiple previously unstudied classes have genomes that encode large enzyme complements for complex carbohydrate degradation. Alternatively, most organisms we detected encode carbohydrate esterases that strip readily accessible methyl and acetyl groups from polymers like pectin and xylan, forming methanol and acetate, the availability of which could explain high prevalences of C1 metabolism and acetate utilization in genomes. Organism abundances among samples collected at three soil depths and under natural and amended rainfall regimes indicate statistically higher associations of inorganic nitrogen metabolism and carbon degradation in deep and shallow soils, respectively. This partitioning decreased in samples under extended spring rainfall indicating long term climate alteration can affect both carbon and nitrogen cycling. Overall, by leveraging natural and experimental gradients with genome-resolved metabolic profiles, we link organisms lacking prior genomic characterization to specific roles in complex carbon, C1, nitrate, and ammonia transformations and constrain factors that impact their distributions in soil.

microbiology

LtrDetector: A modern tool-suite for detecting long terminal repeat retrotransposons de-novo on the genomic scale

Long terminal repeat retrotransposons are the most abundant transposons in plants. They play important roles in alternative splicing, recombination, gene regulation, and genomic evolution. Large-scale sequencing projects for plant genomes are currently underway. Software tools are important for annotating long terminal repeat retrotransposons in these newly available genomes. However, the available tools are not very sensitive to known elements and perform inconsistently on different genomes. Some are hard to install or obsolete. They may struggle to process large plant genomes. None are concurrent or have features to support manual review of new elements. To overcome these limitations, we developed LtrDetector, which uses signal-processing techniques. LtrDetector is easy to install and use. It is not species specific. It utilizes multi-core processors available in personal computers. It is more sensitive than other tools by 14.4%-50.8% while maintaining a low false positive rate on six plant genomes.

bioinformatics

Impact of mutation rate and selection at linked sites on fine-scale DNA variation across the homininae genome

DNA diversity varies across the genome of many species. Variation in diversity across a genome might arise for one of three reasons; regional variation in the mutation rate, selection and biased gene conversion. We show that both non-coding and non-synonymous diversity are correlated to a measure of the mutation rate, the recombination rate and the density of conserved sequences in 50KB windows across the genomes of humans and non-human homininae. We show these patterns persist even when we restrict our analysis to GC-conservative mutations, demonstrating that the patterns are not driven by biased gene conversion. The positive correlation between diversity and our measure of the mutation rate seems to be largely a direct consequence of regions with higher mutation rates having more diversity. However, the positive correlation with recombination rate and the negative correlation with the density of conserved sequences suggests that selection at linked sites affect levels of diversity. This is supported by the observation that the ratio of the number of non-synonymous to non-coding polymorphisms is negatively correlated to a measure of the effective population size across the genome. Furthermore, we find evidence that these genomic variables are better predictors of non-coding diversity in large homininae populations than in small populations, after accounting for statistical power. This is consistent with genetic drift decreasing the impact of selection at linked sites in small populations. In conclusion, our comparative analyses describe for the first time how recombination rate, gene density, mutation rate and genetic drift interact to produce the patterns of DNA diversity that we observe along and between homininae genomes.

evolutionary biology

N6-methyladenosine regulates Influenza A virus mRNA stability yet is rarely found on genomic RNA

Previous studies have found widespread N6-methyladenosine (m6A methylation) on all forms of Influenza A virus (IAV) RNA, with m6A found critical for viral replication, pathogenicity as well as viral RNA packaging. Here we applied the latest quantitative technologies to revisit the methylation landscape on the anti-sense genomic RNA of IAV. Unexpectedly, upon Ultra-Performance Liquid Chromatography-Tandem Mass Spectrometry (UPLC-MS/MS) analysis of IAV virion -extracted genomic RNA, we detected very little m6A regardless of production from human cells or chicken eggs. Concordantly, Nanopore direct RNA sequencing also detected an overall low occurrence and stoichiometry (generally <5%) of m6A across all viral genomic RNA segments, compared with abundant m6A sites on viral mRNAs at ~20-30% m6A. Cross validation with glyoxal- and nitrite-mediated deamination of unmethylated adenosines (GLORI) confirmed multiple m6A sites on viral mRNA yet very few m6A on the genomic RNA. This paucity of m6A on genomic RNA makes it unlikely that m6A contributes to viral RNA packaging. Knockdown or pharmacological inhibition of the m6A methyltransferase METTL3 as well as the reader protein YTHDF2 both reduced viral mRNA levels and infectious viral particle production, with YTHDF2 promoting viral mRNA stability. Thus, the presence of m6A on IAV transcripts is indeed proviral, yet it is the mRNAs instead of genomic RNAs that are methylated at functionally relevant levels. Lastly, we provide proof of concept that a METTL3 small molecule inhibitor can be antiviral, and propose that m6A-targeted antivirals would mainly impact the intracellular gene expression phase of IAV replication.

microbiology

Optical and physical mapping with local finishing enables megabase-scale resolution of agronomically important regions in the wheat genome

BackgroundNumerous scaffold-level sequences for wheat are now being released and, in this context, we report on a strategy for improving the overall assembly to a level comparable to that of the human genome.\n\nResultsUsing chromosome 7A of wheat as a model, sequence-finished megabase scale sections of this chromosome were established by combining a new independent assembly based on a BAC-based physical map, BAC pool paired end sequencing, chromosome arm specific mate-pair sequencing and Bionano optical mapping with the IWGSC RefSeq v1.0 sequence and its underlying raw data. The combined assembly results in 18 super-scaffolds across the chromosome. The value of finished genome regions is demonstrated for two approximately 2.5 Mb regions associated with yield and the grain quality phenotype of fructan carbohydrate grain levels. In addition, the 50 Mb centromere region analysis incorporates cytological data highlighting the importance of non-sequence data in the assembly of this complex genome region.\n\nConclusionsSufficient genome sequence information is shown to be now available for the wheat community to produce sequence-finished releases of each chromosome of the reference genome. The high-level completion identified that an array of seven fructosyl transferase genes underpins grain quality and yield attributes are affected by five f-box-only-protein-ubiquitin ligase domain and four root-specific lipid transfer domain genes. The completed sequence also includes the centromere.

genomics

Improved genome inference in the MHC using a population reference graph

In humans and many other species, while much is known about the extent and structure of genetic variation, such information is typically not used in assembling novel genomes. Rather, a single reference is used against which to map reads, which can lead to poor characterisation of regions of high sequence or structural diversity. Here, we introduce a population reference graph, which combines multiple reference sequences as well as catalogues of SNPs and short indels. The genomes of novel samples are reconstructed as paths through the graph using an efficient hidden Markov Model, allowing for recombination between different haplotypes and variants. By applying the method to the 4.5Mb extended MHC region on chromosome 6, combining eight assembled haplotypes, sequences of known classical HLA alleles and 87,640 SNP variants from the 1000 Genomes Project, we demonstrate, using simulations, SNP genotyping, short-read and longread data, how the method improves the accuracy of genome inference. Moreover, the analysis reveals regions where the current set of reference sequences is substantially incomplete, particularly within the Class II region, indicating the need for continued development of reference-quality genome sequences.

Genomics

Predicting genome sizes and restriction enzyme recognition-sequence probabilities across the eukaryotic tree of life

High-throughput sequencing of reduced representation libraries obtained through digestion with restriction enzymes - generically known as restriction-site associated DNA sequencing (RAD-seq) - is a common strategy to generate genome-wide genotypic and sequence data from eukaryotes. A critical design element of any RAD-seq study is a knowledge of the approximate number of genetic markers that can be obtained for a taxon using different restriction enzymes, as this number determines the scope of a project, and ultimately defines its success. This number can only be directly determined if a reference genome sequence is available, or it can be estimated if the genome size and restriction recognition sequence probabilities are known. However, both scenarios are uncommon for non-model species. Here, we performed systematic in silico surveys of recognition sequences, for diverse and commonly used type II restriction enzymes across the eukaryotic tree of life. Our observations reveal that recognition-sequence frequencies for a given restriction enzyme are strikingly variable among broad eukaryotic taxonomic groups, being largely determined by phylogenetic relatedness. We demonstrate that genome sizes can be predicted from cleavage frequency data obtained with restriction enzymes targeting neutral elements. Models based on genomic compositions are also effective tools to accurately calculate probabilities of recognition sequences across taxa, and can be applied to species for which reduced-representation data is available (including transcriptomes and neutral RAD-seq datasets). The analytical pipeline developed in this study, PredRAD (https://github.com/phrh/PredRAD), and the resulting databases constitute valuable resources that will help guide the design of any study using RAD-seq or related methods.

Genomics

A Comprehensive Assessment of Somatic Mutation Calling in Cancer Genomes

The emergence of next generation DNA sequencing technology is enabling high-resolution cancer genome analysis. Large-scale projects like the International Cancer Genome Consortium (ICGC) are systematically scanning cancer genomes to identify recurrent somatic mutations. Second generation DNA sequencing, however, is still an evolving technology and procedures, both experimental and analytical, are constantly changing. Thus the research community is still defining a set of best practices for cancer genome data analysis, with no single protocol emerging to fulfil this role. Here we describe an extensive benchmark exercise to identify and resolve issues of somatic mutation calling. Whole genome sequence datasets comprising tumor-normal pairs from two different types of cancer, chronic lymphocytic leukaemia and medulloblastoma, were shared within the ICGC and submissions of somatic mutation calls were compared to verified mutations and to each other. Varying strategies to call mutations, incomplete awareness of sources of artefacts, and even lack of agreement on what constitutes an artefact or real mutation manifested in widely varying mutation call rates and somewhat low concordance among submissions. We conclude that somatic mutation calling remains an unsolved problem. However, we have identified many issues that are easy to remedy that are presented here. Our study highlights critical issues that need to be addressed before this valuable technology can be routinely used to inform clinical decision-making.\n\nAbbreviations and Definitionsaligner = mapper, these terms are used interchangeably

Genomics

Mapping bias overestimates reference allele frequencies at the HLA genes in the 1000 Genomes Project phase I data

Next Generation Sequencing (NGS) technologies have become the standard for data generation in studies of population genomics, as the 1000 Genomes Project (1000G). However, these techniques are known to be problematic when applied to highly polymorphic genomic regions, such as the Human Leukocyte Antigen (HLA) genes. Because accurate genotype calls and allele frequency estimations are crucial to population ge-nomics analises, it is important to assess the reliability of NGS data. Here, we evaluate the reliability of genotype calls and allele frequency estimates of the SNPs reported by 1000G (phase I) at five HLA genes (HLA-A, -B, -C, -DRB1, -DQB1). We take advantage of the availability of HLA Sanger sequencing of 930 of the 1,092 1000G samples, and use this as a gold standard to benchmark the 1000G data. We document that 18.6% of SNP genotype calls in HLA genes are incorrect, and that allele frequencies are estimated with an error higher than {+/-}0.1 at approximately 25% of the SNPs in HLA genes. We found a bias towards overestimation of reference allele frequency for the 1000G data, indicating mapping bias is an important cause of error in frequency estimation in this dataset. We provide a list of sites that have poor allele frequency estimates, and discuss the outcomes of including those sites in different kinds of analyses. Since the HLA region is the most polymorphic in the human genome, our results provide insights into the challenges of using of NGS data at other genomic regions of high diversity.\n\nData available in public repositories\n\nhttps://github.com/deboraycb/reliability_hla_1000g

Genomics

Inexpensive Multiplexed Library Preparation for Megabase-Sized Genomes

Whole-genome sequencing has become an indispensible tool of modern biology. However, the cost of sample preparation relative to the cost of sequencing remains high, especially for small genomes where the former is dominant. Here we present a protocol for the rapid and inexpensive preparation of hundreds of multiplexed genomic libraries for Illumina sequencing. By carrying out the Nextera tagmentation reaction in small volumes, replacing costly reagents with cheaper equivalents, and omitting unnecessary steps, we achieve a cost of library preparation of $8 per sample, approximately 6 times cheaper than the widely-used Nextera XT protocol. Furthermore, our procedure takes less than 5 hours for 96 samples and uses nanograms of genomic DNA. Many hundreds of samples can then be pooled on the same HiSeq lane via custom barcodes. Our method is especially useful for re-sequencing of large numbers of full microbial or viral genomes, including those from evolution experiments, genetic screens, and environmental samples.

Genomics

Extensive de novo mutation rate variation between individuals and across the genome of Chlamydomonas reinhardtii

Describing the process of spontaneous mutation is fundamental for understanding the genetic basis of disease, the threat posed by declining population size in conservation biology, and in much evolutionary biology. However, directly studying spontaneous mutation is difficult because of the rarity of de novo mutations. Mutation accumulation (MA) experiments overcome this by allowing mutations to build up over many generations in the near absence of natural selection. In this study, we sequenced the genomes of 85 MA lines derived from six genetically diverse wild strains of the green alga Chlamydomonas reinhardtii. We identified 6,843 spontaneous mutations, more than any other study of spontaneous mutation. We observed seven-fold variation in the mutation rate among strains and that mutator genotypes arose, increasing the mutation rate dramatically in some replicates. We also found evidence for fine-scale heterogeneity in the mutation rate, driven largely by the sequence flanking mutated sites, and by clusters of multiple mutations at closely linked sites. There was little evidence, however, for mutation rate heterogeneity between chromosomes or over large genomic regions of 200Kbp. Using logistic regression, we generated a predictive model of the mutability of sites based on their genomic properties, including local GC content, gene expression level and local sequence context. Our model accurately predicted the average mutation rate and natural levels of genetic diversity of sites across the genome. Notably, trinucleotides vary 17-fold in rate between the most mutable and least mutable sites. Our results uncover a rich heterogeneity in the process of spontaneous mutation both among individuals and across the genome.

Genomics

Distinctive Features of a Saudi Genome

We have fully sequenced the genome of an individual from the region of Saudi Arabia. In order to facilitate comparative analysis, an initial characterization of the new genome was undertaken based on single nucleotide polymorphism (SNP). The SNP data having associated population statistics, essentially the HapMap, served to identify features that were rare by comparison. Methods were developed and applied to tag observed SNPs as different and were extended to identify strings or clusters of difference in the individual relative to comparison populations to effectively increase the significance over single SNP comparison. Difference strings identified in the individual relative to each comparison population showed a genome location pattern with various levels of overlap between the comparison populations. The SNP frequencies from the HapMap population samples Ceu and Yri showed a difference inversion relative to the sample genome. The total SNP difference count was greatest between the individual and the Yri population sample while the number and total span of SNP difference clusters was greatest in comparison with the Ceu population sample. The final pattern of difference clusters has served to define distinctive features in the individual genome toward preliminary characterization.

Genomics

The power of single molecule real-time sequencing technology in the de novo assembly of a eukaryotic genome

Second-generation sequencers (SGS) have been game-changing, achieving cost-effective whole genome sequencing in many non-model organisms. However, a large portion of the genomes still remains unassembled. We reconstructed azuki bean (Vigna angularis) genome using single molecule real-time (SMRT) sequencing technology and achieved the best contiguity and coverage among currently assembled legume crops. The SMRT-based assembly produced 100 times longer contigs with 100 times smaller amount of gaps compared to the SGS-based assemblies. A detailed comparison between the assemblies revealed that the SMRT-based assembly enabled a more comprehensive gene annotation than the SGS-based assemblies where thousands of genes were missing or fragmented. A chromosome-scale assembly was generated based on the high-density genetic map, covering 86% of the azuki bean genome. We demonstrated that SMRT technology, though still needed support of SGS data, achieved a near-complete assembly of a eukaryotic genome.

Genomics

The two-speed genomes of filamentous pathogens: waltz with plants

Fungi and oomycetes include deep and diverse lineages of eukaryotic plant pathogens. The last 10 years have seen the sequencing of the genomes of a multitude of species of these so-called filamentous plant pathogens. Already, fundamental concepts have emerged. Filamentous plant pathogen genomes tend to harbor large repertoires of genes encoding virulence effectors that modulate host plant processes. Effector genes are not randomly distributed across the genomes but tend to be associated with compartments enriched in repetitive sequences and transposable elements. These findings have led to the \"two-speed genome\" model in which filamentous pathogen genomes have a bipartite architecture with gene sparse, repeat rich compartments serving as a cradle for adaptive evolution. Here, we review this concept and discuss how plant pathogens are great model systems to study evolutionary adaptations at multiple time scales. We will also introduce the next phase of research on this topic.

Genomics

The distribution and impact of common copy-number variation in the genome of the domesticated apple, Malus x domestica Borkh.

BackgroundCopy number variation (CNV) is a common feature of eukaryotic genomes, and a growing body of evidence suggests that genes affected by CNV are enriched in processes that are associated with environmental responses. Here we use next generation sequence (NGS) data to detect copy-number variable regions (CNVRs) within the Malus x domestica genome, as well as to examine their distribution and impact.\n\nMethodsCNVRs were detected using NGS data derived from 30 accessions of M. x domestica analyzed using the read-depth method, as implemented in the CNVrd2 software. To improve the reliability of our results, we developed a quality control and analysis procedure that involved checking for organelle DNA, not repeat masking, and the determination of CNVR identity using a permutation testing procedure.\n\nResultsOverall, we identified 876 CNVRs, which spanned 3.5% of the apple genome. To verify that detected CNVRs were not artifacts, we analyzed the B-allele-frequencies (BAF) within a single nucleotide polymorphism (SNP) array dataset derived from a screening of 185 individual apple accessions and found the CNVRs were enriched for SNPs having aberrant BAFs (P < 1e-13, Fishers Exact test). Putative CNVRs overlapped 845 gene models and were enriched for resistance (R) gene models (P < 1e-22, Fishers exact test). Of note was a cluster of resistance gene models on chromosome 2 near a region containing multiple major gene loci conferring resistance to apple scab.\n\nConclusionWe present the first analysis and catalogue of CNVRs in the M. x domestica genome. The enrichment of the CNVRs with R gene models and their overlap with gene loci of agricultural significance draw attention to a form of unexplored genetic variation in apple. This research will underpin further investigation of the role that CNV plays within the apple genome.

Genomics

Whole genome sequence analyses of Western Central African Pygmy hunter-gatherers reveal a complex demographic history and identify candidate genes under positive natural selection

African Pygmies practicing a mobile hunter-gatherer lifestyle are phenotypically and genetically diverged from other anatomically modern humans, and they likely experienced strong selective pressures due to their unique lifestyle in the Central African rainforest. To identify genomic targets of adaptation, we sequenced the genomes of four Biaka Pygmies from the Central African Republic and jointly analyzed these data with the genome sequences of three Baka Pygmies from Cameroon and nine Yoruba famers. To account for the complex demographic history of these populations that includes both isolation and gene flow, we fit models using the joint allele frequency spectrum and validated them using independent approaches. Our two best-fit models both suggest ancient divergence between the ancestors of the farmers and Pygmies, 90,000 or 150,000 years ago. We also find that bi-directional asymmetric gene-flow is statistically better supported than a single pulse of unidirectional gene flow from farmers to Pygmies, as previously suggested. We then applied complementary statistics to scan the genome for evidence of selective sweeps and polygenic selection. We found that conventional statistical outlier approaches were biased toward identifying candidates in regions of high mutation or low recombination rate. To avoid this bias, we assigned P-values for candidates using whole-genome simulations incorporating demography and variation in both recombination and mutation rates. We found that genes and gene sets involved in muscle development, bone synthesis, immunity, reproduction, cell signaling and development, and energy metabolism are likely to be targets of positive natural selection in Western African Pygmies or their recent ancestors.

Genomics