Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Genome-wide genetic data on ~500,000 UK Biobank participants

The UK Biobank project is a large prospective cohort study of ~500,000 individuals from across the United Kingdom, aged between 40-69 at recruitment. A rich variety of phenotypic and health-related information is available on each participant, making the resource unprecedented in its size and scope. Here we describe the genome-wide genotype data (~805,000 markers) collected on all individuals in the cohort and its quality control procedures. Genotype data on this scale offers novel opportunities for assessing quality issues, although the wide range of ancestries of the individuals in the cohort also creates particular challenges. We also conducted a set of analyses that reveal properties of the genetic data - such as population structure and relatedness - that can be important for downstream analyses. In addition, we phased and imputed genotypes into the dataset, using computationally efficient methods combined with the Haplotype Reference Consortium (HRC) and UK10K haplotype resource. This increases the number of testable variants by over 100-fold to ~96 million variants. We also imputed classical allelic variation at 11 human leukocyte antigen (HLA) genes, and as a quality control check of this imputation, we replicate signals of known associations between HLA alleles and many common diseases. We describe tools that allow efficient genome-wide association studies (GWAS) of multiple traits and fast phenome-wide association studies (PheWAS), which work together with a new compressed file format that has been used to distribute the dataset. As a further check of the genotyped and imputed datasets, we performed a test-case genome-wide association scan on a well-studied human trait, standing height.

genetics

Introgression patterns between house mouse subspecies and species reveal genomic windows offrequent exchange

Based on whole genome sequencing data, we have studied the patterns of introgression in a phylogenetically well defined set of populations, sub-species and species of mice (Mus m. domesticus, Mus m. musculus, Mus m. castaneus and Mus spretus). We find that many discrete genomic regions are subject to repeated and mutual introgression and exchange. The majority of these regions code for genes that are involved in parasite defense or genomic conflict. They include genes involved in adaptive immunity, such as the MHC region or antibody coding regions, but also genes involved in innate immune reactions of the epidermis. We find also clusters of KRAB zinc finger proteins that control the spread of transposable elements and genes that are involved in meiotic drive. These findings suggest that even well separated populations and species maintain the capacity to exchange genetic material in a special set of evolutionary active genes.

evolutionary biology

Vibrio cholerae genomic diversity within and between patients

Cholera is a severe, waterborne diarrheal disease caused by toxin-producing strains of the bacterium Vibrio cholerae. Comparative genomics has revealed \"waves\" of cholera transmission and evolution, in which clones are successively replaced over decades and centuries. However, the extent of V. cholerae genetic diversity within an epidemic or even within an individual patient is poorly understood. Here, we characterized V. cholerae genomic diversity at a micro-epidemiological level within and between individual patients from Bangladesh and Haiti. To capture within-patient diversity, we isolated multiple (8 to 20) V. cholerae colonies from each of eight patients, sequenced their genomes and identified point mutations and gene gain/loss events. We found limited but detectable diversity at the level of point mutations within hosts (zero to three single nucleotide variants within each patient), and comparatively higher gene content variation within hosts (at least one gain/loss event per patient, and up to 103 events in one patient). Much of the gene content variation appeared to be due to gain and loss of phage and plasmids within the V. cholerae population, with occasional exchanges between V. cholerae and other members of the gut microbiota. We also show that certain intra-host variants have phenotypic consequences. For example, the acquisition of a Bacteroides plasmid and nonsynonymous mutations in a sensor histidine kinase gene both reduced biofilm formation, an important trait for environmental survival. Together, our results show that V. cholerae is measurably evolving within patients, with possible implications for disease outcomes and transmission dynamics.\n\nAuthor SummaryVibrio cholerae is the etiological agent of cholera, a severe diarrheal disease endemic to Bangladesh and responsible for global outbreaks, including one ongoing in Haiti. Certain bacterial pathogens can evolve and diversify within the human host, often altering virulence and antibiotic resistance. However, most examples of within-host evolution have come from chronic infections, in which the pathogen has sufficient time to mutate and diversify, and little attention has been paid to more acute infections such as the one caused by V. cholerae. The goal of this study was to measure the extent of within-host evolution of V. cholerae within individual infected patients. By sequencing multiple bacterial isolates_from each of eight patients from Bangladesh and Haiti, we found that cholera patients can harbor a diverse population of V. cholerae. As expected for an acute infection, this diversity is limited, ranging from zero to three point mutations (single nucleotide variants) per patient. However, gene gain/loss events are more prevalent than point mutations, occurring in every single patient, and sometimes involving the transfer of dozens of genes on plasmids. Even if rare, point mutations and gene gain/loss events may be maintained by natural selection, and can alter clinically-and environmentally-relevant phenotypes such as biofilm formation. Therefore, within-patient evolution has the potential to impact clinical and epidemiological outcomes. Together, our results demonstrate that within-patient evolution may be a general feature of both acute and chronic infections, and that gene gain/loss may be an important but under-appreciated feature of within-host evolution.

microbiology

C-BERST: Defining subnuclear proteomic landscapes at genomic elements with dCas9-APEX2

Mapping proteomic composition at distinct genomic loci and subnuclear landmarks in living cells has been a long-standing challenge. Here we report that dCas9-APEX2 Biotinylation at genomic Elements by Restricted Spatial Tagging (C-BERST) allows the rapid, unbiased mapping of proteomes near defined genomic loci, as demonstrated for telomeres and centromeres. By combining the spatially restricted enzymatic tagging enabled by APEX2 with programmable DNA targeting by dCas9, C-BERST has successfully identified nearly 50% of known telomere-associated factors and many known centromere-associated factors. We also identified and validated SLX4IP and RPA3 as telomeric factors, confirming C-BERSTs utility as a discovery platform. C-BERST enables the rapid, high-throughput identification of proteins associated with specific sequences, facilitating annotation of these factors and their roles in nuclear and chromosome biology.

molecular biology

Use of whole genome sequencing to investigate an outbreak of gonorrhoea among females in urban New South Wales, Australia, 2012 to 2014.

Increasing rates of gonorrhoea have been observed among urban heterosexuals within the Australian state of New South Wales (NSW). Here, we applied whole genome sequencing (WGS) to better understand transmission dynamics. Ninety-four isolates of a particular N. gonorrhoeae genotype (G122) associated with female patients (years 2012 to 2014) underwent phylogenetic analysis using core single nucleotide polymorphisms (SNPs). Context for genetic variation was provided by including an unbiased selection of 1,870 N. gonorrhoeae genomes from a recent United Kingdom (UK) study. NSW genomes formed a single clade, with the majority of isolates belonging to one of five clusters, and comprised patients of varying age groups. Intra-patient variability was less than 7 core SNPs. Several patients had indistinguishable core SNPs, suggesting a common infection source. These data have provided an enhanced understanding of transmission of N. gonorrhoeae among urban heterosexuals in NSW, Australia, and highlight the value of using WGS in N. gonorrhoeae outbreak investigations.

epidemiology

CHIC: a short read aligner for pan-genomic references

Recently the topic of computational pan-genomics has gained increasing attention, and particularly the problem of moving from a single-reference paradigm to a pan-genomic one. Perhaps the simplest way to represent a pan-genome is to represent it as a set of sequences. While indexing highly repetitive collections has been intensively studied in the computer science community, the research has focused on efficient indexing and exact pattern patching, making most solutions not yet suitable to be used in bioinformatic analysis pipelines.\n\nResultsWe present CHIC, a short-read aligner that indexes very large and repetitive references using a hybrid technique that combines Lempel-Ziv compression with Burrows-Wheeler read aligners.\n\nAvailabilityOur tool is open source and available online at https://gitlab.com/dvalenzu/CHIC

bioinformatics

Genome shuffling in a globalized bacterial plant pathogen: Recombination-mediated evolution in Xanthomonas euvesicatoria and X. perforans

Bacterial recombination and clonality underly the evolution and epidemiology of pathogenic lineages as well as their cosmopolitan spread. While the spread of stable clonal bacterial pathotypes drives disease epidemics, recombination leads to the evolution of new bacterial lineages. Recombinant lineages of plant bacterial pathogens are typically associated with colonization of novel hosts and emergence of new diseases. Here, we show that recombination between evolutionarily and phenotypically distinct plant pathogenic lineages has generated new recombinant lineages with unique combinations of pathogenicity and virulence factors. X. euvesicatoria (Xe) and X. perforans (Xp) are two closely related monophyletic species causing bacterial spot disease on tomato and pepper worldwide. We sequenced the genomes of strains representing populations on tomato in Nigeria and found shuffling of secretion systems and effectors such that these strains contain genes from both Xe and Xp. Multiple strains, from populations in Nigeria, Italy, and Florida, USA, exhibited extensive genomewide homologous recombination and both species exhibited dynamic open pangenomes. Our results show that recombination is generating new lineages of bacterial spot pathogens on tomato with consequences for disease management strategies.\n\nImportanceThe Xanthomonas pathogens that cause bacterial spot of tomato and pepper have been model systems for plant-microbe interactions. Two of these pathogens, X. euvesicatoria and X. perforans, are very closely related. Genome sequences of bacterial spot field strains from Nigeria, Italy, and the United States showed varying levels of homologous recombination that changed the amino sequence of effectors, secretion systems, and other proteins. This shuffling of genome content occurred between X. euvesicatoria and X. perforans, while a Nigerian lineage also contained the lipopolysaccharide cluster of a distantly related Xanthomonas species. Gene content varied among strains and the affected genes are important in the establishment of disease, therefore our findings point to global variation in the host-pathogen interaction driven by gene exchange among evolutionarily distinct lineages.

evolutionary biology

Hemimetabolous genomes reveal molecular basis of termite eusociality

Around 150 million years ago, eusocial termites evolved from within the cockroaches, 50 million years before eusocial Hymenoptera, such as bees and ants, appeared. Here, we report the first, 2GB genome of a cockroach, Blattella germanica, and the 1.3GB genome of the drywood termite, Cryptotermes secundus. We show evolutionary signatures of termite eusociality by comparing the genomes and transcriptomes of three termites and the cockroach against the background of 16 other eusocial and non-eusocial insects. Dramatic adaptive changes in genes underlying the production and perception of pheromones confirm the importance of chemical communication in the termites. These are accompanied by major changes in gene regulation and the molecular evolution of caste determination. Many of these results parallel molecular mechanisms of eusocial evolution in Hymenoptera. However, the specific solutions are remarkably different, thus revealing a striking case of convergence in one of the major evolutionary transitions in biological complexity.

evolutionary biology

Phylogenetic approaches to identifying fragments of the same gene, with application to the wheat genome

As the time and cost of sequencing decrease, the number of available genomes and transcriptomes rapidly increases. Yet the quality of the assemblies and the gene annotations varies considerably and often remains poor, affecting downstream analyses. This is particularly true when fragments of the same gene are annotated as distinct genes and consequently wrongly appear as paralogs. In this study, we introduce two novel phylogenetic tests to infer non-overlapping or partially overlapping genes that are in fact parts of the same gene. One approach collapses branches with low bootstrap support and the other computes a likelihood ratio test. We extensively validated these methods by 1) introducing and recovering fragmentation on the bread wheat, Triticum aestivum cv. Chinese Spring, chromosome 3B; 2) by applying the methods to the low-quality 3B assembly and validating predictions against the high-quality 3B assembly; and 3) by comparing the performance of the proposed methods to the performance of existing methods, namely Ensembl Compara and ESPRIT. Application of this combination to a draft shotgun assembly of the entire bread wheat genome revealed 1221 pairs of genes which are highly likely to be fragments of the same gene. Our approach demonstrates the power of fine-grained evolutionary inferences across multiple species to improving genome assemblies and annotations. An open source software tool is available at https://github.com/DessimozLab/esprit2.

bioinformatics

Molecular Analysis of Parasites in the Choreocolacaceae (Rhodophyta) Reveals a Reduced Harveyella mirabilis Plastid Genome and Supports the Transfer of Genera to the Rhodomelaceae (Rhodophyta)

Parasitism is a life strategy that has repeatedly evolved within the Florideophyceae. Historically, the terms adelphoparasite and alloparasite have been used to distinguish parasites based on the relative phylogenetic relationship of host and parasite. However, analyses using molecular phylogenetics indicate that nearly all red algal parasites infect within their taxonomic family, and a range of relationships exist between host and parasite. To date, all investigated adelphoparasites have lost their plastid, and instead, incorporate a host derived plastid when packaging spores. In contrast, a highly reduced plastid lacking photosynthesis genes was sequenced from the alloparasite Choreocolax polysiphoniae. Here we present the complete Harveyella mirabilis plastid genome, which has also lost genes involved in photosynthesis, and a partial plastid genome from Leachiella pacifica. The H. mirabilis plastid shares more synteny with free-living red algal plastids than that of C. polysiphoniae. Phylogenetic analysis demonstrates that C. polysiphoniae, H. mirabilis, and L. pacifica form a robustly supported clade of parasites, which retain their own plastid genomes, within the Rhodomelaceae. We therefore transfer all three genera from the exclusively parasitic family, Choreocolacaceae, to the Rhodomelaceae. Additionally, we recommend applying the terms archaeplastic parasites (formerly alloparasites), and neoplastic parasites (formerly adelphoparasites) to distinguish red algal parasites using a biological framework rather than taxonomic affiliation with their hosts.

evolutionary biology

Semi-Parametric Covariate-Modulated Local False Discovery Rate For Genome-Wide Association Studies

While genome-wide association studies (GWAS) have discovered thousands of risk loci for heritable disorders, so far even very large meta-analyses have recovered only a fraction of the heritability of most complex traits. Recent work utilizing variance components models has demonstrated that a larger fraction of the heritability of complex phenotypes is captured by the additive effects of SNPs than is evident only in loci surpassing genome-wide significance thresholds, typically set at a Bonferroni-inspired p [≤] 5 x 10-8. Procedures that control false discovery rate can be more powerful, yet these are still under-powered to detect the majority of non-null effects from GWAS. The current work proposes a novel Bayesian semi-parametric two-group mixture model and develops a Markov Chain Monte Carlo (MCMC) algorithm for a covariate-modulated local false discovery rate (cmfdr). The probability of being non-null depends on a set of covariates via a logistic function, and the non-null distribution is approximated as a linear combination of B-spline densities, where the weight of each B-spline density depends on a multinomial function of the covariates. The proposed methods were motivated by work on a large meta-analysis of schizophrenia GWAS performed by the Psychiatric Genetics Consortium (PGC). We show that the new cmfdr model fits the PGC schizophrenia GWAS test statistics well, performing better than our previously proposed parametric gamma model for estimating the non-null density and substantially improving power over usual fdr. Using loci declared significant at cmfdr [≤] 0.20, we perform follow-up pathway analyses using the Kyoto Encyclopedia of Genes and Genomes (KEGG) homo sapiens pathways database. We demonstrate that the increased yield from the cmfdr model results in an improved ability to test for pathways associated with schizophrenia compared to using those SNPs selected according to usual fdr.

bioinformatics

Genome-Wide Association Studies Identify 15 Genetic Markers Associated with Marmite Taste Preference

Marmite is a popular food eaten around the world, to which individuals have commonly considered themselves either \"lovers\" or \"haters\". We aimed to determine whether this food preference has a genetic basis.Weperformed a genome-wide association study (GWAS) for Marmite taste preference using genotype and questionnaire data froma cohort of 261 healthy adults. We found 1 single nucleotide polymorphism (SNP) associated with Marmite taste preferencethat reached genome-wide significance (p<5x10-8) in our GWAS analyses. We found another 4 SNPsassociated with Marmite taste preference that reached genome-wide significance (p<5x10-8) in at leastoneGWAS and/or for at least one phenotype analysed. Moreover, we identified 10 additional SNPs potentially associated with Marmite taste preference through candidate gene analysis. Our results indicate that there is a genetic basis to Marmite taste preference and we have identified 15 genetic markers for this trait. Overall, we conclude that Marmite tastepreference is a complex human trait influenced by multiple genetic markers, as well as the environment.\n\nSummary of Main Results O_TEXTBOXO_LIMarmite taste preference is a complex human trait with many factors influencing whether an individual loves or hates Marmite.\nC_LIO_LIThe relative contribution of genetics versus environment (ie. heritability) for Marmite taste preference is unknown.\nC_LIO_LIThe genetic contribution to Marmite taste preference involves multiple genetic markers each contributing a small amount (ie. the trait is polygenic). There is not one single Marmite gene with a large contribution like in thecase of the TAS2R38gene and bitter taste perception.\nC_LIO_LIWe have found a total of 15 SNPs associated with Marmite taste preference: 5 SNPs by a genetic-association screen atgenome-wide significance, and 10 SNPs by a candidate gene approach at nominal significance.\nC_LIO_LIWe did not find an association between the TAS2R38bitter taste receptor gene and Marmite taste preference.\nC_LIO_LIIt is important to independently replicate the findings of this study in order to validate these genetic markers and get a more accurate idea of their true effect on Marmite taste preference.\nC_LI\n\nC_TEXTBOX

genetics

The number of orphans in yeast and fly is drastically reduced by using combining searches in both proteomes and genomes

The identification of de novo created genes is important as it provides a glimpse on the evolutionary processes of gene creation. Potential de novo created genes are identified by selecting genes that have no homologs outside a particular species, but for an accurate detection this identification needs to be correct.\n\nGenes without any homologs are often referred to as orphans; in addition to de novo created ones, fast evolving genes or genes lost in all related genomes might also be classified as orphans. The identification of orphans is dependent on: (i) a method to detect homologs and (ii) a database including genes from related genomes.\n\nHere, we set out to investigate how the detection of orphans is influenced by these two factors. Using Saccharomyces cerevisiae we identify that best strategy is to use a combination of searching annotated proteins and a six-frame translation of all ORFs from closely related genomes. Using this strategy we obtain a set of 54 orphans in Drosophila melanogaster and 38 in Drosophila pseudoobscura, significantly less than what is reported in some earlier studies.

bioinformatics

Whole genome sequence-based haplotypes reveal single origin of the sickle allele during the Holocene Wet Phase

Five classical designations of sickle haplotypes are based on the presence/absence of restriction sites and named after ethnic groups or geographic regions from which patients originated. Each haplotype is thought to represent an independent occurrence of the sickle mutation. We investigated the origins of the sickle mutation using whole genome sequence data. We identified 156 carriers from the 1000 Genomes Project, the African Genome Variation Project, and Qatar. We defined a new haplotypic classification using 27 polymorphisms in linkage disequilibrium with rs334. Network analysis revealed a common haplotype that differed from the ancestral haplotype only by the derived sickle mutation at rs334 and correlated collectively with the Central African Republic/Bantu, Cameroon, and Arabian/Indian designations. Other haplotypes were derived from this haplotype and fell into two clusters, one comprised of haplotypes correlated with the Senegal designation and the other comprised of haplotypes correlated with both the Benin and Senegal designations. The near-exclusive presence of the original sickle haplotype in the Central African Republic, Kenya, Uganda, and South Africa is consistent with this haplotype predating the Bantu Expansion. Modeling of balancing selection indicated that the heterozygote advantage was 15.2%, an equilibrium frequency of 12.0% was reached after 87 generations, and the selective environment predated the mutation. The posterior distribution of the ancestral recombination graph yielded an age of the sickle mutation of 259 generations, corresponding to 7,300 years and the Holocene Wet Phase. These results clarify the origin of the sickle allele and improve and simplify the classification of sickle haplotypes.

genetics

Theory, practice, and conservation in the age of genomics: the Galapagos giant tortoise as a case study

High-throughput DNA sequencing allows efficient discovery of thousands of single nucleotide polymorphisms (SNPs) in non-model species. Population genetic theory predicts that this large number of independent markers should provide detailed insights into population structure, even when only a few individuals are sampled. Still, sampling design can have a strong impact on such inferences. Here, we use simulations and empirical SNP data to investigate the impacts of sampling design on estimating genetic differentiation among populations that represent three species of Galapagos giant tortoises (Chelonoidis spp.). Though microsatellite and mitochondrial DNA analyses have supported the distinctiveness of these species, a recent study called into question how well these markers matched with data from genomic SNPs, thereby questioning decades of studies in non-model organisms. Using >20,000 genome-wide SNPs from 30 individuals from three Galapagos giant tortoise species, we find distinct structure that matches the relationships described by the traditional genetic markers. Furthermore, we confirm that accurate estimates of genetic differentiation in highly structured natural populations can be obtained using thousands of SNPs and 2-5 individuals, or hundreds of SNPs and 10 individuals, but only if the units of analysis are delineated in a way that is consistent with evolutionary history. We show that the lack of structure in the recent SNP-based study was likely due to unnatural grouping of individuals and erroneous genotype filtering. Our study demonstrates that genomic data enable patterns of genetic differentiation among populations to be elucidated even with few samples per population, and underscores the importance of sampling design. These results have specific implications for studies of population structure in endangered species and subsequent management decisions.\n\n\"Modern molecular techniques provide unprecedented power to understand genetic variation in natural populations. Nevertheless, application of this information requires sound understanding of population genetics theory.\"\n\n- Fred Allendorf (2017, p. 420)

evolutionary biology

A whole genome analysis of the red-crowned crane provides insight into avian longevity

The red-crowned crane (Grus japonensis) is an endangered and large-bodied crane native to East Asia. It is a traditional symbol of longevity and its long lifespan has been confirmed both in captivity and in the wild. Lifespan in birds is positively correlated with body size and negatively correlated with metabolic rate; although the genetic mechanisms for the red-crowned cranes long lifespan have not previously been investigated. Using whole genome sequencing and comparative evolutionary analyses against the grey-crowned crane and other avian genomes, we identified candidate genes that are correlated with longevity. Included among these are positively selected genes with known associations with longevity in metabolism and immunity pathways (NDUFA5, NDUFA8, NUDT12 IL9R, SOD3, NUDT12, PNLIP, CTH, and RPA1). Our analyses provide genetic evidence for low metabolic rate and longevity, accompanied by possible convergent adaptation signatures among distantly related large and long-lived birds. Finally, we identified low genetic diversity in the red-crowned crane, consistent with its listing as an endangered species, and we hope this genome will provide a useful genetic resource for future conservation studies of this rare and iconic species.

bioinformatics

Inferring demographic parameters in bacterial genomic data using Bayesian and hybrid phylogenetic methods

BackgroundRecent developments in sequencing technologies make it possible to obtain genome sequences from a large number of isolates in a very short time. Bayesian phylogenetic approaches can take advantage of these data by simultaneously inferring the phylogenetic tree, evolutionary timescale, and demographic parameters (such as population growth rates), while naturally integrating uncertainty in all parameters. Despite their desirable properties, Bayesian approaches can be computationally intensive, hindering their use for outbreak investigations involving genome data for a large numbers of pathogen isolates. An alternative to using full Bayesian inference is to use a hybrid approach, where the phylogenetic tree and evolutionary timescale are estimated first using maximum likelihood. Under this hybrid approach, demographic parameters are inferred from estimated trees instead of the sequence data, using maximum likelihood, Bayesian inference, or approximate Bayesian computation. This can vastly reduce the computational burden, but has the disadvantage of ignoring the uncertainty in the phylogenetic tree and evolutionary timescale.\n\nResultsWe compared the performance of a fully Bayesian and a hybrid method by analysing six whole-genome SNP data sets from a range of bacteria and simulations. The estimates from the two methods were very similar, suggesting that the hybrid method is a valid alternative for very large datasets. However, we also found that congruence between these methods is contingent on the presence of strong temporal structure in the data (i.e. clocklike behaviour), which is typically verified using a date-randomisation test in a Bayesian framework. To reduce the computational burden of this Bayesian test we implemented a date-randomisation test using a rapid maximum likelihood method, which has similar performance to its Bayesian counterpart.\n\nConclusionsHybrid approaches can produce reliable inferences of evolutionary timescales and phylodynamic parameters in a fraction of the time required for fully Bayesian analyses. As such, they are a valuable alternative in outbreak studies involving a large number of isolates.

bioinformatics

An extensive enhancer-promoter map generated by genome-scale analysis of enhancer and gene activity patterns

Massive efforts have documented hundreds of thousands of putative enhancers in the human genome. A pressing genomic challenge is to identify which of these enhancers are functional and map them to the genes they regulate. We developed a novel method for inferring enhancer-promoter (E-P) links based on correlated activity patterns across many samples. Our method, called FOCS, uses rigorous statistical validation tailored for zero-inflated data, identifying the most important E-P links in each gene model. We applied FOCS to the wide epigenomic and transcriptomic datasets recorded by the ENCODE, Roadmap Epigenomics and FANTOM5 projects, together covering 2,630 samples of human primary cells, tissues and cell lines. In addition, building on expression of enhancer RNAs (eRNAs) as an exquisite mark of enhancer activity and on the robust detection of eRNAs by the GRO-seq technique, we compiled a compendium of eRNA and gene expression profiles based on public GRO-seq data from 245 samples and 23 human cell types. Applying FOCS to this compendium further expanded the coverage of our inferred E-P map. Benchmarking against gold standard E-P links from ChIA-PET and eQTL data, we demonstrate that FOCS prediction of E-P links outperforms extant methods. Collectively, we inferred >300,000 cross-validated E-P links spanning ~16K known genes. Our study presents an improved method for inferring regulatory links between enhancers and promoters, and provides an extensive resource of E-P maps that could greatly assist the functional interpretation of the noncoding regulatory genome. FOCS and our predicted E-P map are publicly available at http://acgt.cs.tau.ac.il/focs.

bioinformatics