Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Correction of autoimmune IL2RA mutations in primary human T cells using non-viral genome targeting

Human T cells are central to physiological immune homeostasis, which protects us from pathogens without collateral autoimmune inflammation. They are also the main effectors in most current cancer immunotherapy strategies1. Several decades of work have aimed to genetically reprogram T cells for therapeutic purposes2-5, but as human T cells are resistant to most standard methods of large DNA insertion these approaches have relied on recombinant viral vectors, which do not target transgenes to specific genomic sites6, 7. In addition, the need for viral vectors has slowed down research and clinical use as their manufacturing and testing is lengthy and expensive. Genome editing brought the promise of specific and efficient insertion of large transgenes into target cells through homology-directed repair (HDR), but to date in human T cells this still requires viral transduction8, 9. Here, we developed a non-viral, CRISPR-Cas9 genome targeting system that permits the rapid and efficient insertion of individual or multiplexed large (>1 kilobase) DNA sequences at specific sites in the genomes of primary human T cells while preserving cell viability and function. We successfully tested the potential therapeutic use of this approach in two settings. First, we corrected a pathogenic IL2RA mutation in primary T cells from multiple family members with monogenic autoimmune disease and demonstrated enhanced signalling function. Second, we replaced the endogenous T cell receptor (TCR) locus with a new TCR redirecting T cells to a cancer antigen. The resulting TCR-engineered T cells specifically recognized the tumour antigen, with concomitant cytokine release and tumour cell killing. Taken together, these studies provide preclinical evidence that non-viral genome targeting will enable rapid and flexible experimental manipulation and therapeutic engineering of primary human immune cells.

genetics

Identifying Pleiotropic Effects: A Two-Stage Approach Using Genome-Wide Association Meta-Analysis Data

Pleiotropic effects occur when a single genetic variant independently influences multiple phenotypes. In genetic epidemiological studies, multiple endo-phenotypes or correlated traits are commonly tested separately in a univariate statistical framework to identify associations with genetic determinants. Subsequently, a simple look-up of overlapping univariate results is applied to identify pleiotropic genetic effects. However, this strategy offers limited power to detect pleiotropy. In contrast, combining correlated traits into a composite test provides a powerful approach for detecting pleiotropic genes. Here, we propose a two-stage approach to identify potential pleiotropic effects by utilizing aggregated results from large-scale genome-wide association (GWAS) meta-analyses. In the first stage, we developed two novel approaches (direct linear combining, dLC; and empirical combining, eLC) combining correlated univariate test statistics to screen potential pleiotropic variants on a genome-wide scale, using either individual-level or aggregated data. Our simulations indicated that dLC and eLC outperform other popular multivariate approaches (such as principal component analysis (PCA), multivariate analysis of variance (MANOVA), canonical correlation (CCA), generalized estimation equations (GEE), linear mixed effects models (LME) and OBrien combining approach). In particular, eLC provides a notable increase in power when the genetic variant exhibits both protective and deleterious effects. In the second stage, we developed a unique approach, conditional pleiotropy testing (cPLT), to examine pleiotropic effects using individual-level data for candidate variants identified in Stage 1. Simulation demonstrated reduced type 1 error for cPLT in identifying pleiotropic genetic variants compared to the typical conditional strategy. We validated our two-stage approach by performing a bivariate GWA study on two correlated quantitative traits, high-density lipoprotein (HDL) and triglycerides (TG), in the Genetic Analysis Workshop 16 (GAW16) simulation dataset. In summary, the proposed two-stage approach allows us to leverage aggregated summary statistics from univariate GWAS and improves the power to identify potential pleiotropy while maintaining valid false-positive rates.\n\nAuthor SummaryPleiotropy, occurring when a single genetic variant contributes to multiple phenotypes, remains difficult to identify in genome-wide association studies (GWAS). To leverage data for multiple phenotypes and incorporate univariate GWAS summary results, we propose a novel two-stage approach for discovering potential pleiotropic variants. In the first stage, two novel combining approaches were developed to screen potential pleiotropic variants on a genome-wide scale. Simulations demonstrated the superior statistical power of these approaches over other multivariate methods. In the second stage, our approach was used to identify potential pleiotropy in the candidate marker sets generated from the first stage. The proposed two-stage approach was applied to the GAW16 simulation dataset to discover pleiotropic variants associated with high-density lipoprotein and triglycerides. In summary, we demonstrate that the proposed two-stage approach can be applied as a viable and robust strategy to accommodate phenotypic and genetic heterogeneity for discovering potential pleiotropy on genome-wide scale.

genetics

Enabling rapid cloud-based analysis of thousands of human genomes via Butler

We present Butler, a computational framework developed in the context of the international Pan-cancer Analysis of Whole Genomes (PCAWG)1 project to overcome the challenges of orchestrating analyses of thousands of human genomes on the cloud. Butler operates equally well on public and academic clouds. This highly flexible framework facilitates management of virtual cloud infrastructure, software configuration, genomics workflow development, and provides unique capabilities in workflow execution management. By comprehensively collecting and analysing metrics and logs, performing anomaly detection as well as notification and cluster self-healing, Butler enables large-scale analytical processing of human genomes with 43% increased throughput compared to prior setups. Butler was key for delivering the germline genetic variant call-sets in 2,834 cancer genomes analysed by PCAWG1.

bioinformatics

Efficient management and analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr

MotivationGenome-wide datasets produced for association studies have dramatically increased in size over the past few years, with modern datasets commonly including millions of variants measured in dozens of thousands of individuals. This increase in data size is a major challenge severely slowing down genomic analyses. Specialized software for every part of the analysis pipeline have been developed to handle large genomic data. However, combining all these software into a single data analysis pipeline might be technically difficult.\n\nResultsHere we present two R packages, bigstatsr and bigsnpr, allowing for management and analysis of large scale genomic data to be performed within a single comprehensive framework. To address large data size, the packages use memory-mapping for accessing data matrices stored on disk instead of in RAM. To perform data pre-processing and data analysis, the packages integrate most of the tools that are commonly used, either through transparent system calls to existing software, or through updated or improved implementation of existing methods. In particular, the packages implement a fast derivation of Principal Component Analysis, functions to remove SNPs in Linkage Disequilibrium, and algorithms to learn Polygenic Risk Scores on millions of SNPs. We illustrate applications of the two R packages by analysing a case-control genomic dataset for the celiac disease, performing an association study and computing Polygenic Risk Scores. Finally, we demonstrate the scalability of the R packages by analyzing a simulated genome-wide dataset including 500,000 individuals and 1 million markers on a single desktop computer.\n\nAvailabilityhttps://privefl.github.io/bigstatsr/ & https://privefl.github.io/bigsnpr/\n\nContactflorian.prive@univ-grenoble-alpes.fr & michael.blum@univ-grenoble-alpes.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Drosophila larval brain neoplasms present tumour-type dependent genome instability

Single nucleotide polymorphisms (SNPs) and copy number variants (CNVs) are found at different rates in human cancer. To determine if these genetic lesions appear in Drosophila tumours we have sequenced the genomes of 17 malignant neoplasms caused by mutations in l(3)mbt, brat, aurA, or lgl. We have found CNVs and SNPs in all the tumours. Tumour-linked CNVs range between 11 and 80 per sample, affecting between 92 and 1546 coding sequences. CNVs are in average less frequent in l(3)mbt than in brat lines. Nearly half of the CNVs fall within the 10 to 100Kb range, all tumour samples contain CNVs larger that 100 Kb and some have CNVs larger than 1Mb. The rates of tumour-linked SNPs change more than 20-fold depending on the tumour type: late stage brat, l(3)mbt, and aurA and lgl lines present median values of SNPs/Mb of exome of 0.16, 0.48, and 3.6, respectively. Higher SNP rates are mostly accounted for by C>A transversions, which likely reflect enhanced oxidative stress conditions in the affected tumours. Both CNVs and SNPs turn over rapidly. We found no evidence for selection of a gene signature affected by CNVs or SNPs in the cohort. Altogether, our results show that the rates of CNVs and SNPs, as well as the distribution of CNV sizes in this cohort of Drosophila tumours are well within the range of those reported for human cancer. Genome instability is therefore inherent to Drosophila malignant neoplastic growth at a variable extent that is tumour type dependent.\n\nAUTHOR SUMMARYDrosophila models of malignant growth can help to understand the molecular mechanisms of malignancy. These models are known to exhibit some of the hallmarks of cancer like sustained growth, immortality, metabolic reprogramming, and others. However, it is currently unclear if these fly models are affected by genome instability, which is another hallmark of many human malignant tumours. To address this issue we have sequenced and analysed the genomes of a cohort of seventeen fly tumour samples. We have found that genome instability is a common trait of Drosophila malignant tumours, which occurs at an extent that is tumour-type dependent, at rates that are similar to those of different human cancers.

cancer biology

MeDEStrand: an improved method to infer genome-wide absolute methylation level from DNA enrichment experiment

BackgroundDNA methylation of dinucleotide CpG is an essential epigenetic modification that plays a key role in transcription. Bisulfite conversion method is a \"gold standard\" for DNA methylation profiling that provides single nucleotide resolution. However, whole-genome bisulfite conversion is very expensive. Alternatively, DNA enrichment-based methods offer high coverage of methylated CpG dinucleotides with the lowest cost per CpG covered genome-wide and have been used widely. They measure the DNA enrichment of methyl-CpG binding, therefore do not directly provide absolute methylation levels. Further, the enrichment is influenced by confounding factors besides the methylation status, e.g., CpG density. Computational models that can accurately derive the absolute methylation levels from the enrichment data are necessary.\n\nResultsWe present MeDEStrand, a method uses sigmoid function to estimate and correct the CpG bias from the numbers of reads that fell within bins that divide the genome. In addition, unlike the previous methods, which estimate CpG bias based on reads mapped at the same genomic loci, MeDEStrand processes the reads for the positive and negative DNA strands separately. We compare the performance of MeDEStrand with three other state-of-the-art methods MEDIPS, BayMeth and QSEA on four independent datasets generated using immortalized cell lines (GM12878 and K562) and human patient primary cells (foreskin fibroblast and mammary epithelial). Based on the comparison between the inferred absolute methylation levels from MeDIP-seq and the corresponding RRBS data, MeDEStrand shows the best performance at high resolution of 25, 50 and 100 base pairs.\n\nConclusions MeDEStrand benefits from the estimation of CpG bias with a sigmoid function and the procedure to process reads mapped to the positive and negative DNA strands separately. MeDEStrand is a tool to infer whole-genome absolute DNA methylation level at the cost of enrichment-based methods with adequate accuracy and resolution. R package MeDEStrand and its tutorial is freely available for download at https://github.com/jxu1234/MeDEStrand.git

bioinformatics

An alignment method for nucleic acid sequences against annotated genomes

MotivationBiological sequence alignment is fundamental to their further interpretation. Current alignment algorithms typically align either nucleic acid or amino acid sequences. Using only nucleic acid sequence similarity, divergent sequences cannot be aligned reliably because of the limited alphabet and genetic saturation. To align divergent coding nucleic acid sequences, one can align using the translated amino acid sequences. This requires the detection of the correct open reading frame, is prone to eventual frame shift errors, and typically requires the treatment of genes separately. It was our motivation to design a nucleic acid sequence alignment algorithm to align a nucleic acid sequence against a (reference) genome sequence, that works equally well for similar and divergent sequences, and produces an optimal alignment considering simultaneously the alignment of all annotated coding sequences.\n\nResultsWe define a genome alignment score for evaluating the quality of an alignment of a nucleic acid query sequence against a reference genome sequence, for which coding sequence features have been annotated (for example in a GenBank record). The genome alignment score combines the a ne gap score for the nucleic acid sequence with an a ne gap score for all amino acid alignments resulting from coding sequences in open reading frames contained within the query sequence. We present a Dynamic Programming algorithm to compute the optimal global or local alignment using this genomic alignment score and provide a formal proof of correctness. This algorithm allows the alignment of nucleic acid sequences from closely related and highly divergent sequences within the same software and using the same parameters, automatically correcting any eventual frame shift errors and produces at the same time the aligned translated amino acid sequences of all relevant coding sequence features.\n\nAvailabilityThe software is available as a web application at http://www.genomedetective.com/app/aga and as command-line application at https://github.com/emweb/aga

bioinformatics

Bayesian Reconstruction of Transmission within Outbreaks using Genomic Variants

Pathogen genome sequencing can reveal details of transmission histories and is a powerful tool in the fight against infectious disease. In particular, within-host pathogen genomic variants identified through heterozygous nucleotide base calls are a potential source of information to identify linked cases and infer direction and time of transmission. However, using such data effectively to model disease transmission presents a number of challenges, including differentiating genuine variants from those observed due to sequencing error, as well as the specification of a realistic model for within-host pathogen population dynamics.\n\nHere we propose a new Bayesian approach to transmission inference, BadTrIP (BAyesian epiDemiological TRansmission Inference from Polymorphisms), that explicitly models evolution of pathogen populations in an outbreak, transmission (including transmission bottlenecks), and sequencing error. BadTrIP enables the inference of host-to-host transmission from pathogen sequencing data and epidemiological data. By assuming that genomic variants are unlinked, our method does not require the computationally intensive and unreliable reconstruction of individual haplotypes. Using simulations we show that BadTrIP is robust in most scenarios and can accurately infer transmission events by efficiently combining information from genetic and epidemiological sources; thanks to its realistic model of pathogen evolution and the inclusion of epidemiological data, BadTrIP is also more accurate than existing approaches. BadTrIP is distributed as an open source package (https://bitbucket.org/nicofmay/badtrip) for the phylogenetic software BEAST2.\n\nWe apply our method to reconstruct transmission history at the early stages of the 2014 Ebola outbreak, showcasing the power of within-host genomic variants to reconstruct transmission events.\n\nAuthor SummaryWe present a new tool to reconstruct transmission events within outbreaks. Our approach makes use of pathogen genetic information, notably genetic variants at low frequency within host that are usually discarded, and combines it with epidemiological information of host exposure to infection. This leads to accurate reconstruction of transmission even in cases where abundant within-host pathogen genetic variation and weak transmission bottlenecks (multiple pathogen units colonising a new host at transmission) would otherwise make inference difficult due to the transmission history differing from the pathogen evolution history inferred from pathogen isolets. Also, the use of within-host pathogen genomic variants increases the resolution of the reconstruction of the transmission tree even in scenarios with limited within-outbreak pathogen genetic diversity: within-host pathogen populations that appear identical at the level of consensus sequences can be discriminated using within-host variants. Our Bayesian approach provides a measure of the confidence in different possible transmission histories, and is published as open source software. We show with simulations and with an analysis of the beginning of the 2014 Ebola outbreak that our approach is applicable in many scenarios, improves our understanding of transmission dynamics, and will contribute to finding and limiting sources and routes of transmission, and therefore preventing the spread of infectious disease.

genetics

MetQy: an R package to query metabolic functions of genes and genomes

SummaryWith the rapid accumulation of sequencing data from genomic and metagenomic studies, there is an acute need for better tools that facilitate their analyses against biological functions. To this end, we developed MetQy, an open-source R package designed for query-based analysis of functional units in [meta]genomes and/or sets of genes using the The Kyoto Encyclopedia of Genes and Genomes (KEGG) database. Furthermore, MetQy contains visualization and analysis tools and facilitates KEGGs flat file manipulation. Thus, MetQy enables better understanding of metabolic capabilities of known genomes or user-specified [meta]genomes by using the available information and can help guide studies in microbial ecology, metabolic engineering and synthetic biology.\n\nAvailability and ImplementationThe MetQy R package is freely available and can be downloaded from our groups website (http://osslab.lifesci.warwick.ac.uk) or GitHub (https://github.com/OSS-Lab/MetQy).\n\nContactO.Soyer@warwick.ac.uk

microbiology

Widespread ancient whole genome duplications in Malpighiales coincide with Eocene global climatic upheaval

Ancient whole genome duplications (WGDs) are important in eukaryotic genome evolution, and are especially prominent in plants. Recent genomic studies from large vascular plant clades, including ferns, gymnosperms, and angiosperms suggest that WGDs may represent a crucial mode of speciation. Moreover, numerous WGDs have been dated to events coinciding with major episodes of global and climatic upheaval, including the mass extinction at the KT boundary (~65 Ma) and during more recent intervals of global aridification in the Miocene (~10-5 Ma). These findings have led to the hypothesis that polyploidization may buffer lineages against the negative consequences of such disruptions. Here, we explore WGDs in the large, and diverse flowering plant clade Malpighiales using a combination of transcriptomes and complete genomes from 42 species. We conservatively identify 22 ancient WGDs, widely distributed across Malpighiales subclades. Our results provide strong support for the hypothesis that WGD is an important mode of speciation in plants. Importantly, we also identify that these events are clustered around the Eocene-Paleocene Transition (~54 Ma), during which time the planet was warmer and wetter than any period in the Cenozoic. These results establish that the Eocene Climate Optimum represents another, previously unrecognized, period of prolific WGDs in plants, and lends support to the hypothesis that polyploidization promotes adaptation and enhances plant survival during major episodes of global change. Malpighiales, in particular, may have been particularly influenced by these events given their predominance in the tropics where Eocene warming likely had profound impacts owing to the relatively tight thermal tolerances of tropical organisms.\n\nSignificance StatementWhole genome duplications (WGDs) are hypothesized to generate adaptive variations during episodes of climate change and global upheaval. Using large-scale phylogenomic assessments, we identify an impressive 22 ancient WGDs in the large, tropical flowering plant clade Malpighiales. This supports growing evidence that ancient WGDs are far more common than has been thought. Additionally, we identify that WGDs are clustered during a narrow window of time, ~54 Ma, when the climate was warmer and more humid than during any period in the last ~65 Ma. This lends support to the hypothesis that WGDs are associated with surviving climatic upheavals, especially for tropical organisms like Malpighiales, which have tight thermal tolerances.

evolutionary biology

GrapeTree: Visualization of core genomic relationships among 100,000 bacterial pathogens

O_LICurrent methods struggle to reconstruct and visualise the genomic relationships of [≥]100,000 bacterial genomes.\nC_LIO_LIGrapeTree facilitates the analyses of allelic profiles from 10,000s of core genomes within a web browser window.\nC_LIO_LIGrapeTree implements a novel minimum spanning tree algorithm to reconstruct genetic relationships despite missing data together with a static \"GrapeTree Layout\" algorithm to render interactive visualisations of large trees.\nC_LIO_LIGrapeTree is a stand-along package for investigating Newick trees plus associated metadata and is also integrated into EnteroBase to facilitate cutting edge navigation of genomic relationships among >160,000 genomes from bacterial pathogens.\nC_LIO_LIThe GrapeTree package was released under the GPL v3.0 Licence.\nC_LI

bioinformatics

Deficiency of global genome nucleotide excision repair explains mutational signature observed in cancer

Nucleotide excision repair (NER) is one of the main DNA repair pathways that protect cells against genomic damage. Disruption of this pathway can contribute to the development of cancer and accelerate aging. Tumors deficient in NER are more sensitive to cisplatin treatment. Characterization of the mutational consequences of NER-deficiency may therefore provide important diagnostic opportunities. Here, we analyzed the somatic mutational profiles of adult stem cells (ASCs) from NER-deficient Ercc1-/{Delta} mice, using whole-genome sequencing analysis of clonally derived organoid cultures. Our results indicate that NER-deficiency increases the base substitution load in liver, but not in small intestinal ASCs, which coincides with a tissue-specific aging-pathology observed in these mice. The mutational landscape changes as a result of NER-deficiency in ASCs of both tissues and shows an increased contribution of Signature 8 mutations, which is a pattern with unknown etiology that is recurrently observed in various cancer types. The scattered genomic distribution of the acquired base substitutions indicates that deficiency of global-genome NER (GG-NER) is responsible for the altered mutational landscape. In line with this, we observed increased Signature 8 mutations in a GG-NER-deficient human organoid culture in which XPC was deleted using CRISPR-Cas9 gene-editing. Furthermore, genomes of NER-deficient breast tumors show an increased contribution of Signature 8 mutations compared with NER-proficient tumors. Elevated levels of Signature 8 mutations may therefore serve as a biomarker for NER-deficiency and could improve personalized cancer treatment strategies.

molecular biology

The origin and remolding of genomic islands of differentiation in the European sea bass

Speciation is a complex process that leads to the progressive establishment of reproductive isolation barriers between diverging populations. Genome-wide comparisons between closely related species have revealed the existence of heterogeneous divergence patterns, dominated by genomic islands of increased divergence supposed to contain reproductive isolation loci. However, this divergence landscape only provides a static picture of the dynamic process of speciation, during which confounding mechanisms unlinked to speciation can interfere. Here, we used haplotype-resolved whole-genome sequences to identify the mechanisms responsible for the formation of genomic islands between Atlantic and Mediterranean sea bass lineages. We show that genomic islands first emerged in allopatry through the effect of linked selection acting on a heterogeneous recombination landscape. Upon secondary contact, preexisting islands were strongly remolded by differential introgression, revealing variable fitness effects among regions involved in reproductive isolation. Interestingly, we found that divergent regions containing ancient polymorphisms conferred the strongest resistance to introgression.

evolutionary biology

Predicting DNA accessibility in the pan-cancer tumor genome using RNA-seq, WGS, and deep learning

DNA accessibility, chromatin regulation, and genome methylation are key drivers of transcriptional events promoting tumor growth. However, understanding the impact of DNA sequence data on transcriptional regulation of gene expression is a challenge, particularly in noncoding regions of the genome. Recently, neural networks have been used to effectively predict DNA accessibility in multiple specific cell types [14]. These models make it possible to explore the impact of mutations on DNA accessibility and transcriptional regulation.\n\nOur work first improved on prior cell-specific accessibility prediction, obtaining a mean receiver operating characteristic (ROC) area under the curve (AUC) = 0.910 and mean precision-recall (PR) AUC = 0.605, compared to the previous mean ROC AUC = 0.895 and mean PR AUC = 0.561 [14].\n\nOur key contribution extended the model to enable accessibility predictions on any new sample for which RNA-seq data is available, without requiring cell-type-specific DNase-seq data for re-training. This new model obtained overall PR AUC = 0.621 and ROC AUC = 0.897 when applied across whole genomes of new samples whose biotypes were held out from training, and PR AUC = 0.725 and ROC AUC = 0.913 on randomly held out new samples whose biotypes were allowed to overlap with training.\n\nMore significantly, we showed that for promoter and promoter flank regions of the genome our model predicts accessibility to high reliability, achieving PR AUC = 0.839 in held out biotypes and PR AUC = 0.911 in randomly held out samples.\n\nThis performance is not sensitive to whether the promoter and flank regions fall within genes used in the input RNA-seq expression vector.\n\nFinally, we utilize this tool to investigate, for the first time, promoter accessibility patterns across several cohorts from The Cancer Genome Atlas (TCGA) [27].

bioinformatics

Identification and analysis of mobile genetic elements in Gibbon genome

Recent sequencing of genome of northern white-cheeked gibbon (Nomascus leucogenys) has provided important insight into fast evolution of gibbons and signatures relevant to gibbon biology. It was revealed that mobile genetic elements (MGE) seems to play major role in gibbon evolution. Here we report that most of the gibbon genome is occupied by the MGEs such as ALUs, MIRs, LINE1, LINE 2, LINE 3, ERVL, ERV-class1, ERV-class II and other DNA elements which include hAT Charlie and TcMar tigger. We provide detailed description and genome wide distribution of all the MGEs present in gibbon genome. Previously, it was reported that gibbon-specific retrotransposon (LAVA) tend to insert into chromosome segregation genes and alter transcription by providing a premature termination site, suggesting a possible molecular mechanism for the genome plasticity of the gibbon lineage. We show that insertion sites of LAVA elements present atypical signals/patterns which are different from typical signals present at insertion sites of Alu elements. This suggests possibility of distinct insertion mechanism used by LAVA elements for their insertions. We also find similarity in signals of LAVA elements insertion sites with atypical signals present at Alus /L1s insertion sites disrupting the genes leading to diseases such as cancer and Duchenne muscular dystrophy. This suggest role of LAVA in premature transcription termination.

bioinformatics

SYSTEMATIC ANALYSES OF AUTOSOMAL RECOMBINATION RATES FROM THE 1000 GENOMES PROJECT UNCOVERS THE GLOBAL RECOMBINATION LANDSCAPE IN HUMANS

BackgroundMeiotic recombination plays an important role in evolution by shuffling different alleles along the chromosomes, thus generating the genetic diversity across generations that is vital for adaptation. The plasticity of recombination rates and presence of hotspots of recombination along the genome has attracted much attention over two decades due to their contribution to the evolution of the genome. Yet, the variation in genome-wide recombination landscape and the differences in the location and strength of hotspots across worldwide human populations remains little explored.\n\nResultsWe make use of the untapped linkage disequilibrium (LD) based genetic maps from the 1000 Genomes Project (1KGP) to perform in-depth analyses of finescale variation in the autosomal recombination rates across 20 human populations to uncover the global recombination landscape. We have generated a detailed map of human recombination landscape comprising of a comprehensive set of 88,841 putative hotspots and 80,129 coldspots with their respective strengths across populations, about 2/3rd of which were previously unknown. We have validated and assessed the number of historical putative hotspots derived from the patterns of LD that are currently active in the contemporary populations using a recently published high-resolution pedigree-based genetic map, constructed and refined using 3.38 million crossovers from various populations. For the first time, we provide statistics regarding the conserved, shared, and unique hotspots across all the populations studied.\n\nConclusionsOur analysis yields clusters of continental groups, reflecting their shared ancestry and genetic similarities in the recombination rates that are linked to the migratory and evolutionary histories of the populations. We provide the genomic locations and strengths of hotspots and coldspots across all the populations studied which are a valuable set of resources arising out our analyses of 1KGP data. The findings are of great importance for further research on human hotspots as we approach the dusk of retiring HapMap-based resources.

evolutionary biology

Comparative Phylogenomic Synteny Network Analysis of Mammalian and Angiosperm Genomes

BackgroundSynteny analysis is a valuable approach for understanding eukaryotic gene and genome evolution, but still relies largely on pairwise or reference-based comparisons. Network approaches can be utilized to expand large-scale phylogenomic microsynteny studies. There is now a wealth of completed mammalian (animal) and angiosperm (plant) genomes, two very important lineages that have evolved and radiated over the last ~170 million years. Genomic organization and conservation differs greatly between these two groups; however, a systematic and comparative characterization of synteny between the two lineages using the same approaches and metrics has not been undertaken.\n\nResultsWe have built complete microsynteny networks for 87 mammalian and 107 angiosperm genomes, which contain 1,464,753 nodes (genes) and 49,426,268 edges (syntenic connections between genes) for mammals, and 2,234,461 nodes and 46,938,272 edges for angiosperms, respectively. Exploiting network statistics, we present the functional characteristics of extremely conserved and diversified gene families. We summarize the features of all syntenic gene clusters and present lineage-wide phylogenetic profiling, revealing intriguing sub-clade lineage-specific clusters. We depict several representative clusters of important developmental genes in humans, such as CENPJ, p53 and NFE2. Finally, we present the complete homeobox gene family networks for both mammals (including Hox and ParaHox gene clusters) and angiosperms.\n\nConclusionsOur results illustrate and quantify overall synteny conservation and diversification properties of all annotated genes for mammals and angiosperms and show that plant genomes are in general more dynamic.

evolutionary biology

Seave: a comprehensive web platform for storing and interrogating human genomic variation

Capability for genome sequencing and variant calling has increased dramatically, enabling large scale genomic interrogation of human disease. However, discovery is hindered by the current limitations in genomic interpretation, which remains a complicated and disjointed process. We introduce Seave, a web platform that enables variants to be easily filtered and annotated with in silico pathogenicity prediction scores and annotations from popular disease databases. Seave stores genomic variation of all types and sizes, and allows filtering for specific inheritance patterns, quality values, allele frequencies and gene lists. Seave is open source and deployable locally, or on a cloud computing provider, and works readily with gene panel, exome and whole genome data, scaling from single labs to multi-institution scale.

bioinformatics