Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Free energy based high-resolution modeling of CTCF-mediated chromatin loops for human genome

A thermodynamic method for computing the stability and dynamics of chromatin loops is proposed. The CTCF-mediated interactions as observed in ChIA-PET experiments for human B-lymphoblastoid cells are evaluated in terms of a polymer model for chain folding physical properties and the experimentally observed frequency of contacts within the chromatin regions. To estimate the optimal free energy and a Boltzmann distribution of suboptimal structures, the approach uses dynamic programming with methods to handle degeneracy and heuristics to compute parallel and antiparallel chain stems and pseudoknots. Moreover, multiple loops mediated by CTCF proteins connected together and forming multimeric islands are simulated using the same model. Based on the thermodynamic properties of those topological three-dimensional structures, we predict the correlation between the relative activity of chromatin loop and the Boltzmann probability, or the minimum free energy, depending also on its genomic length. Segments of chromatin where the structures show a more stable minimum free energy (for a given genomic distance) tend to be inactive, whereas structures that have lower stability in the minimum free energy (with the same genomic distance) tend to be active.

genomics

Hallmarks of early sex-chromosome evolution in the dioecious plant Mercurialis annua revealed by de novo genome assembly, genetic mapping and transcriptome analysis

Suppressed recombination around a sex-determining locus allows divergence between homologous sex chromosomes and the functionality of their genes. Here, we reveal patterns of the earliest stages of sex-chromosome evolution in the diploid dioecious herb Mercurialis annua on the basis of cytological analysis, de novo genome assembly and annotation, genetic mapping, exome resequencing of natural populations, and transcriptome analysis. Both genetic mapping and exome resequencing of individuals across the species range independently identified the largest linkage group, LG1, as the sex chromosome. Although the sex chromosomes of M. annua are karyotypically homomorphic, we estimate that about a third of the Y chromosome has ceased recombining, a region containing 568 transcripts and spanning 22.3 cM in the corresponding female map. Patterns of gene expression hint at the possible role of sexually antagonistic selection in having favored suppressed recombination. In total, the genome assembly contained 34,105 expressed genes, of which 10,076 were assigned to linkage groups. There was limited evidence of Y-chromosome degeneration in terms of gene loss and pseudogenization, but sequence divergence between the X and Y copies of many sex-linked genes was higher than between M. annua and its dioecious sister species M. huetii with which it shares a sex-determining region. The Mendelian inheritance of sex in interspecific crosses, combined with the other observed pattern, suggest that the M. annua Y chromosome has at least two evolutionary strata: a small old stratum shared with M. huetii, and a more recent larger stratum that is probably unique to M. annua and that stopped recombining about one million years ago. Article summaryPlants that evolved separate sexes (dioecy) recently are ideal models for studying the early stages of sex-chromosome evolution. Here, we use karyological, whole genome and transcriptome data to characterize the homomorphic sex chromosomes of the annual dioecious plant Mercurialis annua. Our analysis reveals many typical hallmarks of dioecy and sex-chromosome evolution, including sex-biased gene expression and high X/Y sequence divergence, yet few premature stop codons in Y-linked genes and very little outright gene loss, despite 1/3 of the sex chromosome having ceased recombination in males. Our results confirm that the M. annua species complex is a fertile system for probing early stages in the evolution of sex chromosomes.

genomics

Wild tobacco genomes reveal the evolution of nicotine biosynthesis

Nicotine, the signature alkaloid of Nicotiana species responsible for the addictive properties of human tobacco smoking, functions as a defensive neurotoxin against attacking herbivores. However, the evolution of the genetic features that contributed to the assembly of the nicotine biosynthetic pathway remains unknown. We sequenced and assembled genomes of two wild tobaccos, Nicotiana attenuata (2.5 Gb) and N. obtusifolia (1.5 Gb), two ecological models for investigating adaptive traits in nature. We show that after the Solanaceae whole genome triplication event, a repertoire of rapidly expanding transposable elements (TEs) bloated these Nicotiana genomes, promoted expression divergences among duplicated genes and contributed to the evolution of herbivory-induced signaling and defenses, including nicotine biosynthesis. The biosynthetic machinery that allows for nicotine synthesis in the roots evolved from the stepwise duplications of two ancient primary metabolic pathways: the polyamine and nicotinic acid dinucleotide (NAD) pathways. While the duplication of the former is shared among several Solanaceous genera which produce polyamine-derived tropane alkaloids, the innovation and efficient production of nicotine in the genus Nicotiana required lineage-specific duplications within the NAD pathway and the evolution of root-specific expression of the duplicated Solanaceae-specific ethylene response factor (ERF) that activates the expression of all nicotine biosynthetic genes. Furthermore, TE insertions that incorporated transcription factor binding motifs also likely contributed to the coordinated metabolic flux of the nicotine biosynthetic pathway. Together, these results provide evidence that TEs and gene duplications facilitated the emergence of a key metabolic innovation relevant to plant fitness.

genomics

A genome-wide interactome of DNA-associated proteins in the human liver

Large-scale efforts like the Encyclopedia of DNA Elements (ENCODE) Project have made tremendous progress in cataloging the genomic binding patterns of DNA-associated proteins (DAPs), such as transcription factors (TFs). However most chromatin immunoprecipitation-sequencing (ChIP-seq) analyses have focused on a few immortalized cell lines whose activities and physiology deviate in important ways from endogenous cells and tissues. Consequently, binding data from primary human tissue are essential to improving our understanding of in vivo gene regulation. Here we analyze ChIP-seq data for 20 DAPs assayed in two healthy human liver tissue samples, identifying more than 450,000 binding sites. We integrated binding data with transcriptome and phased whole genome data to investigate allelic DAP interactions and the impact of heterozygous sequence variation on the expression of neighboring genes. We find our tissue-based dataset demonstrates binding patterns more consistent with liver biology than cell lines, and describe uses of these data to better prioritize impactful non-coding variation. Collectively, our rich dataset offers novel insights into genome function in healthy liver tissue and provides a valuable research resource for assessing disease-related disruptions.

genomics

Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome

A primary goal of The Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to develop an African Diaspora Power Chip (ADPC), a genotyping array consisting of tagging SNPs, useful in comprehensively identifying African specific genetic variation. This array is designed based on the novel variation identified in 642 CAAPA samples of African ancestry with high coverage whole genome sequence data (~30x depth). This novel variation extends the pattern of variation catalogued in the 1000 Genomes and Exome Sequencing Projects to a spectrum of populations representing the wide range of West African genomic diversity. These individuals from CAAPA also comprise a large swath of the African Diaspora population and incorporate historical genetic diversity covering nearly the entire Atlantic coast of the Americas. Here we show the results of designing and producing such a microchip array. This novel array covers African specific variation far better than other commercially available arrays, and will enable better GWAS analyses for researchers with individuals of African descent in their study populations. A recent study1 cataloging variation in continental African populations suggests this type of African-specific genotyping array is both necessary and valuable for facilitating large-scale GWAS in populations of African ancestry.

genomics

LARGE-SCALE GENOMIC REORGANIZATIONS OF TOPOLOGICAL DOMAINS (TADs) AT THE HoxD LOCUS

BackgroundThe transcriptional activation of Hoxd genes during mammalian limb development involves dynamic interactions with the two Topologically Associating Domains (TADs) flanking the HoxD cluster. In particular, the activation of the most posterior Hoxd genes in developing digits is controlled by regulatory elements located in the centromeric TAD (C-DOM) through long-range contacts. To assess the structure-function relationships underlying such interactions, we measured compaction levels and TAD discreteness using a combination of chromosome conformation capture (4C-seq) and DNA FISH.\n\nResultsWe challenged the robustness of the TAD architecture by using a series of genomic deletions and inversions that impact the integrity of this chromatin domain and that remodel the long-range contacts. We report multi-partite associations between Hoxd genes and up to three enhancers and show that breaking the native chromatin topology leads to the remodelling of TAD structure.\n\nConclusionsOur results reveal that the re-composition of TADs architectures after severe genomic re-arrangements depends on a boundary-selection mechanism that uses CTCF-mediated gating of long-range contacts in combination with genomic distance and, to a certain extent, sequence specificity.

genomics

Population Genomics And The Evolution Of Virulence In The Fungal Pathogen Cryptococcus neoformans

Cryptococcus neoformans is an opportunistic fungal pathogen that causes approximately 625,000 deaths per year from nervous system infections. Here, we leveraged a unique, genetically diverse population of C. neoformans from sub-Saharan Africa, commonly isolated from mopane trees, to determine how selective pressures in the environment coincidentally adapted C. neoformans for human virulence. Genome sequencing and phylogenetic analysis of 387 isolates, representing the global VNI and African VNB lineages, highlighted a deep, non-recombining split in VNB (herein VNBI and VNBII). VNBII was enriched for clinical samples relative to VNBI, while phenotypic profiling of 183 isolates demonstrated that VNBI isolates were significantly more resistant to oxidative stress and more heavily melanized than VNBII isolates. Lack of melanization in both lineages was associated with loss-of-function mutations in the BZP4 transcription factor. A genome-wide association study across all VNB isolates revealed sequence differences between clinical and environmental isolates in virulence factors and stress response genes. Inositol transporters and catabolism genes, which process sugars present in plants and the human nervous system, were identified as targets of selection in all three lineages. Further phylogenetic and population genomic analyses revealed extensive loss of genetic diversity in VNBI, suggestive of a history of population bottlenecks, along with unique evolutionary trajectories for mating type loci. These data highlight the complex evolutionary interplay between adaptation to natural environments and opportunistic infections, and that selection on specific pathways may predispose isolates to human virulence.

genomics

Accurate and Reproducible Functional Maps in 127 Human Cell Types via 2D Genome Segmentation

The Roadmap Epigenomics consortium has published whole-genome functional annotation maps in 127 human cell types and cancer cell lines by integrating data from multiple epigenetic marks. These maps have thereby been widely used by the community for studying gene regulation in cell type specific contexts and predicting functional impacts of DNA mutations on disease. Here, we present a new map of functional elements produced by a recently published method called IDEAS on the same data set. The IDEAS method has several unique advantages and was shown to outperform existing methods, including the one used by the Roadmap Epigenomics consortium. We further introduce a simple but highly effective pipeline to greatly improve the reproducibility of functional annotation. Using five categories of independent experimental results, we extensively compared the annotation produced by IDEAS and the Roadmap Epigenomics consortium. While the overall concordance between the two maps was high, we observed many differences in the details and in the position-wise consistency of annotation across cell types. We show that the IDEAS annotation was uniformly and often substantially more accurate than the Roadmap Epigenomics result. This study therefore reports on the quality of an existing functional map in 127 human genomes and provides an alternative and better map to be used by the community. The annotation result can be visualized in the UCSC genome browser via the hub at http://bx.psu.edu/~yuzhang/Roadmap_ideas/ideas_hub.txt

genomics

Epistasis in genomic and survival data of cancer patients

Cancer aggressiveness and its effect on patient survival depends on mutations in the tumor genome. Epistatic interactions between the mutated genes may guide the choice of anticancer therapy and set predictive factors of its success. Inhibitors targeting synthetic lethal partners of genes mutated in tumors are already utilized for efficient and specific treatment in the clinic. The space of possible epistatic interactions, how-ever, is overwhelming, and computational methods are needed to limit the experimental effort of validating the interactions for therapy and characterizing their biomarkers. Here, we introduce SurvLRT, a statistical likelihood ratio test for identifying epistatic gene pairs and triplets from cancer patient genomic and survival data. Compared to established approaches, SurvLRT performed favorable in predicting known, experimentally verified synthetic lethal partners of PARP1 from TCGA data. Our approach is the first to test for epistasis between triplets of genes to identify biomarkers of synthetic lethality-based therapy. SurvLRT proved successful in identifying the known gene TP53BP1 as the biomarker of success of the therapy targeting PARP in BRCA1 deficient tumors. Search for other biomarkers for the same interaction revealed a region whose deletion was a more significant biomarker than deletion of TP53BP1. With the ability to detect not only pairwise but twelve different types of triple epistasis, applicability of SurvLRT goes beyond cancer therapy, to the level of characterization of shapes of fitness landscapes.\n\nAuthor SummaryGenomic alterations in tumors affect the fitness of tumor cells, controlling how well they replicate and survive compared to other cells. The landscape of tumor fitness is shaped by epistasis. Epistasis occurs when the contribution of gene alterations to the total fitness is non-linear. The type of epistatic genetic interactions with great potential for cancer therapy is synthetic lethality. Inhibitors targeting synthetic lethal partners of genes mutated in tumors can selectively kill tumor and not normal cells. Therapy based on synthetic lethality is, however, context dependent, and it is crucial to identify its biomarkers. Unfortunately, the space of possible interactions and their biomarkers is overwhelming for experimental validation. Computational pre-selection methods are required to limit the experimental effort. Here, we introduce a statistical approach called SurvLRT, for the identification of epistatic gene pairs and triplets based on patient genomic and survival data. First, we show that using SurvLRT, we can deliver synthetic lethal interactions of pairs of genes that are specific to cancer. Second, we demonstrate the applicability of SurvLRT to identify biomarkers for synthetic lethality, such as mutational status of other genes that can alleviate the synthetic effect.

genomics

Automated Structural Variant Verification In Human Genomes Using Single-Molecule Electronic DNA Mapping

The importance of structural variation in human disease and the difficulty of detecting structural variants larger than 50 base pairs has led to the development of several long-read sequencing technologies and optical mapping platforms. Frequently, multiple technologies and ad hoc methods are required to obtain a consensus regarding the location, size and nature of a structural variant, with no approach able to reliably bridge the gap of variant sizes between the domain of short-read approaches and the largest rearrangements observed with optical mapping.\n\nTo address this unmet need, we have developed a new software package, SV-Verify, which utilizes data collected with the Nabsys High Definition Mapping (HD-Mapping) system, to perform hypothesis-based verification of putative deletions. We demonstrate that whole genome maps, constructed from electronic detection of tagged DNA, hundreds of kilobases in length, can be used effectively to facilitate calling of structural variants ranging in size from 300 base pairs to hundreds of kilobase pairs. SV-Verify implements hypothesis-based verification of putative structural variants using a set of support vector machines and is capable of concurrently testing several thousand independent hypotheses. We describe support vector machine training, utilizing a well-characterized human genome, and application of the resulting classifiers to another human genome, demonstrating high sensitivity and specificity for deletions [≥]300 base pairs.

genomics

Genome-Wide Analysis Of Repetitive Elements Associated With Gene Regulation

Nearly half of the human genome is made up of transposable elements (TEs) and there is evidence that TEs are involved in gene regulation. Here, we have integrated publicly available genomic, epigenetic and transcriptomic data to investigate this in a genome-wide manner. A bootstrapping statistical method was applied to minimize the confounder effects from different repeat types. Our results show that although most TE classes are primarily associated with reduced gene expression, Alu elements are associated with up regulated gene expression. Furthermore, Alu elements had the highest probability of any TE class of contributing to regulatory regions of any type defined by chromatin state. This suggests a general model where clade specific SINEs may contribute more to gene regulation than ancient/ancestral TEs. Finally, non-coding regions were found to have a high probability of TE content within regulatory sequences, most notably in repressors. Our exhaustive analysis has extended and updated our understanding of TEs in terms of their global impact on gene regulation, and suggests that the most recently derived types of TEs, i.e. clade or species specific SINES, have the greatest overall impact on gene regulation.

genomics

Karyotype stability and unbiased fractionation in the paleo-allotetraploid Cucurbita genomes

The Cucurbita genus contains several economically important species in the Cucurbitaceae family. Interspecific hybrids between C. maxima and C. moschata are widely used as rootstocks for other cucurbit crops. We report high-quality genome sequences of C. maxima and C. moschata and provide evidence supporting an allotetraploidization event in Cucurbita. We are able to partition the genome into two homoeologous subgenomes based on different genetic distances to melon, cucumber and watermelon in the Benincaseae tribe. We estimate that the two diploid progenitors successively diverged from Benincaseae around 31 and 26 million years ago (Mya), and the allotetraploidization happened earlier than 3 Mya, when C. maxima and C. moschata diverged. The subgenomes have largely maintained the chromosome structures of their diploid progenitors. Such long-term karyotype stability after polyploidization is uncommon in plant polyploids. The two subgenomes have retained similar numbers of genes, and neither subgenome is globally dominant in gene expression. Allele-specific expression analysis in the C. maxima x C. moschata interspecific F1 hybrid and the two parents indicates the predominance of trans-regulatory effects underlying expression divergence of the parents, and detects transgressive gene expression changes in the hybrid correlated with heterosis in important agronomic traits. Our study provides insights into plant genome evolution and valuable resources for genetic improvement of cucurbit crops.

genomics

Genome reconstruction and characterisation of extensively drug-resistant bacterial pathogens through direct metagenomic sequencing of human faeces

Whole-genome sequencing of microbial pathogens is revolutionising modern approaches to outbreaks of infectious diseases and is reliant upon organism culture. Culture-independent methods have shown promise in identifying pathogens, but high level reconstruction of microbial genomes from microbiologically complex samples for more in-depth analyses remains a challenge. Here, using metagenomic sequencing of a human faecal sample and analysis by tetranucleotide frequency profiling projected onto emergent self-organising maps, we were able to reconstruct the underlying populations of two extensively-drug resistant pathogens, Klebsiella pneumoniae carbapenemase (KPC)-producing Klebsiella pneumoniae and vancomycin-resistant Enterococcus faecium. From these genomes, we were able to ascertain molecular typing results, such as MLST, and identify highly discriminatory mutations in the metagenome to distinguish closely related strains. These proof-of-principle results demonstrate the utility of clinical sample metagenomics to recover sequences of important drug-resistant bacteria and application of the approach in outbreak investigations, independent of the need to culture the organisms.

genomics

Natural selection shaped the rise and fall of passenger pigeon genomic diversity

The extinct passenger pigeon was once the most abundant bird in North America, and possibly the world. While theory predicts that large populations will be more genetically diverse and respond more efficiently to selection, passenger pigeon genetic diversity was surprisingly low. To investigate this we analysed 41 mitochondrial and 4 nuclear genomes from passenger pigeons, and 2 genomes from band-tailed pigeons, passenger pigeons closest living relatives. We find that passenger pigeons large population size allowed for faster adaptive evolution and removal of harmful mutations, but that this drove a huge loss in neutral genetic diversity. These results demonstrate how great an impact selection can have on a vertebrate genome, and invalidate previous results that suggested population instability contributed to this species surprisingly rapid extinction.

genomics

Genome-wide population diversity in Hymenoscyphus fraxineus points to an eastern Russian origin of European Ash dieback.

European forests are experiencing extensive invasion from the Ash pathogen Hymenoscyphus fraxineus, an ecological niche competitor to the non-pathogenic native congener H. albidus. We report the genome-wide diversity and population structure in Asia (native) and Europe (the introduced range). We show H. fraxineus underwent a dramatic bottleneck upon introduction to Europe around 30-40 generations ago, leaving a genomic signature, characterized by long segments of fixation, interspersed with \"diversity islands\" that are identical throughout Europe. This means no effective secondary contact with other populations has occurred. Genome-wide variation is consistently high within sampled locations in Japan and the Russian Far East, and lack of differentiation amongst Russian locations suggests extensive gene flow, similar to Europe. A local ancestry analysis supports Russia as a more likely source population than Japan. Negligible latency, rapid host-range expansion and viability of small founding populations specify strong biosecurity forewarnings against new introductions from outside Europe.

genomics

A Genome-wide Association and Admixture Mapping Study of Bronchodilator Drug Response in African Americans with Asthma

BackgroundShort-acting B2-adrenergic receptor agonists (SABAs) are the most commonly prescribed asthma medications worldwide. Response to SABAs is measured as bronchodilator drug response (BDR), which varies among racial/ethnic groups in the U.S 1, 2. However, the genetic variation that contributes to BDR is largely undefined in African Americans with asthma3\n\nObjectiveTo identify genetic variants that may contribute to differences in BDR in African Americans with asthma.\n\nMethodsWe performed a genome-wide association study of BDR in 949 African American children with asthma, genotyped with the Axiom World Array 4 (Affymetrix, Santa Clara, CA) followed by imputation using 1000 Genomes phase 3 genotypes. We used linear regression models adjusting for age, sex, body mass index and genetic ancestry to test for an association between BDR and genotype at single nucleotide polymorphisms (SNPs). To increase power and distinguish between shared vs. population-specific associations with BDR in children with asthma, we performed a meta-analysis across 949 African Americans and 1,830 Latinos (Total=2,779). Lastly, we performed genome-wide admixture mapping to identify regions whereby local African or European ancestry is associated with BDR in African Americans. Two additional populations of 416 Latinos and 1,325 African Americans were used to replicate significant associations.\n\nResultsWe identified a population-specific association with an intergenic SNP on chromosome 9q21 that was significantly associated with BDR (rs73650726, p=7.69 x 10-9). A trans-ethnic meta-analysis across African Americans and Latinos identified three additional SNPs within the intron of PRKG1 that were significantly associated with BDR (rs7903366, rs7070958, and rs7081864, p[≤]5 x 10-8).\n\nConclusionsOur findings indicate that both population specific and shared genetic variation contributes to differences in BDR in minority children with asthma, and that the genetic underpinnings of BDR may differ between racial/ethnic groups.\n\nKey messagesO_LIA GWAS for BDR in African American children with asthma identified an intergenic population specific variant at 9q21 to be associated with increased bronchodilator drug response (BDR).\nC_LIO_LIA meta-analysis of GWAS across African Americans and Latinos identified shared genetic variants at 10q21 in the intron of PRKG1 to be associated with differences in BDR.\nC_LIO_LIFurther genetic studies need to be performed in diverse populations to identify the full set of genetic variants that contribute to BDR.\nC_LI

genomics

The most developmentally truncated fishes show extensive Hox gene loss and miniaturized genomes

Hox genes play a fundamental role in regulating the embryonic development of all animals. Manipulation of these transcription factors in model organisms has unraveled key aspects of evolution, like the transition from fin to limb. However, by virtue of their fundamental role and pleiotropic effects, simultaneous knockouts of several of these genes pose significant challenges. Here, we report on evolutionary simplification in two species of the dwarf minnow genus Paedocypris using whole genome sequencing. The two species feature unprecedented Hox gene loss and genome reduction in association with their massive developmental truncation. We also show how other genes involved in the development of musculature, nervous system, and skeleton have been lost in Paedocypris, mirroring its highly progenetic phenotype. Further, we identify two mechanisms responsible for genome streamlining: severe intron shortening and reduced repeat content. As a naturally simplified system closely related to zebrafish, Paedocypris provides novel insights into vertebrate development.

genomics

Resolving systematic errors in widely-used enhancer activity assays in human cells enables genome-wide functional enhancer characterization

The identification of transcriptional enhancers in the human genome is a prime goal in biology. Enhancers are typically predicted via chromatin marks, yet their function is primarily assessed with plasmid-based reporter assays. Here, we show that two previous observations relating to plasmid-transfection into human cells render such assays unreliable: (1) the function of the bacterial plasmid origin-of-replication (ORI) as a conflicting core-promoter and (2) the activation of a type I interferon (IFN-I) response. These problems cause strongly confounding false-positives and -negatives in luciferase assays and genome-wide STARR-seq screens. We overcome both problems by directly employing the ORI as a core-promoter and by inhibiting two kinases central to IFN-I induction. This corrects luciferase assays and enables genome-wide STARR-seq screens in human cells. Comprehensive enhancer activity profiles in HeLa-S3 cells uncover strong enhancers, IFN-I-induced enhancers, and enhancers endogenously silenced at the chromatin level. Our findings apply to all episomal enhancer activity assays in mammalian cells, and are key to the characterization of human enhancers.

genomics