Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Genome-scale oscillations in DNA methylation during exit from pluripotency

Pluripotency is accompanied by the erasure of parental epigenetic memory with naive pluripotent cells exhibiting global DNA hypomethylation both in vitro and in vivo. Exit from pluripotency and priming for differentiation into somatic lineages is associated with genome-wide de novo DNA methylation. We show that during this phase, coexpression of enzymes required for DNA methylation turnover, DNMT3s and TETs, promotes cell-to-cell variability in this epigenetic mark. Using a combination of single-cell sequencing and quantitative biophysical modelling, we show that this variability is associated with coherent, genome-scale, oscillations in DNA methylation with an amplitude dependent on CpG density. Analysis of parallel single-cell transcriptional and epigenetic profiling provides evidence for oscillatory dynamics both in vitro and in vivo. These observations provide fresh insights into the emergence of epigenetic heterogeneity during early embryo development, indicating that dynamic changes in DNA methylation might influence early cell fate decisions.\n\nHighlightsO_LICo-expression of DNMT3s and TETs drive genome-scale oscillations of DNA methylation\nC_LIO_LIOscillation amplitude is greatest at a CpG density characteristic of enhancers\nC_LIO_LICell synchronisation reveals oscillation period and link with primary transcripts\nC_LIO_LIMultiomic single-cell profiling provides evidence for oscillatory dynamics in vivo\nC_LI

genomics

Whole genome sequencing, de novo assembly and phenotypic profiling for the new budding yeast species Saccharomyces jurei

Saccharomyces sensu stricto complex consist of yeast species, which are not only important in the fermentation industry but are also model systems for genomic and ecological analysis. Here, we present the complete genome assemblies of Saccharomyces jurei, a newly discovered Saccharomyces sensu stricto species from high altitude oaks. Phylogenetic and phenotypic analysis revealed that S. jurei is a sister-species to S. mikatae, than S. cerevisiae, and S. paradoxus. The karyotype of S. jurei presents two reciprocal chromosomal translocations between chromosome VI/VII and I/XIII when compared to S. cerevisiae genome. Interestingly, while the rearrangement I/XIII is unique to S. jurei, the other is in common with S. mikatae strain IFO1815, suggesting shared evolutionary history of this species after the split between S. cerevisiae and S. mikatae. The number of Ty elements differed in the new species, with a higher number of Ty elements present in S. jurei than in S. cerevisiae. Phenotypically, the S. jurei strain NCYC 3962 has relatively higher fitness than the other strain NCYC 3947T under most of the environmental stress conditions tested and showed remarkably increased fitness in higher concentration of acetic acid compared to the other sensu stricto species. Both strains were found to be better adapted to lower temperatures compared to S. cerevisiae.

genomics

Whole-genome sequences suggest long term declines of spotted owl (Strix occidentalis) (Aves: Strigiformes: Strigidae) populations in California

We analyzed whole-genome data of four spotted owls (Strix occidentalis) to provide a broad-scale assessment of the genome-wide nucleotide diversity across S. occidentalis populations in California. We assumed that each of the four samples was representative of its population and we estimated effective population sizes through time for each corresponding population. Our estimates provided evidence of long-term population declines in all California S. occidentalis populations. We found no evidence of genetic differentiation between northern spotted owl (S. o. caurina) populations in the counties of Marin and Humboldt in California. We estimated greater differentiation between populations at the northern and southern extremes of the range of the California spotted owl (S. o. occidentalis) than between populations of S. o. occidentalis and S. o. caurina in northern California. The San Diego County S. o. occidentalis population was substantially diverged from the other three S. occidentalis populations. These whole-genome data support a pattern of isolation-by-distance across spotted owl populations in California, rather than elevated differentiation between currently recognized subspecies.

genomics

Re-identification of genomic data using long range familial searches

Consumer genomics databases reached the scale of millions of individuals. Recently, law enforcement investigators have started to exploit some of these databases to find distant familial relatives, which can lead to a complete re-identification. Here, we leveraged genomic data of 600,000 individuals tested with consumer genomics to investigate the power of such long-range familial searches. We project that half of the searches with European-descent individuals will result with a third cousin or closer match and will provide a search space small enough to permit re-identification using common demographic identifiers. Moreover, in the near future, virtually any European-descent US person could be implicated by this technique. We propose a potential mitigation strategy based on cryptographic signature that can resolve the issue and discuss policy implications to human subject research.

genomics

Whole Proteome Clustering of 2,307 Genomes Reveals Remarkable Conservation of Four Proteins Among Proteobacteria While Revealing Significant Annotation Issues

To explore the concept of a minimal gene set, we clustered 8.76 M protein sequences deduced from 2,307 completely sequenced Proteobacterial genomes. To our knowledge this is the first study of this scale. Clustering resulted in 707,311 clusters of which 224,442 ranged in size from 2 to 2,894 sequences. The resulting clusters allowed us to ask the question: Is a set of proteins conserved across all Proteobacteria? We chose four essential proteins, the chaperonin GroEL, DNA dependent RNA polymerase subunits beta and beta (RpoB/RpoB), and DNA polymerase I (PolA), representing fundamental cellular functions, and examined their distribution in the clusters. We found these proteins to be remarkably conserved. Although the groEL gene was universally conserved in all the organisms in the study, the protein was not represented in all the deduced proteomes. The genes for RpoB and RpoB were missing from two genomes and merged in 88 genomes, and the sequences were sufficiently divergent that they formed separate clusters for 18 RpoB proteins (seven clusters) and 14 RpoB proteins (three clusters). For PolA, 52 organisms lacked an identifiable sequence, and seven sequences were sufficiently divergent that they formed five separate clusters. Interestingly, organisms lacking an identifiable PolA and those with divergent RpoB/RpoB were almost all endosymbionts. Furthermore, we present a range of examples of annotation issues that caused the deduced proteins to be incorrectly represented in the proteome. These annotation issues represent a significant obstacle for high throughput analyses.

genomics

Demographic inference in a spatially-explicit ecological model from genomic data: a proof of concept for the Mojave Desert Tortoise

In this paper, we study the general problem of extracting information from spatially explicit genomic data to inform inference of ecologically and geographically realistic population models. We describe methods and apply them to simulations motivated by the demography of the Mojave desert tortoise (Gopherus agassizii). The tortoise is an example of a long-lived, threatened species for which we have an excellent understanding of range, habitat preference, and certain aspects of demography, but inadequate information on other life history components that are important for conservation management. We use an individual-based model on a discretized geographic landscape with overlapping generations and age and sex-specific dispersal, fecundity, and mortality to develop and test a method that uses genomic data to infer demographic parameters. We do this by seeking parameters that best match a set of spatial statistics of genomes, which we introduce and discuss. We find that for inferring only overall population density and mean migration distance, a simple statistical learning method performs well using simulated training data, inferring parameters to within 10% accuracy. In the process, we introduce spatial analogues of common population genetics statistics, and discuss how and why they are expected to contain signal about the geography of population dynamics that are key for ecological modeling generally and conservation of endangered taxa.

genomics

Profiling the genome-wide landscape of tandem repeat expansions

Tandem Repeat (TR) expansions have been implicated in dozens of genetic diseases, including Huntingtons Disease, Fragile X Syndrome, and hereditary ataxias. Furthermore, TRs have recently been implicated in a range of complex traits, including gene expression and cancer risk. While the human genome harbors hundreds of thousands of TRs, analysis of TR expansions has been mainly limited to known pathogenic loci. A major challenge is that expanded repeats are beyond the read length of most next-generation sequencing (NGS) datasets. We present GangSTR, a novel algorithm for genome-wide profiling of both normal and expanded TRs. GangSTR extracts information from paired-end reads into a unified model to estimate maximum likelihood TR lengths. We validated GangSTR on real and simulated TR expansions and show that GangSTR outperforms alternative methods. We applied GangSTR to more than 150 individuals to profile the landscape of TR expansions in a healthy population and validated novel expansions using orthogonal technologies. Our analysis revealed that each individual harbors dozens of TR alleles longer than standard read lengths and identified hundreds of potentially mis-annotated TRs in the reference genome. GangSTR is packaged as a standalone tool that will likely enable discovery of novel pathogenic variants not currently accessible from NGS.

genomics

Integration of Dominance and Marker x Environment Interactions into Maize Genomic Prediction Models

Hybrid breeding programs are driven by the potential to explore the heterosis phenomenon in traits with non-additive inheritance. Traditionally, progress has been achieved by crossing lines from different heterotic groups and measuring phenotypic performance of hybrids in multiple environment trials. With the reduction in genotyping prices, genomic selection has become a reality for phenotype prediction and a promising tool to predict hybrid performances. However, its prediction ability is directly associated with models that represent the trait and breeding scheme under investigation. Herein, we assess modelling approaches where dominance effects and multi-environment statistical are considered for genomic selection in maize hybrid. To this end, we evaluated the predictive ability of grain yield and grain moisture collected over three production cycles in different locations. Hybrid genotypes were inferred in silico based on their parental inbred lines using single-nucleotide polymorphism markers obtained via a 500k SNP chip. We considered the importance to decomposes additive and dominance marker effects into components that are constant across environments and deviations that are group-specific. Prediction within and across environments were tested. The incorporation of dominance effect increased the predictive ability for grain production by up to 30% in some scenarios. Contrastingly, additive models yielded better results for grain moisture. For multi-environment modelling, the inclusion of interaction effects increased the predictive ability overall. More generally, we demonstrate that including dominance and genotype by environment interactions resulted in gains in accuracy and hence could be considered for genomic selection implementation in maize breeding programs.

genomics

Genomic signatures of honey bee association in an acetic acid symbiont

Honey bee queens are central to the success and productivity of their colonies; queens are the only reproductive members of the colony, and therefore queen longevity and fecundity can directly impact overall colony health. Recent declines in the health of the honey bee have startled researchers and lay people alike as honey bees are agricultures most important pollinator. Honey bees are important pollinators of many major crops and add billions of dollars annually to the US economy through their services. One factor that may influence queen and colony health is the microbial community. Although honey bee worker guts have a characteristic community of bee-specific microbes, the honey bee queen digestive tracts are colonized by a few bacteria, notably an acetic acid bacterium not seen in worker guts: Bombella apis. This bacterium is related to flower-associated microbes such as Saccharibacter floricola and other species in the genus Saccharibacter, and initial phylogenetic analyses placed it as sister to these environmental bacteria. We used comparative genomics of multiple honey bee-associated strains and the nectar-associated Saccharibacter to identify genomic changes associated with the ecological transition to bee association. We identified several genomic differences in the honey bee-associated strains, including a complete CRISPR/Cas system. Many of the changes we note here are predicted to confer upon them the ability to survive in royal jelly and defend themselves against mobile elements, including phages. Our results are a first step towards identifying potential benefits provided by the honey bee queen microbiota to the colonys matriarch.

genomics

Exploring the unmapped DNA and RNA reads in a songbird genome

BackgroundA widely used approach in next-generation sequencing projects is the alignment of reads to a reference genome. A significant percentage of reads, however, frequently remain unmapped despite improvements in the methods and hardware, which have enhanced the efficiency and accuracy of alignments. Usually unmapped reads are discarded from the analysis process, but significant biological information and insights can be uncovered from this data. We explored the unmapped DNA (normal and bisulfite treated) and RNA sequence reads of the great tit (Parus major) reference genome individual. From the unmapped reads we generated de novo assemblies. The generated sequence contigs were then aligned to the NCBI non-redundant nucleotide database using BLAST, identifying the closest known matching sequence.\n\nResultsMany of the aligned contigs showed sequence similarity to sequences from different bird species and genes that were absent in the great tit reference assembly. Furthermore, there were also contigs that represented known P. major pathogenic species. Most interesting were several species of blood parasites such as Plasmodium and Trypanosoma.\n\nConclusionsOur analyses revealed that meaningful biological information can be found when further exploring unmapped reads. It is possible to discover sequences that are either absent or misassembled in the reference genome and sequences that indicate infection or sample contamination. In this study we also propose strategies to aid the capture and interpretation of this information from unmapped reads.

genomics

A genome-wide association study of mitochondrial DNA copy number in two population-based cohorts

Mitochondrial DNA copy number (mtDNA CN) exhibits interindividual and intercellular variation, but few genome-wide association studies (GWAS) of directly assayed mtDNA CN exist.\n\nWe undertook a GWAS of qPCR-assayed mtDNA CN in the Avon Longitudinal Study of Parents and Children (ALSPAC), and the UK Blood Service (UKBS) cohort. After validating and harmonising data, 5461 ALSPAC mothers (16-43 years at mtDNA CN assay), and 1338 UKBS females (17-69 years) were included in a meta-analysis. Sensitivity analyses restricted to females with white cell-extracted DNA, and adjusted for estimated or assayed cell proportions. Associations were also explored in ALSPAC children, and UKBS males.\n\nA neutrophil-associated locus approached genome-wide significance (rs709591 [MED24], {beta}[SE] -0.084 [0.016], p=1.54e-07) in the main meta-analysis of adult females. This association was concordant in magnitude and direction in UKBS males and ALSPAC neonates. SNPs in and around ABHD8 were associated with mtDNA CN in ALSPAC neonates (rs10424198, {beta}[SE] 0.262 [0.034], p=1.40e-14), but not other study groups. In a meta-analysis of unrelated individuals (N=11253), we replicated a published association in TFAM {beta}[SE] 0.046 [0.017], p=0.006), with an effect size much smaller than that observed in the replication analysis of a previous in silico GWAS.\n\nIn a hypothesis-generating GWAS, we confirm an association between TFAM and mtDNA CN, and present putative loci requiring replication in much larger samples. We discuss the limitations of our work, in terms of measurement error and cellular heterogeneity, and highlight the need for larger studies to better understand nuclear genomic control of mtDNA copy number.

genomics

Assembly of Mb-size genome segments from linked read sequencing of CRISPR DNA targets

We developed a targeted sequencing method for intact high molecular weight (HMW) DNA targets as large as 0.2 Mb. This process uses HMW DNA isolated from intact cells, custom designed Cas9-guide RNA complexes to generate 0.1 - 0.2 Mb DNA targets, electrophoretic isolation of the DNA targets and sequencing with barcode linked reads. We used alignment methods as well as local assembly of the target regions to identify haplotypes and structural variants (SVs) across multi-Megabase genomic regions. To demonstrate the performance of this approach, we designed three assays that covered a 0.2 Mb region surrounding the BRCA1 gene, a set of 40 overlapping 0.2 Mb targets covering the entire 4-Mb MHC locus, and 18 well-characterized structural variants. Using the highly characterized NA12878 genome, we achieved on-target coverage of more than 50X, while overall whole genome coverage was approximately 4X. We generated haplotypes that completely covered each targeted locus, with a maximum size of 4 Mb (for the MHC region). This method detected structural variants such as deletions and inversions with determination of the exact breakpoints and genotypes. Even breakpoints inside highly homologous segmental duplications are precisely determined with our high-quality assemblies. Overall, this is a new method to sequence large DNA segments.

genomics

Refined detection and phasing of structural aberrations in pediatric acute lymphoblastic leukemia by linked-read whole genome sequencing

Structural chromosomal rearrangements that may lead to in-frame gene-fusions represent a leading source of information for diagnosis, risk stratification, and prognosis in pediatric acute lymphoblastic leukemia (ALL). However, short-read whole genome sequencing (WGS) technologies struggle to accurately identify and phase such large-scale chromosomal aberrations in cancer genomes. We therefore evaluated linked-read WGS for detection of chromosomal rearrangements in an ALL cell line (REH) and primary samples of varying DNA quality from 12 patients diagnosed with ALL. We assessed the effect of input DNA quality on phased haplotype block size and the detectability of copy number aberrations (CNAs) and structural variants (SVs). Biobanked DNA isolated by standard column-based extraction methods was sufficient to detect chromosomal rearrangements even at low 10x sequencing coverage. Linked-read WGS enabled precise, allele-specific, digital karyotyping at a base-pair resolution for a wide range of structural variants including complex rearrangements and aneuploidy assessment. With use of haplotype information from the linked-reads, we also identified additional structural variants, such as a compound heterozygous deletion of ERG in a patient with the DUX4-IGH fusion gene. Thus, linked-read WGS allows detection of important pathogenic variants in ALL genomes at a resolution beyond that of traditional karyotyping or short-read WGS.

genomics

Genome-wide association study of multiple yield components in a diversity panel of polyploid sugarcane (Saccharum spp.)

Sugarcane (Saccharum spp.) is an important economic crop, contributes up to 80% of sugar and approximately 60% bio-fuel globally. To meet the increased demand for sugar and bio-fuel supplies, it is critical to breed sugarcane cultivars with robust performance in yield components. Therefore, dissection of causal DNA sequence variants is of great importance by providing genetic resources and fundamental information for crop improvement. In this study, we evaluated and analyzed nine yield components in a sugarcane diversity panel consisting of 308 accessions primarily selected from the \"world collection of sugarcane and related grasses\". By genotyping the diversity panel using target enrichment sequencing, we identified a large number of sequence variants. Genome-wide association study between the markers and traits were conducted with dosages and gene actions taken into consideration. In total, 217 non-redundant markers and 225 candidate genes were identified to be significantly associated with the yield components, which can serve as a comprehensive genetic resource database for future gene identification, characterization, and selection for sugarcane improvement. We further investigated runs of homozygosity (ROH) in the sugarcane diversity panel. We characterized 282 ROHs, and found that the occurrence of ROH in the genome were non-random and probably under selection. ROHs were associated with total weight and dry weight, and high ROHs resulted in decrease of the two traits. This study approved that genomic inbreeding has led to negative impacts on sugarcane yield.

genomics

Systematic dissection of biases in whole-exome and whole-genome sequencing reveals major determinants of coding sequence coverage

Next generation DNA sequencing technologies are rapidly transforming the world of human genomics. Advantages and diagnostic effectiveness of the two most widely used resequencing approaches, whole exome (WES) and whole genome (WGS) sequencing, are still frequently debated. In our study we developed a set of statistical tools to systematically assess coverage of CDS regions provided by several modern WES platforms, as well as PCR-free WGS. Using several novel metrics to characterize exon coverage in WES and WGS, we showed that some of the WES platforms achieve substantially less biased CDS coverage than others, with lower within- and between-interval variation and virtually absent GC-content bias. We discovered that, contrary to a common view, most of the coverage bias in WES stems from mappability limitations of short reads, as well as exome probe design. We identified the ~ 500 kb region of human exome that could not be effectively characterized using short read technology. We also showed that the overall power for SNP and indel discovery in CDS region is virtually indistinguishable for WGS and best WES platforms. Our results indicate that deep WES (100x) using least biased technologies provides similar effective coverage (97% of 10x q10+ bases) and CDS variant discovery to the standard 30x WGS, suggesting that WES remains an efficient alternative to WGS in many applications. Our work could serve as a guide for selection of an up-to-date resequencing approach in human genomic studies.

genomics

Genome-wide DNA methylation profiling identifies convergent molecular signatures associated with idiopathic and syndromic forms of autism in postmortem human brain tissue.

Autism spectrum disorder (ASD) encompasses a collection of complex neuropsychiatric disorders characterized by deficits in social functioning, communication and repetitive behavior. Building on recent studies supporting a role for developmentally moderated regulatory genomic variation in the molecular etiology of ASD, we quantified genome-wide patterns of DNA methylation in 233 post-mortem tissues samples isolated from three brain regions (prefrontal cortex, temporal cortex and cerebellum) dissected from 43 ASD patients and 38 non-psychiatric control donors. We identified widespread differences in DNA methylation associated with idiopathic ASD (iASD), with consistent signals in both cortical regions that were distinct to those observed in the cerebellum. Individuals carrying a duplication on chromosome 15q (dup15q), representing a genetically-defined subtype of ASD, were characterized by striking differences in DNA methylation across a discrete domain spanning an imprinted gene cluster within the duplicated region. In addition to the dramatic cis-effects on DNA methylation observed in dup15q carriers, we identified convergent methylomic signatures associated with both iASD and dup15q, reflecting the findings from previous studies of gene expression and H3K27ac. Cortical co-methylation network analysis identified a number of co-methylated modules significantly associated with ASD that are enriched for genomic regions annotated to genes involved in the immune system, synaptic signalling and neuronal regulation. Our study represents the first systematic analysis of DNA methylation associated with ASD across multiple brain regions, providing novel evidence for convergent molecular signatures associated with both idiopathic and syndromic autism.

genomics

Chiral DNA sequences as commutable reference standards for clinical genomics

Chirality is a geometric property describing any object that is inequivalent to a mirror image of itself. Due to its 5-3 directionality, a DNA sequence is distinct from a mirrored sequence arranged in reverse nucleotide order, and is therefore chiral. A given sequence and its opposing chiral partner sequence share many properties, such as nucleotide composition and sequence entropy. Here we demonstrate that chiral DNA sequence pairs also perform equivalently during molecular and bioinformatic techniques that underpin modern genetic analysis, including PCR amplification, hybridization, whole-genome, target-enriched and nanopore sequencing, sequence alignment and variant detection. Given these shared properties, synthetic DNA sequences that directly mirror clinically relevant and/or analytically challenging regions of the human genome are ideal reference standards for clinical genomics. We show how the addition of chiral DNA standards to patient tumor samples can prevent false-positive and false-negative mutation detection and, thereby, improve diagnosis. Accordingly, we propose that chiral DNA standards can fulfill the unmet need for commutable internal reference standards in precision medicine.

genomics

Genome-wide identification and expression specificity analysis of the DNA methyltransferase gene family under adversity stresses in cotton

DNA methylation is an important epigenetic mode of genomic DNA modification that is an important part of maintaining epigenetic content and regulating gene expression. DNA methyltransferases (MTases) are the key enzymes in the process of DNA methylation. Thus far, there has been no systematic analysis the DNA MTases found in cotton. In this study, the whole genome of cotton C5-Mtase coding genes was identified and analyzed using a bioinformatics method based on information from the cotton genome. In this study, 51 DNA MTase genes were identified, of which 8 belonged to G. raimondii (group D), 9 belonged to G. arboretum L. (group A), 16 belonged to G. hirsutum L. (group AD1) and 18 belonged to G. barbadebse L. (group AD2). Systematic evolutionary analysis divided the 51 genes into four subfamilies, including 7 MET homologous proteins, 25 CMT homologous proteins, 14 DRM homologous proteins and 5 DNMT2 homologous proteins. Further studies showed that the DNA MTases in cotton were more phylogenetically conserved. The comparison of their protein domains showed that the C-terminal functional domain of the 51 proteins had six conserved motifs involved in methylation modification, indicating that the protein has a basic catalytic methylation function and the difference in the N-terminal regulatory domains of the 51 proteins divided the proteins into four classes, MET, CMT, DRM and DNMT2, in which DNMT2 lacks an N-terminal regulatory domain. Gene expression in cotton is not the same under different stress treatments. Different expression patterns of DNA MTases show the functional diversity of the cotton DNA methyltransferase gene family. VIGS silenced Gossypium hirsutum l. in the cotton seedling of DNMT2 family gene GhDMT6, after stress treatment the growth condition was better than the control. The distribution of DNA MTases varies among cotton species. Different DNA MTase family members have different genetic structures, and the expression level changes with different stresses, showing tissue specificity. Under salt and drought stress, G. hirsutum L. TM-1 increased the number of genes more than G. raimondii and G. arboreum L. Shixiya 1. The resistance of Gossypium hirsutum L.TM-1 to cold, drought and salt stress was increased after the plants were silenced with GhDMT6 gene.

genomics