Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Demographic inference in a spatially-explicit ecological model from genomic data: a proof of concept for the Mojave Desert Tortoise

In this paper, we study the general problem of extracting information from spatially explicit genomic data to inform inference of ecologically and geographically realistic population models. We describe methods and apply them to simulations motivated by the demography of the Mojave desert tortoise (Gopherus agassizii). The tortoise is an example of a long-lived, threatened species for which we have an excellent understanding of range, habitat preference, and certain aspects of demography, but inadequate information on other life history components that are important for conservation management. We use an individual-based model on a discretized geographic landscape with overlapping generations and age and sex-specific dispersal, fecundity, and mortality to develop and test a method that uses genomic data to infer demographic parameters. We do this by seeking parameters that best match a set of spatial statistics of genomes, which we introduce and discuss. We find that for inferring only overall population density and mean migration distance, a simple statistical learning method performs well using simulated training data, inferring parameters to within 10% accuracy. In the process, we introduce spatial analogues of common population genetics statistics, and discuss how and why they are expected to contain signal about the geography of population dynamics that are key for ecological modeling generally and conservation of endangered taxa.

genomics

Profiling the genome-wide landscape of tandem repeat expansions

Tandem Repeat (TR) expansions have been implicated in dozens of genetic diseases, including Huntingtons Disease, Fragile X Syndrome, and hereditary ataxias. Furthermore, TRs have recently been implicated in a range of complex traits, including gene expression and cancer risk. While the human genome harbors hundreds of thousands of TRs, analysis of TR expansions has been mainly limited to known pathogenic loci. A major challenge is that expanded repeats are beyond the read length of most next-generation sequencing (NGS) datasets. We present GangSTR, a novel algorithm for genome-wide profiling of both normal and expanded TRs. GangSTR extracts information from paired-end reads into a unified model to estimate maximum likelihood TR lengths. We validated GangSTR on real and simulated TR expansions and show that GangSTR outperforms alternative methods. We applied GangSTR to more than 150 individuals to profile the landscape of TR expansions in a healthy population and validated novel expansions using orthogonal technologies. Our analysis revealed that each individual harbors dozens of TR alleles longer than standard read lengths and identified hundreds of potentially mis-annotated TRs in the reference genome. GangSTR is packaged as a standalone tool that will likely enable discovery of novel pathogenic variants not currently accessible from NGS.

genomics

Integration of Dominance and Marker x Environment Interactions into Maize Genomic Prediction Models

Hybrid breeding programs are driven by the potential to explore the heterosis phenomenon in traits with non-additive inheritance. Traditionally, progress has been achieved by crossing lines from different heterotic groups and measuring phenotypic performance of hybrids in multiple environment trials. With the reduction in genotyping prices, genomic selection has become a reality for phenotype prediction and a promising tool to predict hybrid performances. However, its prediction ability is directly associated with models that represent the trait and breeding scheme under investigation. Herein, we assess modelling approaches where dominance effects and multi-environment statistical are considered for genomic selection in maize hybrid. To this end, we evaluated the predictive ability of grain yield and grain moisture collected over three production cycles in different locations. Hybrid genotypes were inferred in silico based on their parental inbred lines using single-nucleotide polymorphism markers obtained via a 500k SNP chip. We considered the importance to decomposes additive and dominance marker effects into components that are constant across environments and deviations that are group-specific. Prediction within and across environments were tested. The incorporation of dominance effect increased the predictive ability for grain production by up to 30% in some scenarios. Contrastingly, additive models yielded better results for grain moisture. For multi-environment modelling, the inclusion of interaction effects increased the predictive ability overall. More generally, we demonstrate that including dominance and genotype by environment interactions resulted in gains in accuracy and hence could be considered for genomic selection implementation in maize breeding programs.

genomics

Genomic signatures of honey bee association in an acetic acid symbiont

Honey bee queens are central to the success and productivity of their colonies; queens are the only reproductive members of the colony, and therefore queen longevity and fecundity can directly impact overall colony health. Recent declines in the health of the honey bee have startled researchers and lay people alike as honey bees are agricultures most important pollinator. Honey bees are important pollinators of many major crops and add billions of dollars annually to the US economy through their services. One factor that may influence queen and colony health is the microbial community. Although honey bee worker guts have a characteristic community of bee-specific microbes, the honey bee queen digestive tracts are colonized by a few bacteria, notably an acetic acid bacterium not seen in worker guts: Bombella apis. This bacterium is related to flower-associated microbes such as Saccharibacter floricola and other species in the genus Saccharibacter, and initial phylogenetic analyses placed it as sister to these environmental bacteria. We used comparative genomics of multiple honey bee-associated strains and the nectar-associated Saccharibacter to identify genomic changes associated with the ecological transition to bee association. We identified several genomic differences in the honey bee-associated strains, including a complete CRISPR/Cas system. Many of the changes we note here are predicted to confer upon them the ability to survive in royal jelly and defend themselves against mobile elements, including phages. Our results are a first step towards identifying potential benefits provided by the honey bee queen microbiota to the colonys matriarch.

genomics

Exploring the unmapped DNA and RNA reads in a songbird genome

BackgroundA widely used approach in next-generation sequencing projects is the alignment of reads to a reference genome. A significant percentage of reads, however, frequently remain unmapped despite improvements in the methods and hardware, which have enhanced the efficiency and accuracy of alignments. Usually unmapped reads are discarded from the analysis process, but significant biological information and insights can be uncovered from this data. We explored the unmapped DNA (normal and bisulfite treated) and RNA sequence reads of the great tit (Parus major) reference genome individual. From the unmapped reads we generated de novo assemblies. The generated sequence contigs were then aligned to the NCBI non-redundant nucleotide database using BLAST, identifying the closest known matching sequence.\n\nResultsMany of the aligned contigs showed sequence similarity to sequences from different bird species and genes that were absent in the great tit reference assembly. Furthermore, there were also contigs that represented known P. major pathogenic species. Most interesting were several species of blood parasites such as Plasmodium and Trypanosoma.\n\nConclusionsOur analyses revealed that meaningful biological information can be found when further exploring unmapped reads. It is possible to discover sequences that are either absent or misassembled in the reference genome and sequences that indicate infection or sample contamination. In this study we also propose strategies to aid the capture and interpretation of this information from unmapped reads.

genomics

A genome-wide association study of mitochondrial DNA copy number in two population-based cohorts

Mitochondrial DNA copy number (mtDNA CN) exhibits interindividual and intercellular variation, but few genome-wide association studies (GWAS) of directly assayed mtDNA CN exist.\n\nWe undertook a GWAS of qPCR-assayed mtDNA CN in the Avon Longitudinal Study of Parents and Children (ALSPAC), and the UK Blood Service (UKBS) cohort. After validating and harmonising data, 5461 ALSPAC mothers (16-43 years at mtDNA CN assay), and 1338 UKBS females (17-69 years) were included in a meta-analysis. Sensitivity analyses restricted to females with white cell-extracted DNA, and adjusted for estimated or assayed cell proportions. Associations were also explored in ALSPAC children, and UKBS males.\n\nA neutrophil-associated locus approached genome-wide significance (rs709591 [MED24], {beta}[SE] -0.084 [0.016], p=1.54e-07) in the main meta-analysis of adult females. This association was concordant in magnitude and direction in UKBS males and ALSPAC neonates. SNPs in and around ABHD8 were associated with mtDNA CN in ALSPAC neonates (rs10424198, {beta}[SE] 0.262 [0.034], p=1.40e-14), but not other study groups. In a meta-analysis of unrelated individuals (N=11253), we replicated a published association in TFAM {beta}[SE] 0.046 [0.017], p=0.006), with an effect size much smaller than that observed in the replication analysis of a previous in silico GWAS.\n\nIn a hypothesis-generating GWAS, we confirm an association between TFAM and mtDNA CN, and present putative loci requiring replication in much larger samples. We discuss the limitations of our work, in terms of measurement error and cellular heterogeneity, and highlight the need for larger studies to better understand nuclear genomic control of mtDNA copy number.

genomics

Assembly of Mb-size genome segments from linked read sequencing of CRISPR DNA targets

We developed a targeted sequencing method for intact high molecular weight (HMW) DNA targets as large as 0.2 Mb. This process uses HMW DNA isolated from intact cells, custom designed Cas9-guide RNA complexes to generate 0.1 - 0.2 Mb DNA targets, electrophoretic isolation of the DNA targets and sequencing with barcode linked reads. We used alignment methods as well as local assembly of the target regions to identify haplotypes and structural variants (SVs) across multi-Megabase genomic regions. To demonstrate the performance of this approach, we designed three assays that covered a 0.2 Mb region surrounding the BRCA1 gene, a set of 40 overlapping 0.2 Mb targets covering the entire 4-Mb MHC locus, and 18 well-characterized structural variants. Using the highly characterized NA12878 genome, we achieved on-target coverage of more than 50X, while overall whole genome coverage was approximately 4X. We generated haplotypes that completely covered each targeted locus, with a maximum size of 4 Mb (for the MHC region). This method detected structural variants such as deletions and inversions with determination of the exact breakpoints and genotypes. Even breakpoints inside highly homologous segmental duplications are precisely determined with our high-quality assemblies. Overall, this is a new method to sequence large DNA segments.

genomics

Refined detection and phasing of structural aberrations in pediatric acute lymphoblastic leukemia by linked-read whole genome sequencing

Structural chromosomal rearrangements that may lead to in-frame gene-fusions represent a leading source of information for diagnosis, risk stratification, and prognosis in pediatric acute lymphoblastic leukemia (ALL). However, short-read whole genome sequencing (WGS) technologies struggle to accurately identify and phase such large-scale chromosomal aberrations in cancer genomes. We therefore evaluated linked-read WGS for detection of chromosomal rearrangements in an ALL cell line (REH) and primary samples of varying DNA quality from 12 patients diagnosed with ALL. We assessed the effect of input DNA quality on phased haplotype block size and the detectability of copy number aberrations (CNAs) and structural variants (SVs). Biobanked DNA isolated by standard column-based extraction methods was sufficient to detect chromosomal rearrangements even at low 10x sequencing coverage. Linked-read WGS enabled precise, allele-specific, digital karyotyping at a base-pair resolution for a wide range of structural variants including complex rearrangements and aneuploidy assessment. With use of haplotype information from the linked-reads, we also identified additional structural variants, such as a compound heterozygous deletion of ERG in a patient with the DUX4-IGH fusion gene. Thus, linked-read WGS allows detection of important pathogenic variants in ALL genomes at a resolution beyond that of traditional karyotyping or short-read WGS.

genomics

Genome-wide association study of multiple yield components in a diversity panel of polyploid sugarcane (Saccharum spp.)

Sugarcane (Saccharum spp.) is an important economic crop, contributes up to 80% of sugar and approximately 60% bio-fuel globally. To meet the increased demand for sugar and bio-fuel supplies, it is critical to breed sugarcane cultivars with robust performance in yield components. Therefore, dissection of causal DNA sequence variants is of great importance by providing genetic resources and fundamental information for crop improvement. In this study, we evaluated and analyzed nine yield components in a sugarcane diversity panel consisting of 308 accessions primarily selected from the \"world collection of sugarcane and related grasses\". By genotyping the diversity panel using target enrichment sequencing, we identified a large number of sequence variants. Genome-wide association study between the markers and traits were conducted with dosages and gene actions taken into consideration. In total, 217 non-redundant markers and 225 candidate genes were identified to be significantly associated with the yield components, which can serve as a comprehensive genetic resource database for future gene identification, characterization, and selection for sugarcane improvement. We further investigated runs of homozygosity (ROH) in the sugarcane diversity panel. We characterized 282 ROHs, and found that the occurrence of ROH in the genome were non-random and probably under selection. ROHs were associated with total weight and dry weight, and high ROHs resulted in decrease of the two traits. This study approved that genomic inbreeding has led to negative impacts on sugarcane yield.

genomics

Systematic dissection of biases in whole-exome and whole-genome sequencing reveals major determinants of coding sequence coverage

Next generation DNA sequencing technologies are rapidly transforming the world of human genomics. Advantages and diagnostic effectiveness of the two most widely used resequencing approaches, whole exome (WES) and whole genome (WGS) sequencing, are still frequently debated. In our study we developed a set of statistical tools to systematically assess coverage of CDS regions provided by several modern WES platforms, as well as PCR-free WGS. Using several novel metrics to characterize exon coverage in WES and WGS, we showed that some of the WES platforms achieve substantially less biased CDS coverage than others, with lower within- and between-interval variation and virtually absent GC-content bias. We discovered that, contrary to a common view, most of the coverage bias in WES stems from mappability limitations of short reads, as well as exome probe design. We identified the ~ 500 kb region of human exome that could not be effectively characterized using short read technology. We also showed that the overall power for SNP and indel discovery in CDS region is virtually indistinguishable for WGS and best WES platforms. Our results indicate that deep WES (100x) using least biased technologies provides similar effective coverage (97% of 10x q10+ bases) and CDS variant discovery to the standard 30x WGS, suggesting that WES remains an efficient alternative to WGS in many applications. Our work could serve as a guide for selection of an up-to-date resequencing approach in human genomic studies.

genomics

Genome-wide DNA methylation profiling identifies convergent molecular signatures associated with idiopathic and syndromic forms of autism in postmortem human brain tissue.

Autism spectrum disorder (ASD) encompasses a collection of complex neuropsychiatric disorders characterized by deficits in social functioning, communication and repetitive behavior. Building on recent studies supporting a role for developmentally moderated regulatory genomic variation in the molecular etiology of ASD, we quantified genome-wide patterns of DNA methylation in 233 post-mortem tissues samples isolated from three brain regions (prefrontal cortex, temporal cortex and cerebellum) dissected from 43 ASD patients and 38 non-psychiatric control donors. We identified widespread differences in DNA methylation associated with idiopathic ASD (iASD), with consistent signals in both cortical regions that were distinct to those observed in the cerebellum. Individuals carrying a duplication on chromosome 15q (dup15q), representing a genetically-defined subtype of ASD, were characterized by striking differences in DNA methylation across a discrete domain spanning an imprinted gene cluster within the duplicated region. In addition to the dramatic cis-effects on DNA methylation observed in dup15q carriers, we identified convergent methylomic signatures associated with both iASD and dup15q, reflecting the findings from previous studies of gene expression and H3K27ac. Cortical co-methylation network analysis identified a number of co-methylated modules significantly associated with ASD that are enriched for genomic regions annotated to genes involved in the immune system, synaptic signalling and neuronal regulation. Our study represents the first systematic analysis of DNA methylation associated with ASD across multiple brain regions, providing novel evidence for convergent molecular signatures associated with both idiopathic and syndromic autism.

genomics

Chiral DNA sequences as commutable reference standards for clinical genomics

Chirality is a geometric property describing any object that is inequivalent to a mirror image of itself. Due to its 5-3 directionality, a DNA sequence is distinct from a mirrored sequence arranged in reverse nucleotide order, and is therefore chiral. A given sequence and its opposing chiral partner sequence share many properties, such as nucleotide composition and sequence entropy. Here we demonstrate that chiral DNA sequence pairs also perform equivalently during molecular and bioinformatic techniques that underpin modern genetic analysis, including PCR amplification, hybridization, whole-genome, target-enriched and nanopore sequencing, sequence alignment and variant detection. Given these shared properties, synthetic DNA sequences that directly mirror clinically relevant and/or analytically challenging regions of the human genome are ideal reference standards for clinical genomics. We show how the addition of chiral DNA standards to patient tumor samples can prevent false-positive and false-negative mutation detection and, thereby, improve diagnosis. Accordingly, we propose that chiral DNA standards can fulfill the unmet need for commutable internal reference standards in precision medicine.

genomics

Genome-wide identification and expression specificity analysis of the DNA methyltransferase gene family under adversity stresses in cotton

DNA methylation is an important epigenetic mode of genomic DNA modification that is an important part of maintaining epigenetic content and regulating gene expression. DNA methyltransferases (MTases) are the key enzymes in the process of DNA methylation. Thus far, there has been no systematic analysis the DNA MTases found in cotton. In this study, the whole genome of cotton C5-Mtase coding genes was identified and analyzed using a bioinformatics method based on information from the cotton genome. In this study, 51 DNA MTase genes were identified, of which 8 belonged to G. raimondii (group D), 9 belonged to G. arboretum L. (group A), 16 belonged to G. hirsutum L. (group AD1) and 18 belonged to G. barbadebse L. (group AD2). Systematic evolutionary analysis divided the 51 genes into four subfamilies, including 7 MET homologous proteins, 25 CMT homologous proteins, 14 DRM homologous proteins and 5 DNMT2 homologous proteins. Further studies showed that the DNA MTases in cotton were more phylogenetically conserved. The comparison of their protein domains showed that the C-terminal functional domain of the 51 proteins had six conserved motifs involved in methylation modification, indicating that the protein has a basic catalytic methylation function and the difference in the N-terminal regulatory domains of the 51 proteins divided the proteins into four classes, MET, CMT, DRM and DNMT2, in which DNMT2 lacks an N-terminal regulatory domain. Gene expression in cotton is not the same under different stress treatments. Different expression patterns of DNA MTases show the functional diversity of the cotton DNA methyltransferase gene family. VIGS silenced Gossypium hirsutum l. in the cotton seedling of DNMT2 family gene GhDMT6, after stress treatment the growth condition was better than the control. The distribution of DNA MTases varies among cotton species. Different DNA MTase family members have different genetic structures, and the expression level changes with different stresses, showing tissue specificity. Under salt and drought stress, G. hirsutum L. TM-1 increased the number of genes more than G. raimondii and G. arboreum L. Shixiya 1. The resistance of Gossypium hirsutum L.TM-1 to cold, drought and salt stress was increased after the plants were silenced with GhDMT6 gene.

genomics

Dating genomic variants and shared ancestry in population-scale sequencing data

The origin and fate of new mutations within species is the fundamental process underlying evolution. However, while much attention has been focused on characterizing the presence, frequency, and phenotypic impact of genetic variation, the evolutionary histories of most variants are largely unexplored. We have developed a non-parametric approach for estimating the date of origin of genetic variants in large-scale sequencing data sets. The accuracy and robustness of the approach is demonstrated through simulation. Using data from two publicly available human genomic diversity resources, we estimated the age of more than 45 million single nucleotide polymorphisms (SNPs) in the human genome and release the Atlas of Variant Age as a public online database. We characterize the relationship between variant age and frequency in different geographical regions, and demonstrate the value of age information in interpreting variants of functional and selective importance. Finally, we use allele age estimates to power a rapid approach for inferring the ancestry shared between individual genomes, to quantify genealogical relationships at different points in the past, as well as describe and explore the evolutionary history of modern human populations.

genomics

Predicting mRNA abundance directly from genomic sequence using deep convolutional neural networks

Algorithms that accurately predict gene structure from primary sequence alone were transformative for annotating the human genome. Can we also predict the expression levels of genes based solely on genome sequence? Here we sought to apply deep convolutional neural networks towards this goal. Surprisingly, a model that includes only promoter sequences and features associated with mRNA stability explains 59% and 71% of variation in steady-state mRNA levels in human and mouse, respectively. This model, which we call Xpresso, more than doubles the accuracy of alternative sequence-based models, and isolates rules as predictive as models relying on ChIP-seq data. Xpresso recapitulates genome-wide patterns of transcriptional activity and predicts the influence of enhancers, heterochromatic domains, and microRNAs. Model interpretation reveals that promoter-proximal CpG dinucleotides strongly predict transcriptional activity. Looking forward, we propose the accurate prediction of cell type-specific gene expression based solely on primary sequence as a grand challenge for the field.

genomics

Long-read based assembly and annotation of a Drosophila simulans genome

Long-read sequencing technologies enable high-quality, contiguous genome assemblies. Here we used SMRT sequencing to assemble the genome of a Drosophila simulans strain originating from Madagascar, the ancestral range of the species. We generated 8 Gb of raw data (~50x coverage) with a mean read length of 6,410 bp, a NR50 of 9,125 bp and the longest subread at 49 kb. We benchmarked six different assemblers and merged the best two assemblies from Canu and Falcon. Our final assembly was 127.41 Mb with a N50 of 5.38 Mb and 305 contigs. We anchored more than 4 Mb of novel sequence to the major chromosome arms, and significantly improved the assembly of peri-centromeric and telomeric regions. Finally, we performed full-length transcript sequencing and used this data in conjunction with short-read RNAseq data to annotate 13,422 genes in the genome, improving the annotation in regions with complex, nested gene structures.

genomics

Genomic prediction of autotetraploids; influence of relationship matrices, allele dosage, and continuous genotyping calls in phenotype prediction

Estimation of allele dosage in autopolyploids is challenging and current methods often result in the misclassification of genotypes. Here we propose and compare the use of next generation sequencing read depth as continuous parameterization for autotetraploid genomic prediction of breeding values, using blueberry (Vaccinium corybosum spp.) as a model. Additionally, we investigated the influence of different sources of information to build relationship matrices in phenotype prediction; no relationship, pedigree, and genomic information, considering either diploid or tetraploid parameterizations. A real breeding population composed of 1,847 individuals was phenotyped for eight yield and fruit quality traits over two years. Analyses were based on extensive pedigree (since 1908) and high-density marker data (86K markers). Our results show that marker-based matrices can yield significantly better prediction than pedigree for most of the traits, based on model fitting and expected genetic gain. Continuous genotypic based models performed as well as the current best models and presented a significantly better goodness-of-fit for all traits analyzed. This approach also reduces the computational time required for marker calling and avoids problems associated with misclassification of genotypic classes when assigning dosage in polyploid species. Accuracies are encouraging for application of genomic selection (GS) for blueberry breeding. Conservatively, GS could reduce the time for cultivar release by three years. GS could increase the genetic gain per cycle by 86% on average when compared to phenotypic selection, and 32% when compared with pedigree-based selection.

genomics

Genome-wide transcription during early wheat meiosis is independent of synapsis, ploidy level and the Ph1 locus

Polyploidization is a fundamental process in plant evolution. One of the biggest challenges faced by a new polyploid is meiosis, particularly discriminating between multiple related chromosomes so that only homologous chromosomes synapse and recombine to ensure regular chromosome segregation and balanced gametes. Despite its large genome size, high DNA repetitive content and similarity between homoeologous chromosomes, hexaploid wheat completes meiosis in a shorter period than diploid species with a much smaller genome. Therefore, during wheat meiosis, mechanisms additional to the classical model based on DNA sequence homology, must facilitate more efficient homologous recognition. One such mechanism could involve exploitation of differences in chromosome structure between homologues and homoeologues at the onset of meiosis. In turn, these chromatin changes, can be expected to be linked to transcriptional gene activity. In this study, we present an extensive analysis of a large RNA-Seq data derived from six different genotypes: wheat, wheat-rye hybrids and newly synthesized octoploid triticale, both in the presence and absence of the Ph1 locus. Plant material was collected at early prophase, at the transition leptotene-zygotene, when the telomere bouquet is forming and synapsis between homologues is beginning. The six genotypes exhibit different levels of synapsis and chromatin structure at this stage; therefore, recombination and consequently segregation, are also different. Unexpectedly, our study reveals that neither synapsis, whole genome duplication nor the absence of the Ph1 locus are associated with major changes in gene expression levels during early meiotic prophase. Overall wheat transcription at this meiotic stage is therefore highly resilient to such alterations, even in the presence of major chromatin structural changes. This suggests that post-transcriptional and post-translational processes are likely to be more important. Thus, further studies will be required to reveal whether these observations are specific to wheat meiosis, and whether there are significant changes in post-transcriptional and post-translational modifications in wheat and other polyploid species associated with their polyploidisation.

genomics