Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

High quality whole genome sequence of an abundant Holarctic odontocete, the harbour porpoise (Phocoena phocoena)

The harbour porpoise (Phocoena phocoena) is a highly mobile cetacean found in waters across the Northern hemisphere. It occurs in coastal water and inhabits water basins that vary broadly in salinity, temperature, and food availability. These diverse habitats could drive differentiation among populations. Here we report the first harbour porpoise genome, assembled de novo from a Swedish Kattegat individual. The genome is one of the most complete cetacean genomes currently available, with a total size of 2.7 Gb and 50% of the total length found in just 34 scaffolds. Using the largest 122 scaffolds, we were able to validate a high level of homology to the chromosome-level genome assembly of the closest related species for which such resource was available, the domestic cattle (Bos taurus). The draft annotation comprises 22,154 predicted gene models, which we further annotated through matches to the NCBI nucleotide database, GO categorization, and motif prediction. To infer the adaptive abilities of this species, as well as their population history, we performed a Bayesian skyline analysis, and produced results that are concordant with the demographic history of this species, including expansion and fragmentation events. Overall, this genome assembly, together with the draft annotation, represents a crucial addition to the limited genetic markers currently available for the study of porpoises and Phocoenidae conservation, phylogeny, and evolution.

genomics

A high-quality sequence of Rosa chinensis to elucidate genome structure and ornamental traits

Rose is the worlds most important ornamental plant with economic, cultural and symbolic value. Roses are cultivated worldwide and sold as garden roses, cut flowers and potted plants. Rose has a complex genome with high heterozygosity and various ploidy levels. Our objectives were (i) to develop the first high-quality reference genome sequence for the genus Rosa by sequencing a doubled haploid, combining long and short read sequencing, and anchoring to a high-density genetic map and (ii) to study the genome structure and the genetic basis of major ornamental traits.\n\nWe produced a haploid rose line from R. chinensis Old Blush and generated the first rose genome sequence at the pseudo-molecule scale (512 Mbp with N50 of 3.4 Mb and L75 of 97). The sequence was validated using high-density diploid and tetraploid genetic maps. We delineated hallmark chromosomal features including the pericentromeric regions through annotation of TE families and positioned centromeric repeats using FISH. Genetic diversity was analysed by resequencing eight Rosa species. Combining genetic and genomic approaches, we identified potential genetic regulators of key ornamental traits, including prickle density and number of flower petals. A rose APETALA2 homologue is proposed to be the major regulator of petals number in rose. This reference sequence is an important resource for studying polyploidisation, meiosis and developmental processes as we demonstrated for flower and prickle development. This reference sequence will also accelerate breeding through the development of molecular markers linked to traits, the identification of the genes underlying them and the exploitation of synteny across Rosaceae.

genomics

Whole genome sequence of an edible and potential medicinal fungus, Cordyceps guangdongensis

Cordyceps guangdongensis is an edible fungus which has been approved as a Novel Food by the Chinese Ministry of Public Health in 2013. It also has a broad application prospect in pharmaceutical industries with many medicinal activities. In this study, the whole genome of C. guangdongensis GD15, a single spore isolate from a wild strain, was sequenced and assembled with Illumina and PacBio sequencing technology. The generated genome is 29.05 Mb in size, comprising 9 scaffolds with an average GC content of 57.01%. It is predicted to contain a total of 9150 protein-coding genes. Sequence identification and comparative analysis indicated that the assembled scaffolds contained two complete chromosomes and four single-end chromosomes, showing a high level assembly. Gene annotation revealed a diversity of transporters that could contribute to the genome size and evolution. Besides, approximately 15.49% and 13.70% genes involved in metabolic processes were annotated by KEGG and COG respectively. Genes belonging to CAZymes accounted for a proportion of 2.84% of the total genes. In addition, 435 transcription factors (TFs) were identified, which were involved in various biological processes. Among the identified TFs, the fungal transcription regulatory proteins (18.39%) and fungal-specific TFs (19.77%) represented the two largest classes of TFs. These data provided a much needed genomic resource for studying C. guangdongensis, laying a solid foundation for further genetic and biological studies, especially for elucidating the genome evolution and exploring the regulatory mechanism of fruiting body development.

genomics

Insights from deconvolution of cell subtype proportions enhance the interpretation of functional genomic data.

Cell subtype proportional differences between samples significantly contribute to variation of functional genomic properties such as gene expression or DNA methylation. Current analytical approaches typically deal with cell subtype proportion influences as a nuisance variable to be eliminated. Here we demonstrate how harvesting information about cell subtype proportions from functional genomics data provides insights into the cellular events in human phenotypes. We note a striking concordance between cell subtype proportions estimated from orthogonal genome-wide assays, and demonstrate the potential for single-cell RNA-seq data to be used in tissues for which reference cell subtype functional genomic datasets are not available. Taken together, our results confirm the importance of estimating cell subtype proportions when testing a model of cellular reprogramming in human phenotypic association studies, and the value of simultaneously testing for systematic cell subtype proportional alterations as a separate phenotypic association, gaining extra insights from functional genomic studies.

genomics

A meta-analysis of the diagnostic sensitivity and clinical utility of genome sequencing, exome sequencing and chromosomal microarray in children with suspected genetic diseases

IMPORTANCEGenetic diseases are a leading cause of childhood mortality. Whole genome sequencing (WGS) and whole exome sequencing (WES) are relatively new methods for diagnosing genetic diseases.\n\nOBJECTIVESCompare the diagnostic sensitivity (rate of causative, pathogenic or likely pathogenic genotypes in known disease genes) and rate of clinical utility (proportion in whom medical or surgical management was changed by diagnosis) of WGS, WES, and chromosomal microarrays (CMA) in children with suspected genetic diseases.\n\nDATA SOURCES AND STUDY SELECTIONSystematic review of the literature (January 2011 - August 2017) for studies of diagnostic sensitivity and/or clinical utility of WGS, WES, and/or CMA in children with suspected genetic diseases. 2% of identified studies met selection criteria.\n\nDATA EXTRACTION AND SYNTHESISTwo investigators extracted data independently following MOOSE/PRISMA guidelines.\n\nMAIN OUTCOMES AND MEASURESPooled rates and 95% Cl were estimated with a random-effects model. Metaanalysis of the rate of diagnosis was based on test type, family structure, and site of testing.\n\nRESULTSIn 36 observational series and one randomized control trial, comprising 20,068 children, the diagnostic sensitivity of WGS (0.41, 95% Cl 0.34-0.48, I2=44%) and WES (0.35, 95% Cl 0.31-0.39, I2=85%) were qualitatively greater than CMA (0.10, 95% Cl 0.08-0.12, I2=81%). Subgroup meta-analyses showed that the diagnostic sensitivity of WGS was significantly greater than CMA in studies published in 2017 (P<.0001, I2=13% and I2=40%, respectively), and the diagnostic sensitivity of WES was significantly greater than CMA in studies featuring within-cohort comparisons (P<001, I2=36%). Evidence for a significant difference in the diagnostic sensitivity of WGS and WES was lacking. In studies featuring within-cohort comparisons of singleton and trio WGS/WES, the likelihood of diagnosis was significantly greater for trios (odds ratio 2.04, 95% Cl 1.62-2.56, I2=12%; P<.0001). The diagnostic sensitivity of WGS/WES with hospital-based interpretation (0.41, 95% Cl 0.38-0.45, I2=50%) was qualitatively higher than that of reference laboratories (0.28, 95% Cl 0.24-0.32, I2=81%); this difference was significant in meta-analysis of studies published in 2017 (P=.004, I2=34% and I2=26%, respectively). The rates of clinical utility of WGS (0.27, 95% Cl 0.17-0.40, I2=54%) and WES (0.18, 95% Cl 0.13-0.24, I2-77%) were higher than CMA (0.06, 95% Cl 0.05-0.07, I2=42%); this difference was significant in meta-analysis of WGS vs CMA (P<.0001).\n\nCONCLUSIONS AND RELEVANCEIn children with suspected genetic diseases, the diagnostic sensitivity and rate of clinical utility of WGS/WES were greater than CMA. Subgroups with higher WGS/WES diagnostic sensitivity were trios and those receiving hospital-based interpretation. WGS/WES should be considered a first-line genomic test for children with suspected genetic diseases.\n\nKey PointsO_ST_ABSQuestionC_ST_ABSWhat is the relative diagnostic sensitivity and clinical utility of different genome tests in children with suspected genetic diseases?\n\nFindingsWhole genome sequencing had greater diagnostic sensitivity and clinical utility than chromosomal microarrays. Testing parent-child trios had greater diagnostic sensitivity than proband singletons. Hospital-based testing had greater diagnostic sensitivity than reference laboratories.\n\nMeaningTrio genomic sequencing is the most sensitive diagnostic test for children with suspected genetic diseases.

genomics

Robustness of Transposable Element regulation but no genomic shock observed in interspecific Arabidopsis hybrids

The merging of two divergent genomes in a hybrid is believed to trigger a \"genomic shock\", disrupting gene regulation and transposable element (TE) silencing. Here, we tested this expectation by comparing the pattern of expression of transposable elements in their native and hybrid genomic context. For this, we sequenced the transcriptome of the Arabidopsis thaliana genotype Col-0, the A. lyrata genotype MN47 and their F1 hybrid. Contrary to expectations, we observe that the level of TE expression in the hybrid is strongly correlated to levels in the parental species. We detect that at most 1.1% of expressed transposable elements belonging to two specific subfamilies change their expression level upon hybridization. Most of these changes, however, are of small magnitude. We observe that the few hybrid-specific modifications in TE expression are more likely to occur when TE insertions are close to genes. In addition, changes in epigenetic histone marks H3K9me2 and H3K27me3 following hybridization do not coincide with TEs with changed expression. Finally, we further examined TE expression in parents and hybrids exposed to severe dehydration stress. Despite the major reorganization of gene and TE expression by stress, we observe that hybridization does not lead to increased disorganization of TE expression in the hybrid. We conclude that TE expression is globally robust to hybridization and that the term \"genomic shock\" is no longerappropriate to describe the anticipated consequences of merging divergent genomes in a hybrid.

genomics

CONSTRUCTION OF WHOLE GENOMES FROM SCAFFOLDS USING SINGLE CELL STRAND-SEQ DATA

Accurate reference genome sequences provide the foundation for modern molecular biology and genomics as the interpretation of sequence data to study evolution, gene expression and epigenetics depends heavily on the quality of the genome assembly used for its alignment. Correctly organising sequenced fragments such as contigs and scaffolds in relation to each other is a critical and often challenging step in the construction of robust genome references. We previously identified misoriented regions in the mouse and human reference assemblies using Strand-seq, a single cell sequencing technique that preserves DNA directionality1, 2. Here we demonstrate the ability of Strand-seq to build and correct full-length chromosomes, by identifying which scaffolds belong to the same chromosome and determining their correct order and orientation, without the need for overlapping sequences. We demonstrate that Strand-seq exquisitely maps assembly fragments into large related groups and chromosome-sized clusters without using new assembly data. Using template strand inheritance as a bi-allelic marker, we employ genetic mapping principles to cluster scaffolds that are derived from the same chromosome and order them within the chromosome based solely on directionality of DNA strand inheritance. We prove the utility of our approach by generating improved genome assemblies for several model organisms including the ferret, pig, Xenopus, zebrafish, Tasmanian devil and the Guinea pig.

genomics

Evaluation of Whole Exome Sequencing as an Alternative of BeadChip and Whole Genome Sequencing in Human Population Genetic Analysis

Understanding the underlying genetic structure of human populations is of fundamental interest to both biological and social sciences. Advances in high-throughput genotyping technology have markedly improved our understanding of global patterns of human genetic variation. The most widely used methods for collecting variant information at the DNA-level include whole genome sequencing, which continues to remain costly, and the more economical solution of array-based techniques, as these are capable of simultaneously genotyping a pre-selected set of variable DNA sites in the human genome. The largest publicly accessible set of human genomic sequence data available today originates from exome sequencing that comprises around 1.2% of the whole genome (approximately 30 million base pairs). In this study, we compared the application of the exome dataset to the array-based dataset and to the gold standard whole genome dataset using the same population genetic analysis methods. Our results draw attention to some of the inherent problems that arise from using pre-selected SNP sets for population genetic analysis. Additionally, we demonstrate that exome sequencing provides a better alternative to the array-based methods for population genetic analysis. In this study, we propose a strategy for unbiased variant collection from exome data and offer a bioinformatics protocol for proper data processing.

genomics

Association of whole-genome and NETRIN1 signaling pathway-derived polygenic risk scores for Major Depressive Disorder and thalamic radiation white matter microstructure in UK Biobank

BackgroundMajor Depressive Disorder (MDD) is a clinically heterogeneous psychiatric disorder with a polygenic architecture. Genome-wide association studies have identified a number of risk-associated variants across the genome, and growing evidence of NETRIN1 pathway involvement. Stratifying disease risk by genetic variation within the NETRIN1 pathway may provide an important route for identification of disease mechanisms by focusing on a specific process excluding heterogeneous risk-associated variation in other pathways. Here, we sought to investigate whether MDD polygenic risk scores derived from the NETRIN1 signaling pathway (NETRIN1-PRS) and the whole genome excluding NETRIN1 pathway genes (genomic-PRS) were associated with white matter integrity.\n\nMethodsWe used two diffusion tensor imaging measures, fractional anisotropy (FA) and mean diffusivity (MD), in the most up-to-date UK Biobank neuroimaging data release (FA: N = 6,401; MD: N = 6,390).\n\nResultsWe found significantly lower FA in the superior longitudinal fasciculus ({beta} = -0.035, pcorrected = 0.029) and significantly higher MD in a global measure of thalamic radiations ({beta} = 0.029, pcorrected = 0.021), as well as higher MD in the superior ({beta} = 0.034, pcorrected = 0.039) and inferior ({beta} = 0.029, pcorrected = 0.043) longitudinal fasciculus and in the anterior ({beta} = 0.025, pcorrected = 0.046) and superior ({beta} = 0.027, pcorrected = 0.043) thalamic radiation associated with NETRIN1-PRS. Genomic-PRS was also associated with lower FA and higher MD in several tracts.\n\nConclusionsOur findings indicate that variation in the NETRIN1 signaling pathway may confer risk for MDD through effects on thalamic radiation white matter microstructure.

genetics

Immuno-genomic PanCancer Landscape Reveals Diverse Immune Escape Mechanisms and Immuno-Editing Histories

Immune reactions in the tumor micro-environment are one of the cancer hallmarks and emerging immune therapies have been proven effective in many types of cancer. To investigate cancer genome-immune interactions and the role of immuno-editing or immune escape mechanisms in cancer development, we analyzed 2,834 whole genomes and RNA-seq datasets across 31 distinct tumor types from the PanCancer Analysis of Whole Genomes (PCAWG) project with respect to key immuno-genomic aspects. We show that selective copy number changes in immune-related genes could contribute to immune escape. Furthermore, we developed an index of the immuno-editing history of each tumor sample based on the information of mutations in exonic regions and pseudogenes. Our immuno-genomic analyses of pan-cancer analyses have the potential to identify a subset of tumors with immunogenicity and diverse background or intrinsic pathways associated with their immune status and immuno-editing history.

genomics

Three invariant Hi-C interaction patterns: applications to genome assembly

Assembly of reference-quality genomes from next-generation sequencing data is a key challenge in genomics. Recently, we and others have shown that Hi-C data can be used to address several outstanding challenges in the field of genome assembly. This principle has since been developed in academia and industry, and has been used in the assembly of several major genomes. In this paper, we explore the central principles underlying Hi-C-based assembly approaches, by quantitatively defining and characterizing three invariant Hi-C interaction patterns on which these approaches can build: Intrachromosomal interaction enrichment, distance-dependent interaction decay and local interaction smoothness. Specifically, we evaluate to what degree each invariant pattern holds on a single locus level in different species, cell types and Hi-C map resolutions. We find that these patterns are generally consistent across species and cell types but are affected by sequencing depth, and that matrix balancing improves consistency of loci with all three invariant patterns. Finally, we overview current Hi-C-based assembly approaches in light of these invariant patterns and demonstrate how local interaction smoothness can be used to easily detect scaffolding errors in extremely sparse Hi-C maps. We suggest that simultaneously considering all three invariant patterns may lead to better Hi-C-based genome assembly methods.

genomics

EquCab3, an Updated Reference Genome for the Domestic Horse

EquCab2, a high-quality reference genome for the domestic horse, was released in 2007. Since then, it has served as the foundation for nearly all genomic work done in equids. Recent advances in genomic sequencing technology and computational assembly methods have allowed scientists to improve reference assemblies of large animal and plant genomes in terms of contiguity and composition. In 2014, the equine genomics research community began a project to improve the reference sequence for the horse, building upon the solid foundation of EquCab2 and incorporating new short-read data, long-read data, and proximity ligation data. The result, EquCab3, is presented here. The count of non-N bases in the incorporated chromosomes is improved from 2.33Gb in EquCab2 to 2.41Gb from EquCab3. Contiguity has also been improved nearly 40-fold with a contig N50 of 4.5Mb and scaffold contiguity enhanced to where all but one of the 32 chromosomes is comprised of a single scaffold.

genomics

TSA-Seq Mapping of Nuclear Genome Organization

While nuclear compartmentalization is an essential feature of three-dimensional genome organization, no genomic method exists for measuring chromosome distances to defined nuclear structures. Here we describe TSA-Seq, a new mapping method able to estimate mean chromosomal distances from nuclear speckles genome-wide and predict several Mbp chromosome trajectories between nuclear compartments without sophisticated computational modeling. Ensemble-averaged results reveal a clear nuclear lamina to speckle axis correlated with a striking spatial gradient in genome activity. This gradient represents a convolution of multiple, spatially separated nuclear domains, including two types of transcription \"hot-zones\". Transcription hot-zones protruding furthest into the nuclear interior and positioning deterministically very close to nuclear speckles have higher numbers of total genes, the most highly expressed genes, house-keeping genes, genes with low transcriptional pausing, and super-enhancers. Our results demonstrate the capability of TSA-Seq for genome-wide mapping of nuclear structure and suggest a new model for nuclear spatial organization of transcription.

genomics

Genomic prediction using individual-level data and summary statistics from multiple populations

This study presents a method for genomic prediction that uses individual-level data and summary statistics from multiple populations. Genome-wide markers are nowadays widely used to predict complex traits, and genomic prediction using multi-population data is an appealing approach to achieve higher prediction accuracies. However, sharing of individual-level data across populations is not always possible. We present a method that enables integration of summary statistics from separate analyses with the available individual-level data. The data can either consist of individuals with single or multiple (weighted) phenotype records per individual. We developed a method based on a hypothetical joint analysis model and absorption of population specific information. We show that population specific information is fully captured by estimated allele substitution effects and the accuracy of those estimates, i.e. the summary statistics. The method gives identical result as the joint analysis of all individual-level data when complete summary statistics are available. We provide a series of easy-to-use approximations that can be used when complete summary statistics are not available or impractical to share. Simulations show that approximations enables integration of different sources of information across a wide range of settings yielding accurate predictions. The method can be readily extended to multiple-traits. In summary, the developed method enables integration of genome-wide data in the individual-level or summary statistics form from multiple populations to obtain more accurate estimates of allele substitution effects and genomic predictions.

genomics

Livestock genome annotation: transcriptome and chromatin structure profiling in cattle, goat, chicken and pig.

BackgroundFunctional annotation of livestock genomes is a critical step to decipher the genotype-to-phenotype relationship underlying complex traits. As part of the Functional Annotation of Animal Genomes (FAANG) action, the FR-AgENCODE project (http://www.fragencode.org) aimed to profile the landscape of transcription (RNA-seq), chromatin accessibility (ATAC-seq) and conformation (Hi-C) in four livestock species representing ruminants (cattle, goat), monogastrics (pig) and birds (chicken), using three target samples related to metabolism (liver) and immunity (CD4+ and CD8+ T cells).\n\nResultsRNA-seq assays considerably extended the available catalog of annotated transcripts and identified differentially expressed genes with unknown function, including new syntenic lncRNAs. ATAC-seq highlighted an enrichment for transcription factor binding sites in differentially accessible regions of the chromatin. Comparative analyses revealed a core set of conserved regulatory regions across species. Topologically Associating Domains (TADs) and epigenetic A/B compartments annotated from Hi-C data were consistent with RNA-seq and ATAC-seq data. Multi-species comparisons showed that conserved TAD boundaries had stronger insulation properties than species-specific ones and that the genomic distribution of orthologous genes in A/B compartments was significantly conserved across species.\n\nConclusionsWe report the first multi-species and multi-assay genome annotation results obtained by a FAANG project. Beyond the generation of reference annotations and the confirmation of previous findings on model animals, the integrative analysis of data from multiple assays and species sheds a new light on the multi-scale selective pressure shaping genome organization from birds to mammals. Overall, these results emphasize the value of FAANG for research on domesticated animals and reinforces the importance of future meta-analyses of the reference datasets being generated by this community on different species.

genomics

Viral Taxonomy Derived From Evolutionary Genome Relationships

We describe a new genome alignment-based model for classification of viruses based on evolutionary genetic relationships. This approach uses information theory and a physical model to determine the information shared by the genes in two genomes. Pairwise comparisons of genes from the viruses are created from alignments using NCBI BLAST, and their match scores are combined to produce a metric between genomes, which is in turn used to determine a global classification using the 5,817 viruses on RefSeq. In cases where there is no measurable alignment between any genes, the method falls back to a coarser measure of genome relationship: the mutual information of k-mer frequency. This results in a principled model which depends only on the genome sequence, which captures many interesting relationships between viral families, and which creates clusters which correlate well with both the Baltimore and ICTV classifications. The incremental computational cost of classifying a novel virus is low and therefore newly discovered viruses can be quickly identified and classified.

genomics

The African Bullfrog (Pyxicephalus adspersus) genome unites the two ancestral ingredients formaking vertebrate sex chromosomes

Heteromorphic sex chromosomes have evolved repeatedly among vertebrate lineages despite largely deleterious reductions in gene dose. Understanding how this gene dose problem is overcome is hampered by the lack of genomic information at the base of tetrapods and comparisons across the evolutionary history of vertebrates. To address this problem, we produced a chromosome-level genome assembly for the African Bullfrog (Pyxicephalus adspersus)--an amphibian with heteromorphic ZW sex chromosomes--and discovered that the Bullfrog Z is surprisingly homologous to substantial portions of the human X. Using this new reference genome, we identified ancestral synteny among the sex chromosomes of major vertebrate lineages, showing that non-mammalian sex chromosomes are strongly associated with a single vertebrate ancestral chromosome, while mammals are associated with another that displays increased haploinsufficiency. The sex chromosomes of the African Bullfrog however, share genomic blocks with both humans and non-mammalian vertebrates, connecting the two ancestral chromosome sequences that repeatedly characterize vertebrate sex chromosomes. Our results highlight the consistency of sex-linked sequences despite sex determination system lability and reveal the repeated use of two major genomic sequence blocks during vertebrate sex chromosome evolution.

genomics

Extended regions of suspected mis-assembly in the rat reference genome

We performed whole-genome sequencing for eight inbred rat strains commonly used in genetic mapping studies, and they are the founders of the NIH heterogeneous stock (HS) outbred colony. We provide their sequences and variant calls to the rat genomics community. When analyzing the variant calls we identified regions with unusually high heterozygosity. We show that these regions are consistent across the eight inbred strains, including the BN strain, which was the basis of the rat reference genome. These regions show significantly higher read depth than other regions in the genome. The evidence suggests that these regions may contain segmental duplications that are incorrectly overlaid in the reference genome. We provide masks for these suspected regions of mis-assembly as a resource for the community to flag potentially false interpretations of mapping results or functional data.

genomics