Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

First draft genome assembly of an iconic clownfish species (Amphiprion frenatus)

Clownfishes (or anemonefishes) form an iconic group of coral reef fishes, particularly known for their mutualistic interaction with sea anemones. They are characterized by particular life history traits, such as a complex social structure and mating system involving sequential hermaphroditism, coupled with an exceptionally long lifespan. Additionally, clownfishes are considered to be one of the rare group to have experienced an adaptive radiation in the marine environment.\n\nHere, we assembled and annotated the first genome of a clownfish species, the tomato clownfish (Amphiprion frenatus). We obtained a total of 17,801 assembled scaffolds, containing a total of 26,917 genes. The completeness of the assembly and annotation was satisfying, with 96.5% of the Actinopterygii BUSCOs (Benchmarking Universal Single-Copy Orthologs) being retrieved in A. frenatus assembly. The quality of the resulting assembly is comparable to other bony fish assemblies.\n\nThis resource is valuable for the advancing of studies of the particular life-history traits of clownfishes, as well as being useful for population genetic studies and the development of new phylogenetic markers. It will also open the way to comparative genomics. Indeed, future genomic comparison among closely related fishes may provide means to identify genes related to the unique adaptations to different sea anemone hosts, as well as better characterize the genomic signatures of an adaptive radiation.

genomics

Leveraging Transcriptomics Data for Genomic Prediction Models in Cassava

BackgroundGenomic prediction models were, in principle, developed to include all the available marker information; with this approach, these models have shown in various crops moderate to high predictive accuracies. Previous studies in cassava have demonstrated that, even with relatively small training populations and low-density GBS markers, prediction models are feasible for genomic selection. In the present study, we prioritized SNPs in close proximity to genome regions with biological importance for a given trait. We used a number of strategies to select variants that were then included in single and multiple kernel GBLUP models. Specifically, our sources of information were transcriptomics, GWAS, and immunity-related genes, with the ultimate goal to increase predictive accuracies for Cassava Brown Streak Disease (CBSD) severity.\n\nResultsWe used single and multi-kernel GBLUP models with markers imputed to whole genome sequence level to accommodate various sources of biological information; fitting more than one kinship matrix allowed for differential weighting of the individual marker relationships. We applied these GBLUP approaches to CBSD phenotypes (i.e., root infection and leaf severity three and six months after planting) in a Ugandan Breeding Population (n = 955). Three means of exploiting an established RNAseq experiment of CBSD-infected cassava plants were used. Compared to the biology-agnostic GBLUP model, the accuracy of the informed multi-kernel models increased the prediction accuracy only marginally (1.78% to 2.52%).\n\nConclusionsOur results show that markers imputed to whole genome sequence level do not provide enhanced prediction accuracies compared to using standard GBS marker data in cassava. The use of transcriptomics data and other sources of biological information resulted in prediction accuracies that were nominally superior to those obtained from traditional prediction models.

genomics

Analysis Of The Genomic Basis Of Functional Diversity In Dinoflagellates Using A Transcriptome-Based Sequence Similarity Network

Dinoflagellates are one of the most abundant and functionally diverse groups of eukaryotes. Despite an overall scarcity of genomic information for dinoflagellates, constantly emerging high-throughput sequencing resources can be used to characterize and compare these organisms. We assembled de novo and processed 46 dinoflagellate transcriptomes and used a sequence similarity network (SSN) to compare the underlying genomic basis of functional features within the group. This approach constitutes the most comprehensive picture to date of the genomic potential of dinoflagellates. A core proteome composed of 252 connected components (CCs) of putative conserved protein domains (pCDs) was identified. Of these, 206 were novel and 16 lacked any functional annotation in public databases. Integration of functional information in our network analyses allowed investigation of pCDs specifically associated to functional traits. With respect to toxicity, sequences homologous to those of proteins involved in toxin biosynthesis pathways (e.g. sxtA1-4 and sxtG) were not specific to known toxin-producing species. Although not fully specific to symbiosis, the most represented functions associated with proteins involved in the symbiotic trait were related to membrane processes and ion transport. Overall, our SSN approach led to identification of 45,207 and 90,794 specific and constitutive pCDs of respectively the toxic and symbiotic species represented in our analyses. Of these, 56% and 57% respectively (i.e. 25,393 and 52,193 pCDs) completely lacked annotation in public databases. This stresses the extent of our lack of knowledge, while emphasizing the potential of SSNs to identify candidate pCDs for further functional genomic characterization.

genomics

Insights into platypus population structure and history from whole-genome sequencing

The platypus is an egg-laying mammal which, alongside the echidna, occupies a unique place in the mammalian phylogenetic tree. Despite widespread interest in its unusual biology, little is known about its population structure or recent evolutionary history. To provide new insights into the dispersal and demographic history of this iconic species, we sequenced the genomes of 57 platypuses from across the whole species range in eastern mainland Australia and Tasmania. Using a highly-improved reference genome, we called over 6.7M SNPs, providing an informative genetic data set for population analyses. Our results show very strong population structure in the platypus, with our sampling locations corresponding to discrete groupings between which there is no evidence for recent gene flow. Genome-wide data allowed us to establish that 28 of the 57 sampled individuals had at least a third-degree relative amongst other samples from the same river, often taken at different times. Taking advantage of a sampled family quartet, we estimated the de novo mutation rate in the platypus at 7.0x10-9/bp/generation (95% CI 4.1x10-9 - 1.2x10-8/bp/generation). We estimated effective population sizes of ancestral populations and haplotype sharing between current groupings, and found evidence for bottlenecks and long-term population decline in multiple regions, and early divergence between populations in different regions. This study demonstrates the power of whole-genome sequencing for studying natural populations of an evolutionarily important species.

genomics

Expressed Exome Capture Sequencing (EecSeq): a method for cost-effective exome sequencing for all organisms with or without genomic resources

Exome capture is an effective tool for surveying the genome for loci under selection. However, traditional methods require annotated genomic resources. Here, we present a method for creating cDNA probes from expressed mRNA, which are then used to enrich and capture genomic DNA for exon regions. This approach, called \"EecSeq\", eliminates the need for costly probe design and synthesis. We tested EecSeq in the eastern oyster, Crassostrea virginica, using a controlled exposure experiment. Four adult oysters were heat shocked at 36{degrees} C for 1 hour along with four control oysters kept at 14{degrees} C. Stranded mRNA libraries were prepared for two individuals from each treatment and pooled. Half of the combined library was used for probe synthesis and half was sequenced to evaluate capture efficiency. Genomic DNA was extracted from all individuals, enriched via captured probes, and sequenced directly. We found that EecSeq had an average capture sensitivity of 86.8% across all known exons and had over 99.4% sensitivity for exons with detectable levels of expression in the mRNA library. For all mapped reads, over 47.9% mapped to exons and 37.0% mapped to expressed targets, which is similar to previously published exon capture studies. EecSeq displayed relatively even coverage within exons (i.e. minor \"edge effects\") and even coverage across exon GC content. We discovered 5,951 SNPs with a minimum average coverage of 80X, with 3,508 SNPs appearing in exonic regions. We show that EecSeq provides comparable, if not superior, specificity and capture efficiency compared to costly, traditional methods.

genomics

Whole genome hybrid assembly and protein-coding gene annotation of the entirely black native Korean chicken breed Yeonsan Ogye

Yeonsan Ogye (YO), an indigenous Korean chicken breed (gallus gallus domesticus), has entirely black external features and internal organs. In this study, the draft genome of YO was assembled using a hybrid de novo assembly method that takes advantage of high-depth Illumina short-reads (232.2X) and low-depth PacBio long-reads (11.5X). Although the contig and scaffold N50s (defined as the shortest contig or scaffold length at 50% of the entire assembly) of the initial de novo assembly were 53.6Kbp and 10.7Mbp, respectively, additional and pseudo-reference-assisted assemblies extended the assembly to 504.8Kbp for contig N50 (pseudo-contig) and 21.2Mbp for scaffold N50, which included 551 structural variations including the Fibromelanosis (FM) locus duplication, compared to galGal4 and 5. The completeness (97.6%) of the draft genome (Ogye_1) was evaluated with single copy orthologous genes using BUSCO, and found to be comparable to the current chicken reference genome (galGal5; 97.4%), which was assembled with a long read-only method, and superior to other avian genomes (92~93%), assembled with short read-only and hybrid methods. To comprehensively reconstruct transcriptome maps, RNA sequencing (RNA-seq) and representation bisulfite sequencing (RRBS) data were analyzed from twenty different tissues, including black tissues. The maps included 15,766 protein-coding and 6,900 long non-coding RNA genes, many of which were expressed in the tissue-specific manner, closely related with the DNA methylation pattern in the promoter regions.

genomics

Deep-coverage whole genome sequences and blood lipids among 16,324 individuals

Deep-coverage whole genome sequencing at the population level is now feasible and offers potential advantages for locus discovery, particularly in the analysis rare mutations in non-coding regions. Here, we performed whole genome sequencing in 16,324 participants from four ancestries at mean depth >29X and analyzed correlations of genotypes with four quantitative traits - plasma levels of total cholesterol, low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol, and triglycerides. We conducted a discovery analysis including common or rare variants in coding as well as non-coding regions and developed a framework to interpret genome sequence for dyslipidemia risk. Common variant association yielded loci previously described with the exception of a few variants not captured earlier by arrays or imputation. In coding sequence, rare variant association yielded known Mendelian dyslipidemia genes and, in non-coding sequence, we detected no rare variant association signals after application of four approaches to aggregate variants in non-coding regions. We developed a new, genome-wide polygenic score for LDL-C and observed that a high polygenic score conferred similar effect size to a monogenic mutation (~30 mg/dl higher LDL-C for each); however, among those with extremely high LDL-C, a high polygenic score was considerably more prevalent than a monogenic mutation (23% versus 2% of participants, respectively).

genomics

Decoding the chromatin proteome of a single genomic locus by DNA sequencing

Transcription, replication and repair involve interactions of specific genomic loci with many different proteins. How these interactions are orchestrated at any given location and under changing cellular conditions is largely unknown because systematically measuring protein-DNA interactions at a specific locus in the genome is challenging. To address this problem, we developed Epi-Decoder, a Tag-ChIP-Barcode-Seq technology in budding yeast to identify and quantify in an unbiased and systematic manner the proteome of an individual genomic locus. Epi-Decoder is orthogonal to proteomics approaches because it does not rely on mass spectrometry but instead takes advantage of DNA sequencing. Analysis of the proteome of a transcribed locus proximal to an origin of replication revealed more than 400 proteins. Moreover, replication stress induced changes in local chromatin-proteome composition prior to local origin firing, affecting replication proteins as well as transcription proteins. Epi-Decoder will enable the delineation of complex and dynamic protein-DNA interactions across many regions of the genome.

genomics

Discovery and characterization of coding and non-coding driver mutations in more than 2,500 whole cancer genomes

Discovery of cancer drivers has traditionally focused on the identification of protein-coding genes. Here we present a comprehensive analysis of putative cancer driver mutations in both protein-coding and non-coding genomic regions across >2,500 whole cancer genomes from the Pan-Cancer Analysis of Whole Genomes (PCAWG) Consortium. We developed a statistically rigorous strategy for combining significance levels from multiple driver discovery methods and demonstrate that the integrated results overcome limitations of individual methods. We combined this strategy with careful filtering and applied it to protein-coding genes, promoters, untranslated regions (UTRs), distal enhancers and non-coding RNAs. These analyses redefine the landscape of non-coding driver mutations in cancer genomes, confirming a few previously reported elements and raising doubts about others, while identifying novel candidate elements across 27 cancer types. Novel recurrent events were found in the promoters or 5UTRs of TP53, RFTN1, RNF34, and MTG2, in the 3UTRs of NFKBIZ and TOB1, and in the non-coding RNA RMRP. We provide evidence that the previously reported non-coding RNAs NEAT1 and MALAT1 may be subject to a localized mutational process. Perhaps the most striking finding is the relative paucity of point mutations driving cancer in non-coding genes and regulatory elements. Though we have limited power to discover infrequent non-coding drivers in individual cohorts, combined analysis of promoters of known cancer genes show little excess of mutations beyond TERT.

genomics

Firefly genomes illuminate the origin and evolution of bioluminescence

Fireflies are among the best-studied of the bioluminescent organisms. Despite longterm interest in the biochemistry, neurobiology, and evolution of firefly flash signals and the widespread biotechnological applications of firefly luciferase, only a limited set of genes related to this complex trait have been described. To investigate the genetic basis of firefly bioluminescence, we generated a high-quality reference genome for the Big Dipper firefly Photinus pyralis, from which the first laboratory luciferase was cloned, using long-read (PacBio), short-read (Illumina), and Hi-C sequencing technologies. To facilitate comparative genomics, we also generated short-read genome assemblies for a Japanese firefly Aquatica lateralis and a bioluminescent click beetle, Ignelater luminosus. Analyses of these genomic datasets supports at least two independent gains of luminescence in beetles, and provides new insights into the evolution of beetle bioluminescence and chemical defenses that likely co-evolved over their 100 million years of evolution.

genomics

Epigenomic and genomic landscape of Drosophila melanogaster heterochromatic genes

Heterochromatin is associated with transcriptional repression. In contrast, several genes in the pericentromeric regions of Drosophila melanogaster are dependent on this heterochromatic environment for their expression. Heterochromatic genes encode proteins involved in various developmental processes. Several studies have shown that a variety of epigenetic modifications is associated with these genes. Here we present a comprehensive analysis of the epigenetic landscape of heterochromatic genes across all the developmental stages of Drosophila using the available histone modification and expression data from modENCODE. We find that heterochromatic genes exhibit combinations of active and inactive histone marks that correspond to their level of expression during development. Thus, we classified these genes into three groups based on the combinations of histone modifications present. We also looked for potential regulatory DNA sequence elements in the genomic neighborhood of these genes. Our results show that Nuclear Matrix Associated Regions (MARs) are prominently present in the intergenic regions of heterochromatic genes during embryonic stages suggesting their plausible role in pericentromeric genome organization. We also find that the intergenic sequences in the heterochromatic regions have binding sites for transcription factors known to modulate epigenetic status. Taken together, our meta-analysis of the various genomic datasets suggest that the epigenomic and genomic landscape of the heterochromatic genes are distinct from that of euchromatic genes. These features could be contributing to the unusual regulatory status of the heterochromatic genes as opposed to the surrounding heterochromatin, which is repressive in nature.

genomics

Comparison of two African rice species through a new pan-genomic approach on massive data

Pangenome theory implies that individuals from a given group/species share only a given part of their genome (core-genome), the remaining part being the dispensable one. Domestication process implies a small number of founder individuals, and thus a large core-genome compared to dispensable at the first steps of domestication. We sequenced at high depth 120 cultivated African rice Oryza glaberrima and of 74 wild relatives O. barthii, and mapped them on the external reference from Asian rice O. sativa. We then use a novel DepthOfCoverage approach to identif missing genes. After comparing the two species, we shown that the cultivated species has a smaller core-genome than the wild one, as well as an expected smaller dispensable one. This unexpected output however replaces in perspective the inadequacy of cultivated crops to wilderness.

genomics

The Juicebox Assembly Tools module facilitates de novo assembly of mammalian genomes with chromosome-length scaffolds for under $1000

Hi-C contact maps are valuable for genome assembly (Lieberman-Aiden, van Berkum et al. 2009; Burton et al. 2013; Dudchenko et al. 2017). Recently, we developed Juicebox, a system for the visual exploration of Hi-C data (Durand, Robinson et al. 2016), and 3D-DNA, an automated pipeline for using Hi-C data to assemble genomes (Dudchenko et al. 2017). Here, we introduce \"Assembly Tools,\" a new module for Juicebox, which provides a point-and-click interface for using Hi-C heatmaps to identify and correct errors in a genome assembly. Together, 3D-DNA and the Juicebox Assembly Tools greatly reduce the cost of accurately assembling complex eukaryotic genomes. To illustrate, we generated de novo assemblies with chromosome-length scaffolds for three mammals: the wombat, Vombatus ursinus (3.3Gb), the Virginia opossum, Didelphis virginiana (3.3Gb), and the raccoon, Procyon lotor (2.5Gb). The only inputs for each assembly were Illumina reads from a short insert DNA-Seq library (300 million Illumina reads, maximum length 2x150 bases) and an in situ Hi-C library (100 million Illumina reads, maximum read length 2x150 bases), which cost <$1000.

genomics

Germline DNA replication timing shapes mammalian genome composition

Mammalian DNA is replicated in a highly organized and regulated manner. Large, Mb-sized regions are replicated at defined times along S phase. DNA Replication Timing (RT) has been suggested to play an important role in shaping the mammalian genome by affecting mutation rates. Previous analyses relied on somatic DNA RT profiles, while to fully understand the influences of RT on the mammalian genome, germ cell RT information is necessary, as only germline mutations are passed to offspring and thus affect genomic composition. Using an improved RT mapping technique that allows mapping the RT from limited amounts of cells, we measured RT from two stages in the mouse germline - primordial germ cells (PGCs) and spermatogonial stem cells (SSCs). The germ cell RT profiles were distinct from those of both somatic and embryonic tissues. The correlations between RT and both mutation rate and recombination hotspots were not only confirmed in the germline tissues, but were shown to be stronger compared to correlations with RT of somatic tissues, emphasizing the importance of using RT profiles from the correct tissue of origin. Expanding the analysis to additional genetic features such as GC content, transposable elements (SINEs and LINEs) and gene density, also revealed a stronger correlation with the germ cell RT maps. GC content stratification along with multiple regression analysis revealed the independent contribution of RT to SINE, gene, mutation and recombination hotspot densities. Taken together, our results point to the centrality of RT in shaping multiple levels of mammalian genome composition.

genomics

Sequencing Metrics of Human Genomes Extracted from Single Cancer Cells Individually Isolated in a Valveless Microfluidic Device

Sequencing the genomes of individual cells enables the direct determination of genetic heterogeneity amongst cells within a population. We have developed an injection-moulded valveless microfluidic device in which single cells from colorectal cell (LS174T, LS180 and RKO) lines and fresh colorectal cancers are individually trapped, their genomes extracted and prepared for sequencing, using multiple displacement amplification (MDA). Ninety nine percent of the DNA sequences obtained mapped to a reference human genome, indicating that there was effectively no contamination of these samples from non-human sources. In addition, most of the reads are correctly paired, with a low percentage of singletons (0.17 {+/-} 0.06 %) and we obtain genome coverages approaching 90%. To achieve this high quality, our device design and process shows that amplification can be conducted in microliter volumes as long as extraction is in sub-nanoliter volumes. Our data also demonstrates that high quality single cell sequencing can be achieved using a relatively simple, inexpensive and scalable device.

genomics

How to make use of ordination methods to identify local adaptation: a comparison of genome scans based on PCA and RDA

Ordination is a common tool in ecology that aims at representing complex biological information in a reduced space. In landscape genetics, ordination methods such as principal component analysis (PCA) have been used to detect adaptive variation based on genomic data. Taking advantage of environmental data in addition to genotype data, redundancy analysis (RDA) is another ordination approach that is useful to detect adaptive variation. This paper aims at proposing a test statistic based on RDA to search for loci under selection. We compare redundancy analysis to pcadapt, which is a nonconstrained ordination method, and to a latent factor mixed model (LFMM), which is a univariate genotype-environment association method. Individual-based simulations identify evolutionary scenarios where RDA genome scans have a greater statistical power than genome scans based on PCA. By constraining the analysis with environmental variables, RDA performs better than PCA in identifying adaptive variation when selection gradients are weakly correlated with population structure. Additionally, we show that if RDA and LFMM have a similar power to identify genetic markers associated with environmental variables, the RDA-based procedure has the advantage to identify the main selective gradients as a combination of environmental variables. To give a concrete illustration of RDA in population genomics, we apply this method to the detection of outliers and selective gradients on an SNP data set of Populus trichocarpa (Geraldes et al., 2013). The RDA-based approach identifies the main selective gradient contrasting southern and coastal populations to northern and continental populations in the northwestern American coast.

genomics

Genomes of trombidid mites reveal novel predicted allergens and laterally-transferred genes associated with secondary metabolism

BackgroundTrombidid mites have a unique lifecycle in which only the larval stage is ectoparasitic. In the superfamily Trombiculoidea (\"chiggers\"), the larvae feed preferentially on vertebrates, including humans. Species in the genus Leptotrombidium are vectors of a potentially fatal bacterial infection, scrub typhus, which affects 1 million people annually. Moreover, chiggers can cause pruritic dermatitis (trombiculiasis) in humans and domesticated animals. In the Trombidioidea (velvet mites), the larvae feed on other arthropods and are potential biological control agents for agricultural pests. Here, we present the first trombidid mites genomes, obtained both for a chigger, Leptotrombidium deliense, and for a velvet mite, Dinothrombium tinctorium.\n\nResultsSequencing was performed using Illumina technology. A 180 Mb draft assembly for D. tinctorium was generated from two paired-end and one mate-pair library using a single adult specimen. For L. deliense, a lower-coverage draft assembly (117 Mb) was obtained using pooled, engorged larvae with a single paired-end library. Remarkably, both genomes exhibited evidence of ancient lateral gene transfer from soil-derived bacteria or fungi. The transferred genes confer functions that are rare in animals, including terpene and carotenoid synthesis. Thirty-seven allergenic protein families were predicted in the L. deliense genome, of which nine were unique. Preliminary proteomic analyses identified several of these putative allergens in larvae.\n\nConclusionsTrombidid mite genomes appear to be more dynamic than those of other acariform mites. A priority for future research is to determine the biological function of terpene synthesis in this taxon and its potential for exploitation in disease control.

genomics

Spatial Chromatin Architecture Alteration by Structural Variations in Human Genomes at Population Scale

This genome-wide study is focused on the impact of structural variants identified in individuals from 26 human populations onto three-dimensional structures of their genomes. We assess the tendency of structural variants to accumulate in spatially interacting genomic segments and design a high-resolution computational algorithm to model the 3D conformational changes resulted by structural variations. We show that differential gene transcription is closely linked to variation in chromatin interaction networks mediated by RNA polymerase II. We also demonstrate that CTCF-mediated interactions are well conserved across population, but enriched with disease-associated SNPs. Altogether, this study assesses the critical impact of structural variants on the higher order organization of chromatin folding and provides unique insight into the mechanisms regulating gene transcription at the population scale, among which the local arrangement of chromatin loops seems to be the leading one. It is the first insight into the variability of the human 3D genome at the population scale.

genomics