Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

High-depth whole genome sequencing of a large population-specific reference panel: Enhancing sensitivity, accuracy, and imputation

BackgroundWhile increasingly large reference panels for genome-wide imputation have been recently made available, the degree to which imputation accuracy can be enhanced by population-specific reference panels remains an open question. In the present study, we sequenced at full-depth ([&ge;]30x) a moderately large (n=738) cohort of samples drawn from the Ashkenazi Jewish population across two platforms (Illumina X Ten and Complete Genomics, Inc.). We developed and refined a series of quality control steps to optimize sensitivity, specificity, and comprehensiveness of variant calls in the reference panel, and then tested the accuracy of imputation against target cohorts drawn from the same population.\n\nResultsFor samples sequenced on the Illumina X Ten platform, quality thresholds were identified that permitted highly accurate calling of single nucleotide variants across 94% of the genome. The Complete Genomics, Inc. platform was more conservative (fewer variants called) compared to the Illumina platform, but also demonstrated relatively greater numbers of false positives that needed to be filtered. Quality control procedures also permitted detection of novel genome reads that are not mapped to current reference or alternate assemblies. After stringent quality control, the population-specific reference panel produced more accurate and comprehensive imputation results relative to publicly available, large cosmopolitan reference panels. The population-specific reference panel also permitted enhanced filtering of clinically irrelevant variants from personal genomes.\n\nConclusionsOur primary results demonstrate enhanced accuracy of a population-specific imputation panel relative to cosmopolitan panels, especially in the range of infrequent (<5% non-reference allele frequency) and rare (<1% non-reference allele frequency) variants that may be most critical to further progress in mapping of complex phenotypes.

genomics

Genomic landscape of oxidative DNA damage and repair reveals regioselective protection from mutagenesis

DNA is subject to constant chemical modification and damage, which eventually results in variable mutation rates throughout the genome. Although detailed molecular mechanisms of DNA damage and repair are well-understood, damage impact and execution of repair across a genome remains poorly defined. To bridge the gap between our understanding of DNA repair and mutation distributions we developed a novel method, AP-seq, capable of mapping apurinic sitesand 8-oxo-7,8-dihydroguanine bases at [~]300bp resolution on a genome-wide scale. We directly demonstrate that the accumulation rate of oxidative damage varies widely across the genome, with hot spots acquiring many times more damage than cold spots. Unlike SNVs in cancers, damage burden correlates with marks for open chromatin notably H3K9ac and H3K4me2. Oxidative damage is also highly enriched in transposable elements and other repetitive sequences. In contrast, we observe decreased damage at promoters, exons and termination sites, but not introns, in a seemingly transcription-independent manner. Leveraging cancer genomic data, we also find locally reduced SNV rates in promoters, genes and other functional elements. Taken together, our study reveals that oxidative DNA damage accumulation and repair differ strongly across the genome, but culminate in a previously unappreciated mechanism that safe-guards the regulatory sequences and the coding regions of genes from mutations.

genomics

Genome-reconstruction for eukaryotes from complex natural microbial communities

Microbial eukaryotes are integral components of natural microbial communities and their inclusion is critical for many ecosystem studies yet the majority of published metagenome analyses ignore eukaryotes. In order to include eukaryotes in environmental studies we propose a method to recover eukaryotic genomes from complex metagenomic samples. A key step for genome recovery is separation of eukaryotic and prokaryotic fragments. We developed a kmer-based strategy, EukRep, for eukaryotic sequence identification and applied it to environmental samples to show that it enables genome recovery, genome completeness evaluation and prediction of metabolic potential. We used this approach to test the effect of addition of organic carbon on a geyser-associated microbial community and detected a substantial change of the community metabolism, with selection against almost all candidate phyla bacteria and archaea and for eukaryotes. Near complete genomes were reconstructed for three fungi placed within the eurotiomycetes and an arthropod. While carbon fixation and sulfur oxidation were important functions in the geyser community prior to carbon addition, the organic carbon impacted community showed enrichment for secreted proteases, secreted lipases, cellulose targeting CAZymes, and methanol oxidation. We demonstrate the broader utility of EukRep by reconstructing and evaluating relatively high quality fungal, protist, and rotifer genomes from complex environmental samples. This approach opens the way for cultivation-independent analyses of whole microbial communities.

genomics

A Mechanism of Cohesin-Dependent Loop Extrusion Organizes Zygotic Genome Architecture

Fertilization triggers assembly of higher-order chromatin structure from a naive genome to generate a totipotent embryo. Chromatin loops and domains are detected in mouse zygotes by single-nucleus Hi-C (snHi-C) but not bulk Hi-C. We resolve this discrepancy by investigating whether a mechanism of cohesin-dependent loop extrusion generates zygotic chromatin conformations. Using snHi-C of mouse knockout embryos, we demonstrate that the zygotic genome folds into loops and domains that depend on Scc1-cohesin and are regulated in size by Wapl. Remarkably, we discovered distinct effects on maternal and paternal chromatin loop sizes, likely reflecting loop extrusion dynamics and epigenetic reprogramming. Polymer simulations based on snHi-C are consistent with a model where cohesin locally compacts chromatin and thus restricts inter-chromosomal interactions by active loop extrusion, whose processivity is controlled by Wapl. Our simulations and experimental data provide evidence that cohesin-dependent loop extrusion organizes mammalian genomes over multiple scales from the one-cell embryo onwards.\n\nHighlightsO_LIZygotic genomes are organized into cohesin-dependent chromatin loops and TADs\nC_LIO_LILoop extrusion leads to different loop strengths in maternal and paternal genomes\nC_LIO_LICohesin restricts inter-chromosomal interactions by altering chromosome surface area\nC_LIO_LILoop extrusion organizes chromatin at multiple genomic scales\nC_LI

genomics

De novo assembly and phasing of dikaryotic genomes from two isolates of Puccinia coronata f. sp. avenae, the causal agent of oat crown rust

Oat crown rust, caused by the fungus Puccinia coronata f. sp. avenae (Pca), is a devastating disease that impacts worldwide oat production. For much of its life cycle, Pca is dikaryotic, with two separate haploid nuclei that may vary in virulence genotype, highlighting the importance of understanding haplotype diversity in this species. We generated highly contiguous de novo genome assemblies of two Pca isolates, 12SD80 and 12NC29, from long-read sequences. In total, we assembled 603 primary contigs for a total assembly length of 99.16 Mbp for 12SD80 and 777 primary contigs with a total length of 105.25 Mbp for 12NC29, and approximately 52% of each genome was assembled into alternate haplotypes. This revealed structural variation between haplotypes in each isolate equivalent to more than 2% of the genome size, in addition to about 260,000 and 380,000 heterozygous single-nucleotide polymorphisms in 12SD80 and 12NC29, respectively. Transcript-based annotation identified 26,796 and 28,801 coding sequences for isolates 12SD80 and 12NC29, respectively, including about 7,000 allele pairs in haplotype-phased regions. Furthermore, expression profiling revealed clusters of co-expressed secreted effector candidates, and the majority of orthologous effectors between isolates showed conservation of expression patterns. However, a small subset of orthologs showed divergence in expression, which may contribute to differences in virulence between 12SD80 and 12NC29. This study provides the first haplotype-phased reference genome for a dikaryotic rust fungus as a foundation for future studies into virulence mechanisms in Pca.\n\nImportanceDisease management strategies for oat crown rust are challenged by the rapid evolution of Puccinia coronata f. sp. avenae (Pca), which renders resistance genes in oat varieties ineffective. Despite the economic importance of understanding Pca, resources to study the molecular mechanisms underpinning pathogenicity and emergence of new virulence traits are lacking. Such limitations are partly due to the obligate biotrophic lifestyle of Pca as well as the dikaryotic nature of the genome, features that are also shared with other important rust pathogens. This study reports the first release of a haplotype-phased genome assembly for a dikaryotic fungal species and demonstrates the amenability of using emerging technologies to investigate genetic diversity in populations of Pca.

genomics

The whole-genome panorama of cancer drivers

The advance of personalized cancer medicine requires the accurate identification of the mutations driving each patients tumor. However, to date, we have only been able to obtain partial insights into the contribution of genomic events to tumor development. Here, we design a comprehensive approach to identify the driver mutations in each patients tumor and obtain a whole-genome panorama of driver events across more than 2,500 tumors from 37 types of cancer. This panorama includes coding and non-coding point mutations, copy number alterations and other genomic rearrangements of somatic origin, and potentially predisposing germline variants. We demonstrate that genomic events are at the root of virtually all tumors, with each carrying on average 4.6 driver events. Most individual tumors harbor a unique combination of drivers, and we uncover the most frequent co-occurring driver events. Half of all cancer genes are affected by several types of driver mutations. In summary, the panorama described here provides answers to fundamental questions in cancer genomics and bridges the gap between cancer genomics and personalized cancer medicine.

cancer biology

Differential distribution of Neandertal genomic signatures in human mitochondrial haplogroups

Genetic contributions of Neanderthals to the modern human genome have been evidenced by comparison of present-day human genomes with paleogenomes suggesting that the Neanderthal introgression is higher in Asians and Europeans and lower in Africans. Neanderthal signatures in extant human genomes are attributed to intercrosses between Neanderthals and archaic Anatomically Modern Humans (AMH). Although Neanderthal signatures are well documented in the nuclear genome, it has been proposed that there is no contribution of Neanderthal mitochondrial DNA to contemporary human genomes. Here we show that modern human mitochondrial genomes contain potential 66 Neanderthal signatures, or Neanderthal single nucleotide variants (N-SNVs) being 36 in coding regions of which 7 are nonsynonymous. Also, 7 N-SNVs are associated with traits such as cycling vomiting syndrome, Alzheimers disease, Parkinsons disease and 2 N-SNVs are associated with intelligence quotient. Based on recombination tests, Principal Component Analysis (PCA) and the complete absence of these N-SNVs in 41 archaic AMH mitogenomes we conclude that convergent evolution due to homoplasy and not recombination, explains the presence of N-SNVs in present-day human mitogenomes.

genomics

Symbiodinium genomes reveal adaptive evolution of functions related to symbiosis

Symbiosis between dinoflagellates of the genus Symbiodinium and reef-building corals forms the trophic foundation of the worlds coral reef ecosystems. Here we present the first draft genome of Symbiodinium goreaui (Clade C, type C1: 1.03 Gbp), one of the most ubiquitous endosymbionts associated with corals, and an improved draft genome of Symbiodinium kawagutii (Clade F, strain CS-156: 1.05 Gbp), previously sequenced as strain CCMP2468, to further elucidate genomic signatures of this symbiosis. Comparative analysis of four available Symbiodinium genomes against other dinoflagellate genomes led to the identification of 2460 nuclear gene families that show evidence of positive selection, including genes involved in photosynthesis, transmembrane ion transport, synthesis and modification of amino acids and glycoproteins, and stress response. Further, we identified extensive sets of genes for meiosis and response to light stress. These draft genomes provide a foundational resource for advancing our understanding Symbiodinium biology and the coral-algal symbiosis.

genomics

Genome-wide characterization of copy number variants in epilepsy patients

Epilepsy will affect nearly 3% of people at some point during their lifetime. Previous copy number variants (CNVs) studies of epilepsy have used array-based technology and were restricted to the detection of large or exonic events. In contrast, whole-genome sequencing (WGS) has the potential to more comprehensively profile CNVs but existing analytic methods suffer from limited accuracy. We show that this is in part due to the non-uniformity of read coverage, even after intra-sample normalization. To improve on this, we developed PopSV, an algorithm that uses multiple samples to control for technical variation and enables the robust detection of CNVs. Using WGS and PopSV, we performed a comprehensive characterization of CNVs in 198 individuals affected with epilepsy and 301 controls. For both large and small variants, we found an enrichment of rare exonic events in epilepsy patients, especially in genes with predicted loss-of-function intolerance. Notably, this genome-wide survey also revealed an enrichment of rare non-coding CNVs near previously known epilepsy genes. This enrichment was strongest for non-coding CNVs located within 100 Kbp of an epilepsy gene and in regions associated with changes in the gene expression, such as expression QTLs or DNase I hypersensitive sites. Finally, we report on 21 potentially damaging events that could be associated with known or new candidate epilepsy genes. Our results suggest that comprehensive sequence-based profiling of CNVs could help explain a larger fraction of epilepsy cases.\n\nAuthor summaryEpilepsy is a common neurological disorder affecting around 3% of the population. In some cases, epilepsy is caused by brain trauma or other brain anomalies but there are often no clear causes. Genetic factors have been associated with epilepsy in the past such as rare genetic variations found by linkage studies as well as common genetic variations found by genome-wide association studies and large copy-number variants. We sequenced the genome of[~] 200 epilepsy patients and[~] 300 healthy controls and compared the distribution of deletion (loss of a copy) and duplication (additional copy) of genomic regions. Thanks to the sequencing technology and a new method that takes advantage of the large sample size, we could compare the distribution of small copy- number variants between epilepsy patients and controls. Overall, we found that small variants are also associated with epilepsy. Indeed, the genome of epilepsy patients had more exonic copy- number variants, especially when rare or affecting genes with predicted loss-of-function intolerance. Focusing on regions around genes that have been previously associated with epilepsy, we also found more non-coding variants in epilepsy patients, especially deletions or variants in regulatory regions. Finally, we provide a list of 21 regions in which we found likely pathogenic variants.

genomics

Genomics in healthcare: GA4GH looks to 2022

The Global Alliance for Genomics and Health (GA4GH), the standards-setting body in genomics for healthcare, aims to accelerate biomedical advancement globally. We describe the differences between healthcare- and research-driven genomics, discuss the implications of global, population-scale collections of human data for research, and outline mission-critical considerations in ethics, regulation, technology, data protection, and society. We present a crude model for estimating the rate of healthcare-funded genomes worldwide that accounts for the preparedness of each country for genomics, and infers a progression of cancer-related sequencing over time. We estimate that over 60 million patients will have their genome sequenced in a healthcare context by 2025. This represents a large technical challenge for healthcare systems, and a huge opportunity for research. We identify eight major practical, principled arguments to support the position that virtual cohorts of 100 million people or more would have tangible research benefits.

genomics

Draft genomes of the fungal pathogen Phellinus noxius in Hong Kong

The fungal pathogen Phellinus noxius is the underlying cause of brown root rot, a disease with causing tree mortality globally, causing extensive damage in urban areas and crop plants. This disease currently has no cure, and despite the global epidemic, little is known about the pathogenesis and virulence of this pathogen.\n\nUsing Ion Torrent PGM, Illumina MiSeq and PacBio RSII sequencing platforms with various genome assembly methods, we produced the draft genome sequences of four P. noxius strains isolated from infected trees in Hong Kong to further understand the pathogen and identify the mechanisms behind the aggressive nature and virulence of this fungus. The resulting genomes ranged from 30.8Mb to 31.8Mb in size, and of the four sequences, the YTM97 strain was chosen to produce a high-quality Hong Kong strain genome sequence, resulting in a 31Mb final assembly with 457 scaffolds, an N50 length of 275,889 bp and 96.2% genome completeness. RNA-seq of YTM97 using Illumina HiSeq400 was performed for improved gene prediction. AUGUSTUS and Genemark-ES prediction programs predicted 9,887 protein-coding genes which were annotated using GO and Pfam databases. The encoded carbohydrate active enzymes revealed large numbers of lignolytic enzymes present, comparable to those of other white-rot plant pathogens. In addition, P. noxius also possessed larger numbers of cellulose, xylan and hemicellulose degrading enzymes than other plant pathogens. Searches for virulence genes was also performed using PHI-Base and DFVF databases revealing a host of virulence-related genes and effectors. The combination of non-specific host range, unique carbohydrate active enzyme profile and large amount of putative virulence genes could explain the reasons behind the aggressive nature and increased virulence of this plant pathogen.\n\nThe draft genome sequences presented here will provide references for strains found in Hong Kong. Together with emerging research, this information could be used for genetic diversity and epidemiology research on a global scale as well as expediting our efforts towards discovering the mechanisms of pathogenicity of this devastating pathogen.

genomics

Genomic characterisation and conservation genetics of the indigenous Irish Kerry cattle breed

Kerry cattle are an endangered landrace heritage breed of cultural importance to Ireland. In the present study we have used genome-wide SNP data (Illumina(R) BovineSNP50 array) to evaluate genomic diversity within the Kerry cattle population and between Kerry cattle and other European cattle breeds. Visualisation of patterns of genetic differentiation and gene flow among cattle breeds using phylogenetic trees with ancestry graphs highlighted, in particular, historical gene flow from the British Shorthorn breed into the ancestral population of modern Kerry cattle. Principal component analysis (PCA) and genetic clustering emphasised the genetic distinctiveness of Kerry cattle relative to comparator British and European cattle breeds. Modelling of genetic effective population size (Ne) revealed a demographic trend of diminishing Ne over time and that recent estimated Ne values for the Kerry breed may be less than the threshold for sustainable genetic conservation. In addition, analysis of genome-wide autozygosity (FROH) showed that genomic inbreeding has increased significantly during the 20 years between 1992 and 2012. Finally, signatures of selection revealed genomic regions subject to natural and artificial selection as Kerry cattle adapted to the climate, physical geography and agro-ecology of southwest Ireland.\n\nNote 1: This is an Associate Editor (D.E.M) Inaugural Article submission to Frontiers in Genetics: Livestock Genomics\n\nNote 2: British English language style preferred for publication of this article.

genomics

Whole-genome sequencing of Nicotiana glauca

BackgroundNicotiana glauca (tree tobacco) is a member of the Solanaceae family, which includes important crops (potato, tomato, eggplant, pepper) and many medicinal plants. This diploid plant is native to South America and is one of the first Nicotiana species with Agrobacterium cellular T-DNA (cT-DNA). Its cT-DNA is a partial, inverted repeat, called gT. Tree tobacco belongs to the section Noctiflorae. Sequencing of the genomes of N. tomentosiformis and N. otophora (section Tomentosae) and N. tabacum (section Nicotiana) allowed the detection of previously unknown multiple cT-DNAs, raising the question whether there are other T-DNA insertions in the N. glauca. NGS data can help answer this question. Besides, N. glauca contains a profile of alkaloids different from N. tabacum. The plant is used for medicinal purposes. Comparative analysis of genomic data of phylogenetically distant tobacco species will provide valuable information on the genetic basis for various traits, especially secondary metabolism.\n\nFindingsWe report a high-depth sequencing and de novo assembly of N. glauca full genome, which was obtained from 210 Gb Illumina HiSeq data. The final draft genome is 3.2 Gb, with N50 size of 31.1 kbp. T-DNA analysis confirmed the presence of the previously described gT insertion and the absence of other ones.\n\nConclusionWe provide the first comprehensive de novo full genome assembly of three tobacco, and a cT-DNA insertion analysis. These genome data could be used in pharmacological and in phylogenetic studies.

genomics

Dot2dot: Accurate Whole-Genome Tandem Repeats Discovery

The advent of sequencing technologies and the consequent computational analysis of genomes has confirmed the evidence that DNA sequences contain a relevant amount of repetitions. A particularly important category of repeating sequences is that of tandem repeats (TRs). TRs are short, almost identical sequences that lie adjacent to each other. The abundance of TRs in eukaryotic genomes has suggested that they play a role in many cellular processes and, indeed, are also involved in the onset and progress of several genetic disorders.\n\nBuilding upon the idea that similar sequences can be easily displayed using graphical methods, we formalized the structure that TRs induce in dot plot matrices where a sequence is compared with itself. We further observed that a compact representation of these matrices can be built and searched in linear time in the size of the input sequence. Exploiting this observation, we developed an algorithm fast enough to be suitable for whole-genome discovery of tandem repeats.\n\nWe compared our algorithm with seven state of the art methods using as a gold standard five collections of tandem repeats: pathology-linked, forensic, for population analysis, genealogic-oriented, and variable TRs in regulatory regions. In addition, we run our algorithm on seven reference genomes to test the suitability of our approach for whole-genome analysis. Experiments show that our method: is always more accurate than the other methods, and completes the analysis of the biggest available reference genome in about one day running at a rate of 0.98Gbp/h on a standard workstation.

genomics

Genome-wide epigenetic profiling of 5-hydroxymethylcytosine by long-read optical mapping

The epigenetic mark 5-hydroxymethylcytosine (5-hmC) is a distinct product of active enzymatic demethylation that is linked to gene regulation, development and disease. Genome-wide 5-hmC profiles generated by short-read next-generation sequencing are limited in providing long-range epigenetic information relevant to highly variable genomic regions, such as the 3.7 Mbp disease-related Human Leukocyte Antigen (HLA) region. We present a long-read, single-molecule mapping technology that generates hybrid genetic/epigenetic profiles of native chromosomal DNA. The genome-wide distribution of 5- hmC in human peripheral blood cells correlates well with 5-hmC DNA immunoprecipitation (hMeDIP) sequencing. However, the long read length of 100 kbp-1Mbp produces 5-hmC profiles across variable genomic regions that failed to showup in the sequencing data. In addition, optical 5-hmC mapping shows strong correlation between the 5-hmC density in gene bodies and the corresponding level of gene expression. The single molecule concept provides information on the distribution and coexistence of 5-hmC signals at multiple genomic loci on the same genomic DNA molecule, revealing long-range correlations and cell-to-cell epigenetic variation.

genomics

Rapid low-cost assembly of the Drosophila melanogaster reference genome using low-coverage, long-read sequencing

Accurate and comprehensive characterization of genetic variation is essential for deciphering the genetic basis of diseases and other phenotypes. A vast amount of genetic variation stems from large-scale sequence changes arising from the duplication, deletion, inversion, and translocation of sequences. In the past 10 years, high-throughput short reads have greatly expanded our ability to assay sequence variation due to single nucleotide polymorphisms. However, a recent de novo assembly of a second Drosophila melanogaster reference genome has revealed that short read genotyping methods miss hundreds of structural variants, including those affecting phenotypes. While genomes assembled using high-coverage long reads can achieve high levels of contiguity and completeness, concerns about cost, errors, and low yield have limited widespread adoption of such sequencing approaches. Here we resequenced the reference strain of D. melanogaster (ISO1) on a single Oxford Nanopore MinION flow cell run for 24 hours. Using only reads longer than 1 kb or with at least 30x coverage, we assembled a highly contiguous de novo genome. The addition of inexpensive paired reads and subsequent scaffolding using an optical map technology achieved an assembly with completeness and contiguity comparable to the D. melanogaster reference assembly. Comparison of our assembly to the reference assembly of ISO1 uncovered a number of structural variants (SVs), including novel LTR transposable element insertions and duplications affecting genes with developmental, behavioral, and metabolic functions. Collectively, these SVs provide a snapshot of the dynamics of genome evolution. Furthermore, our assembly and comparison to the D. melanogaster reference genome demonstrates that high-quality de novo assembly of reference genomes and comprehensive variant discovery using such assemblies are now possible by a single lab for under $1,000 (USD).

genomics

Constructing and visualizing cancer genomic maps in 3D spatial context by phenotype-based high-throughput laser-aided isolation and sequencing (PHLI-seq)

A spatially resolved analysis of the heterogeneous cancer genome, in which the data are connected to the three-dimensional space of a tumour, is crucial to understand cancer biology and the clinical impact of cancer heterogeneity on patients. However, despite recent progress in spatially resolved transcriptomics, spatial mapping of genomic data in a high-throughput and high-resolution manner has been challenging due to current technical limitations. Here, we describe a novel approach, phenotype-based high-throughput laser-aided isolation and sequencing (PHLI-seq), which enables high-throughput isolation of a single-cell or a small number of cells and their genome-wide sequence analysis to construct genomic maps within cancer tissue in relation to the phenotypes of the cells. By applying PHLI-seq, we reveal the heterogeneity of breast cancer tissues at a high resolution and map the genomic landscape of the cells to their corresponding spatial locations and phenotypes in the tumour mass. Additionally, with different staining modalities, the genotypes of the cells can be connected to corresponding phenotypic information of the tissue. Together with the spatially resolved genomic analysis, we can infer the histories of heterogeneous cancer cells in two or three dimensions, providing significant insight into cancer biology and precision medicine.

genomics

Finding Nemo’s Genes: A chromosome-scale reference assembly of the genome of the orange clownfish Amphiprion percula

The iconic orange clownfish, Amphiprion percula, is a model organism for studying the ecology and evolution of reef fishes, including patterns of population connectivity, sex change, social organization, habitat selection and adaptation to climate change. Notably, the orange clownfish is the only reef fish for which a complete larval dispersal kernel has been established and was the first fish species for which it was demonstrated that anti-predator responses of reef fishes could be impaired by ocean acidification. Despite its importance, molecular resources for this species remain scarce and until now it lacked a reference genome assembly. Here we present a de novo chromosome-scale assembly of the genome of the orange clownfish Amphiprion percula. We utilized single-molecule real-time sequencing technology from Pacific Biosciences to produce an initial polished assembly comprised of 1,414 contigs, with a contig N50 length of 1.86 Mb. Using Hi-C based chromatin contact maps, 98% of the genome assembly were placed into 24 chromosomes, resulting in a final assembly of 908.8 Mb in length with contig and scaffold N50s of 3.12 and 38.4 Mb, respectively. This makes it one of the most contiguous and complete fish genome assemblies currently available. The genome was annotated with 26,597 protein coding genes and contains 96% of the core set of conserved actinopterygian orthologs. The availability of this reference genome assembly as a community resource will further strengthen the role of the orange clownfish as a model species for research on the ecology and evolution of reef fishes.

genomics