Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genome-wide and single-base resolution DNA methylomes of the Sea Lamprey (Petromyzon marinus) Reveal Gradual Transition of the Genomic Methylation Pattern in Early Vertebrates

In eukaryotes, cytosine methylation is a primary heritable epigenetic modification of the genome that regulates many cellular processes. While the whole-genome methylation pattern has been generally conserved in different eukaryotic groups, invertebrates and vertebrates exhibit two distinct patterns. Whereas almost all CpG sites are methylated in most vertebrates, with the exception of short unmethylated regions call CpG islands, the most frequent pattern in invertebrate animals is mosaic methylation, comprising domains of heavily methylated DNA interspersed with domains that are methylation free. The mechanism by which the genome methylation pattern transited from a mosaic to a global pattern and the role of the one or two-round whole-genome duplication in this transition remain largely elusive, partly owing to the lack of methylome data from early vertebrates. In this study, we used the whole-genome bisulfite-sequencing technology to investigate the genome-wide methylation in three tissues (heart, muscle, and sperm) from the sea lamprey, an extant Agarthan vertebrate. Analyses of methylation level and the extent of CpG dinucleotide depletion of geneencoding, intergenic and promoter regions revealed a gradual increase in the methylation level from invertebrates to vertebrates, with the sea lamprey exhibiting an intermediate position. In addition, the methylation level of the majority of CpGs was intermediate in each sea lamprey tissue, indicating a high level of heterogeneity of methylation status between individual cells. In this regard, we defined the genomic methylation pattern of sea lamprey as \"global genomic DNA intermediate methylation\". The methylation features in different genomic regions, such as the transcription start site (TSS) region of the gene body, exon-intron boundaries, transposons, as well as genes grouping with different expression levels, supported the gradual methylation transition hypothesis. We further discussed that the copy number difference in DNA methylation transferases and the loss of the PWWP domain and/or DNTase domain in DNMT3 sub-family enzymes may have contributed to the methylation pattern transition in early vertebrates. These findings demonstrate an intermediate genomic methylation pattern between invertebrates and jawed vertebrates, providing evidence that supports the hypothesis that methylation patterns underwent a gradual transition from invertebrates (mosaic) to vertebrates (global).

Evolutionary Biology

dRep: A tool for fast and accurate genome de-replication that enables tracking of microbial genotypes and improved genome recovery from metagenomes

The number of microbial genomes sequenced each year is expanding rapidly, in part due to genome-resolved metagenomic studies that routinely recover hundreds of draft-quality genomes. Rapid algorithms have been developed to comprehensively compare large genome sets, but they are not accurate with draft-quality genomes. Here we present dRep, a program that sequentially applies a fast, inaccurate estimation of genome distance and a slow but accurate measure of average nucleotide identity to reduce the computational time for pair-wise genome set comparisons by orders of magnitude. We demonstrate its use in a study where we separately assembled each metagenome from time series datasets. Groups of essentially identical genomes were identified with dRep, and the best genome from each set was selected. This resulted in recovery of significantly more and higher-quality genomes compared to the set recovered using the typical co-assembly method. Documentation is available at http://drep.readthedocs.io/en/master/ and source code is available at https://github.com/MrOlm/drep.

bioinformatics

seq-seq-pan: Building a computational pan-genome data structure on whole genome alignment

BackgroundThe increasing application of next generation sequencing technologies has led to the availability of thousands of reference genomes, often providing multiple genomes for the same or closely related species. The current approach to represent a species or a population with a single reference sequence and a set of variations cannot represent their full diversity and introduces bias towards the chosen reference. There is a need for the representation of multiple sequences in a composite way that is compatible with existing data sources for annotation and suitable for established sequence analysis methods. At the same time, this representation needs to be easily accessible and extendable to account for the constant change of available genomes.\n\nResultsWe introduce seq-seq-pan, a framework that provides methods for adding or removing new genomes from a set of aligned genomes and uses these to construct a whole genome alignment. Throughout the sequential workflow the alignment is optimized for generating a representative linear presentation of the aligned set of genomes, that enables its usage for annotation and in downstream analyses.\n\nConclusionsBy providing dynamic updates and optimized processing, our approach enables the usage of whole genome alignment in the field of pan-genomics. In addition, the sequential workflow can be used as a fast alternative to existing whole genome aligners. seq-seq-pan is freely available at https://gitlab.com/groups/rki_bioinformatics

bioinformatics

Imaging-Genomics Study Of Head-Neck Squamous Cell Carcinoma: Associations Between Radiomic Phenotypes And Genomic Mechanisms Via Integration Of TCGA And TCIA

PurposeRecent data suggest that imaging radiomics features for a tumor could predict important genomic biomarkers. Understanding the relationship between radiomic and genomic features is important for basic cancer research and future patient care. For Head and Neck Squamous Cell Carcinoma (HNSCC), we perform a comprehensive study to discover the imaging-genomics associations and explore the potential of predicting tumor genomic alternations using radiomic features.\n\nMethodsOur retrospective study integrates whole-genome multi-omics data from The Cancer Genome Atlas (TCGA) with matched computed tomography imaging data from The Cancer Imaging Archive (TCIA) for the same set of 126 HNSCC patients. Linear regression analysis and gene set enrichment analysis are used to identify statistically significant associations between radiomic imaging features and genomic features. Random forest classifier is used to predict two key HNSCC molecular biomarkers, the status of human papilloma virus (HPV) and disruptive TP53 mutation, based on radiomic features.\n\nResultsWide-spread and statistically significant associations are discovered between genomic features (including miRNA expressions, protein expressions, somatic mutations, and transcriptional activities, copy number variations, and promoter region DNA methylation changes of pathways) and radiomic features characterizing the size, shape, and texture of tumor. Prediction of HPV and TP53 mutation status using radiomic features achieves an area under the receiver operating characteristics curve (AUC) of 0.71 and 0.641, respectively.\n\nConclusionOur analysis suggests that radiomic features are associated with genomic characteristics in HNSCC and provides justification for continued development of radiomics as biomarkers for relevant genomic alterations in HNSCC.

cancer biology

Genome-wide features of introns are evolutionary decoupled among themselves and from genome size throughout Eukarya

The impact of spliceosomal introns on genome and organismal evolution remains puzzling. Here, we investigated the correlative associations among genome-wide features of introns from protein-coding genes (e.g., size, density, genome-content, repeats), genome size and multicellular complexity on 461 eukaryotes. Thus, we formally distinguished simple from complex multicellular organisms (CMOs), and developed the program GenomeContent to systematically estimate genomic traits. We performed robust phylogenetic controlled analyses, by taking into account significant uncertainties in the tree of eukaryotes and variation in genome size estimates. We found that changes in the variation of some intron features (such as size and repeat composition) are only weakly, while other features measuring intron abundance (within and across genes) are not, scaling with changes in genome size at the broadest phylogenetic scale. Accordingly, the strength of these associations fluctuates at the lineage-specific level, and changes in the length and abundance of introns within a genome are found to be largely evolving independently throughout Eukarya. Thereby, our findings are in disagreement with previous estimations claiming a concerted evolution between genome size and introns across eukaryotes. We also observe that intron features vary homogeneously (with low repetitive composition) within fungi, plants and stramenophiles; but they vary dramatically (with higher repetitive composition) within holozoans, chlorophytes, alveolates and amoebozoans. We also found that CMOs and their closest ancestral relatives are characterized by high intron-richness, regardless their genome size. These patterns contrast the narrow distribution of exon features found across eukaryotes. Collectively, our findings unveil spliceosomal introns as a dynamically evolving non-coding DNA class and strongly argue against both, a particular intron feature as key determinant of eukaryotic gene architecture, as well as a major mechanism (adaptive or non-adaptive) behind the evolutionary dynamics of introns over a large phylogenetic scale. We hypothesize that intron-richness is a pre-condition to evolve complex multicellularity.

evolutionary biology

A Model for Genome-First Care: Returning Secondary Genomic Findings to Participants and Their Healthcare Providers in a Large Research Cohort

BackgroundResearch cohorts with linked genomic data exist, or are being developed, at many research centers. Within any such \"sequenced cohort\" of more than 100 participants, it is likely that there are participants with previously undisclosed risk for life-threatening monogenic diseases that could be identified with targeted analysis of their existing data. Identification of such disease-associated findings are not usually primary to the enrollment research goals. At Geisinger Health System, MyCode(R) Community Health Initiative (MyCode) participants represent one such large sequenced cohort. Since 2013, MyCode participants in discovery research have been consented for secondary analysis of their existing research genomic sequences to allow delivery of medically actionable findings to them and their healthcare providers. This return of genomic results program was developed to manage an anticipated 3.5% of MyCode participants who will receive clinically confirmed genomic variants from an approved gene list out of more than 150,000 total participants. Risk-associated DNA sequences alone without any clinical parameter, prompt \"genome-first\" follow-up encounters.\n\nMethodsThis article describes our process for generating clinical grade results from research-based genomic sequencing data, delivering results to patients and their providers, facilitating targeted clinical evaluations of patients and promoting cascade testing of at-risk relatives. We also summarize our early data about the results generated during this process and our ability to contact patients and their providers to disclose the information.\n\nResultsThis process has been used to generate 343 results on 339 patients. 93% of patients with a result have been successfully contacted about their results as evidenced by direct interaction about their result with the research team or a healthcare provider. 222 healthcare providers have been notified of a result on one or more patient through this result delivery process.\n\nConclusionsHere we describe the existing GHS model to deliver genomic data into the electronic medical record and the clinical interactions that are prompted and supported. Elements of this genome-first care model can be applied in other healthcare settings and in national efforts, such as \"All of Us\", that wish to establish programs for returning genomic results to research participants.

genomics

Genome Build Information Is An Essential Part Of Genomic Track Files

Genomic locations are represented as coordinates on a specific genome build version, but the build information is frequently missing when coordinates are provided. It is essential to correctly interpret and analyse the genomic intervals contained in genomic track files. Here, we demonstrate that this crucial metadatum (or rather datum) is often isolated from the genomic track files in public repositories and journal articles, which could be a major time thief. We propose best practices to ensure that genome build version is always carried along with genomic track files. Although not a substitute to the best practices, we also provide a tool to predict the genome build version of genomic track files.

bioinformatics

Identification of meiotic recombination through gamete genome reconstruction using whole genome linked-reads

Meiotic recombination (MR), which transmits exchanged genetic materials between homologous chromosomes to offspring, plays a crucial role in shaping genomic diversity in eukaryotic organisms. In humans, thousands of meiotic recombination hotspots have been mapped by population genetics approaches. However, direct identification of MR events for individuals is still challenging due to the difficulty in resolving the haplotypes of homologous chromosomes and reconstructing the gamete genome. Whole genome linked-read sequencing (lrWGS) can generate haplotype sequences of mega-base pairs (N50 ~2.5Mb) after computational phasing. However, the haplotype information is still isolated in a large number of fragmented genomic regions and limited by switch errors, impeding its further application in the chromosome-scale analysis. In this study, we developed a tool MRLR (Meiotic Recombination identification by Linked-Read sequencing) for the analysis of individual MR events. By leveraging trio pedigree information with lrWGS haplotypes, our pipeline is sufficient to reconstruct the whole human gamete genome with 99.8% haplotyping accuracy. By analyzing the haplotype exchange between homologous chromosomes, MRLR identified 462 high-resolution MR events in 6 human trio samples from the Genome In A Bottle (GIAB) and the Human Genome Structural Variation Consortium (HGSVC). In three datasets of the HGSVC, our results recapitulated 149 (92%) previously identified high-confident MR events and discovered 85 novel events. About half (40) of the new events are supported by single-cell template strand sequencing (Strand-seq) results. We found that 332 (71.9%) MR events co-localize with recombination hotspots (>10 cM/Mb) in human populations, and MR breakpoint regions are enriched in PRDM9 and DMC1 binding sites. In addition, 48% (221) breakpoint regions were detected inside a gene, indicating these MRs can directly affect the haplotype diversity of genic regions. Taken together, our approach provides new opportunities in the haplotype-based genomic analysis of individual meiotic recombination. The MRLR software is implemented in Perl and is freely available at https://github.com/ChongLab/MRLR.

genomics

Ethnically relevant consensus Korean reference genome towards personal reference genomes

Human genomes are routinely compared against a universal reference. However, this strategy could miss population-specific or personal genomic variations, which may be detected more efficiently using an ethnically-relevant and/or a personal reference. Here we report a hybrid assembly of Korean reference (KOREF) as a pilot case for constructing personal and ethnic references by combining sequencing and mapping methods. KOREF is also the first consensus variome reference, providing information on millions of variants from additional ethnically homogeneous personal genomes. We found that this ethnically-relevant consensus reference was beneficial for efficiently detecting variants. Systematic comparison of KOREF with previously established human assemblies showed the importance of assembly quality, suggesting the necessity of using new technologies to comprehensively map ethnic and personal genomic structure variations. In the era of large-scale population genome projects, the leveraging of ethnicity-specific genome assemblies as well as the human reference genome will accelerate mapping all human genome diversity.

Genomics

An improved assembly and annotation of the allohexaploid wheat genome identifies complete families of agronomic genes and provides genomic evidence for chromosomal translocations.

Advances in genome sequencing and assembly technologies are generating many high quality genome sequences, but assemblies of large, repeat-rich polyploid genomes, such as that of bread wheat, remain fragmented and incomplete. We have generated a new wheat whole-genome shotgun sequence assembly using a combination of optimised data types and an assembly algorithm designed to deal with large and complex genomes. The new assembly represents more than 78% of the genome with a scaffold N50 of 88.8kbp that has a high fidelity to the input data. Our new annotation combines strand-specific Illumina RNAseq and PacBio full-length cDNAs to identify 104,091 high confidence protein-coding genes and 10,156 non-coding RNA genes. We confirmed three known and identified one novel genome rearrangements. Our approach enables the rapid and scalable assembly of wheat genomes, the identification of structural variants, and the definition of complete gene models, all powerful resources for trait analysis and breeding of this key global crop. [Supplemental material is available for this article.]

genomics

Heterogeneity Among Estimates Of The Core Genome And Pan-Genome In Different Pneumococcal Populations

BackgroundUnderstanding the structure of a bacterial population is essential in order to understand bacterial evolution, or which genetic lineages cause disease, or the consequences of perturbations to the bacterial population. Estimating the core genome, the genes common to all or nearly all strains of a species, is an essential component of such analyses. The size and composition of the core genome varies by dataset, but our hypothesis was that variation between different collections of the same bacterial species should be minimal. To test this, the genome sequences of 3,121 pneumococci recovered from healthy individuals in Reykjavik (Iceland), Southampton (United Kingdom), Boston (USA) and Maela (Thailand) were analysed.\n\nResultsThe analyses revealed a supercore genome (genes shared by all 3,121 pneumococci) of only 303 genes, although 461 additional core genes were shared by pneumococci from Reykjavik, Southampton and Boston. Overall, the size and composition of the core genomes and pan-genomes among pneumococci recovered in Reykjavik, Southampton and Boston were very similar, but pneumococci from Maela were distinctly different. Inspection of the pan-genome of Maela pneumococci revealed several >25 Kb sequence regions that were homologous to genomic regions found in other bacterial species.\n\nConclusionsSome subsets of the global pneumococcal population are highly heterogeneous and thus our hypothesis was rejected. This is an essential point of consideration before generalising the findings from a single dataset to the wider pneumococcal population.

genomics

Entering the era of conservation genomics: Cost-effective assembly of the African wild dog genome using linked long reads

A high-quality reference genome assembly is a valuable tool for the study of non- model organisms across disciplines. Genomic techniques can provide important insights about past population sizes, local adaptation, and even aid in the development of breeding management plans. This information can be particularly important for fields like conservation genetics, where endangered species require critical and immediate attention. However, funding for genomic-based methods can be sparse for conservation projects, as costs for general species management can consume budgets. Here we report the generation of high-quality reference genomes for the African wild dog (Lycaon pictus) at a low cost, thereby facilitating future studies of this endangered canid. We generated assemblies for three individuals from whole blood samples using the linked-read 10x Genomics Chromium system. The most continuous assembly had a scaffold N50 of 21 Mb, a contig N50 of 83 Kb, and completely reconstructed 95% of conserved mammalian genes as reported by BUSCO v2, indicating a high assembly quality. Thus, we show that 10x Genomics Chromium data can be used to effectively generate high-quality genomes of mammal species from Illumina short-read data of intermediate coverage ([~]25-50x). Interestingly, the African wild dog shows a much higher heterozygosity than other species of conservation concern, possibly as a result of its behavioral ecology. The availability of reference genomes for non-model organisms will facilitate better genetic monitoring of threatened species such as the African wild dog. At the same time, they can help researchers and conservationists to better understand the ecology and adaptability of those species in a changing environment.

genomics

Fine-scale characterization of genomic structural variation in the human genome reveals adaptive and biomedically relevant hotspots

Genomic structural variants (SVs) are distributed nonrandomly across the human genome. These \"hotspots\" have been implicated in critical evolutionary innovations, as well as serious medical conditions. However, the evolutionary and biomedical features of these hotspots remain incompletely understood. In this study, we analyzed data from 2,504 genomes from the 1000 Genomes Project Consortium and constructed a refined map of 1,148 SV hotspots in human genomes. By studying the genomic architecture of these hotspots, we found that both nonallelic homologous recombination and non-homologous mechanisms act as mechanistic drivers of SV formation. We found that the majority of SV hotspots are within gene-poor regions and evolve under relaxed negative selection or neutrality. However, we found that a small subset of SV hotspots harbor genes that are enriched for anthropologically crucial functions, including blood oxygen transport, olfaction, synapse assembly, and antigen binding. We provide evidence that balancing selection may have maintained these SV hotspots, which include two independent hotspots on different chromosomes affecting alpha and beta hemoglobin gene clusters. Biomedically, we found that the SV hotspots coincide with breakpoints of clinically relevant, large de novo SVs, significantly more often than genome-wide expectations. As an example, we showed that the breakpoints of multiple large de novo SVs, which lead to idiopathic short stature, coincide with SV hotspots. As such, the mutational instability in SV hotpots likely enables chromosomal breaks that lead to pathogenic structural variation formations. Our study contributes to a better understanding of the mutational landscape of the genome and implicates both mechanistic and adaptive forces in the formation and maintenance of SV hotspots.

genomics

High coverage genome sequencing and identification of genomic variants in Bengal tiger (Panthera tigris tigris)

Bengal tiger (Panthera tigris tigris), one of six extant tiger subspecies, occurs solely in the Indian subcontinent. Although endangered and threatened by various extinction risks, this is the most populous tiger subspecies with the highest genetic diversity and strongest chance of survival in the wild. Availability of high quality genomic information on this animal will help us understand its ability to adapt to different habitats and environmental changes, in addition to comparative studies with other subspecies. Here we report high coverage sequencing of the Bengal tiger genome and its mapping to the Amur tiger genome in order to discover single nucleotide to large structural variants. A total of 345 Gb, roughly equivalent to 144X coverage of the genome, was generated from 1,149,381,669 raw read pairs. Further, 990,060,729 clean read pairs, again equivalent to 115X coverage, were retained from the raw read data and considered for comparative analysis with the Amur tiger genome. This alignment showed that 97.35% of the bases mapped at 5X depth, 97.26% at 10X and 90.44% at 50X depth. We identified a total of 3,601,882 single nucleotide variants, 948 structural variants, 56,649 copy number variants and 1,760,347 simple sequence repeats. We report the first high coverage genome sequence of Bengal tiger with an overview of its genomic variants when compared to the Amur tiger genome. Of the several variants identified, we further have to assess and validate variants potentially associated with the ability of the animal to adapt to environmental changes, disease susceptibility and other important biological phenomena.

genomics

EuMicrobedbLite: A lightweight genomic resource and analytic platform for draft oomycete genomes

We have developed EuMicrobedbLite - A light weight comprehensive genome resource and sequence analysis platform for oomycete organisms. EuMicrobedbLite is a successor of the VBI Microbial Database (VMD) that was built using the Genome Unified Schema (GUS). In this version, the GUS schema has been greatly simplified with removal of many obsolete modules and redesign of others to incorporate contemporary data. Several dependencies such as perl object layers used for data loading in VMD have been replaced with independent light weight scripts. EumicrobedbLite now runs on a powerful annotation engine developed at our lab called \"Genome Annotator Lite\". Currently this database has 26 publicly available genomes and 10 EST datasets of oomycete organisms. The browser page has dynamic tracks presenting comparative genomics analyses, coding and non-coding data, tRNA genes, repeats and EST alignments. In addition, we have defined 44,777 core conserved proteins from twelve oomycete organisms that form 2974 clusters. Synteny viewing is enabled by incorporation of the Genome Synteny Viewer (GSV) tool. The user interface has undergone major changes for ease of browsing. Queryable comparative genomics information, conserved orthologous genes and pathways are among the new key features updated in this database. The browser has been upgraded to enable user upload of GFF files for quick view of genome annotation comparisons. The toolkit page integrates the EMBOSS package and has a gene prediction tool. Annotations for the organisms are updated once every six months to ensure quality. The database resource is available at www.eumicrobedb.org.

bioinformatics

Comparative Genomics Of Apomictic Root-Knot Nematodes: Hybridization, Ploidy, And Dynamic Genome Change

The Root-Knot Nematodes (RKN; genus Meloidogyne) are important plant parasites causing substantial agricultural losses. The Meloidogyne incognita group (MIG) of species, most of which are obligatory apomicts (mitotic parthenogens), are extremely polyphagous and important problems for global agriculture. While understanding the genomic basis for their variable success on different crops could benefit future agriculture, analyses of their genomes pose challenges due to complex evolutionary histories that may incorporate hybridization, ploidy changes, and chromosomal fragmentation. Here we sequence 19 genomes, representing five species of key RKN collected from different geographic origins. We show that a hybrid origin that predated speciation within the MIG has resulted in each species possessing two divergent genomic copies. Additionally, the MIG apomicts are hypotriploids, with a proportion of one genome present in a second copy, and this proportion varies among species. The evolutionary history of the MIG genomes is revealed to be very dynamic, with non-crossover recombination both homogenising the genomic copies, and acting as a mechanism for generating divergence between species. Interestingly, the automictic MIG species M. floridensis differs from the apomict species in that it has become homozygous throughout much of its genome.

evolutionary biology

Genome-Wide Comparison Of Toxigenic And Non-Toxigenic Corynebacterium diphtheriae Isolates Identifies Differences In The Pan Genomes Between Respiratory And Cutaneous Strains

ObjectivesCorynebacterium diphtheriae is the main etiological agent of diphtheria, a global disease causing life-threatening infections, particularly in infants and children. Vaccination with diphtheria toxoid protects against infection with potent toxin producing strains. However a growing number of apparently non-toxigenic but potentially invasive C. diphtheriae strains are identified in countries with low prevalence of diphtheria, raising key questions about genomic structures and population dynamics of the species.\n\nMethodsThis study examined genomic diversity among 47 C. diphtheriae isolates collected in Australia over a 10-year period using whole genome sequencing. Phylogeny was determined using SNP-based mapping and genome wide analysis.\n\nResultsC. diphtheriae sequence type (ST) 32, a non-toxigenic ST with evidence of enhanced virulence that is also circulating in Europe, appears to be endemic in Australia. Isolates from temporospatially related patients displayed the same ST and similarity in their core genomes. The genome-wide analysis highlighted a role of pilins, adhesion factors and iron utilization in infections caused by toxigenic as well as non-toxigenic strains.\n\nConclusionsThe genomic diversity of toxigenic and non-toxigenic strains of C. diphtheriae in Australia suggests multiple local and overseas sources of infection and colonisation. Our findings suggest that regular genomic surveillance of co-circulating toxigenic and non-toxigenic C. diphtheriae can deliver highly nuanced data in order to inform targeted public health actions and policy for predicting the future impact of this highly successful pathogen.

microbiology

A European whitefish linkage map and its implications for understanding genome-wide synteny between salmonids following whole genome duplication

Genomic datasets continue to increase in size and ease of production for a wider selection of species including non-model organisms. For many of these species highly contiguous and well-annotated genomes are unavailable due to their prohibitive complexity and cost. As a result, a common starting point for genomic work in non-model species is the production of a linkage map, which involves the grouping and relative ordering of genetic markers along the genome. Dense linkage maps facilitate the analysis of genomic data in a variety of ways, from broad scale observations regarding genome structure e.g. chromosome number and type or sex-related structural differences, to fine scale patterns e.g. recombination rate variation and co-localisation of differentiated regions. Here we present both a sex-averaged and sex-specific linkage maps for Coregonus sp. \"Albock\" containing 5395 single nucleotide polymorphism (SNP) loci across 40 linkage groups to facilitate future investigation into the genomic basis of whitefish adaptation and speciation. The map was produced using restriction-site associated digestion (RAD) sequencing data from two wild-caught parents and 156 F1 offspring in Lep-MAP3. We discuss the differences between our sex-avagerated and sex-specific maps and identify synteny between C. sp. \"Albock\" linkage groups and the Atlantic salmon (Salmo salar) genome. Our synteny analysis confirms that many patterns of homology observed between Atlantic salmon and Oncorhynchus and Salvelinus species are also shared by members of the Coregoninae subfamily.

evolutionary biology