Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Molecular evolutionary trends and feeding ecology diversification in the Hemiptera, anchored by the milkweed bug genome

BackgroundThe Hemiptera (aphids, cicadas, and true bugs) are a key insect order, with high diversity for feeding ecology and excellent experimental tractability for molecular genetics. Building upon recent sequencing of hemipteran pests such as phloem-feeding aphids and blood-feeding bed bugs, we present the genome sequence and comparative analyses centered on the milkweed bug Oncopeltus fasciatus, a seed feeder of the family Lygaeidae.\n\nResultsThe 926-Mb Oncopeltus genome is well represented by the current assembly and official gene set. We use our genomic and RNA-seq data not only to characterize the protein-coding gene repertoire and perform isoform-specific RNAi, but also to elucidate patterns of molecular evolution and physiology. We find ongoing, lineage-specific expansion and diversification of repressive C2H2 zinc finger proteins. The discovery of intron gain and turnover specific to the Hemiptera also prompted evaluation of lineage and genome size as predictors of gene structure evolution. Furthermore, we identify enzymatic gains and losses that correlate with feeding biology, particularly for reductions associated with derived, fluid-nutrition feeding.\n\nConclusionsWith the milkweed bug, we now have a critical mass of sequenced species for a hemimetabolous insect order and close outgroup to the Holometabola, substantially improving the diversity of insect genomics. We thereby define commonalities among the Hemiptera and delve into how hemipteran genomes reflect distinct feeding ecologies. Given Oncopeltus's strength as an experimental model, these new sequence resources bolster the foundation for molecular research and highlight technical considerations for the analysis of medium-sized invertebrate genomes.

genomics

Comparative analysis examining patterns of genomic differentiation across multiple episodes of population divergence in birds

Heterogeneous patterns of genomic differentiation are commonly documented between closely related populations and there is considerable interest in identifying factors that contribute to their formation. These factors could include genomic features (e.g., areas of low recombination) that promote processes like linked selection (positive or purifying selection that affects linked neutral sites) at specific genomic regions. Examinations of repeatable patterns of differentiation across population pairs can provide insight into the role of these factors. Birds are well suited for this work, as genome structure is conserved across this group. Accordingly, we re-estimated relative (FST) and absolute (dXY) differentiation between eight sister pairs of birds that span a broad taxonomic range using a common pipeline. Across pairs, there were modest but significant correlations in window-based estimates of differentiation (up to 3% of variation explained for FST and 26% for dXY), supporting a role for processes at conserved genomic features in generating heterogeneous patterns of differentiation. This suggestion was reinforced by linear models identifying several genomic features (e.g., gene densities) as significant predictors of FST and dXY repeatability. FST repeatability was higher among pairs that were further along the speciation continuum (i.e., more reproductively isolated), suggesting that early stages of speciation may be dominated by positive selection that is different between pairs and replaced by processes acting according to shared genomic features as speciation proceeds.

genomics

Amplification-free, CRISPR-Cas9 Targeted Enrichment and SMRT Sequencing of Repeat-Expansion Disease Causative Genomic Regions

Targeted sequencing has proven to be an economical means of obtaining sequence information for one or more defined regions of a larger genome. However, most target enrichment methods require amplification. Some genomic regions, such as those with extreme GC content and repetitive sequences, are recalcitrant to faithful amplification. Yet, many human genetic disorders are caused by repeat expansions, including difficult to sequence tandem repeats.\n\nWe have developed a novel, amplification-free enrichment technique that employs the CRISPR-Cas9 system for specific targeting multiple genomic loci. This method, in conjunction with long reads generated through Single Molecule, Real-Time (SMRT) sequencing and unbiased coverage, enables enrichment and sequencing of complex genomic regions that cannot be investigated with other technologies. Using human genomic DNA samples, we demonstrate successful targeting of causative loci for Huntingtons disease (HTT; CAG repeat), Fragile X syndrome (FMR1; CGG repeat), amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (C9orf72; GGGGCC repeat), and spinocerebellar ataxia type 10 (SCA10) (ATXN10; variable ATTCT repeat). The method, amenable to multiplexing across multiple genomic loci, uses an amplification-free approach that facilitates the isolation of hundreds of individual on-target molecules in a single SMRT Cell and accurate sequencing through long repeat stretches, regardless of extreme GC percent or sequence complexity content. Our novel targeted sequencing method opens new doors to genomic analyses independent of PCR amplification that will facilitate the study of repeat expansion disorders.

genomics

Ultraconserved elements occupy specific arenas of three-dimensional mammalian genome organization

This study explores the relationships between three-dimensional genome organization and the ultraconserved elements (UCEs), an enigmatic set of DNA elements that show very high DNA sequence conservation between vertebrate reference genomes. Examining both human and mouse genomes, we interrogate the relationship of UCEs to three features of chromosome organization derived from Hi-C studies. Firstly, we report that UCEs are enriched within contact domains and, further, that the UCEs that fall into domains shared across diverse cell types are linked to kidney-related and neuronal processes. In boundaries, UCEs are generally depleted, with those that do overlap boundaries being overrepresented in exonic UCEs. Regarding loop anchors, UCEs are neither over- nor under-represented, with those present in loop anchors being enriched for splice sites compared to all UCEs. Finally, as all of the relationships we observed between UCEs and genomic features are conserved in the mouse genome, our findings suggest that UCEs contribute to interspecies conservation of genome organization and, thus, genome stability.

genomics

Multiple laboratory mouse reference genomes define strain specific haplotypes and novel functional loci

The most commonly employed mammalian model organism is the laboratory mouse. A wide variety of genetically diverse inbred mouse strains, representing distinct physiological states, disease susceptibilities, and biological mechanisms have been developed over the last century. We report full length draft de novo genome assemblies for 16 of the most widely used inbred strains and reveal for the first time extensive strain-specific haplotype variation. We identify and characterise 2,567 regions on the current Genome Reference Consortium mouse reference genome exhibiting the greatest sequence diversity between strains. These regions are enriched for genes involved in defence and immunity, and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. Several immune related loci, some in previously identified QTLs for disease response have novel haplotypes not present in the reference that may explain the phenotype. We used these genomes to improve the mouse reference genome resulting in the completion of 10 new gene structures, and 62 new coding loci were added to the reference genome annotation. Notably this high quality collection of genomes revealed a previously unannotated gene (Efcab3-like) encoding 5,874 amino acids, one of the largest known in the rodent lineage. Interestingly, Efcab3-like-/- mice exhibit severe size anomalies in four regions of the brain suggesting a mechanism of Efcab3-like regulating brain development.

genomics

Genome-wide Meta-analysis of 158,000 Individuals of European Ancestry Identifies Three Loci Associated with Chronic Back Pain

OBJECTIVESTo conduct a genome-wide association study (GWAS) meta-analysis of chronic back pain (CBP).\n\nMETHODSAdults of European ancestry were included from 16 cohorts in Europe and North America. CBP cases were defined as those reporting back pain present for >3-6 months; non-cases were included as comparisons (\"controls\"). Each cohort conducted genotyping using commercially available arrays followed by imputation. GWAS used logistic regression models with additive genetic effects, adjusting for age, sex, study-specific covariates, and population substructure. The threshold for genome-wide significance in the fixed-effect inverse-variance weighted meta-analysis was p<5x10-8. Suggestive (p<5x10-7) and genome-wide significant (p<5x10-8) variants were carried forward for replication or further investigation in an independent sample.\n\nRESULTSThe discovery sample was comprised of 158,025 individuals, including 29,531 CBP cases. A genome-wide significant association was found for the intronic variant rs12310519 in SOX5 (OR 1.08, p=7.2x10-10). This was subsequently replicated in an independent sample of 283,752 subjects, including 50,915 cases (OR 1.06, p=5.3x10-11), and exceeded genome-wide significance in joint meta-analysis (0R=1.07, p=4.5x10-19). We found suggestive associations at three other loci in the discovery sample, two of which exceeded genome-wide significance in joint meta-analysis: an intergenic variant, rs7833174, located between CCDC26 and GSDMC (OR 1.05, p=4.4x10-13), and an intronic variant, rs4384683, in DCC (OR 0.97, p=2.4x10-10).\n\nDISCUSSIONIn this first reported meta-analysis of GWAS for CBP, we identified and replicated a genetic locus associated with CBP (SOX5). We also identified 2 other loci that reached genome-wide significance in a 2-stage joint meta-analysis (CCDC26/GSDMC and DCC).

genomics

An improved mitochondrial reference genome for Arabidopsis thaliana Col-0

Arabidopsis thaliana remains the foremost model system for plant genetics and genomics, and researchers rely on the accuracy of its genomic resources. The first completely sequenced angiosperm mitochondrial genome was obtained from A. thaliana C24 (Unseld et al., 1997), and more recent efforts have produced additional A. thaliana reference genomes, including one for Col-0, the most widely used ecotype (Davila et al., 2011). These studies were based on older DNA sequencing methods, making them subject to errors associated with lower levels of sequencing coverage or the extremely short read lengths produced by early-generation Illumina technologies. Indeed, although the more recently published A. thaliana mitochondrial reference genome sequences made substantial progress in improving upon earlier versions, they still have high error rates. By comparing publicly available Illumina sequence data to the A. thaliana Col-0 reference genome, we found that it contains a sequence error every 2.4 kb on average, including 57 SNPs, 96 indels (up to 901 bp in size), and a large repeat-mediated rearrangement. Most of these errors appear to have been carried over from the original A. thaliana mitochondrial genome sequence by reference-based assembly approaches, which has misled subsequent studies of plant mitochondrial mutation and molecular evolution by giving the false impression that the errors are naturally occurring variants present in multiple ecotypes. Building on the progress made by previous researchers, we provide a corrected reference sequence that we hope will serve as a useful community resource for future investigations in the field of plant mitochondrial genetics.

genomics

Culture-free generation of microbial genomes from human and marine microbiomes

Our understanding of natural microbial communities is shaped by the careful investigation of a relatively small number of isolated and cultured organisms, and by analysis of genomic sequences obtained by culture-free metagenomic sequencing approaches. Metagenomic shotgun sequencing has facilitated partial reconstruction of strain-level community structure and functional repertoire. Unfortunately, it remains difficult to cost-effectively produce high quality genome drafts for individual microbes without isolation and culture. Recent molecular techniques that partition long DNA fragments and then barcode short fragments derived from them produce \"read clouds\", which are short-read sequences containing long-range information. Here, we present a novel application of a read cloud technique to microbiome samples, as well as Athena, a de novo assembler that uses these barcodes to produce improved metagenomic assemblies. We apply our approach to sequence human stool samples from two healthy individuals, and compare it to existing short read and synthetic long read metagenomic sequencing approaches. We find that read cloud metagenomic sequencing and Athena assembly produce the most complete individual genome drafts. These genome drafts are also highly contiguous (>200kb N50, <10 contigs), even for bacteria that have relatively low (20x) raw short read sequence coverage. We also apply this approach to a significantly more complex marine sediment sample and obtain 23 genome drafts with valuable 16S ribosomal RNA taxonomic marker sequences, nine of which are complete genome drafts. Read cloud metagenomic sequencing allows culture-free generation of high quality microbial genome drafts using only a single shotgun experiment.

genomics

Genomic Exploration of Within-Host Microevolution Reveals a Distinctive Molecular Signature of Persistent Staphylococcus aureus Bacteraemia

BackgroundLarge-scale genomic studies of within-host evolution during Staphylococcus aureus bacteraemia (SAB) are needed to understanding bacterial adaptation underlying persistence and thus refining the role of genomics in management of SAB. However, available comparative genomic studies of sequential SAB isolates have tended to focus on selected cases of unusually prolonged bacteraemia, where secondary antimicrobial resistance has developed. To understand the bacterial genomic evolution during SAB more broadly, we applied whole genome sequencing to a large collection of sequential isolates obtained from patients with persistent or relapsing bacteraemia.\n\nResultsWe show that, while adapation pathways are heterogenous and episode-specific, isolates from persistent bacteraemia have a distinctive molecular signature, characterised by a low mutation frequency and high proportion of non-silent mutations. By performing an extensive analysis of structural genomic variants in addition to point mutations, we found that these often overlooked genetic events are commonly acquired during SAB. We discovered that IS256 insertion may represent the most effective driver of within-host microevolution in selected lineages, with up to three new insertion events per isolate even in the absence of other mutations. Genetic mechanisms resulting in significant phenotypic changes, such as increases in vancomycin resistance, development of small colony phenotypes, and decreases in cytotoxicity, included mutations in key genes (rpoB, stp, agrA) and an IS256 insertion upstream of the walKR operon.\n\nConclusionsThis study provides for the first time a large-scale analysis of within-host evolution during invasive S. aureus infection and describes specific patterns of adaptation that will be informative for both understanding S. aureus pathoadaptation and utilising genomics for management of complicated S. aureus infections.

genomics

A Freeloader?: The Highly Eroded Yet Big-Genomed Serratia symbiotica symbiont of Cinara strobi

Genome reduction is pervasive among maternally-inherited bacterial endosymbionts. This genome reduction can eventually lead to serious deterioration of essential metabolic pathways, thus rendering an obligate endosymbiont unable to provide essential nutrients to its host. This loss of essential pathways can lead to either symbiont complementation (sharing of the nutrient production with a novel co-obligate symbiont) or symbiont replacement (complete takeover of nutrient production by the novel symbiont). However, the process by which these two evolutionary events happen remains somewhat enigmatic by the lack of examples of intermediate stages of this process. Cinara aphids (Hemiptera: Aphididae) typically harbour two obligate bacterial symbionts: Buchnera and Serratia symbiotica. However, the latter has been replaced by different bacterial taxa in specific lineages, and thus species within this aphid lineage could provide important clues into the process of symbiont replacement. In the present study, using 16S rRNA high-throughput amplicon sequencing, we determined that the aphid Cinara strobi harbours not two, but three fixed bacterial symbionts: Buchnera aphidicola, a Sodalis sp., and S. symbiotica. Through genome assembly and genome-based metabolic inference, we have found that only the first two symbionts (Buchnera and Sodalis) actually contribute to the hosts supply of essential nutrients while S. symbiotica has become unable to contribute towards this task. We found that S. symbiotica has a rather large and highly eroded genome which codes only for a few proteins and displays extensive pseudogenisation. Thus, we propose an ongoing symbiont replacement within C. strobi, in which a once competent\" S. symbiotica does no longer contribute towards the beneficial association. These results suggest that in dual symbiotic systems, when a substitute co-symbiont is available, genome deterioration can precede genome reduction and a symbiont can be maintained despite the apparent lack of benefit to its host.

genomics

The Global State of Genome Editing

Genome editing technologies hold great promise in fundamental biomedical research, development of treatments for animal and plant diseases, and engineering biological organisms for food and industrial applications. Therefore, a global understanding of the growth of the field is needed to identify challenges, opportunities and biases that could shape the impact of the technology. To address this, this work applies automated literature mining of scientific publications on genome editing in the past year to infer research trends in 2 key genome editing technologies-CRISPR/Cas systems and TALENs. The study finds that genome editing research is disproportionately distributed between and within countries, with researchers in the US and China accounting for 50% of authors in the field whereas countries across Africa are underrepresented. Furthermore, genome editing research is also disproportionately being explored on diseases such as cancer, Duchene Muscular Dystrophy, sickle cell disease and malaria. Gender biases are also evident in genome editing research with considerably fewer women as principal investigators. The results of this study suggest that automated mining of scientific literature could help identify biases in genome editing research as a means to mitigate future inequalities and tap the full potential of the technology.

genomics

Feasibility of constructing multi-dimensional genomic maps of juvenile idiopathic arthritis

BackgroundJuvenile idiopathic arthritis (JIA) is one of the most common chronic conditions of childhood. Like many common chronic human illnesses, JIA likely involves complex interactions between genes and the environment, mediated by the epigenome. Such interactions are best understood through multi-dimensional genomic maps that identify critical genetic and epigenetic components of the disease. However, constructing such maps in a cost-effective way is challenging, and this challenge is further complicated by the challenge of obtaining biospecimens from pediatric patients at time of disease diagnosis, prior to therapy, as well as the limited quantity of biospecimen that can be obtained from children,particularly those who are unwell. In this paper, we demonstrate the feasibility and utility of creating multi-dimensional genomic maps for JIA from limited sample numbers.\n\nMethodsTo accomplish our aims, we used an approach similar to that used in the ENCODE and Roadmap Epigenomics projects, which used only 2 replicates for each component of the genomic maps. We used genome-wide DNA methylation sequencing, whole genome sequencing on the Illumina 10x platform, RNA sequencing, and chromatin immunoprecipitation-sequencing for informative histone marks (H3K4me1 and H3K27ac) to construct a multi-dimensional map of JIA neutrophils, a cell we have shown to be important in the pathobiology of JIA.\n\nResultsThe epigenomes of JIA neutrophils display numerous differences from those from healthy children. DNA methylation changes, however, had only a weak effect on differential gene expression. In contrast, H3K4me1 and H3K27ac, commonly associated with enhancer functions, strongly correlated with gene expression. Furthermore, although unique/novel enhancer marks were associated with insertion-deletion events (indels) identified on whole genome sequencing, we saw no strong association between epigenetic changes and underlying genetic variation. The initiation of treatment in JIA is associated with a re-ordering of both DNA methylation and histone modifications, demonstrating the plasticity of the epigenome in this setting.\n\nConclusionsThese findings, generated from a small number of patient samples, demonstrate how multidimensional genomic studies may yield new understandings of biology of JIA and provide insight into how therapy alters gene expression patterns.

genomics

The draft genome sequence of mandrill (Mandrillus sphinx)

BackgroundMandrill (Mandrillus sphinx) is a primate species which belong to Old World monkey (Cercopithecidae) family. It is closely related to human, serving as model for some human diseases researches. However, genetic researches and genomic resources of mandrill were limited, especially comparing to other primate species.\n\nFindingsHere we sequenced 284 Gb data, providing 96-fold coverage (considering the estimate genome size of 2.9 Gb), to construct a reference genome for mandrill. The assembled draft genome was 2.79 Gb with contig N50 of 20.48 Kb and scaffold N50 of 3.56 Mb. We annotated the mandrill genome to find 43.83% repeat elements, as well as 21,906 protein coding genes. We found good quality of the draft genome and gene annotation by BUSCO analysis which revealed 98% coverage of the BUSCOs.\n\nConclusionsWe established the first draft genome sequence of mandrill, which is valuable resource for future evolutionary and human diseases studies.

genomics

CRISPR-bind: a simple, custom CRISPR/dCas9-mediated labeling of genomic DNA for mapping in nanochannel arrays

Bionano genome mapping is a robust optical mapping technology used for de novo construction of whole genomes using ultra-long DNA molecules, able to efficiently interrogate genomic structural variation. It is also used for functional analysis such as epigenetic analysis and DNA replication mapping and kinetics. Genomic labeling for genome mapping is currently specified by a single strand nicking restriction enzyme followed by fluorophore incorporation by nick-translation (NLRS), or by a direct label and stain (DLS) chemistry which conjugates a fluorophore directly to an enzyme-defined recognition site. Although these methods are efficient and produce high quality whole genome mapping data, they are limited by the number of available enzymes--and thus the number of recognition sequences--to choose from. The ability to label other sequences can provide higher definition in the data and may be used for countless additional applications. Previously, custom labeling was accomplished via the nick-translation approach using CRISPR-Cas9, leveraging Cas9 mutant D10A which has one of its cleavage sites deactivated, thus effectively converting the CRISPR-Cas9 complex into a nickase with customizable target sequences. Here we have improved upon this approach by using dCas9, a nuclease-deficient double knockout Cas9 with no cutting activity, to directly label DNA with a fluorescent CRISPR-dCas9 complex (CRISPR-bind). Unlike labeling with CRISPR-Cas9 D10A nickase, in which nicking, labeling, and repair by ligation, all occur as separate steps, the new assay has the advantage of labeling DNA in one step, since the CRISPR-dCas9 complex itself is fluorescent and remains bound during imaging. CRISPR-bind can be added directly to a sample that has already been labeled using DLS or NLRS, thus overlaying additional information onto the same molecules. Using the dCas9 protein assembled with custom target crRNA and fluorescently labeled tracrRNA, we demonstrate rapid labeling of repetitive DUF1220 elements. We also combine NLRS-based whole genome mapping with CRISPR-bind labeling targeting Alu loci. This rapid, convenient, non-damaging, and cost-effective technology is a valuable tool for custom labeling of any CRISPR-Cas9 amenable target sequence.

genomics

CRISPR/Cas9 Targeted Capture Of Mammalian Genomic Regions For Characterization By NGS

The robust detection of structural variants in mammalian genomes remains a challenge. It is particularly difficult in the case of genetically unstable Chinese hamster ovary (CHO) cell lines with only draft genome assemblies available. We explore the potential of the CRISPR/Cas9 system for the targeted capture of genomic loci containing integrated vectors in CHO-K1-based cell lines, and compare it to popular target-enrichment methods and to whole genome sequencing (WGS). The CRISPR/Cas9-based techniques allow for amplification-free capture of genomic regions, which reduces the possibility of sequencing artifacts. Other advantages of these methods are the ease of bioinformatics analysis, potential for multiplexing, and the production of longer sequencing templates for real-time sequencing. The utility of these protocols has been proven by identification of transgene integration sites and flanking sequences in a number of CHO cell lines. However, data produced by these and other targeted capture methods are not always sufficient to analyze complex genomic rearrangements (CGRs) or unexpected sequences introduced into genome by vector integration events. In contrast, WGS provides complete information about vector integration sites, vector copy number, CGRs, and foreign DNA-but despite these benefits, WGS is not easily implemented due to the cost and complexity of the analysis.

genomics

Fast and general-purpose linear mixed models for genome-wide genetics

Linear mixed effect models are powerful tools used to account for population structure in genome-wide association studies (GWASs) and estimate the genetic architecture of complex traits. However, fully-specified models are computationally demanding and common simplifications often lead to reduced power or biased inference. We describe Grid-LMM (https://github.com/deruncie/GridLMM), an extendable algorithm for repeatedly fitting complex linear models that account for multiple sources of heterogeneity, such as additive and non-additive genetic variance, spatial heterogeneity, and genotype-environment interactions. Grid-LMM can compute approximate (yet highly accurate) frequentist test statistics or Bayesian posterior summaries at a genome-wide scale in a fraction of the time compared to existing general-purpose methods. We apply Grid-LMM to two types of quantitative genetic analyses. The first is focused on accounting for spatial variability and non-additive genetic variance while scanning for QTL; and the second aims to identify gene expression traits affected by non-additive genetic variation. In both cases, modeling multiple sources of heterogeneity leads to new discoveries.\n\nAuthor summaryThe goal of quantitative genetics is to characterize the relationship between genetic variation and variation in quantitative traits such as height, productivity, or disease susceptibility. A statistical method known as the linear mixed effect model has been critical to the development of quantitative genetics. First applied to animal breeding, this model now forms the basis of a wide-range of modern genomic analyses including genome-wide associations, polygenic modeling, and genomic prediction. The same model is also widely used in ecology, evolutionary genetics, social sciences, and many other fields. Mixed models are frequently multi-faceted, which is necessary for accurately modeling data that is generated from complex experimental designs. However, most genomic applications use only the simplest form of linear mixed methods because the computational demands for model fitting can be too great. We develop a flexible approach for fitting linear mixed models to genome scale data that greatly reduces their computational burden and provides flexibility for users to choose the best statistical paradigm for their data analysis. We demonstrate improved accuracy for genetic association tests, increased power to discover causal genetic variants, and the ability to provide accurate summaries of model uncertainty using both simulated and real data examples.

genomics

Evaluating the quality of the 1000 Genomes Project data

Data from the 1000 Genomes project is quite often used as a reference for human genomic analysis. However, its accuracy needs to be assessed to understand the quality of predictions made using this reference. We present here an assessment of the genotype, phasing, and imputation accuracy data in the 1000 Genomes project. We compare the phased haplotype calls from the 1000 Genomes project to experimentally phased haplotypes for 28 of the same individuals sequenced using the 10X Genomics platform. We observe that phasing and imputation for rare variants are unreliable, which likely reflects the limited sample size of the 1000 Genomes project data. Further, it appears that using a population specific reference panel does not improve the accuracy of imputation over using the entire 1000 Genomes data set as a reference panel. We also note that the error rates and trends depend on the choice of definition of error, and hence any error reporting needs to take these definitions into account.

genomics

Chromosome-scale assemblies reveal the structural evolution of African cichlid genomes

BackgroundAfrican cichlid fishes are well known for their rapid radiations and are a model system for studying evolutionary processes. Here we compare multiple, high-quality, chromosome-scale genome assemblies to understand the genetic mechanisms underlying cichlid diversification and study how genome structure evolves in rapidly radiating lineages.\n\nResultsWe re-anchored our recent assembly of the Nile tilapia (Oreochromis niloticus) genome using a new high-density genetic map. We developed a new de novo genome assembly of the Lake Malawi cichlid, Metriaclima zebra, using high-coverage PacBio sequencing, and anchored contigs to linkage groups (LGs) using four different genetic maps. These new anchored assemblies allow the first chromosome-scale comparisons of African cichlid genomes.\n\nLarge intra-chromosomal structural differences (~2-28Mbp) among species are common, while inter-chromosomal differences are rare (< 10Mbp total). Placement of the centromeres within chromosome-scale assemblies identifies large structural differences that explain many of the karyotype differences among species. Structural differences are also associated with unique patterns of recombination on sex chromosomes. Structural differences on LG9, LG11 and LG20 are associated with reductions in recombination, indicative of inversions between the rock- and sand-dwelling clades of Lake Malawi cichlids. M. zebra has a larger number of recent transposable element (TE) insertions compared to O. niloticus, suggesting that several TE families have a higher rate of insertion in the haplochromine cichlid lineage.\n\nConclusionThis study identifies novel structural variation among East African cichlid genomes and provides a new set of genomic resources to support research on the mechanisms driving cichlid adaptation and speciation.

genomics