Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Screening for the ancient polar bear mitochondrial genome reveals low integration of mitochondrial pseudogenes (numts) in bears

Phylogenetic analyses of nuclear and mitochondrial genomes have shown that polar bears captured the mitochondrial genome of brown bears some 160,00 years ago. This hybridization event likely led to an extinction of the original polar bear mitochondrial genome. However, parts of the mitochondrial DNA occasionally integrates into the nuclear genome, forming pseudogenes called numts (nuclear mitochondrial integrations). Screening the polar bear genome for numts, we identified only 13 such integrations. Analyses of whole-genome sequences from additional polar bears, brown and American black bears as well as the giant panda indicates that the discovered numts entered the bear lineage before the initial ursid radiation some 14 million years ago. Our findings suggests a low integration rate of numts in the bear lineage and a complete loss of the original polar bear mitochondrial genome.

evolutionary biology

McClintock: An integrated pipeline for detecting transposable element insertions in whole genome shotgun sequencing data.

BackgroundTransposable element (TE) insertions are among the most challenging type of variants to detect in genomic data because of their repetitive nature and complex mechanisms of replication. Nevertheless, the recent availability of large resequencing datasets has spurred the development of many new methods to detect TE insertions in whole genome shotgun sequences. These methods generate output in diverse formats and have a large number of software and data dependencies, making their comparative evaluation challenging for potential users.\n\nResultsHere we develop an integrated bioinformatics pipeline for the detection of TE insertions in whole genome shotgun data, called McClintock (https://github.com/bergmanlab/mcclintock), that automatically runs and generates standardized output for multiple TE detection methods. We demonstrate the utility of the McClintock system by performing comparative evaluation of six TE detection methods using simulated and real genome data from the model microbal eukaryote, Saccharomyces cerevisiae. We find substantial variation among McClintock component methods in their ability to detect non-reference insertions in the yeast genome, but show that non-reference TEs at nearly all biologically-realistic locations can be detected in simulated data by combining multiple methods that use split-read and read-pair evidence. In general, our results reveal that split-read methods detect fewer non-reference TE insertions than read-pair methods, but generally have much higher positional accuracy. Analysis of a large sample of real yeast genomes reveals that most, but not all, McClintock component methods can recover known aspects of TE biology in yeast such as the transpositional activity status of families, tRNA gene target preferences, and target site duplication structure, albeit with varying levels of positional accuracy.\n\nConclusionsOur results suggest that no single TE detection method currently provides comprehensive detection of non-reference TEs, even in the context of a simplified model eukaryotic genome like S. cerevisiae. In spite of these limitations, the McClintock system provides a framework for testing, developing and integrating results from multiple TE detection methods to achieve this ultimate aim, as well as useful guidance for yeast researchers to select appropriate TE detection tools.

bioinformatics

PGS: a dynamic and automated population-based genome structure software

Hi-C technologies are widely used to investigate the spatial organization of genomes. However, the structural variability of the genome is a great challenge to interpreting ensemble-averaged Hi-C data, particularly for long-range/interchromosomal interactions. We pioneered a probabilistic approach for generating a population of distinct diploid 3D genome structures consistent with all the chromatin-chromatin interaction probabilities from Hi-C experiments. Each structure in the population is a physical model of the genome in 3D. Analysis of these models yields new insights into the causes and the functional properties of the genomes organization in space and time. We provide a user-friendly software package, called PGS, that runs on local machines and high-performance computing platforms. PGS takes a genome-wide Hi-C contact frequency matrix and produces an ensemble of 3D genome structures entirely consistent with the input. The software automatically generates an analysis report, and also provides tools to extract and analyze the 3D coordinates of specific domains.

bioinformatics

Habitat generalists or specialists, insights from comparative genomic analyses of Thermosipho lineages

Thermosipho species inhabit various extreme environments such as marine hydrothermal vents, petroleum reservoirs and terrestrial hot springs. A 16S rRNA phylogeny of available Thermosipho spp. sequences suggested habitat specialists adapted to living in hydrothermal vents only, and habitat generalists inhabiting oil reservoirs, hydrothermal vents and hotsprings. Comparative genomics and recombination analysis of the genomes of 15 Thermosipho isolates separated them into three species with different habitat distributions, the widely distributed T. africanus and the more specialized, T. melanesiensis and T. affectus. The three Thermosipho species can also be differentiated on the basis of genome content. For instance the T. africanus genomes had the largest repertoire of carbohydrate metabolism, which could explain why these isolates were obtained from ecologically more divergent habitats. The three species also show different capacities for defense against foreign DNA. T. melanesiensis and T. africanus both had a complete RM system, while this was missing in T. affectus. These observations also correlated with Pacbio sequencing, which revealed a methylated T. melanesiensis BI431 genome, while no methylation was detected among two T. affectus isolates. All the genomes carry CRISPR arrays accompanied by more or less complete CRISPR-cas systems. Interestingly, some isolates of both T. melanesiensis and T. africanus carry integrated prophage elements, with spacers matching these in their CRISPR arrays. Taken together, the comparative genomic analyses of Thermosipho spp. revealed genetic variation allowing habitat differentiation within the genus as well as differentiation with respect to invading mobile DNA that is present in subsurface ecosystems.

microbiology

Pan-genome and phylogeny of Bacillus cereus sensu lato

BackgroundBacillus cereus sensu lato (s. l.) is an ecologically diverse bacterial group of medical and agricultural significance. In this study, I use publicly available genomes to characterize the B. cereus s. l. pan-genome and perform the largest phylogenetic and population genetic analyses of this group to date in terms of the number of genes and taxa included. With these fundamental data in hand, I identify genes associated with particular phenotypic traits (i.e., \"pan-GWAS\" analysis), and quantify the degree to which taxa sharing common attributes are phylogenetically clustered.\n\nMethodsA rapid k-mer based approach (Mash) was used to create reduced representations of selected Bacillus genomes, and a fast distance-based phylogenetic analysis of this data (FastME) was performed to determine which species should be included in B. cereus s. l. The complete genomes of eight B. cereus s. l. species were annotated de novo with Prokka, and these annotations were used by Roary to produce the B. cereus s. l. pan-genome. Scoary was used to associate gene presence and absence patterns with various phenotypes. The orthologous protein sequence clusters produced by Roary were filtered and used to build HaMStR databases of gene models that were used in turn to construct phylogenetic data matrices. Phylogenetic analyses used RAxML, DendroPy, ClonalFrameML, PAUP*, and SplitsTree. Bayesian model-based population genetic analysis assigned taxa to clusters using hierBAPS. The genealogical sorting index was used to quantify the phylogenetic clustering of taxa sharing common attributes.\n\nThe B. cereus s. l. pan-genome currently consists of {approx}60,000 genes, {approx}600 of which are \"core\" (common to at least 99% of taxa sampled). Pan-GWAS analysis revealed genes associated with phenotypes such as isolation source, oxygen requirement, and ability to cause diseases such as anthrax or food poisoning. Extensive phylogenetic analyses using an unprecedented amount of data produced phylogenies that were largely concordant with each other and with previous studies. Phylogenetic support as measured by bootstrap probabilities increased markedly when all suitable pan-genome data was included in phylogenetic analyses, as opposed to when only core genes were used. Bayesian population genetic analysis recommended subdividing the three major clades of B. cereus s. l. into nine clusters. Taxa sharing common traits and species designations exhibited varying degrees of phylogenetic clustering.

evolutionary biology

Properties Of Genomic Relationships For Estimating Current Genetic Variances Within And Genetic Correlations Between Populations

Different methods are available to calculate multi-population genomic relationship matrices. Since those matrices differ in base population, it is anticipated that the method used to calculate the genomic relationship matrix affect the estimate of genetic variances, covariances and correlations. The aim of this paper is to define a multi-population genomic relationship matrix to estimate current genetic variances within and genetic correlations between populations. The genomic relationship matrix containing two populations consists of four blocks, one block for population 1, one block for population 2, and two blocks for relationships between the populations. It is known, based on literature, that current genetic variances are estimated when the current population is used as base population of the relationship matrix. In this paper, we theoretically derived the properties of the genomic relationship matrix to estimate genetic correlations and validated it using simulations. When the scaling factors of the genomic relationship matrix fulfill the property [Formula], the genetic correlation is estimated even though estimated variance components are not necessarily related to the current population. When this property is not met, the correlation based on estimated variance components should be multiplied by [Formula] to rescale the genetic correlation. In this study we present a genomic relationship matrix which directly results in current genetic variances as well as genetic correlations between populations.

genetics

BasePlayer: Versatile Analysis Software For Large-Scale Genomic Variant Discovery

Next-generation sequencing (NGS) is being routinely applied in life sciences and clinical practice, where the interpretation of the resulting massive data has become a critical challenge. Computational workflows, such as the Broad GATK, have been established to take raw sequencing data and produce processed data for downstream analyses. Consequently, results of these computationally demanding workflows, consisting of e.g. sequence alignment and variant calling, are increasingly being provided for customers by sequencing and bioinformatics facilities. However, downstream variant analysis, whole-genome level in particular, has been lacking a multi-purpose tool, which could take advantage of rapidly growing genomic information and integrate genetic variant, sequence, genomic annotation and regulatory (e.g. ENCODE) data interactively and in a visual fashion. Here we introduce a highly efficient and user-friendly software, BasePlayer (http://baseplayer.fi), for biological discovery in large-scale NGS data. BasePlayer enables tightly integrated comparative variant analysis and visualization of thousands of NGS data samples and millions of variants, with numerous applications in disease, regulatory and population genomics. Although BasePlayer has been designed primarily for whole-genome and exome sequencing data, it is well-suited to various study settings, diseases and organisms by supporting standard and upcoming file formats. BasePlayer transforms an ordinary desktop computer into a large-scale genomic research platform, enabling also a non-technical user to perform complex comparative variant analyses, population frequency filtering and genome level annotations under intuitive, scalable and highly-responsive user interface to facilitate everyday genetic research as well as the search of novel discoveries.

bioinformatics

Demography And Mating System Shape The Genome-Wide Impact Of Purifying Selection In Arabis alpina

Plant mating systems have profound effects on levels and structuring of genetic variation, and can affect the impact of natural selection. While theory predicts that intermediate outcrossing rates may allow plants to prevent accumulation of deleterious alleles, few studies have empirically tested this prediction using genomic data. Here, we study the effect of mating system on purifying selection by conducting population genomic analyses on whole-genome resequencing data from 38 European individuals of the arctic-alpine crucifer Arabis alpina. We find that outcrossing and mixed-mating populations maintain genetic diversity at similar levels, whereas highly self-fertilizing Scandinavian A. alpina show a strong reduction in genetic diversity, most likely as a result of a postglacial colonization bottleneck. We further find evidence for accumulation of genetic load in highly self-fertilizing populations, whereas the genome-wide impact of purifying selection does not differ greatly between mixed-mating and outcrossing populations. Our results demonstrate that intermediate levels of outcrossing may allow efficient selection against harmful alleles whereas demographic effects can be important for relaxed purifying selection in highly selfing populations. Thus, both mating system and demography shape the impact of purifying selection on genomic variation in A. alpina. These results are important for an improved understanding of the evolutionary consequences of mating system variation and the maintenance of mixed-mating strategies.\n\nSignificanceIntermediate outcrossing rates are theoretically predicted to maintain effective selection against harmful alleles, but few studies have empirically tested this prediction using genomic data. We used whole-genome resequencing data from alpine rock-cress to study how genetic variation and purifying selection vary with mating system. We find that populations with intermediate outcrossing rates have similar levels of genetic diversity as outcrossing populations, and that purifying selection against harmful alleles is efficient in mixed-mating populations. In contrast, self-fertilizing populations from Scandinavia have strongly reduced genetic diversity, and accumulate harmful mutations, likely as a result of demographic effects of postglacial colonization. Our results suggest that mixed-mating populations can avoid the negative evolutionary consequences of high self-fertilization rates.

evolutionary biology

Exploring the Diversity of Bacillus whole genome sequencing projects using Peasant, the Prokaryotic Assembly and Annotation Tool

BackgroundThe persistent decrease in cost and difficulty of whole genome sequencing of microbial organisms has led to a dramatic increase in the number of species and strains characterized from a wide variety of environments. Microbial genome sequencing can now be conducted by small laboratories and as part of undergraduate curriculum. While sequencing is routine in microbiology, assembly, annotation and downstream analyses still require computational resources and expertise, often necessitating familiarity with programming languages. To address this problem, we have created a light-weight, user-friendly tool for the assembly and annotation of microbial sequencing projects.\n\nResultsThe Prokaryotic Assembly and Annotation Tool, Peasant, automates the processes of read quality control, genome assembly, and annotation for microbial sequencing projects. High-quality assemblies and annotations can be generated by Peasant without the need of programming expertise or high-performance computing resources. Furthermore, statistics are calculated so that users can evaluate their sequencing project. To illustrate the computational speed and accuracy of Peasant, the SRA records of 322 Illumina platform whole genome sequencing assays for Bacillus species were retrieved from NCBI, assembled and annotated on a single desktop computer. From the assemblies and annotations produced, a comprehensive analysis of the diversity of over 200 high-quality samples was conducted, looking at both the 16S rRNA phylogenetic marker as well as the Bacillus core genome.\n\nConclusionsPeasant provides an intuitive solution for high-quality whole genome sequence assembly and annotation for users with limited programing experience and/or computational resources. The analysis of the Bacillus whole genome sequencing projects exemplifies the utility of this tool. Furthermore, the study conducted here provides insight into the diversity of the species, the largest such comparison conducted to date.

bioinformatics

Patterns of divergence across the geographic and genomic landscape of a butterfly hybrid zone associated with a climatic gradient

Hybrid zones are a valuable tool for studying the process of speciation and for identifying the genomic regions undergoing divergence and the ecological (extrinsic) and non-ecological (intrinsic) factors involved. Here, we explored the genomic and geographic landscape of divergence in a hybrid zone between Papilio glaucus and Papilio canadensis. Using a genome scan of 28,417 ddRAD SNPs, we identified genomic regions under possible selection and examined their distribution in the context of previously identified candidate genes for ecological adaptations. We showed that differentiation was genome-wide, including multiple candidate genes for ecological adaptations, particularly those involved in seasonal adaptation and host plant detoxification. The Z-chromosome and four autosomes showed a disproportionate amount of differentiation, suggesting genes on these chromosomes play a potential role in reproductive isolation. Cline analyses of significantly differentiated genomic SNPs, and of species diagnostic genetic markers, showed a high degree of geographic coincidence (81%) and concordance (80%) and were associated with the geographic distribution of a climate-mediated developmental threshold (length of the growing season). A relatively large proportion (1.3%) of the outliers for divergent selection were not associated with candidate genes for ecological adaptations and may reflect the presence of previously unrecognized intrinsic barriers between these species. These results suggest that exogenous (climate-mediated) and endogenous (unknown) clines may have become coupled and act together to reinforce reproductive isolation. This approach of assessing divergence across both the genomic and geographic landscape can provide insight about the interplay between the genetic architecture of reproductive isolation and endogenous and exogenous selection.

evolutionary biology

The Reconstruction of 2,631 Draft Metagenome-Assembled Genomes from the Global Oceans

Microorganisms play a crucial role in mediating global biogeochemical cycles in the marine environment. By reconstructing the genomes of environmental organisms through metagenomics, researchers are able to study the metabolic potential of Bacteria and Archaea that are resistant to isolation in the laboratory. Utilizing the large metagenomic dataset generated from 234 samples collected during the Tara Oceans circumnavigation expedition, we were able to assemble 102 billion paired-end reads into 562 million contigs, which in turn were co-assembled and consolidated in to 7.2 million contigs [≥]2kb in length. Approximately 1 million of these contigs were binned to reconstruct draft genomes. In total, 2,631 draft genomes with an estimated completion of [≥]50% were generated (1,491 draft genomes >70% complete; 603 high-quality genomes >90% complete). A majority of the draft genomes were manually assigned phylogeny based on sets of concatenated phylogenetic marker genes and/or 16S rRNA gene sequences. The draft genomes are now publically available for the research community at-large.

microbiology

Evolutionary rewiring of the human regulatory network by waves of genome expansion

Genome expansion is believed to be an important driver of the evolution of gene regulation. To investigate the role of newly arising sequence in rewiring the regulatory network we estimated the age of each region of the human genome by applying maximum parsimony to genome-wide alignments with 100 vertebrates. We then studied the age distribution of several types of functional regions, with a focus on regulatory elements. The age distribution of regulatory elements reveals the extensive use of newly formed genomic sequence in the evolution of regulatory interactions. Many transcription factors have expanded their repertoire of targets through waves of genomic expansions that can be traced to specific evolutionary times. Repeated elements contributed a major part of such expansion: many classes of such elements are enriched in binding sites of one or a few specific transcription factors, whose binding sites are localized in specific portions of the element and characterized by distinctive motif words. These features suggest that the binding sites were available as soon as the new sequence entered the genome, rather than being created later by accumulation of point mutations. By comparing the age of regulatory regions to the evolutionary shift in expression of nearby genes we show that rewiring through genome expansion played an important role in shaping the human regulatory network.

evolutionary biology

The genomic rate of adaptation in the fungal wheat pathogen Zymoseptoria tritici

Antagonistic host-pathogen co-evolution is a determining factor in the outcome of infection and shapes genetic diversity at the population level of both partners. While the molecular function of an increasing number of genes involved in pathogenicity is being uncovered, little is known about the molecular bases and genomic impact of hst-pathogen coevolution and rapid adaptation. Here, we apply a population genomic approach to infer genome-wide patterns of selection among thirteen isolates of the fungal pathogen Zymoseptoria tritici. Using whole genome alignments, we characterize intragenic polymorphism, and we apply different test statistics based on the distribution of non-synonymous and synonymous polymorphisms (pN/pS) and substitutions (dN/dS) to (1) characterise the selection regime acting on each gene, (2) estimate rates of adaptation and (3) identify targets of selection. We correlate our estimates with different genome variables to identify the main determinants of past and ongoing adaptive evolution, as well as purifying and balancing selection. We report a negative relationship between pN/pS and fine-scale recombination rate and a strong positive correlation between the rate of adaptive non-synonymous substitutions ({omega}a) and recombination rate. This result suggests a pervasive role of Hill-Robertson interference even in a species with an exceptionally high recombination rate (60 cM/Mb). Moreover, we report that the genome-wide fraction of adaptive non-synonymous substitutions () is ~ 44%, however in genes encoding determinants of pathogenicity we find a mean value of alpha ~ 68% demonstrating a considerably faster rate of adaptive evolution in this class of genes. We identify 787 candidate genes under balancing selection with an enrichment of genes involved in secondary metabolism and host infection, but not predicted effectors. This suggests that different classes of pathogenicity-related genes evolve according to distinct selection regimes. Overall our study shows that sexual recombination is a main driver of genome evolution in this pathogen.

evolutionary biology

Divergent genome evolution caused by regional variation in DNA gain and loss between human and mouse

The forces driving the accumulation and removal of non-coding DNA and ultimately the evolution of genome size in complex organisms are intimately linked to genome structure and organisation. Our analysis provides a novel method for capturing the regional variation of lineage-specific DNA gain and loss events in their respective genomic contexts. To further understand this connection we used comparative genomics to identify genome-wide individual DNA gain and loss events in the human and mouse genomes. Focusing on the distribution of DNA gains and losses, relationships to important structural features and potential impact on biological processes, we found that in autosomes, DNA gains and losses both followed separate lineage-specific accumulation patterns. However, in both species chromosome X was particularly enriched for DNA gain, consistent with its high L1 retrotransposon content required for X inactivation. We found that DNA loss was associated with gene-rich open chromatin regions and DNA gain events with gene-poor closed chromatin regions. Additionally, we found that DNA loss events tended to be smaller than DNA gain events suggesting that they were more tolerated in open chromatin regions. GO term enrichment in human gain hotspots showed terms related to cell cycle/metabolism, human loss hotspots were enriched for terms related to gene silencing, and mouse gain hotspots were enriched for terms related to transcription regulation. Interestingly, mouse loss hotspots were strongly enriched for terms related to developmental processes, suggesting that DNA loss in mouse is associated with phenotypic changes in mouse morphology. This is consistent with a model in which DNA gain and loss results in turnover or \"churning\" of regulatory regions that are then subjected to selection, resulting in the differences we now observe, both genomic and phenotypic/morphological.

evolutionary biology

Phylogenomics reveals an extensive history of genome duplication in diatoms (Bacillariophyta)

Premise of the studyDiatoms are one of the most species-rich lineages of microbial eukaryotes. Similarities in clade age, species richness, and contributions to primary production motivate comparisons to flowering plants, whose genomes have been inordinately shaped by whole genome duplication (WGD). These events that have been linked to speciation and increased rates of lineage diversification, identifying WGDs as a principal driver of angiosperm evolution. We synthesized a relatively large but scattered body of evidence that, taken together, suggests that polyploidy may be common in diatoms.\n\nMethodsWe used data from gene counts, gene trees, and patterns of synonymous divergence to carry out the first large-scale phylogenomic analysis of genome-scale duplication histories for a phylogenetically diverse set of 37 diatom taxa.\n\nKey resultsSeveral methods identified WGD events of varying age across diatoms, though determining the exact number and placement of events and, more broadly, inferences of WGD at all, were greatly impacted by gene-tree uncertainty. Gene-tree reconciliations supported allopolyploidy as the predominant mode of polyploid formation, with particularly strong evidence for ancient allopolyploid events in the thalassiosiroid and pennate diatom clades.\n\nConclusionsWhole genome duplication appears to have been an important driver of genome evolution in diatoms. Denser taxon sampling will better pinpoint the timing of WGDs and likely reveal many more of them. We outline potential challenges in reconstructing paleopolyploid events in diatoms that, together with these results, offer a framework for understanding the evolutionary roles of genome duplication in a group that likely harbors substantial genomic diversity.

evolutionary biology

GVC: A superfast and universal genomic variant caller

Germline and somatic variant detection from human and cancer whole-genome sequencing data is a challenge task for genome-wide association study and cancer genomics in precision medicine. Many confounding factors contribute the difficulties including complexity of variant, sequencing and alignment error, tumor clonality and sample purity etc. Current genomic variant callers are too time-consuming to meet the requirement of clinical application in precision medicine. We developed superfast and universal Genomic Variant Caller (GVC), which can simultaneously detect various genomic variants including SNV, sINDEL and SV from personal and normal-cancer paired whole-genome/exome sequencing data within fifteen minutes. Whats more, it achieved higher sensitivity and precision than popular variant callers including GATK4, Mutect, NovoBreak in germline and somatic variant detection from NA12878 and ICGC-TCGA Dream Challenge Datasets respectuvely. It is worth mentioning that GVC achieved comparable performance in variant detection from NA12878 sequenced by three different high-throughput sequencing platforms including Illumina HiSeq2000, NovaSeq and BGISEQ-500.

bioinformatics

Comparison of single genome and allele frequency data reveals discordant demographic histories

Inference of demographic history from genetic data is a primary goal of population genetics of model and non-model organisms. Whole genome-based approaches such as the Pairwise/Multiple Sequentially Markovian Coalescent (PSMC/MSMC) methods use genomic data from one to four individuals to infer the demographic history of an entire population, while site frequency spectrum (SFS)-based methods use the distribution of allele frequencies in a sample to reconstruct the same historical events. Although both methods are extensively used in empirical studies and perform well on data simulated under simple models, there have been only limited comparisons of them in more complex and realistic settings. Here we use published demographic models based on data from three human populations (Yoruba (YRI), descendants of northwest-Europeans (CEU), and Han Chinese (CHB)) as an empirical test case to study the behavior of both inference procedures. We find that several of the demographic histories inferred by the whole genome-based methods do not predict the genome-wide distribution of heterozygosity nor do they predict the empirical SFS. However, using simulated data, we also find that the whole genome methods can reconstruct the complex demographic models inferred by SFS-based methods, suggesting that the discordant patterns of genetic variation are not attributable to a lack of statistical power, but may reflect unmodeled complexities in the underlying demography. More generally, our findings indicate that demographic inference from a small number of genomes, routine in genomic studies of nonmodel organisms, should be interpreted cautiously, as these models cannot recapitulate other summaries of the data.

evolutionary biology

Virtual Genome Walking: Generating gene models for the salamander Ambystoma mexicanum

Large repeat rich genomes present challenges for assembly and identification of gene models with short read technologies. Here we present a method we call Virtual Genome Walking which uses an iterative assembly approach to first identify exons from de-novo assembled transcripts and assemble whole genome reads against each exon. This process is iterated allowing the extension of exons. These linked assemblies are refined to generate gene models including upstream and downstream genomic sequence as well as intronic sequence. We test this method using a 20X genomic read set for the axolotl, the genome of which is estimated to be 30 Gb in size. These reads were previously reported to be effectively impossible to assemble. Here we provide almost 1 Gb of assembled sequence describing over 19,000 gene models for the axolotl. Gene models stop assembling either due to localised low coverage in the genomic reads, or the presence of repeats. We validate our observations by comparison with previously published axolotl bacterial artificial chromosome (BAC) sequences. In addition we analysed axolotl intron length, intron-exon structure, repeat content and synteny. These gene-models, sequences and annotations are freely available for download from https://tinyurl.com/y8gydc6n. The software pipeline including a docker image is available from https://github.com/LooseLab/iterassemble. These methods will increase the value of low coverage sequencing of understudied model systems.

bioinformatics