Search bioRxivSearch

Biology subjects

Vinnere Pettersson, O.

Publications and source records attributed to Vinnere Pettersson, O..

3 recordsLinked to original sources

A hybrid de novo genome assembly of the honeybee, Apis mellifera, with chromosome-length scaffolds

BackgroundThe ability to generate long sequencing reads and access long-range linkage information is revolutionizing the quality and completeness of genome assemblies. Here we use a hybrid approach that combines data from four genome sequencing and mapping technologies to generate a new genome assembly of the honeybee Apis mellifera. We first generated contigs based on PacBio sequencing libraries, which were then merged with linked-read 10x Chromium data followed by scaffolding using a BioNano optical genome map and a Hi-C chromatin interaction map, complemented by a genetic linkage map.\n\nResultsEach of the assembly steps reduced the number of gaps and incorporated a substantial amount of additional sequence into scaffolds. The new assembly (Amel_HAv3) is significantly more contiguous and complete than the previous one (Amel_4.5), based mainly on Sanger sequencing reads. N50 of contigs is 120-fold higher (5.381 Mbp compared to 0.053 Mbp) and we anchor >98% of the sequence to chromosomes. All of the 16 chromosomes are represented as single scaffolds with an average of three sequence gaps per chromosome. The improvements are largely due to the inclusion of repetitive sequence that was unplaced in previous assemblies. In particular, our assembly is highly contiguous across centromeres and telomeres and includes hundreds of AvaI and AluI repeats associated with these features.\n\nConclusionsThe improved assembly will be of utility for refining gene models, studying genome function, mapping functional genetic variation, identification of structural variants, and comparative genomics.

genomics

A comprehensive model of DNA fragmentation for the preservation of High Molecular Weight DNA

For long-read sequencing applications, shearing of DNA is a significant issue as it limits the read-lengths generated by sequencing. During extraction and storage of DNA the DNA polymers are susceptible to physical and chemical shearing. In particular, the mechanisms of physical shearing are poorly understood in most laboratories as they are of little relevance to commonly used short-read sequencing technologies. This study draws upon lessons learned in a diverse set of research fields to create a comprehensive theoretical framework for obtaining high molecular weight DNA (HMW-DNA) to support improved quality management in laboratories and biobanks for long-read sequencing applications.\n\nUnder common laboratory conditions physical and chemical shearing yields DNA fragments of 5-35 kilobases (kb) in length. This fragment length is sufficient for DNA sequencing using short-read technologies but for Nanopore sequencing, linked reads and single molecular real time sequencing (SMRT) poorly preserved DNA will limit the length of the reads generated.\n\nThe shearing process can be divided into physical and chemical shearing which generates different patterns of fragmentation. Exposure to physical shearing creates a characteristic fragment length where the main cause of shearing is shear stress induced by turbulence. The characteristic fragment length is several thousand base pairs longer than the reads produced by short-read sequencing as the shear stress imposed on short DNA fragments is insufficient to shear the DNA. This characteristic length can be measured using gel electrophoresis or instruments for DNA fragment analysis. Chemical shearing generates randomly distributed fragment lengths visible as a smear of DNA below the peak fragment length. By measuring the peak of the DNA fragment length distribution and the proportion of very short DNA fragments, both sources of shearing can be measured using commonly used laboratory techniques, providing a suitable quantification of DNA integrity of DNA for sequencing with long-read technologies.

molecular biology

Amplicon sequencing of the 16S-ITS-23S rRNA operon with long-read technology for improved phylogenetic classification of uncultured prokaryotes

Amplicon sequencing of the 16S rRNA gene is the predominant method to quantify microbial compositions of environmental samples and to discover previously unknown lineages. Its unique structure of interspersed conserved and variable regions is an excellent target for PCR and allows for classification of reads at all taxonomic levels. However, the relatively few phylogenetically informative sites prevent confident phylogenetic placements of novel lineages that are deep branching relative to reference taxa. This problem is exacerbated when only short 16S rRNA gene fragments are sequenced. To resolve their placement, it is common practice to gather more informative sites by combining multiple conserved genes into concatenated datasets. This however requires genomic data which may be obtained through relatively expensive metagenome sequencing and computationally demanding analyses. Here we develop a protocol that amplifies a large part of 16S and 23S rRNA genes within the rRNA operon, including the ITS region, and sequences the amplicons with PacBio long-read technology. We tested our method with a synthetic mock community and developed a read curation pipeline that reduces the overall error rate to 0.18%. Applying our method on four diverse environmental samples, we were able to capture near full-length rRNA operon amplicons from a large diversity of prokaryotes. Phylogenetic trees constructed with these sequences showed an increase in statistical support compared to trees inferred with shorter, Illumina-like sequences using only the 16S rRNA gene (250 bp). Our method is a cost-effective solution to generate high quality, near full-length 16S and 23S rRNA gene sequences from environmental prokaryotes.

microbiology