Search bioRxivSearch

Biology subjects

Maliepaard, C.

Publications and source records attributed to Maliepaard, C..

6 recordsLinked to original sources

Family-based haplotype estimation and allele dosage correction for polyploids using short sequence reads

DNA sequence reads contain information about the genomic variants located on a single chromosome. By extracting and extending this information (using the overlaps of the reads), the haplotypes of an individual can be obtained. Adding parent-offspring relationships to the read information in a population can considerably improve the quality of the haplotypes obtained from short reads, as pedigree information can compensate for spurious overlaps (due to sequencing errors) and insufficient overlaps (due to shallow coverage). This improvement is especially beneficial for polyploid organisms, which have more than two copies of each chromosome and are therefore more difficult to be haplotyped compared to diploids. We develop a novel method, PopPoly, to estimate polyploid haplotypes in an F1-population from short sequence data by considering the transmission of the haplotypes from the parents to the offspring. In addition, PopPoly employs this information to improve genotype dosage estimation and to call missing genotypes in the population. Through realistic simulations, we compare PopPoly to other haplotyping methods and show its better performance in terms of phasing accuracy and the accuracy of phased genotypes. We apply PopPoly to estimate the parental and offspring haplotypes for a tetraploid potato cross with 10 offspring, using Illumina HiSeq sequence data of 9 genomic regions involved in plant maturity and tuberisation.

bioinformatics

A high-quality sequence of Rosa chinensis to elucidate genome structure and ornamental traits

Rose is the worlds most important ornamental plant with economic, cultural and symbolic value. Roses are cultivated worldwide and sold as garden roses, cut flowers and potted plants. Rose has a complex genome with high heterozygosity and various ploidy levels. Our objectives were (i) to develop the first high-quality reference genome sequence for the genus Rosa by sequencing a doubled haploid, combining long and short read sequencing, and anchoring to a high-density genetic map and (ii) to study the genome structure and the genetic basis of major ornamental traits.\n\nWe produced a haploid rose line from R. chinensis Old Blush and generated the first rose genome sequence at the pseudo-molecule scale (512 Mbp with N50 of 3.4 Mb and L75 of 97). The sequence was validated using high-density diploid and tetraploid genetic maps. We delineated hallmark chromosomal features including the pericentromeric regions through annotation of TE families and positioned centromeric repeats using FISH. Genetic diversity was analysed by resequencing eight Rosa species. Combining genetic and genomic approaches, we identified potential genetic regulators of key ornamental traits, including prickle density and number of flower petals. A rose APETALA2 homologue is proposed to be the major regulator of petals number in rose. This reference sequence is an important resource for studying polyploidisation, meiosis and developmental processes as we demonstrated for flower and prickle development. This reference sequence will also accelerate breeding through the development of molecular markers linked to traits, the identification of the genes underlying them and the exploitation of synteny across Rosaceae.

genomics

polymapR: linkage analysis and genetic map construction from F1 populations of outcrossing polyploids

MotivationPolyploid species carry more than two copies of each chromosome, a condition found in many of the worlds most important crops. Genetic mapping in polyploids is more complex than in diploid species, resulting in a lack of available software tools. These are needed if we are to realise all the opportunities offered by modern genotyping platforms for genetic research and breeding in polyploid crops.\n\nResultspolymapR is an R package for genetic linkage analysis and integrated genetic map construction from bi-parental populations of outcrossing autopolyploids. It can currently analyse triploid, tetraploid and hexaploid marker datasets and is applicable to various crops including potato, leek, alfalfa, blueberry, chrysanthemum, sweet potato or kiwifruit. It can detect, estimate and correct for preferential chromosome pairing, and has been tested on high-density marker datasets from potato, rose and chrysanthemum, generating high-density integrated linkage maps in all of these crops.\n\nAvailability and ImplementationpolymapR is freely available under the general public license from the Comprehensive R Archive Network (CRAN) at http://cran.r-project.org/packages=polymapR.\n\nContactChris Maliepaard chris.maliepaard@wur.nl or Roeland E. Voorrips roeland.voorrips@wur.nl

genetics

Mapfuser: an integrative toolbox for consensus map construction and Marey maps

MotivationWhere standalone tools for genetic map visualisation, consensus map construction, and Marey map analysis offer useful analyses for research and breeding, no single tool combines the three. Manual data curation is part of each of these analyses, which is difficult to standardize and consequently error prone.\n\nResultsMapfuser provides a high-level interface for common analyses in breeding programs and quantitative genetics. Combined with interactive visualisations and automated quality control, mapfuser provides a standardized and flexible toolbox and is available as Shiny app for biologists. Reproducible research in the R package is facilitated by storage of raw data, function parameters, and results in an R object. In the shiny app a rmarkdown report is available.\n\nAvailabilityMapfuser is available as R package at https://github.com/dmuijen/mapfuser under the GPL-3 License and is available for public use as Shiny application at https://plantbreeding.shinyapps.io/mapfuser. Documentation for the R package is available as package vignette. Shiny app documentation is integrated in the application.\n\nContactdennis.vanmuijen@wur.nl

genetics

TriPoly: a haplotype estimation approach for polyploids using sequencing data of related individuals

Knowledge of \"haplotypes\", i.e. phased and ordered marker alleles on a chromosome, is essential to answer many questions in genetics and genomics. By generating short pieces of DNA sequence, high-throughput modern sequencing technologies make estimation of haplotypes possible for single individuals. In polyploids, however, haplotype estimation methods usually require deep coverage to achieve sufficient accuracy. This often renders sequencing-based approaches too costly to be applied to large populations needed in studies of Quantitative Trait Loci (QTL).\n\nWe propose a novel haplotype estimation method for polyploids, TriPoly, that combines sequencing data with Mendelian inheritance rules to infer haplotypes in parent-offspring trios. Using realistic simulations of short- read sequencing data for potato (Solanum tuberosum) and banana (Musa acuminata) trios, we show that TriPoly yields more accurate progeny haplotypes at low coverages compared to the existing methods that work on single individuals.

bioinformatics

Exploiting Next Generation Sequencing to solve the Haplotyping puzzle in Polyploids: a Simulation study

Haplotypes are the units of inheritance in an organism, and many genetic analyses depend on their precise determination. Methods for haplotyping single individuals use the phasing information available in Next Generation sequencing reads, by matching overlapping sNPs while penalizing post hoc nucleotide corrections made. Haplotyping diploids is relatively easy, but the complexity of the problem increases drastically for polyploid genomes, which are found in both model organisms and in economically relevant plant and animal species. While a number of tools are available for haplotyping polyploids, the effects of the genomic makeup and the sequencing strategy followed on the accuracy of these methods have hitherto not been thoroughly evaluated.\n\nWe developed the simulation pipeline haplosim to evaluate the performance of haplotype estimation algorithms for polyploids: HapCompass, HapTree and SDhaP, in settings varying in sequencing approach, ploidy levels and genomic diversity, using tetraploid potato as the model. Our results show that sequencing depth is the major determinant of haplotype estimation quality, that 1kb PacBio CCS reads and Illumina reads with large insert-sizes are competitive, and that all methods fail to produce good haplotypes when ploidy levels increase. Comparing the three methods, HapTree produces the most accurate estimates, but also consumes the most resources. There is clearly room for improvement in polyploid haplotyping algorithms.

bioinformatics