Search bioRxivSearch

Biology subjects

Vidalis, A.

Publications and source records attributed to Vidalis, A..

4 recordsLinked to original sources

Association mapping identified novel candidate loci affecting wood formation in Norway spruce

[tpltrtarr] Norway spruce (Picea abies) is an important boreal forest tree species of significant ecological and economic importance. Hence there is a strong imperative to dissect the genetics controlling important wood quality traits in the species.\n[tpltrtarr]We performed a functional genome-wide association mapping of 17 wood traits in Norway spruce using 178101 single-nucleotide polymorphisms (SNPs) generated from exome genotyping of 517 mother trees. The wood traits were defined using functional modelling of wood properties across annual growth rings.\n[tpltrtarr]Association mapping was performed using a multilocus LASSO penalized regression method and we detected a total of 51 significant SNPs from 39 candidate genes that are involved in wood formation.\n[tpltrtarr]Our study represents the first functional multi-locus genome-wide association mapping (AM) in Norway spruce. The results advance our understanding of the genetics influencing wood traits, identify novel candidate genes for further functional studies and support current Norway spruce breeding efforts.

genetics

An ultra-dense haploid genetic map for evaluating the highly fragmented genome assembly of Norway spruce (Picea abies)

Norway spruce (Picea abies (L.) Karst.) is a conifer species of substanital economic and ecological importance. In common with most conifers, the P. abies genome is very large ([~]20 Gbp) and contains a high fraction of repetitive DNA. The current P. abies genome assembly (v1.0) covers approximately 60% of the total genome size but is highly fragmented, consisting of >10 million scaffolds. The genome annotation contains 66,632 gene models that are at least partially validated (www.congenie.org), however, the fragmented nature of the assembly means that there is currently little information available on how these genes are physically distributed over the 12 P. abies chromosomes. By creating an ultra-dense genetic linkage map, we anchored and ordered scaffolds into linkage groups, which complements the fine-scale information available in assembly contigs. Our ultra-dense haploid consensus genetic map consists of 21,056 markers derived from 14,336 scaffolds that contain 17,079 gene models (25.6% of the validated gene models) that we have anchored to the 12 linkage groups. We used data from three independent component maps, as well as comparisons with previously published Picea maps to evaluate the accuracy and marker ordering of the linkage groups. We demonstrate that approximately 3.8% of the anchored scaffolds and 1.6% of the gene models covered by the consensus map have likely assembly errors as they contain genetic markers that map to different regions within or between linkage groups. We further evaluate the utility of the genetic map for the conifer research community by using an independent data set of unrelated individuals to assess genome-wide variation in genetic diversity using the genomic regions anchored to linkage groups. The results show that our map is sufficiently dense to enable detailed evolutionary analyses across the P. abies genome.

genetics

Design and evaluation of a large sequence-capture probe set and associated SNPs for diploid and haploid samples of Norway spruce (Picea abies)

Massively parallel sequencing has revolutionized the field of genetics by providing comparatively high-resolution insights into whole genomes for large number of species so far. However, whole-genome resequencing of many conspecific individuals remains cost-prohibitive for most species. This is especially true for species with very large genomes with extensive genomic redundancy, such as the genomes of coniferous trees. The genome assembly for the conifer Norway spruce (Picea abies) was the first published draft genome assembly for any gymnosperm. Our goal was to develop a dense set of genome-wide SNP markers for Norway spruce to be used for assembly improvement and population studies. From 80,000 initial probe candidates, we developed two partially-overlapping sets of sequence capture probes: one developed against 56 haploid megagametophytes, to aid assembly improvement; and the other developed against 6 diploid needle samples, to aid population studies. We focused probe development within genes, as delineated via the annotation of ~67,000 gene models accompanying P. abies assembly version 1.0. The 31,277 probes developed against megagametophytes covered 19,268 gene models (mean 1.62 probes/model). The 40,018 probes developed against diploid tissue covered 26,219 gene modules (mean 1.53 probes/model). Analysis of read coverage and variant quality around probe sites showed that initial alignment of captured reads should be done against the whole genome sequence, rather than a subset of probe-containing scaffolds, to overcome occasional capture of sequences outside of designed regions. All three probe sets, anchored to the P. abies 1.0 genome assembly and annotation, are available for download.

genomics

METHimpute: Imputation-guided construction of complete methylomes from WGBS data

Whole-genome Bisulfite sequencing (WGBS) has become the standard method for interrogating plant methylomes at base resolution. However, deep WGBS measurements remain cost prohibitive for large, complex genomes and for population-level studies. As a result, most published plant methylomes are sequenced far below saturation, with a large proportion of cytosines having either missing data or insufficient coverage. Here we present METHimpute, a Hidden Markov Model (HMM) based imputation algorithm for the analysis of WGBS data. Unlike existing methods, METHimpute enables the construction of complete methylomes by inferring the methylation status and level of all cytosines in the genome regardless of coverage. Application of METHimpute to maize, rice and Arabidopsis shows that the algorithm infers cytosine-resolution methylomes with high accuracy from data as low as 6X, compared to data with 60X, thus making it a cost-effective solution for large-scale studies. Although METHimpute has been extensively tested in plants, it should be broadly applicable to other species.

bioinformatics