Search bioRxiv⌕ Search

Biology subjects

Kujala, S. T.

Publications and source records attributed to Kujala, S. T..

3 recordsLinked to original sources

Optimizing Exome Captures in Species with Large Genomes Using Species-specific Repetitive DNA Blocker

Large and highly repetitive genomes are common. However, research interests usually lie within the non-repetitive parts of the genome, as they are more likely functional, and can be used to answer questions related to adaptation, selection, and evolutionary history. Exome capture is a cost-effective method for providing sequencing data from protein-coding parts of the genes. C0t-1 DNA blockers consist of repetitive DNA and are used in exome captures to prevent the hybridization of repetitive DNA sequences to capture baits or bait-bound genomic DNA. Universal blockers target repetitive regions shared by many species, while species-specific c0t-1 DNA is prepared from the DNA of the studied species, thus perfectly matching the repetitive DNA contents of the species. So far the use of species-specific c0t-1 DNA has been limited to a few model species. Here, we evaluated the performance of blocker treatments in exome captures of Pinus sylvestris, a widely distributed conifer species with a large (> 20 Gbp) and highly repetitive genome. We compared treatment with a commercial universal blocker to treatments with species-specific c0t-1 (30,000 ng and 60,000 ng). Species-specific c0t-1 captured more unique exons than the initial set of targets, reduced sequencing of tandem repeats, and produced more target regions with high read coverage and narrower depth distribution than the universal blocker. Based on our results, we recommend optimizing exome captures by using at least 60,000 ng species-specific c0t-1 DNA. It is relatively easy and fast to prepare and can also be used with existing bait set designs.

genomics↗

Does the seed fall far from the tree? - weak fine scale genetic structure in a continuous Scots pine population

Knowledge of fine-scale spatial genetic structure, i.e., the distribution of genetic diversity at short distances, is important in evolutionary research and in practical applications such as conservation and breeding programs. In trees, related individuals often grow close to each other due to limited seed and/or pollen dispersal. The extent of seed dispersal also limits the speed at which a tree species can spread to new areas. We studied the fine-scale spatial genetic structure of Scots pine (Pinus sylvestris) in two naturally regenerated sites located 20 km from each other in continuous south-eastern Finnish forest. We genotyped almost 500 adult trees for 150k SNPs using a custom made Affymetrix array. We detected some pairwise relatedness at short distances, but the average relatedness was low and decreased with increasing distance, as expected. Despite the clustering of related individuals, the sampling sites were not differentiated (FST = 0.0005). According to our results, Scots pine has a large neighborhood size (Nb = 1680- 3210), but a relatively short gene dispersal distance ({sigma}g = 36.5-71.3 m). Knowledge of Scots pine fine-scale spatial genetic structure can be used to define suitable sampling distances for evolutionary studies and practical applications. Detailed empirical estimates of dispersal are necessary both in studying post-glacial recolonization and predicting the response of forest trees to climate change.

evolutionary biology↗

Taming the massive genome of Scots pine with PiSy50k, a new genotyping array for conifer research

Scots pine (Pinus sylvestris) is the most widespread coniferous tree in the boreal forests of Eurasia and has major economic and ecological importance. However, its large and repetitive genome presents a challenge for conducting genome-wide analyses such as association studies and genomic selection. We present a new 50K SNP genotyping array for Scots pine research, breeding programs, and other applications. To select the SNP set, we first genotyped 480 Scots pine samples on a 407 540 SNP screening array, and identified 47 712 high-quality SNPs for the final array (called PiSy50k). Here, we provide details of the design and testing, as well as allele frequency estimates from the discovery panel, functional annotation, tissue-specific expression patterns, and expression level information for the SNPs or corresponding genes, when available. We validated the performance of the PiSy50k array using samples from breeding populations from Finland and Scotland. Overall, 39 678 (83.2%) SNPs showed low error rates (mean = 0.92%). Relatedness estimates based on array genotypes were consistent with the expected pedigrees, and the amount of Mendelian error was negligible. In addition, array genotypes successfully discriminate Scots pine populations from different geographic origins. The PiSy50k array will be a valuable tool for future genetic studies and forestry applications. Significance statementScots pine is an evolutionary, economically and ecologically impressive coniferous species but its gigantic genome has limited studying e.g. the genetic basis of its functional trait variation. We have developed a genotyping array that facilitates Scots pine genetic research and linking its trait variation to genetic polymorphisms and gene expression levels across the genome.

genomics↗