Search bioRxivSearch

Biology subjects

Jannink, J.-L.

Publications and source records attributed to Jannink, J.-L..

9 recordsLinked to original sources

A low resolution epistasis mapping approach to identify chromosome arm interactions in allohexaploid wheat

1Epistasis is an important contributor to genetic variance, even in inbred populations where it is present as additive by additive interactions. Testing for epistasis presents a multiple testing problem as the search space for modest numbers of markers is large. Additionally, single markers do not necessarily track functional units of interacting chromatin as well as haplotype based methods do. To harness the power of multiple markers while drastically minimizing the number of tests conducted, we present a low resolution test for epistatic interactions across whole chromosome arms. Two additive genetic covariance matrices are constructed from markers on two different chromosome arms. The Hadamard product of these additive covariance matrices is then used to produce the additive by additive epistasis covariance matrix between the two chromosome arms. The covariance matrices are subsequently used to estimate an epistatic interaction variance parameter in a mixed model framework, while correcting for background additive and epistatic effects. We find significant epistatic interactions for 2% of interactions tested for four agronomic traits in a population of winter wheat. Interactions across homeologous chromosome arms were identified, but were less abundant than other interaction chromosome arm pairs. Of these, homeologous chromosome arm pair 4BL and 4DL showed a strong relationship between the product of their additive effects and the interaction effect that may be indicative of functional redundancy. Several chromosome arms were involved in many interactions across the genome, suggesting that they may contain important large effect regulatory factors. The differential patterns of epistasis across different traits suggests that detection of epistatic interactions is robust when correcting for background additive and epistatic effects in the population. The low resolution epistasis mapping method presented here identifies important epistatic interactions with a limited number of statistical tests at the cost of relatively lower precision.

genetics

A subfunctionalization epistasis model to evaluate homeologous gene interactions in allopolyploid wheat

1Hybridization between related species results in the formation of an allopolyploid with multiple subgenomes. These subgenomes will each contain complete, yet evolutionarily divergent, sets of genes. Like a diploid hybrid, allopolyploids will have two versions, or homeoalleles, for every gene. Partial functional redundancy between homeologous genes should result in a deviation from additivity. These epistatic interactions between homeoalleles are analogous to dominance effects, but are fixed across subgenomes through self pollination. An allopolyploid can be viewed as an immortalized hybrid, with the opportunity to select and fix favorable homeoallelic interactions within inbred varieties. We present a subfunctionalization epistasis model to estimate the degree of functional redundancy between homeoallelic loci and a statistical framework to determine their importance within a population. We provide an example using the homeologous dwarfing genes of allohexaploid wheat, Rht-1, and search for genome-wide patterns indicative of homeoallelic subfunctionalization in a breeding population. Using the IWGSC RefSeq vl.0 sequence, 23,796 homeoallelic gene sets were identified and anchored to the nearest DNA marker to form 10,172 homeologous marker sets. Interaction predictors constructed from products of marker scores were used to fit the homeologous main and interaction effects, as well as estimate whole genome genetic values. Some traits displayed a pattern indicative of homeoallelic subfunctionalization, while other traits showed a less clear pattern or were not affected. Using genomic prediction accuracy to evaluate importance of marker interactions, we show that homeologous interactions explain a portion of the non-additive genetic signal, but are less important than other epistatic interactions.

genetics

RNA polymerase mapping in plants identifies enhancers enriched in causal variants

Promoter-proximal pausing and divergent transcription at promoters and enhancers, which are prominent features in animals, have been reported to be absent in plants based on a study of Arabidopsis thaliana. Here, our PRO-Seq analysis in cassava (Manihot esculenta) identified peaks of transcriptionally-engaged RNA polymerase II (Pol2) at both 5 and 3 ends of genes, consistent with paused or slowly-moving Pol2, and divergent transcription at potential intragenic enhancers. A full genome search for bi-directional transcription using an algorithm for enhancer detection developed in mammals (dREG) identified many enhancer candidates. These sites show distinct patterns of methylation and nucleotide variation based on genomic evolutionary rate profiling characteristic of active enhancers. Maize GRO-Seq data showed RNA polymerase occupancy at promoters and enhancers consistent with cassava but not Arabidopsis. Furthermore, putative enhancers in maize identified by dREG significantly overlapped with sites previously identified on the basis of open chromatin, histone marks, and methylation. We show that SNPs within these divergently transcribed intergenic regions predict significantly more variation in fitness and root composition than SNPs in chromosomal segments randomly ascertained from the same intergenic distribution, suggesting a functional importance of these sites on cassava. The findings shed new light on plant transcription regulation and its impact on development and plasticity.

genomics

Prediction of subgenome additive and interaction effects in allohexaploid wheat

1Whole genome duplications have played an important role in the evolution of angiosperms. These events often occur through hybridization between closely related species, resulting in an allopolyploid with multiple subgenomes. With the availability of affordable genotyping and a reference genome to locate markers, breeders of allopolyploids now have the opportunity to manipulate subgenomes independently. This also presents a unique opportunity to investigate epistatic interactions between homeologous orthologs across subgenomes. We present a statistical framework for partitioning genetic variance to the subgenomes of an allopolyploid, predicting breeding values for each subgenome, and determining the importance of inter-genomic epistasis. We demonstrate using an allohexaploid wheat breeding population evaluated in Ithaca, NY and an important wheat dataset previously shown to demonstrate non-additive genetic variance. Subgenome covariance matrices were constructed and used to calculate subgenome interaction covariance matrices across subgenomes for variance component estimation and genomic prediction. We propose a method to extract population structure from all subgenomes at once before covariances are calculated to reduce collinearity between subgenome estimates. Variance parameter estimation was shown to be reliable for additive subgenome effects, but was less reliable for subgenome interaction components. Predictive ability was equivalent to current genomic prediction methods. Including only inter-genomic interactions resulted in the same increase in accuracy as modeling all pairwise marker interactions. Thus, we provide a new tool for breeders of allopolyploid crops to characterize the genetic architecture of existing populations, determine breeding goals, and develop new strategies for selection of additive effects and fixation of inter-genomic epistasis.

genetics

Leveraging Transcriptomics Data for Genomic Prediction Models in Cassava

BackgroundGenomic prediction models were, in principle, developed to include all the available marker information; with this approach, these models have shown in various crops moderate to high predictive accuracies. Previous studies in cassava have demonstrated that, even with relatively small training populations and low-density GBS markers, prediction models are feasible for genomic selection. In the present study, we prioritized SNPs in close proximity to genome regions with biological importance for a given trait. We used a number of strategies to select variants that were then included in single and multiple kernel GBLUP models. Specifically, our sources of information were transcriptomics, GWAS, and immunity-related genes, with the ultimate goal to increase predictive accuracies for Cassava Brown Streak Disease (CBSD) severity.\n\nResultsWe used single and multi-kernel GBLUP models with markers imputed to whole genome sequence level to accommodate various sources of biological information; fitting more than one kinship matrix allowed for differential weighting of the individual marker relationships. We applied these GBLUP approaches to CBSD phenotypes (i.e., root infection and leaf severity three and six months after planting) in a Ugandan Breeding Population (n = 955). Three means of exploiting an established RNAseq experiment of CBSD-infected cassava plants were used. Compared to the biology-agnostic GBLUP model, the accuracy of the informed multi-kernel models increased the prediction accuracy only marginally (1.78% to 2.52%).\n\nConclusionsOur results show that markers imputed to whole genome sequence level do not provide enhanced prediction accuracies compared to using standard GBS marker data in cassava. The use of transcriptomics data and other sources of biological information resulted in prediction accuracies that were nominally superior to those obtained from traditional prediction models.

genomics

Genome-wide association mapping and genomic prediction unravels CBSD resistance in a Manihot esculenta breeding population

Cassava (Manihot esculenta Crantz), a key carbohydrate dietary source for millions of people in Africa, faces severe yield loses due to two viral diseases: cassava brown streak disease (CBSD) and cassava mosaic disease (CMD). The completion of the cassava genome sequence and the whole genome marker profiling of clones from African breeding programs (www.nextgencassava.org) provides cassava breeders the opportunity to deploy additional breeding strategies and develop superior varieties with both farmer and industry preferred traits. Here the identification of genomic segments associated with resistance to CBSD foliar symptoms and root necrosis as measured in two breeding panels at different growth stages and locations is reported. Using genome-wide association mapping and genomic prediction models we describe the genetic architecture for CBSD severity and identify loci strongly associated on chromosomes 4 and 11. Moreover, the significantly associated region on chromosome 4 colocalises with a Manihot glaziovii introgression segment and the significant SNP markers on chromosome 11 are situated within a cluster of nucleotide-binding site leucine-rich repeat (NBS-LRR) genes previously described in cassava. Overall, predictive accuracy values found in this study varied between CBSD severity traits and across GS models with Random Forest and RKHS showing the highest predictive accuracies for foliar and root CBSD severity scores.

genetics

Accuracies Of Univariate And Multivariate Genomic Prediction Models In African Cassava.

List of abbreviations\n\nAbstractO_ST_ABSBackgroundC_ST_ABSGenomic selection (GS) promises to accelerate genetic gain in plant breeding programs especially for long cycle crops like cassava. To practically implement GS in cassava breeding, it is useful to evaluate different GS models and to develop suitable models for an optimized breeding pipeline.\n\nMethodsWe compared prediction accuracies from a single-trait (uT) and a multi-trait (MT) mixed model for single environment genetic evaluation (Scenario 1) while for multi-environment evaluation accounting for genotype-by-environment interaction (Scenario 2) we compared accuracies from a univariate (uE) and a multivariate (ME) multi-environment mixed model. We used sixteen years of data for six target cassava traits for these analyses. All models for Scenario 1 and Scenario 2 were based on the one-step approach. A 5-fold cross validation scheme with 10-repeat cycles were used to assess model prediction accuracies.\n\nResultsIn Scenario 1, the MT models had higher prediction accuracies than the uT models for most traits and locations analyzed amounting to 32 percent better prediction accuracy on average. However for Scenario 2, we observed that the ME model had on average (across all locations and traits) 12 percent better predictive power than the uE model.\n\nConclusionWe recommend the use of multivariate mixed models (MT and ME) for cassava genetic evaluation. These models may be useful for other plant species.

genetics

Prospects for genomic selection in cassava breeding

Cassava (Manihot esculenta Crantz) is a clonally propagated staple food crop in the tropics. Genomic selection (GS) reduces selection cycle times by the prediction of breeding value for selection of unevaluated lines based on genome-wide marker data. GS has been implemented at three breeding programs in sub-Saharan Africa. Initial studies provided promising estimates of predictive abilities in single populations using standard prediction models and scenarios. In the present study we expand on previous analyses by assessing the accuracy of seven prediction models for seven traits in three prediction scenarios: (1) cross-validation within each population, (2) cross-population prediction and (3) cross-generation prediction. We also evaluated the impact of increasing training population size by phenotyping progenies selected either at random or using a genetic algorithm. Cross-validation results were mostly consistent across breeding programs, with non-additive models like RKHS predicting an average of 10% more accurately. Accuracy was generally associated with heritability. Cross-population prediction accuracy was generally low (mean 0.18 across traits and models) but prediction of cassava mosaic disease severity increased up to 57% in one Nigerian population, when combining data from another related population. Accuracy across-generation was poorer than within (cross-validation) as expected, but indicated that accuracy should be sufficient for rapid-cycling GS on several traits. Selection of prediction model made some difference across generations, but increasing training population (TP) size was more important. In some cases, using a genetic algorithm, selecting one third of progeny could achieve accuracy equivalent to phenotyping all progeny. Based on the datasets analyzed in this study, it was apparent that the size of a training population (TP) has a significant impact on prediction accuracy for most traits. We are still in the early stages of GS in this crop, but results are promising, at least for some traits. The TPs need to continue to grow and quality phenotyping is more critical than ever. General guidelines for successful GS are emerging. Phenotyping can be done on fewer individuals, cleverly selected, making for trials that are more focused on the quality of the data collected.\n\nAbbreviations

genetics

Genome-wide association mapping of correlated traits in cassava: dry matter and total carotenoid content

Cassava (Manihot esculenta (L.) Crantz) is a starchy root crop cultivated in the tropics for fresh consumption and commercial processing. Dry matter content and micronutrient density, particularly of provitamin A - traits that are negatively correlated - are among the primary selection objectives in cassava breeding. This study aimed at identifying genetic markers associated with these traits and uncovering the potential underlying cause of their negative correlation - whether linkage and/or pleiotropy. A genome-wide association mapping using 672 clones genotyped at 72,279 SNP loci was carried out. Root yellowness was used indirectly to assess variation in carotenoid content. Two major loci for root yellowness was identified on chromosome 1 at positions 24.1 and 30.5 Mbp. A single locus for dry matter content that co-located with the 24.1 Mbp peak for carotenoid content was identified. Haplotypes at these loci explained a large proportion of the phenotypic variability. Evidence of mega-base-scale linkage disequilibrium around the major loci of the two traits and detection of the major dry matter locus in independent analysis for the white- and yellow-root subpopulations suggests that physical linkage rather that pleiotropy is more likely to be the cause of the negative correlation between the target traits. Moreover, candidate genes for carotenoid (phytoene synthase) and starch biosynthesis (UDP-glucose pyrophosphorylase and sucrose synthase) occurred in the vicinity of the identified locus at 24.1 Mbp. These findings elucidate on the genetic architecture of carotenoids and dry matter in cassava and provides an opportunity to accelerate genetic improvement of these traits.\n\nCORE IDEASO_LICassava, a starchy root crop, is a major source of dietary calories in the tropics.\nC_LIO_LIMost varieties consumed are poor in micronutrients, including pro-vitamin A.\nC_LIO_LIThese two traits are governed by few major loci on chromosome one.\nC_LIO_LIGenetic linkage, rather than pleiotropy, is the most likely cause of their negative correlation.\nC_LI

genetics