Search bioRxiv⌕ Search

Biology subjects

Merrick, L. F.

Publications and source records attributed to Merrick, L. F..

4 recordsLinked to original sources

Classification and Regression Models for Genomic Selection of Skewed Phenotypes: A Case for Disease Resistance in Winter Wheat (Triticum aestivum L.)

Most genomic prediction models are linear regression models that assume continuous and normally distributed phenotypes, but responses to diseases such as stripe rust (caused by Puccinia striiformis f. sp. tritici) are commonly recorded in ordinal scales and percentages. Disease severity (SEV) and infection type (IT) data in germplasm screening nurseries generally do not follow these assumptions. On this regard, researchers may ignore the lack of normality, transform the phenotypes, use generalized linear models, or use supervised learning algorithms and classification models with no restriction on the distribution of response variables, which are less sensitive when modeling ordinal scores. The goal of this research was to compare classification and regression genomic selection models for skewed phenotypes using stripe rust SEV and IT in winter wheat. We extensively compared both regression and classification prediction models using two training populations composed of breeding lines phenotyped in four years (2016-2018, and 2020) and a diversity panel phenotyped in four years (2013-2016). The prediction models used 19,861 genotyping-by-sequencing single-nucleotide polymorphism markers. Overall, square root transformed phenotypes using rrBLUP and support vector machine regression models displayed the highest combination of accuracy and relative efficiency across the regression and classification models. Further, a classification system based on support vector machine and ordinal Bayesian models with a 2-Class scale for SEV reached the highest class accuracy of 0.99. This study showed that breeders can use linear and non-parametric regression models within their own breeding lines over combined years to accurately predict skewed phenotypes.

genetics↗

Comparison of Single-Trait and Multi-Trait Genome-Wide Association Models and Inclusion of Correlated Traits in the Dissection of the Genetic Architecture of a Complex Trait in a Breeding Program

Traits with an unknown genetic architecture make it difficult to create a useful bi-parental mapping population to characterize the genetic basis of the trait due to a combination of complex and pleiotropic effects. Seedling emergence of wheat (Triticum aestivum L.) from deep planting is a vital factor affecting stand establishment and grain yield, has a poorly understood genetic architecture, and is historically correlated with coleoptile length. The creation of bi-parental mapping populations can be overcome by using genome-wide association studies (GWAS). This study aimed to dissect the genetic architecture of seedling emergence while accounting for correlated traits using one multi-trait GWAS model (MT-GWAS) and three single-trait GWAS models (ST-GWAS) with the inclusion of covariates for correlated traits. The ST-GWAS models included one single locus model (MLM), and two multiple loci models (FarmCPU and BLINK). We conducted the GWAS using two populations, the first consisting of 473 varieties from a diverse association mapping panel (DP) phenotyped from 2015-2019, and the other population used as a validation population consisting of 279 breeding lines (BL) phenotyped in 2015 in Lind, WA, with 40,368 markers. We also compared the inclusion of coleoptile length and markers associated with reduced height as covariates in our ST-GWAS models for the DP. ST-GWAS found 107 significant markers across 19 chromosomes, while MT-GWAS found 82 significant markers across 14 chromosomes. MT-GWAS models were able to identify large-effect markers on chromosome 5A. FarmCPU and BLINK models were able to identify many small effect markers, and the inclusion of covariates helped to identify the large effect markers on chromosome 5A. Therefore, by using multi-locus models combined with pleiotropic covariates, breeding programs can uncover the complex nature of traits to help identify candidate genes and the underlying architecture of a trait, such as seedling emergence of deep-sown winter wheat.

genetics↗

Breeding with Major and Minor Genes: Genomic Selection for Quantitative Disease Resistance

Most disease resistance in plants is quantitative, with both major and minor genes controlling resistance. This research aimed to optimize genomic selection (GS) models for use in breeding programs needing to select both major and minor genes for resistance. In this experiment, stripe rust (Puccinia striiformis Westend. f. sp. tritici Erikss.) of wheat (Triticum aestivum L.) was used as a model for quantitative disease resistance. The quantitative nature of stripe rust is usually phenotyped with two disease traits, infection type and disease severity. We compared two types of training populations composed of 2,630 breeding lines phenotyped in single plot trials from four years (2016-2020) and 475 diversity panel lines from four years (2013-2016), both across two locations. We also compared the accuracy of models with four different major gene markers and genome-wide association (GWAS) markers as fixed effects. The prediction models used 31,975 markers replicated 50 times using 5-fold cross-validation. We then compared the GS models with marker-assisted selection to compare the prediction accuracy of the markers alone and in combination. The GS models had higher accuracies than marker-assisted selection and reached an accuracy of 0.72 for disease severity. The major gene and GWAS markers had only a small to zero increase in prediction accuracy over the base GS model, with the highest accuracy increase of 0.03 for major markers and 0.06 for GWAS markers. There was a statistical increase in accuracy by using the disease severity trait, the breeding lines, population type, and by combing years. There was also a statistical increase in accuracy using major markers within the validation sets as the mean accuracy decreased. The inclusion of fixed effects in low prediction scenarios increased accuracy up to 0.06 for GS models using significant GWAS markers. Our results indicate that GS can accurately predict quantitative disease resistance in the presence of major and minor genes.

genetics↗

Comparison of Genomic Selection Models for Exploring Predictive Ability of Complex Traits in Breeding Programs

Traits with a complex unknown genetic architecture are common in breeding programs. However, they pose a challenge for selection due to a combination of complex environmental and pleiotropic effects that impede the ability to create mapping populations to characterize the traits genetic basis. One such trait, seedling emergence of wheat (Triticum aestivum L.) from deep planting, presents a unique opportunity to explore the best method to use and implement GS models to predict a complex trait. 17 GS models were compared using two training populations, consisting of 473 genotypes from a diverse association mapping panel (DP) phenotyped from 2015-2019 and the other training population consisting of 643 breeding lines phenotyped in 2015 and 2020 in Lind, WA with 40,368 markers. There were only a few significant differences between GS models, with support vector machines reaching the highest accuracy of 0.56 in a single breeding line trial using cross-validations. However, the consistent moderate accuracy of cBLUP and other parametric models indicates no need to implement computationally demanding non-parametric models for complex traits. There was an increase in accuracy using cross-validations from 0.40 to 0.41 and independent validations from 0.10 to 0.17 using diversity panels lines to breeding lines. The environmental effects of complex traits can be overcome by combining years of the same populations. Overall, our study showed that breeders can accurately predict and implement GS for a complex trait by using parametric models within their own breeding programs with increased accuracy as they combine training populations over the years.

genetics↗