Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.06.23.661178

Improving Multi-Trait Genomic Prediction Efficiency Through The Incorporation Of Synthetic Traits Selected Based on Co-heritability

Abstract

Genomic prediction (GP) is an essential tool in the field of plant breeding to accelerate the cultivar development pipeline by predicting the performance of unphenotyped lines. The precision of prediction is constrained by the heritability of the target trait when applying a single-trait genomic prediction model. To overcome this limitation, a multi-trait genomic prediction model leveraging high-heritability secondary traits co-heritable with the target trait can boost predictive ability for the target trait. However, this is practically challenging because it requires additional phenotyping effort and prior knowledge of trait co-heritability. This study aimed to assess the efficiency of multi-trait genomic prediction models powered by secondary traits derived from high-throughput phenotyping data when predicting important leaf functional target traits, i.e., nitrogen (N) content and specific leaf area (SLA) in diverse sorghum accessions. Since these traits can be predicted from hyperspectral reflectance data, there is significant potential for other wavelengths within the existing dataset to meet the criteria needed to improve prediction accuracy using multi-trait approaches. Therefore, experiments were performed on traditional direct measures of leaf N content and SLA, plus partial least squares regression predictions of them (Leaf N-PLSR, SLA-PLSR), i.e., four target traits in total. Three secondary, "synthetic traits" (S1, S2, S3), each a ratio of two wavelengths within the hyperspectral data, were identified based on high co-heritability with a given target trait. Single-trait GBLUP (Genomic Best Linear Unbiased Predictor) was fitted as a baseline model, followed by three multi-trait GBLUP models using synthetic traits and target traits together. Model performance was assessed using k-fold (k=5) cross-validation (CV), which consisted of single-trait, CV1, and CV2 schemes. The synthetic traits high genetic correlation and heritability met the requirements for their use as secondary traits. There was a significant increase in accuracy when synthetic traits were used in the multi-trait genomic prediction model compared to a single trait alone for all four target traits. It improved prediction accuracy while using secondary traits derived from hyperspectral high-throughput phenotyping data in the multi-trait genomic prediction model, suggesting that this approach could be broadly applied in a post-hoc fashion to many datasets without any additional phenotyping effort. Our analysis highlights a practical approach to improve multi-trait genomic prediction model performance using synthetic traits with no intrinsic biological meaning selected through co-heritability estimation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Upadhyay, A., Azam, R., Yuan, M., Ivanovic, S., Ferguson, J. N., Paul, R. E., Koyejo, O., El-Kebir, M., Lipka, A. E., Leakey, A. D., Fernandes, S. B.. 2025-06-27. Improving Multi-Trait Genomic Prediction Efficiency Through The Incorporation Of Synthetic Traits Selected Based on Co-heritability. https://doi.org/10.1101/2025.06.23.661178

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Large language model-based bibliometric evaluation of population descriptors in human genetics

As the use of population descriptors such as race, ethnicity, and ancestry have become increasingly common in modern genetics research, there have been growing calls to critically examine their use. Most notably, in 2023, the National Academies of Science, Engineering, and Medicine (NASEM) published a report titled Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field, which included eight specific and actionable recommendations for researchers to implement the ethical and accurate use of population descriptors in genetic research. Here, we use the 2023 NASEM report as a benchmark to analyze the use of population descriptors in genome-wide association studies (GWAS). We develop a general toolkit for large language model-based bibliometrics, operationalize the report's recommendations into an evaluation framework, and apply this framework to evaluate all 4,007 papers from the GWAS Catalog published between 2007 and 2025 with full text available on PubMedCentral. We find significant improvements in adherence to NASEM report recommendations over time. However, most improvements predate the publication of the NASEM report itself, suggesting the report functioned primarily as a synthesis of existing best practices rather than a catalyst for change. We conclude by highlighting opportunities for growth in the field of human genetics.

genetics↗

Mitigating biases of rescaling in forward-in-time population genetic simulations

Forward-in-time population genetic simulations are widely used in evolutionary analyses, but simulating large populations and long genomic regions remains computationally demanding. To reduce this cost, parameter rescaling is widely employed, in which the original evolutionary process is approximated by one with a smaller population size and fewer generations. Recently, several studies using the SLiM simulator have raised concerns about the accuracy of this rescaling approach. In this study, we show that many of the biases reported in these studies can be mitigated by using a different simulation algorithm. These results reveal that the accuracy of parameter rescaling depends on how well the simulation algorithm preserves diffusion-limit properties under rescaling.

genetics↗

OPA1 controls mitochondrial dysfunction-driven liver fibrosis in MASLD

Progressive hepatic fibrosis is the principal determinant of morbidity and mortality in metabolic dysfunction-associated steatotic liver disease and steatohepatitis (MASLD/MASH). Mitochondrial dysfunction is a hallmark of MASH, and the release of mitochondrial damage-associated molecular patterns (mito-DAMPs) from injured hepatocytes can promote fibrosis. However, how mitochondrial dynamics and quality control shape the fibrotic response in MASLD/MASH remains unclear. Here, through large-scale genomic analyses of mitochondrial genes governing mitophagy, fusion and fission in human MASLD, with a power-equivalent sample size of approximately 700,000 individuals, we identify a strong association between hepatic fibrosis and the mitochondrial fusion factor dynamin-like GTPase optic atrophy 1 (OPA1). OPA1 transcripts and protein abundance in the liver epithelium were progressively dysregulated with advancing fibrosis. In mice, hepatocyte-specific OPA1 loss alone was sufficient to induce hepatic stellate cell activation and fibrosis in zone 3, promoted the release of mito-DAMPs into the circulation and exacerbated fibrosis in experimental MASH. These findings identify OPA1 as a central regulator of the hepatic fibrotic response and connect defective mitochondrial homeostasis to mito-DAMP release, hepatic stellate cell activation and fibrosis in MASLD.

genetics↗