Search bioRxiv⌕ Search

Biology subjects

Tribout, T.

Publications and source records attributed to Tribout, T..

3 recordsLinked to original sources

Deep Generative Models for Discrete Genotype Simulation

Deep generative models open new avenues for simulating realistic genomic data while preserving privacy and addressing data accessibility constraints. While previous studies have primarily focused on generating gene expression or haplotype data, this study explores generating genotype data in both unconditioned and phenotype-conditioned settings, which is inherently more challenging due to the discrete nature of genotype data. In this work, we developed and evaluated commonly used generative models, including Variational Autoencoders (VAEs), Diffusion Models, and Generative Adversarial Networks (GANs), and proposed adaptation tailored to discrete genotype data. We conducted extensive experiments on large-scale datasets, including all chromosomes from cow and multiple chromosomes from human. Model performance was assessed using a well-established set of metrics drawn from both deep learning and quantitative genetics literature. Our results show that these models can effectively capture genetic patterns and preserve genotype-phenotype association. Our findings provide a comprehensive comparison of these models and offer practical guidelines for future research in genotype simulation. We have made our code publicly available at https://github.com/SihanXXX/DiscreteGenoGen.

bioinformatics↗

Unravelling the genetic architecture of persistence in production, quality, and efficiency traits in laying hens at late production stages

BackgroundThe laying hen industry aims to extend production for the economic and environmental benefits it offers, while at the same time facing the challenges of declining egg production and quality in aging hens. To explore trait persistence, we studied 998 Rhode Island Red purebred hens from the Novogen nucleus. We recorded daily egg production from 70 to 92 weeks of age and measured individual feed intake twice a week for three weeks, starting at 70, 80, and 90 weeks, as well as body weight at the start and at the end of each feed intake recording period. Random regression models were used to study trait trajectories over time, and PCA and hierarchical clustering were applied to identify groups of hens based on estimated breeding values for the intercept and slope of traits trajectory. ResultsResults showed different aging trajectories among traits. Daily body weight variation, Feed conversion ratio, Haugh unit and yolk percentage showed persistence (i.e., stability) over the measured period. On the contrary, daily feed intake, residual feed intake, laying rate, egg mass, eggshell breaking strength and stiffness decreased over time, while body weight, mean egg weight and eggshell colour increased. To assess the feasibility of selecting for trait persistence, we estimated the genetic variance of the slope and its correlation with the intercept. We found that, for egg weight and eggshell colour, genetic variance of the slope was negligible, indicating that selection for persistence on these traits requires other means. On the contrary, the slope for other traits such as laying rate and residual feed intake showed significant additive genetic variance. Strong genetic correlations between trait estimates at different ages were also observed and heritabilities estimates were low to high depending of the traits and period. ConclusionThe study explores hens trait persistence from 70 to 92 weeks, suggesting potential for improved egg production persistence. Challenges arise from low genetic variances impacting the efficiency of the potential selection on persistence. Clustering analysis reveals distinctive response patterns to elongation of production and underlined that selecting for enhanced persistence of different traits will necessitate compromises in breeding goals.

genetics↗

Predicting nonlinear genetic relationships between traits in multi-trait evaluations by using a GBLUP-assisted Deep Learning model

BackgroundGenomic prediction aims to predict the breeding values of multiple complex traits, usually assumed to be normally distributed by the largely used statistical methods, thus imposing linear genetic correlations between traits. While statistical methods are of great value for genomic prediction, these methods do not account for nonlinear genetic relationships between traits. If such relationships exist, although statistical models do perform a fair linear approximation, their prediction accuracy is limited due to the nonlinearity. Deep learning (DL) is a promising methodology for predicting multiple complex traits, in scenarios where nonlinear genetic relationships are present, due to its capacity to capture complex and nonlinear patterns in large data. We proposed a novel hybrid DLGBLUP model which uses the output of the traditional GBLUP, and enhances its PGV by accounting for nonlinear genetic relationships between traits using DL. Using simulated data, we compared the accuracy of the PGV obtained with the proposed hybrid DLGBLUP model, a DL model, and the traditional GBLUP model - the latter being our baseline reference. ResultsWe found that both DL and DLGBLUP models either outperformed GBLUP, or presented equally accurate PGV, with a particular greater accuracy for traits presenting a strongly characterized nonlinear genetic relationship. Overall, DLGBLUP presented the highest prediction accuracy, up to 0.2 points higher than GBLUP, and smallest mean squared error of the PGV for all traits. Additionally, we evolved a base population over seven generations and compared the genetic progress when selecting individuals based on the additive PGV obtained by either DL, DLGBLUP or GBLUP. For all traits with a nonlinear genetic relationship, after the fourth generation, the observed genetic gain when selection was based on the additive PGV from GBLUP was always inferior to the one achieved from either DL or DLGBLUP. ConclusionsThe integration of DL into genomic prediction enables the possibility of modeling nonlinear relationships between traits. Moreover, by identifying these nonlinear genetic relationships, our DL and DLGBLUP models improved prediction accuracy, when compared to GBLUP. The possibility of nonlinear relationships between traits offers a different perspective into multi-trait evaluations and prediction, as well as into the traits evolution over generations, with potential to further improve selection strategies in commercial livestock breeding programs. Moreover, DLGBLUP shows that DL can be used as a complement to statistical methods, by enhancing their performance.

genomics↗