Search bioRxiv⌕ Search

Biology subjects

Aranzana, M. J.

Publications and source records attributed to Aranzana, M. J..

4 recordsLinked to original sources

GenoDrawing: An autoencoder framework for image prediction from SNP markers

Advancements in genome sequencing have facilitated whole genome characterization of numerous plant species, providing an abundance of genotypic data for genomic analysis. Genomic selection and neural networks, particularly deep learning, have been developed to predict complex traits from dense genotypic data. Autoencoders, a neural network model to extract features from images in an unsupervised manner, has proven to be useful for plant phenotyping. This study introduces an autoencoder framework, GenoDrawing, for predicting and retrieving apple images from a low-depth single nucleotide polymorphism (SNP) array, potentially useful in predicting traits that are difficult to define. GenoDrawing demonstrated proficiency in its task while using a small dataset of shape-related SNPs, and multiple experiments were conducted to evaluate the impact of SNP selection and shape relation. Results indicated that the correct relationship of SNPs with visual traits had a significant impact on the generated images, consistent with biological interpretation. While using significant SNPs is crucial, incorporating additional, unrelated SNPs results in performance degradation for simple NN architectures that cannot easily identify the most important inputs. The proposed GenoDrawing method is a practical framework for exploring genomic prediction in fruit tree phenotyping, particularly beneficial for small to medium breeding companies to predict economically significant heritable traits. Although GenoDrawing has limitations, it sets the groundwork for future research in image prediction from genomic markers. Future studies should focus on using stronger models for image reproduction, SNP information extraction, and improved dataset balance in terms of shape for more precise outcomes.

plant biology↗

An LTR retrotransposon in the promoter of a PsMYB10.2 gene associated with the regulation of fruit flesh color in Japanese plum

Japanese plums exhibit wide diversity of fruit coloration. The red to black hues are caused by the accumulation of anthocyanins, while their absence results in yellow, orange or green fruits. In Prunus, MYB10 genes are determinants for anthocyanin accumulation. In peach, QTLs for red plant organ traits map in an LG3 region with three MYB10 copies (PpMYB10.1, PpMYB10.2 and PpMYB10.3). In Japanese plum the gene copy number in this region differs with respect to peach, with at least three copies of PsMYB10.1. Polymorphisms in one of these copies correlate with fruit skin color. The objective of this study was to determine a possible role of LG3-PsMYB10 genes in the natural variability of the flesh color trait and to develop a molecular marker for marker-assisted selection (MAS). We explored LG3-PsMYB10 variability, including the analysis of long-range sequences obtained in previous studies through CRISPR-Cas9 enrichment sequencing. We found that the PsMYB10.2 gene was only expressed in red flesh fruits. Its role in promoting anthocyanin biosynthesis was validated by transient overexpression in Japanese plum fruits. The analysis of long-range sequences identified an LTR retrotransposon in the promoter of the expressed PsMYB10.2 gene that explained the trait in 93.1% of the 145 individuals analyzed. We hypothesize that the LTR retrotransposon may promote the PsMYB10.2 expression and activate the anthocyanin biosynthesis pathway. We provide a molecular marker for the red flesh trait which, together with that for skin color, will serve for the early selection of fruit color in breeding programs.

molecular biology↗

An efficient CRISPR-Cas9 enrichment sequencing strategy for characterizing complex and highly duplicated genomic regions. A case study in the Prunus salicina LG3-MYB10 genes cluster.

Genome complexity is largely linked to diversification and crop innovation. Examples of regions with duplicated genes with relevant roles in agricultural traits are found in many crops. In both duplicated and non-duplicated genes, much of the variability in agronomic traits is caused by large as well as small and middle scale structural variants (SVs), which highlights the relevance of the identification and characterization of complex variability between genomes for plant breeding. Here we improve and demonstrate the use of CRISPR-Cas9 enrichment combined with long-read sequencing technology to resolve the MYB10 region in the linkage group 3 (LG3) of Japanese plum (Prunus salicina), which has a length from 90 kb to 271 kb according to the P. salicina genomes available. We demonstrate the high complexity of this region, with homology levels between Japanese plum varieties comparable to those between Prunus species. We cleaved MYB10 genes in five plum varieties using the Cas9 enzyme guided by a pool of crRNAs. The barcoded fragments were then pooled and sequenced in a single MinION Oxford Nanopore Technologies (ONT) run, yielding 194 Mb of sequence. The enrichment was confirmed by aligning the long reads to the plum reference genomes, with a mean read on-target value of 4.5% and a depth per sample of 11.9x. From the alignment, 3,261 SNPs and 287 SVs were called and phased. A de novo assembly was constructed for each variety, which also allowed detection, at the haplotype level, of the variability in this region. CRISPR-Cas9 enrichment is a versatile and powerful tool for long-read targeted sequencing even on highly duplicated and/or polymorphic genomic regions, being especially useful when a reference genome is not available. Potential uses of this methodology as well as its limitations are further discussed.

molecular biology↗

Genetic architecture and genomic prediction accuracy of apple quantitative traits across environments

Implementation of genomic tools is desirable to increase the efficiency of apple breeding. The apple reference population (apple REFPOP) proved useful for rediscovering loci, estimating genomic prediction accuracy, and studying genotype by environment interactions (GxE). Here we show contrasting genetic architecture and genomic prediction accuracies for 30 quantitative traits across up to six European locations using the apple REFPOP. A total of 59 stable and 277 location-specific associations were found using GWAS, 69.2% of which are novel when compared with 41 reviewed publications. Average genomic prediction accuracies of 0.18-0.88 were estimated using single-environment univariate, single-environment multivariate, multi-environment univariate, and multi-environment multivariate models. The GxE accounted for up to 24% of the phenotypic variability. This most comprehensive genomic study in apple in terms of trait-environment combinations provided knowledge of trait biology and prediction models that can be readily applied for marker-assisted or genomic selection, thus facilitating increased breeding efficiency.

genomics↗