Search bioRxiv⌕ Search

Biology subjects

Yildiz, G.

Publications and source records attributed to Yildiz, G..

4 recordsLinked to original sources

Integrative epigenomic analysis uncovers asymmetry of enhancer activity in Brassica napus

Non-coding regulatory regions are essential to the determination of gene expression and plant phenotypes. In this work, we investigated the cis-regulatory landscape of a winter type rapeseed, Express617, across multiple sample types. Combining chromatin accessibility, DNA methylation and gene expression, we annotated thousands of novel regulatory elements in the Brassica napus genome. Among those regions, we discovered and functionally characterized super-enhancers, observing an asymmetrical distribution of these regulatory elements favoring the Cn subgenome. Super-enhancer (SE) associated genes were found enriched in tissue identity and responses to stimuli related processes. We further establish and apply an in-silico validation pipeline for super-enhancers, integrating population-level expression analysis and machine learning (ML) models predicting gene expression levels. Almost 50% of the newly identified SE-associated genes had an observed expression higher than the expression levels predicted by the ML model. Moreover, structural variants disrupting super-enhancer elements correlate with a reduction of expression in the associated genes, both consistent with the positive effect of these regulatory regions. These results greatly expand the functional annotation of rapeseed and contribute to a better understanding of the link between regulatory elements and their target genes and processes, providing novel insights and targets for B. napus (epi)-genome editing strategies.

plant biology↗

Machine learning and multi-omic analysis reveal contrasting recombination landscape of A and C subgenomes of winter oilseed rape

Meiotic recombination is essential for generating genetic diversity, driving plant evolution, and enabling crop improvement, yet its uneven distribution across genomes constrains breeding efforts. Here, we investigated the multi-omic landmarks that shape the recombination landscape in Brassica napus by integrating epigenomic, genomic and transcriptomic data with recombination maps derived from large multi-parental rapeseed populations. Predictive machine-learning accurately predicted recombination rates and hotspot location using only feature information. Recombination was generally suppressed in centromeres and other repeat-rich, methylated regions and enriched in gene-dense, transcriptionally active domains. Proxies for chromatin configuration--such as DNA methylation, transposable elements or genes-- consistently achieved the highest predictive power with the random forest algorithm. We discovered distinct recombination landscape patterns between subgenomes, with crossovers clustering near subtelomeres in the A subgenome and more evenly spread across the C subgenome. Models trained on A-subgenome data outperformed those based on the C subgenome, although combining both subgenomes improved overall accuracy.

bioinformatics↗

Graphical pangenomics-enabled characterisation of structural variant impact on gene expression in Brassica napus

Structural variants (SVs, eg. insertions and deletions) are genomic variations > 50 bp that are known to be associated with a range of crop traits, from yield to flowering behaviour and stress responses. Recently, pangenome graphs have emerged as a powerful framework for analysing genomic data by encoding population- or species-level diversity in one data structure. Pangenome graphs have the potential to serve as unbiased references for downstream applications, including SV genotyping and pan-transcriptomic analyses. In this work, we hypothesized that extensive variation affects transcript quantification and expression quantitative trait locus (eQTL) analysis when relying on a single reference, and that using pangenome graphs can mitigate reference sequence bias. We combined long and short read whole genome sequencing data with expression profiling of Brassica napus (oilseed rape) to assess the impact of SVs on gene expression regulation and explored the utility of pangenome graphs for eQTL analysis. We demonstrate that pangenome graphs provides a superior framework for eQTL analysis by eliminating single reference bias in gene expression quantification. Combined with the graph-based genotyping of SVs, we identified 240 eQTL-SVs found in close proximity of target loci. These SVs affect expression of genes related to important traits, are often not in linkage with SNPs and represent diversity unaccounted for in classical SNP-based analyses. This study highlights the multiple advantages of graph-based approaches in population-scale studies and provides novel insight into gene expression regulation in an important crop.

plant biology↗

Benchmarking Oxford Nanopore Read Alignment-Based Structural Variant Detection Tools in Crop Plant Genomes

Structural variations (SVs) are larger polymorphisms (>50 bp in length), which consist of insertions, deletions, inversions, duplications, and translocations. They can have a strong impact on agronomical traits and play an important role in environmental adaptation. The development of long-read sequencing technologies, including Oxford Nanopore, allows for comprehensive SV discovery and characterization even in complex polyploid crop genomes. However, many of the SV discovery pipeline benchmarks do not include complex plant genome datasets. In this study, we benchmarked popular long-read alignment-based SV detection tools for crop plant genomes. We used real and simulated Oxford Nanopore reads for two crops, allotetraploid Brassica napus (oilseed rape) and diploid Solanum lycopersicum (tomato), and evaluated several read aligners and SV callers across 5x, 10x, and 20x coverages typically used in re-sequencing studies. Our benchmarks provide a useful guide for designing Oxford Nanopore re-sequencing projects and SV discovery pipelines for crop plants.

bioinformatics↗