Search bioRxiv⌕ Search

Biology subjects

Sanchez, M.-P.

Publications and source records attributed to Sanchez, M.-P..

6 recordsLinked to original sources

Integrating Structural Variants into Sequence-Based GWAS Using a Pangenome and Imputation Framework in French Dairy Cattle

BackgroundStructural variants (SVs) are most effectively identified using long-read (LR) sequenc-ing. However, such data remain scarce, and sequenced samples often lack associated phenotypic information. To overcome this limitation, we integrated pangenome-based (variation graph-based) and imputation approaches to enable large-scale SV association studies in the three main French dairy cattle breeds. ResultsA variation graph was constructed using 69,892 deletions, 89,900 insertions, and 17,402 duplications detected in 176 LR samples. We subsequently genotyped 939 samples for each SV in the panel by realigning their short read (SR) sequences to the graph. Validation analyses showed high genotype concordance rates for deletions (0.79) and insertions (0.79); however, concordance for duplications was low (0.14), leading to their exclusion from further analyses. The retained SVs were combined with single nucleotide variants (SNVs) to build a sequence-level imputation reference panel. Using SNP genotyping array data, we imputed SVs and SNVs for 11,902 Holstein, 3,753 Montbeliarde, and 3,053 Normande bulls. After quality control, more than 14 million SNVs and 40 thousand SVs were retained for within-breed genome-wide association studies (GWAS) us-ing daughter yield deviations for stature and four milk production and composition traits. The GWAS results reveled genetic architectures consistent with previous findings and identified 40 genome-wide significant associations between structural variant and key phenotypes. Conditional analyses showed that ten of these SVs as strong candidates associated with milk fat and protein contents, as well as stature. ConclusionsBy integrating LR, SR, and SNP genotyping data within a unified pangenome and imputation framework, we demonstrate a scalable strategy to systematically interrogate the contribution of SVs to complex traits. The resulting genetic architectures were highly consistent with previous findings, validating both the robustness and transferability of our approach. Our findings highlight the added value of integrating SVs into routine genomic analyses and provide a scalable framework for incorporating SVs into genomic selection in dairy cattle.

genomics↗

Assembly of a pangenome uncovers novel non-reference unique insertion sequences in cattle highlighting their genetic diversity

BackgroundThe current cattle reference genome, derived from a single Hereford cow, does not capture the full spectrum of genetic diversity present within the species. Moreover, detecting structural variations (SVs [≥] 50 nucleotides long) remains challenging using only standard approaches of either short or long-read sequence approaches against a linear reference genome. Recent advances in long-read sequencing technologies and graph-based assembly now enable the construction of breed-specific pangenomes, revealing previously uncharacterized genomic regions that may contribute to important agricultural traits. ResultsIn this study we constructed a cattle pangenome graph using 16 high-quality haplotype-resolved genome assemblies originating from nine breeds representing the diversity of French cattle populations, and including Yak (Bos grunniens) as a close outgroup species. Using a trio-based strategy combined with complementary sequencing technologies and bioinformatics methods, we identified and characterized 101,219 structural variations. Of these, 33,634 were classified as non-reference unique insertions (NRUIs), adding several megabases of novel genomic sequences absent from the current Hereford reference genome. Analysis of the distribution of these NRUIs revealed significant genome-wide enrichment within QTL regions associated with milk production and morphological traits, suggesting their contribution to the genetic basis of economically relevant phenotypes. Furthermore, their functional annotation highlighted two NRUIs located within the intronic regions of ARMH3 and EPHA5, both specific to the Normande breed and significantly associated with milk production and morphological traits, respectively. ConclusionsOur findings demonstrate the value of pangenome approaches to uncover functionally relevant SVs, particularly NRUIs, that are systematically not in the current reference genome. By linking these variants to economically important traits, our work underscores the need to incorporate breed diversity into future genomic analyses and reference-building efforts in cattle.

genetics↗

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

genomics↗

Comprehensive detection of structural variations in long and short reads dataset of French cattle

Structural variants (SVs) correspond to different types of genomic variants larger than 50 bp. Many findings suggest the use of long rather than short reads to improve the accuracy of SV detection. Here, we present the results of an in-depth analysis for detection of SVs, mainly large insertions and deletions, in 14 French bovine breeds, based on whole-genome data comprising 176 long-read and 571 short-read samples, with 154 individuals having both long- and short-read data available. We first investigated possible biases on the performances of well-known SV detection tools, namely CUTESV, PBSV, and SNIFFLES, using long reads from different technologies, including PacBio HiFi, Oxford ONT, and PacBio CLR. We subsequently highlighted the abilities of tools for detecting SVs (DELLY, LUMPY, and MANTA) and for genotyping known SVs (GRAPHTYPER, SVTYPER, PARAGRAPH, and VG toolkit) using short-read data. We then show how the incremental composition of samples in the reference panel affected the SV genotyping for six validation individuals sequenced in short reads. We then searched for the optimal parameters and created the final SV reference panel consisting of 25,191 deletions and 30,118 insertions. Finally, we emphasized the landscape of the genotyped SVs segregating across 571 short-read individuals of 14 breeds.

genomics↗

Application of a French cattle pangenome, from structural variant discovery to association studies on key phenotypes

BackgroundThe current cattle reference genome assembly, a pseudo-linear sequence produced using sequences from a single Hereford cow, represent a limit when performing genetic studies, especially when investigating the whole spectrum of genetic variations within the species. Detecting structural variations (SVs) poses significant challenges when relying solely on conventional methods of short or long-read sequence mapping to the current bovine genome assembly. ResultsIn this study, we used long-reads (LR) and bioinformatic tools to construct a comprehensive bovine pangenome incorporating genetic diversity of 64 good quality de novo genome assemblies representing 14 French dairy and beef cattle breeds. Using a combination of complementary approaches, we explored the pangenome graph and identified 2.563 Gb of sequences common to all samples, and cumulated 0.295 Gb of variable sequences. Notably, we discovered 0.159 Gb of novel sequences not present in the current Hereford reference genome assembly. Our analysis also revealed 109,275 SVs, of which 84,612 were bi-allelic, including 21,840 insertions and 21,340 deletions. Genome-wide association studies using SNPs and a panel of 221 SVs, shared between the pangenome and the EuroGMD chip, revealed several well-known QTLs across the genome for the Holstein, Montbeliarde and Normande breeds. Among those, a QTL on chromosome 11 presents an SV with a highly significant effect on stature in the Holstein breed. This SV is a 6.2 kb deletion affecting the 5UTR, first exon and part of first intron of MATN3 gene, suggesting a potential regulatory and coding effect. ConclusionsOur study provides new insights into the genetic diversity of 14 French dairy and beef breeds and highlights the utility of pangenome graphs in capturing structural variation. The identified SV associated with stature highlights the importance of integrating SVs into GWAS for a more comprehensive understanding of complex traits.

genetics↗

Characterization of bovine vaginal microbiota and its relationship with host fertility, health, and production

BackgroundBecause of its potential influence on the hosts phenotype, increasing attention is paid to organ-specific microbiota in several animal species, including cattle. However, ecosystems other than those related to the digestive tract remain largely understudied. In particular, little is known about the vaginal microbiota of ruminants despite the importance of the reproductive functions of cows in a livestock context, where fertility disorders represent one of the primary reasons for culling. ResultsIn the present study, we aimed at better characterizing the vaginal microbiota of dairy cows through 16S rRNA sequencing, using a large cohort of Holstein cows from Northern France. Our results allowed to define a core microbiota of the dairy cows vagina, and highlighted that 90% of the sequences belonged to the Firmicutes, the Proteobacteria, and the Bacteroidetes phyla. The core microbiota was composed of four phyla, 16 families, 14 genera and only one amplicon sequence variant (ASV), supporting the idea of the high diversity of vaginal microbiota within the studied population. This variability was partly explained by various environmental factors such as the herd, the sampling season, the lactation rank and the lactation stage. In addition, we investigated potential associations between the diversity and the composition of the vaginal microbiota and several health-, performance-, and fertility-related phenotypes. Our analyses highlighted significant associations between the and {beta}- diversities and several traits including the first insemination outcome, the productive longevity, and the culling. Besides, relevant phenotypes were correlated with the abundance of several genera, some of which, such as Leptotrichia, Streptobacillus, Methylobacterium-Methylorubrum, or Negativibacillus, were linked to multiple traits. ConclusionConsidering the large number of samples, which were collected in commercial farms, and the diversity of the phenotypes considered, this study represents a first step towards a better understanding of the close relationship between the vaginal and the dairy cows phenotypes.

microbiology↗