Search bioRxiv⌕ Search

Biology subjects

Damaris, B. F.

Publications and source records attributed to Damaris, B. F..

2 recordsLinked to original sources

Cis non-coding genetic variation drives gene expression changes in the E. coli and P. aeruginosa pangenomes

Bacteria use gene regulation to dynamically adapt to changes in their environment, including resistance to stress and the occupation of new niches. Gene expression is known to vary within a species pangenome, but the extent to which these changes could be explained by genetic variants in cis non-coding regions has so far been poorly investigated. Statistical genetics offers a hypothesis-free approach to this problem, as opposed to mechanistic models, which can be used only for reference isolates that are not representative of the whole species. In this study, we assembled two genomic and transcriptomic datasets for Escherichia coli (N=117) and Pseudomonas aeruginosa (N=413) and identified associations between genetic variants in cis non-coding regions and recorded gene expression variation. We identified at least one associated variant in up to 39% of the tested genes in both species. We partly validated the associations in-silico and in-vitro for E. coli, reinforcing the difficulty of identifying a single mechanism generating gene expression diversity. We then investigated the relevance of non-coding variants in explaining the variability in antimicrobial resistance in both species using two additional publicly available datasets, identifying a large number of these variants across antimicrobial compounds. This work confirms the role of genetic variation in often overlooked regions of bacterial genomes in influencing molecular and clinically relevant phenotypes.

microbiology↗

microGWAS: a computational pipeline to perform large scale bacterial genome-wide association studies

Identifying genetic variants associated with bacterial phenotypes, such as virulence, host preference, and antimicrobial resistance, has great potential for a better understanding of the mechanisms involved in these traits. The availability of large collections of bacterial genomes has made genome-wide association studies (GWAS) a common approach for this purpose. The need to employ multiple software tools for data pre- and post-processing limits the application of these methods by experienced bioinformaticians. To address this issue, we have developed a pipeline to perform bacterial GWAS from a set of assemblies and annotations, with multiple phenotypes as targets. The associations are run using five sets of genetic variants: unitigs, gene presence/absence, rare variants (i.e. gene burden test), gene cluster specific k-mers, and all unitigs jointly. All variants passing the association threshold are further annotated to identify overrepresented biological processes and pathways. The results can be further augmented by generating a phylogenetic tree and by predicting the presence of antimicrobial resistance and virulence associated genes. We tested the microGWAS pipeline on a previously reported dataset on E. coli virulence, successfully identifying the causal variants, and providing further interpretation on the association results. The microGWAS pipeline integrates the state-of-the-art tools to perform bacterial GWAS into a single, user-friendly, and reproducible pipeline, allowing for the democratization of these analyses. The pipeline can be accessed, together with its documentation, at: https://github.com/microbial-pangenomes-lab/microGWAS.

bioinformatics↗