Search bioRxiv⌕ Search

Biology subjects

Lubega, J.

Publications and source records attributed to Lubega, J..

3 recordsLinked to original sources

EffectorGeneP: accurate gene annotation in pathogen genomes from infection transcriptomes

Accurate gene annotation is crucial for inference of biological knowledge from genomes. However, non-canonical genes such as orphan or single-exon genes as well as those residing in rapidly evolving regions are routinely dismissed in annotation pipelines. In filamentous pathogen genomes, this disproportionately affects the annotation of genes encoding disease-promoting effector proteins. We introduce EffectorGeneP, a machine learning tool that self-trains on transcript data, predicts the most likely coding sequence from transcripts and effectively separates bona fide genes from transcriptional noise. EffectorGeneP annotates over 95% of known effectors correctly, while other state-of-the-art methods annotate 15%-78%. We show that EffectorGeneP expands the predicted secretome of pathogens by over 50% and that high-throughput screening of an effector library in plant protoplasts uncovers the previously poorly annotated AvrSr26 gene family in the wheat stem rust fungus. EffectorGeneP decodes genomes at unprecedented resolution and will enable the study of biological processes in important pathogen species.

microbiology↗

Haplotype-phased genomes of the barley leaf rust pathogen reveal evidence of repeat element expansion and somatic hybridization

Barley leaf rust disease, caused by Puccinia hordei, leads to substantial yield losses and diminished malting quality of barley across temperate growing regions worldwide. To address the paucity of high-resolution genomic resources for this pathogen, we generated haplotype-phased, chromosome-scale assemblies for ten globally distributed isolates using PacBio HiFi and Hi-C sequencing. Phylogenomic analysis revealed seven distinct lineages of P. hordei, including evidence of nuclear exchange, with a shared nuclear haplotype detected between two US lineages. Nuclear genome sizes ranged from [~]140-147Mbp, with the exception of isolate 90ISR03 from Israel ([~]163Mbp), which also harbored a 6.2Mbp supernumerary scaffold in one nucleus exhibiting chromosomal characteristics. Consistent with its larger genome, P. hordei had a higher repeat content ([~]70%) than related cereal rust fungi, driven primarily by the proliferation of LTR retroelements and DNA transposons. Across the global pan-genome of 13 unique nuclear haplotypes, approximately one-third of all protein orthogroups were conserved across all isolates. Only 18% of predicted effector orthogroups were shared between all haplotypes, reflecting the highly dynamic and variable nature of the effector repertoire. The long-term propagation of clonal P. hordei lineages is apparent both within the US and globally, and nuclear exchange is important for generating novel diversity and virulence profiles. The high degree of genome plasticity is evident in extensive structural variation, including large-scale translocations and inversions as well as a putative accessory chromosome. These chromosome-level, haplotype-resolved genomes provide a foundational resource for exploring the evolution, diversity, and avirulence gene repertoire of P. hordei.

genomics↗

Haplotype-resolved genomes of diverse oat crown rust isolates reveal both global dispersion of long-lived clonal haplotypes and limited recombination between haplotypes.

Genetic diversity of pathogen populations plays an important role in host adaptation. While single high quality genome references are a valuable resource, a compilation of genome references from individuals increases the breadth of genetic variation within a species that is captured. Puccinia coronata f. sp. avenae (Pca), a fungal pathogen causing crown rust of oat, demonstrates rapid virulence evolution and adaptation to newly released cultivars. To broaden the geographic and temporal distribution of available Pca genomes, we generated nuclear haplotype-resolved genome references for ten isolates from Europe, Africa and the Middle East and compared these with existing references for USA and Australian isolates. Of the full collection of 52 haplotypes, 40 were unique. Importantly, the presence of a nearly identical haplotype in a UK isolate collected in 1984 and in USA isolates from 1990 and 2017 supports the existence of long-lived clonal haplotypes in the global population that have been exchanged between lineages. Taken together with the identification of infrequent recombination between haplotypes from geographically dispersed isolates, this evidence reflects a globally mobile population of Pca that is mostly comprised of persistent clonal lineages with some influence from rare recombination events. Analysis of the core and non-core proteome suggests that while the core proteome is enriched for predicted secreted and effector proteins, sequence and expression variation are most prevalent in non-core orthogroups. We anticipate that this expanded collection of haplotypes will facilitate the development of new surveillance technologies and identification of virulence loci.

genomics↗