Search bioRxiv⌕ Search

Biology subjects

Sielemann, K.

Publications and source records attributed to Sielemann, K..

2 recordsLinked to original sources

plASgraph - using graph neural networks to detect plasmid contigs from an assembly graph

Identification of plasmids from sequencing data is an important and challenging problem related to antimicrobial resistance spread and other One-Health issues. In our work, we provide a new architecture for identifying plasmid contigs in fragmented genome assemblies built from short-read data. Unlike previous machine-learning approaches for this problem, which classify individual contigs separately, we employ graph neural networks (GNNs) to include information from the assembly graph. Propagation of information from nearby nodes in the graph allows accurate classification of even short contigs that are difficult to classify based on sequence features or database searches alone. Our new species-agnostic software tool plASgraph outperforms recently developed PlasForest, which uses database searches to supplement sequence-based features. Since our tool does not rely on existing plasmid databases, it is more suitable for classification of contigs in novel species and discovery of previously unknown plasmid sequences. Our tool can also be trained on a specific species, and in that scenario it outperforms mlplasmids trained on the same species. On one hand, our work provides a new, accurate, and easy to use tool for plasmid classification; on the other hand, it serves as a motivation for more widespread use of GNNs in bioinformatics, such as in pangenome sequence analysis, where sequence graphs serve as a fundamental data structure. Availabilityhttps://github.com/cchauve/plASgraph

bioinformatics↗

Complete pan-plastome sequences enable high resolution phylogenetic classification of sugar beet and closely related crop wild relatives

BackgroundAs the major source of sugar in moderate climates, sugar-producing beets (Beta vulgaris subsp. vulgaris) have a high economic value. However, the low genetic diversity within cultivated beets requires introduction of new traits, for example to increase their tolerance and resistance attributes - traits that often reside in the crop wild relatives. For this, genetic information of wild beet relatives and their phylogenetic placements to each other are crucial. To answer this need, we sequenced and assembled the complete plastome sequences from a broad species spectrum across the beet genera Beta and Patellifolia, both embedded in the Betoideae (order Caryophyllales). This pan-plastome dataset was then used to determine the wild beet phylogeny in high-resolution. ResultsWe sequenced the plastomes of 18 closely related accessions representing 11 species of the Betoideae subfamily and provided high-quality plastome assemblies which represent an important resource for further studies of beet wild relatives and the diverse plant order Caryophyllales. Their assembly sizes range from 149,723 bp (Beta vulgaris subsp. vulgaris) to 152,816 bp (Beta nana), with most variability in the intergenic sequences. Combining plastome-derived phylogenies with read-based treatments based on mitochondrial information, we were able to suggest a unified and highly confident phylogenetic placement of the investigated Betoideae species. Our results show that the genus Beta can be divided into the two clearly separated sections Beta and Corollinae. Our analysis confirms the affiliation of B. nana with the other Corollinae species, and we argue against a separate placement in the Nanae section. Within the Patellifolia genus, the two diploid species Patellifolia procumbens and Patellifolia webbiana are, regarding the plastome sequences, genetically more similar to each other than to the tetraploid Patellifolia patellaris. Nevertheless, all three Patellifolia species are clearly separated. ConclusionIn conclusion, our wild beet plastome assemblies represent a new resource to understand the molecular base of the beet germplasm. Despite large differences on the phenotypic level, our pan-plastome dataset is highly conserved. For the first time in beets, our whole plastome sequences overcome the low sequence variation in individual genes and provide the molecular backbone for highly resolved beet phylogenomics. Hence, our plastome sequencing strategy can also guide genomic approaches to unravel other closely related taxa.

genomics↗