Search bioRxiv⌕ Search

Biology subjects

Breimann, S.

Publications and source records attributed to Breimann, S..

2 recordsLinked to original sources

AAclust: k-optimized clustering for selecting redundancy-reduced sets of amino acid scales

SummaryAmino acid scales are crucial for sequence-based protein prediction tasks, yet no gold standard scale set or simple scale selection methods exist. We developed AAclust, a wrapper for clustering models that require a pre-defined number of clusters k, such as k-means. AAclust obtains redundancy-reduced scale sets by clustering and selecting one representative scale per cluster, where k can either be optimized by AAclust or defined by the user. The utility of AAclust scale selections was assessed by applying machine learning models to 24 protein benchmark datasets. We found that top-performing scale sets were different for each benchmark dataset and significantly outperformed scale sets used in previous studies. Notably, model performance showed a strong positive correlation with the scale set size. AAclust enables a systematic optimization of scale-based feature engineering in machine learning applications. Availability and implementationThe AAclust algorithm is part of AAanalysis, a Python-based framework for interpretable sequence-based protein prediction, which will be made freely accessible in a forthcoming publication. ContactStephan Breimann (Stephan.Breimann@dzne.de) and Dmitrij Frishman (dimitri.frischmann@tum.de) Supplementary informationFurther details on methods and results are provided in Supplementary Material.

bioinformatics↗

AAontology: An ontology of amino acid scales for interpretable machine learning

Amino acid scales are crucial for protein prediction tasks, many of them being curated in the AAindex database. Despite various clustering attempts to organize them and to better understand their relationships, these approaches lack the fine-grained classification necessary for satisfactory interpretability in many protein prediction problems. To address this issue, we developed AAontology--a two-level classification for 586 amino acid scales (mainly from AAindex) together with an in-depth analysis of their relations--using bag-of-word-based classification, clustering, and manual refinement over multiple iterations. AAontology organizes physicochemical scales into 8 categories and 67 subcategories, enhancing the interpretability of scale-based machine learning methods in protein bioinformatics. Thereby it enables researchers to gain a deeper biological insight. We anticipate that AAontology will be a building block to link amino acid properties with protein function and dysfunctions as well as aid informed decision-making in mutation analysis or protein drug design.

bioinformatics↗