Search bioRxiv⌕ Search

Biology subjects

Ghelich, R.

Publications and source records attributed to Ghelich, R..

2 recordsLinked to original sources

RulePep: Interpretable ESM-Guided Neural-Symbolic Peptide Classification

Peptides are increasingly explored as therapeutic candidates, delivery vectors, and functional biomolecules, but experimental screening of peptide activity and safety remains costly because the sequence space is vast and small sequence changes can alter functionality. Computational peptide classification can therefore help prioritize candidates. However, many protein-language-model-based classifiers achieve strong performance using opaque prediction heads, making it difficult to determine which learned evidence supports or opposes a prediction. We present RulePep, an ESM-2-guided neural-symbolic classifier for peptide-function prediction. RulePep maps frozen ESM-2 sequence representation to learned latent predicates, polarity-constrained differentiable rules, and an additive symbolic logit whose components can be inspected at the case level. We evaluate RulePep on three biologically distinct peptide classification tasks: blood-brain barrier penetration, hemolytic potency, and anticancer activity. On the BBPpredict, HemoPI3, and AntiCP 2.0 alternate benchmark datasets, RulePep achieved AUROC/MCC values of 0.8869/0.6850, 0.9155/0.6820, and 0.9765/0.8633, respectively. Ablation experiments supported the contributions of multi-layer representation pooling, rule polarity, mined-rule initialization, symbolic capacity, and rule-derived aggregation. RulePep combines competitive predictive performance with additive logit reconstruction, rule-level evidence reporting, and predicate-suppression auditing, providing a transparent sequence-based framework for peptide candidate prioritization.

bioinformatics↗

Modularity, Limited Sparsity, and Extensive Pleiotropy in a Genotype-Phenotype-Fitness Map of Drug Resistance

Understanding how the myriad molecular impacts of mutation percolate to influence higher-order traits and ultimately fitness requires compressing a many-to-many mapping into something tractable. Decades of theoretical work suggest this may be possible because biological systems are modular: effects of perturbation are often funneled through particular pathways or subsystems rather than propagating freely through the organism. Yet few empirical systems have been able to demonstrate such modularity at scale. Here we show that, even across 774 diverse yeast lineages, fitness variation across 12 drug environments is organized by a strikingly low-dimensional structure defined by only a few inferred phenotypic axes that capture the main patterns of variation. Lineages drawn from multiple evolutionary histories reveal more of these phenotypic axes than those derived from a single selection pressure. Consistent with many of these lineages having evolved under strong selection pressure, their mutations often exhibit broad pleiotropy, affecting nearly all inferred phenotypic axes. However, fitness in any given drug depends on a much sparser subset of the phenotypic modules these axes reflect. By compressing many-to-many relationships, this low-dimensional framework exposes the modular phenotypic space, as well as the context-dependent contribution of each phenotypic module to fitness that can constrain the pleiotropic effects of adaptive mutations. It also highlights that the apparent complexity of genotype-phenotype-fitness maps depends not only on environmental context but also on the diversity of mutations through which they are observed, laying the groundwork for identifying the key phenotypic modules that matter for fitness.

evolutionary biology↗