bioRxiv · 10.64898/2026.05.14.725067
Protein solubility depends on centrifugation: Aiki-Sol, a per-regime predictor for E. coli
Abstract
MotivationSequence-based predictors of recombinant protein solubility in Escherichia coli have plateaued (NESG independent-test AUC 0.760 [->]~ 0.80 over eight years of protein-language-model variants). The plateau hides a latent confound: the centrifugation regime used to separate the soluble from the insoluble fraction is a hidden variable collapsed into a single binary "soluble" label. The proteins biochemistry does not change between regimes; what changes is which fraction of the lysate is recovered as soluble. Existing predictors treat the regime as label noise rather than a feature, and sequence overlap between training and test partitions masks the resulting failure mode. ResultsWe release the Aiki-Sol Dataset, a tiered E. coli solubility corpus: a ~ 85K stringency-annotated benchmark, an Apache-licensed ~ 147K extension adding binary-only-labelled proteins, and a ~ 229K research-tier pool incorporating non-commercially-licensed sources. On the ~ 85K benchmark, scored on sequence-cluster-disjoint partitions, the strongest published binary comparator falls below chance on the 32,000 x g stratum (AUC 0.491 {+/-} 0.020); a fine-tuned ESM-2 650M backbone with five protocol-matched out-puts lifts pooled AUC by +0.108 (paired-bootstrap CI lower bound +0.090). The gain is curation, not architecture: structure-aware predictors given ESMFold structures do not outperform the sequence-only frame, and capacity scaled to 3B parameters does not exceed the conditioned 650M backbone. The released model, Aiki-Sol, jointly supervises five per-stringency outputs alongside a marginal output for stringency-unknown proteins; on five external cohorts it lifts cohort-mean AUC from 0.69-0.70 to 0.825, with a [≥] +0.10-0.16 lift on the three cohorts at measurably-zero training-pool overlap. Availability and implementationAiki-Sol model weights (Apache 2.0), the 147K-row license-clean training pool of the deployment checkpoint (CC BY 4.0), the cluster-disjoint per-stringency 5-fold partition assignments, per-cohort prediction CSVs, and source code for training, inference, and figure reproduction are available at https://github.com/aikium-public/aiki-sol and archived at Zenodo 10.5281/zenodo.20151817. The research-tier 229K checkpoint is released under CC-BY-NC-ND 4.0 (inheriting the most-restrictive upstream-source tier); its training CSV and the 84,809-protein stringency-annotated bench-mark of [§]2.1 mix non-commercial-tier upstream sources and are not redistributed verbatim. Upstream sources are documented in Data availability and SI [§]S1. The deployment artefact is distributed as a Python package (pip install aikisol) with a predict(seq) entry point. Contactvenkatesh@aikium.com. Supplementary informationSupplementary text, figures, and tables are available at Bioinformatics online.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rajagopalan, R., Meda, R. S., Shastry, S., Mysore, V.. 2026-05-14. Protein solubility depends on centrifugation: Aiki-Sol, a per-regime predictor for E. coli. https://doi.org/10.64898/2026.05.14.725067
Cite the original work for its findings. Save a collection to share your selection of sources.