bioRxiv · 10.64898/2026.08.11.744303
Towards a Physiological Scaling Law: Model Quality vs. Cohort Size for Stochastic Sequence Data
Abstract
Scaling laws help determine the optimal data size for training large models but are established in domains where the target is deterministic. Physiological signals are different: heartbeat sequences are stochastic, so part of the error is irreducible even with large amounts of data. Metrics such as MAE do not account for non-deterministic behavior, and therefore assessing scaling requires evaluating distributional calibration (measuring how well predicted probability densities capture true conditional characteristics). We formulate a scaling law metric(n) = E + A n- and evaluate it with five metrics: accuracy (MAE, RMSE), distributional calibration (KS distance, goodness-of-fit), and training objective (negative log loss) using a neural temporal point process trained on a cohort of four-ECG datasets. The law fits all five metrics. While point accuracy is near saturation at n = 183, KS distance and goodness-of-fit improve by 6% and 12% respectively when extrapolated to 10,000 subjects, showing that scaling decisions in stochastic domains must be guided by distributional calibration rather than point accuracy.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sunil, G., Kumar, B. R., Ramsundar, B., Subramanian, S.. 2026-08-20. Towards a Physiological Scaling Law: Model Quality vs. Cohort Size for Stochastic Sequence Data. https://doi.org/10.64898/2026.08.11.744303
Cite the original work for its findings. Save a collection to share your selection of sources.