LearnStop — when learned exit rules beat confidence thresholds
Reasoning models waste compute on easy problems. LearnStop learns when to stop mid-inference by probing hidden states at checkpoint intervals and predicting correctness from answer confidence, entropy, vote stability, and other online signals.
Tested across 18 task-model pairs (GSM8K, MATH-500, MMLU-Pro, AIME-90, GPQA, Qwen3, DeepSeek-R1 distillations) — learned stopping rules beat simple baselines on free-form math but remain task-dependent. The efficiency gains vary by problem type and model.