Researchers tested an alternative to hardcoding clinical calculators: the model writes case-specific Python that a restricted local executor runs as a deterministic solver. On MedCalc-Bench Verified (1,100 cases, 55 calculators), with formulas and gold variables supplied, the approach helped Qwen2.5-32B-AWQ (90.53% vs 83.47%, +7.05 points, 95% CI [0.47, 14.60]) but not Qwen2.5-7B (75.31% vs 72.02%, +3.29 points, 95% CI [-3.49, 10.38]).

A hand-written 22-calculator library was exact on its 440 supported cases but abstained elsewhere, scoring 40.0% overall. The authors audited the benchmark's formulas against current clinical guidelines and flagged 16 of 55 with version, use or coefficient concerns.