MODEL EVALUATION / LOG-UNIFORM VALUES
What changes when the model gets a calculator?
Direct and calculator-assisted answers compared across the log-uniform calculation set by accuracy, operation type, and scientific scale.
Model under testGPT-4.1 mini
gpt-4.1-miniThe same model is used in both evaluation modes.
OpenAI model details and pricingLoading comparison…