MODEL EVALUATION / LOG-UNIFORM VALUES

What changes when the model gets a calculator?

Direct and calculator-assisted answers compared across the log-uniform calculation set by accuracy, operation type, and scientific scale.

Model under testGPT-4.1 minigpt-4.1-mini

The same model is used in both evaluation modes.

Intelligence Index15Artificial Analysis
Input price$0.40per 1M tokens
Output price$1.60per 1M tokens
OpenAI model details and pricing
Loading comparison…