Back to blog
AI News

OpenAI reports two API settings that tripled ARC AGI 3 scores, what UK businesses should test next

A model performance update described by OpenAI says two API settings boosted GPT 5.6 results on ARC AGI 3. If you run reasoning heavy workflows, the practical next step is to run a controlled evaluation, compare quality and efficiency, and update your production settings only after you see the same gains on your own tasks.

What changed in the model API settings

OpenAI describes that adjusting two API settings for GPT 5.6 materially improved performance on the ARC AGI 3 benchmark. The update focuses on keeping reasoning active and allowing compaction, which together drove higher scores and improved efficiency on that benchmark.

Why this matters for business teams

If your business uses AI for tasks where the model needs structured reasoning, small changes in API settings can shift both quality and cost. OpenAI signals that the same direction can also improve efficiency, which means you may be able to reach the same outcome with less spend if your workload benefits from the updated behavior.

The practical adoption plan, test before you roll out

Before changing production, run a controlled evaluation against a representative slice of your real customer and operations tasks. Compare your current configuration to the updated settings described by OpenAI, then measure both outcome quality and operational efficiency. Only roll changes forward when the improvements hold up on your own metrics and user experience.

  • Select a test set that mirrors your real workflow inputs and outputs
  • Run side by side trials using your current settings and the two settings described by OpenAI
  • Track both quality and efficiency using the same evaluation framework your team already trusts
  • Update production prompts and routing only after results are stable and explainable to stakeholders

What to document for risk control

Treat settings changes as operational changes. Document the rationale for the two settings, the evaluation results you used, and the rollback plan. This helps you manage model risk and keeps product, operations, and compliance aligned when you adjust how reasoning is invoked.

OpenAI reports two API settings that tripled ARC AGI 3 scores, what UK businesses should test next | New Era AI