GPT-6 Astra produced two very different results on the same benchmark.
ARC Prize reports that Astra scored 62.7% on ARC-AGI-3 with its Standard harness. With OpenAI’s Provider Adapter harness, it scored 99.9%.[1]
The higher result should not be dismissed. ARC Prize describes both results as state of the art.[1]
The gap changes what buyers should evaluate. The model is no longer the whole product.
The Provider Adapter preserved opaque reasoning state between requests and used compaction for longer conversations. ARC Prize says it let the model reuse earlier work, while its Standard harness passed forward notes instead.[1] OpenAI also disclosed that its ARC-AGI-3 evaluation used a Responses API harness with two settings changed to better reflect what it considers real-world performance.[2]
So the result came from a system: model, state management, context handling, harness and test conditions working together.
That is the useful lesson for business operators.
Your company will not deploy a model in isolation
An SME does not buy intelligence as a number on a leaderboard. It buys an outcome inside a real workflow.
Perhaps the job is to qualify a sales lead, reconcile an invoice, update a CRM record, prepare a management report or flag a purchasing exception. In each case, the model is only one component.
The result will also depend on:
- what data the agent can access;
- what context it retains between steps;
- which tools it can use;
- what permissions it receives;
- how errors and exceptions are handled;
- what the run costs and how long it takes; and
- where a human must review or approve the work.
Change those conditions and you may change the outcome substantially, even when the underlying model stays the same.
Simon Willison highlighted this distinction in his early review of Astra, noting the 99.9% result with the custom adapter and the 62.7% result with the default harness. He also made the sensible point that he had not yet tested the model himself.[3]
That is good evaluation discipline: separate reported results from direct experience, then test the configuration that matters to you.
A benchmark is evidence, not a business case
ARC-AGI-3 is designed to test how an AI system explores unfamiliar interactive environments, works out their mechanics and plans its actions. It is valuable evidence about a particular kind of adaptive performance.[1]
It does not tell you whether an agent will process your supplier invoices correctly.
It does not know the quality of your ERP data, the exceptions in your approval policy or the way your team records customer commitments.
And it does not prove that the configuration available to your business is identical to the one that produced the headline result.
This is where buyers often make the wrong comparison. They line up model scores, choose the highest number and then discover that the production outcome depends more on integration, context and controls than on the ranking.
The benchmark should start the conversation, not end the buying decision.
Five questions to ask before you buy or scale
1. How does the system manage state and context?
Ask what the agent remembers across steps, sessions and tools.
Does it carry forward full reasoning state, a compacted summary, selected notes or nothing at all? What happens when the task exceeds the context window? Can your team inspect and correct the retained business context?
A strong model with weak context management may repeat work, lose constraints or make inconsistent decisions across a long process.
2. Which tools can it use, and what is it allowed to change?
A useful digital coworker needs tools. It may need to read the CRM, check inventory, create a draft invoice or update a ticket.
But access should be bounded. Ask which systems are read-only, which actions require approval and how the agent is prevented from wandering into unrelated data or operations.
Tools turn chat into execution. Permissions determine the blast radius when execution goes wrong.
3. What are the real cost and latency for the complete workflow?
Do not evaluate only the model’s token price.
Include tool calls, retries, retrieval, compaction, validation and human review. Measure the time and cost to complete the whole task at the quality you require.
ARC Prize’s published Astra results are a useful reminder that the harness can affect both performance and resource use.[1] Your own workflow economics may look very different.
4. Has the exact deployment configuration been tested on your work?
A polished demonstration is not enough.
Give the system a bounded set of representative tasks, including incomplete data, common exceptions and a few known failure cases. Define success before the test begins. Compare the result with the current human process on accuracy, cycle time, intervention rate and recovery effort.
Test the version you will actually deploy—not a vendor’s ideal laboratory setup.
5. How will the system be monitored, reviewed and stopped?
Ask what evidence the agent records, how errors are detected and who owns the decision when confidence is low.
For routine work, sampling may be enough. For payments, customer commitments, compliance or sensitive records, keep explicit human approval at the judgment point.
A production-ready agent needs an audit trail, escalation path and rollback plan. “The model is very capable” is not a control.
Orchestrate the complete system
The practical mistake is to debate which model is smartest while leaving the operating environment undefined.
A capable model without trusted context may act on the wrong facts. A capable model with excessive permissions may create avoidable risk. A capable model without verification may produce work that nobody can defend later.
The opposite is also true. A well-designed harness can help a model retain useful state, call the right tools and complete longer work more consistently. That is not a trick around the model. It is part of the product.
For SME leaders, the next step is not to reproduce a frontier benchmark.
Choose one real, bounded workflow. Define the data, tools, permissions, expected output, approval point and recovery path. Then test the complete deployment configuration from start to finish.
Treat the headline score as a starting point. Buy, or build, the operating system that can deliver the outcome safely, repeatedly and at a cost your business can sustain.
Sources
[1] https://arcprize.org/blog/astra — OpenAI's GPT-6 Astra on ARC-AGI-3 [2] https://openai.com/index/gpt-6-astra — GPT-6 Astra: A new generation of intelligence [3] https://simonwillison.net/2026/Sep/3/gpt6-astra — GPT-6 Astra — Simon Willison