The whole GPT-5.6 family, that is Sol, Terra and Luna, now works inside Kiro, AWS's spec-driven coding agent. OpenAI and AWS tuned the environment together. The number they put front and center isn't about quality. It's about cost.
- The whole GPT-5.6 family (Sol, Terra, Luna) is now in Kiro, AWS's agent that turns intent into requirements, a design and tasks.
- By OpenAI's own figures, on Terminal-Bench 2.1 the Terra variant completes successful tasks at around 82 percent lower cost.
- Verifying execution through property-based tests is among the listed capabilities.
Eighty-two percent. That's the number OpenAI put front and center when it announced that GPT-5.6 works inside Kiro.
Not speed. Not accuracy. Cost.
Whose number is it
Theirs. Measured by OpenAI, in an environment tuned by OpenAI and AWS, on a task they chose. There's no independent verification. So I read it as an advert with a precise number attached, not an engineering fact.
The logic underneath, though, holds up. The spec feeds the model clear requirements and context from the start. Less wandering, fewer wasted attempts, fewer tokens burned in the wrong direction.
For a team paying the bill at the end of the month, that's more interesting than any quality benchmark. And it's easy to check for yourself: look at your own invoice a month from now, not their chart.