New-ZZZ
RU / EN
AI Models 9 July 2026

OpenAI launches GPT-5.6 with faster, cheaper frontier models

N
New-ZZZ desk
OpenAI Blog · 4 weeks ago

OpenAI is presenting GPT-5.6 as a new model family built around a simple promise: more useful work from each token, stronger results for each dollar spent, and extra capability when a task is difficult enough to justify more compute. The family includes three models: Sol, the new flagship; Terra, a balanced model aimed at everyday work; and Luna, the lowest-cost option. The announcement frames Sol as the main leap forward, especially in coding, knowledge work, cybersecurity, science, tool use, and design-oriented collaboration.

The central claim is that GPT-5.6 Sol improves both raw intelligence and operating efficiency. OpenAI says it can outperform earlier and competing frontier models while using fewer tokens and costing less in estimated total spend. That matters because advanced AI systems are increasingly judged not only by benchmark scores, but by how much useful work they can complete within a practical budget. The headline shift is performance per dollar: GPT-5.6 is described as doing more successful work for the same spend, or reaching similar outcomes at lower total cost.

OpenAI says GPT-5.6 was trained to extract more value from every token, which means the model should need fewer words, tool calls, or repeated attempts to complete complex work. On Agents' Last Exam, a benchmark covering long-running professional workflows across 55 fields, GPT-5.6 Sol reportedly reaches 53.6, setting a new high and beating Claude Fable 5 with adaptive reasoning by 13.1 points. Even at a medium reasoning setting, Sol is said to outperform Fable 5 by 11.4 points while costing roughly one quarter as much. OpenAI also claims the smaller Terra and Luna models outperform Fable 5 at about one sixteenth of the cost, positioning the family as not just a flagship upgrade but a broader cost-efficiency push.

The announcement also cites the Artificial Analysis Intelligence Index, which measures a mix of agentic work, coding, scientific reasoning, and general capability. There, GPT-5.6 Sol with maximum reasoning reportedly comes within one point of Fable 5 while finishing tasks 61% faster and at roughly half the estimated cost. In practical terms, OpenAI is arguing that Sol may not always need to win every raw-score comparison to be more attractive: if a model is nearly as capable, much faster, and cheaper to run, it can be better suited for real production use.

A major part of the release is the new “ultra” capability setting. OpenAI describes ultra as its highest-capability mode, intended for demanding tasks where spending more tokens and compute is worthwhile. Unlike normal single-agent execution, ultra coordinates multiple agents across parallel workstreams, with four agents running by default. The goal is to let the system explore alternatives, run checks, split subtasks, and converge faster on stronger answers. OpenAI says parallel agents improve the score-latency tradeoff on benchmarks such as BrowseComp, SEC-Bench Pro, and Terminal-Bench 2.1, meaning the system can reach better results in less elapsed time.

GPT-5.6 also expands what OpenAI calls Programmatic Tool Calling in the Responses API. Instead of forcing developers to script every step or send every tool result back through the model, the model can write and run lightweight programs that coordinate tools, filter intermediate data, monitor progress, and decide what to do next. This is important for tool-heavy workflows because large tool outputs can be expensive and distracting. The model is being positioned less as a chatbot and more as an agentic work system that can manage intermediate steps without constant human or developer supervision.

Coding is one of the strongest claims in the announcement. OpenAI calls GPT-5.6 Sol its best coding model so far. On the Artificial Analysis Coding Agent Index, Sol with maximum reasoning reportedly scores 80, which OpenAI says is a new state of the art and 2.8 points above Fable 5. The company adds that Sol uses less than half the output tokens, takes less than half the time, and costs about one third less. The advantage is also said to extend to smaller models: Terra performs slightly above Fable 5, while Luna beats Opus 4.8, with both completing coding tasks in roughly one third of the time, using about half the output tokens, and costing around one quarter as much.

OpenAI also highlights Terminal-Bench 2.1 and DeepSWE, which test more realistic engineering skills than simple coding puzzles. Terminal-Bench focuses on complex command-line workflows, while DeepSWE evaluates longer software engineering tasks in real codebases. These benchmarks are important because many practical coding jobs require navigating files, running commands, interpreting errors, updating an approach, and checking results, not merely writing a function in isolation. If the reported gains hold up in real use, GPT-5.6 could be more useful for software teams that want AI help with implementation, debugging, refactoring, and multi-step engineering tasks.

The release also emphasizes safety. OpenAI says GPT-5.6 launches with its most robust safeguards to date, designed to resist determined and adaptive misuse while avoiding broad restrictions on legitimate work. Before general availability, the company says it ran an extensive evaluation period combining human red teaming with large-scale automated testing. During the preview, OpenAI worked with expert organizations and trusted partners to pressure-test defenses. The resulting safety system combines protections trained into the model with real-time checks, monitoring, and access controls calibrated to trust and risk.

The larger strategic message is that OpenAI wants GPT-5.6 to serve different levels of ambition and budget. Luna is meant to make advanced intelligence cheaper and more abundant. Terra is aimed at routine everyday work where balance matters. Sol is the flagship for high-value tasks. Max reasoning gives Sol more time to think, check, and revise; ultra goes further by coordinating multiple agents in parallel. The release is therefore not just a model upgrade, but a packaging of intelligence across cost, latency, reasoning depth, and multi-agent execution.

Why it matters

  • GPT-5.6 is framed as a major efficiency upgrade, promising stronger AI performance at lower estimated cost.
  • The new ultra mode pushes OpenAI further into multi-agent workflows for difficult professional tasks.
  • Coding, tool use, and long-running agentic work are central to the release, signaling where frontier models are being optimized.

Key facts

  • The GPT-5.6 family includes Sol, Terra, and Luna, covering flagship, balanced, and cost-efficient use cases.
  • OpenAI says GPT-5.6 Sol scores 53.6 on Agents' Last Exam, beating Claude Fable 5 by 13.1 points.
  • Sol reportedly reaches a state-of-the-art score of 80 on the Artificial Analysis Coding Agent Index.
  • The new ultra setting coordinates four agents in parallel by default for demanding tasks.
  • Programmatic Tool Calling lets the model run lightweight programs to coordinate tools and reduce unnecessary model round trips.
Read the original

The full text is in the original source. Here we provide a brief summary and key facts.

/ related