New-ZZZ
RU / EN
LLM 26 June 2026

GPT-5.6 Sol Launches: New Era of AI Capability with Subagents and Enhanced Safety

N
New-ZZZ desk
OpenAI Blog · 1 month ago

The announcement details the limited preview of the GPT-5.6 series, introducing three distinct models: Sol, Terra, and Luna. GPT-5.6 Sol is positioned as the flagship, most capable model, while Terra offers a balanced, cost-effective alternative, and Luna provides strong capability at the lowest cost. This tiered approach aims to provide broad access to advanced AI capabilities across different use cases and budgets.

GPT-5.6 Sol represents a significant leap in AI performance, particularly in complex, multi-step reasoning tasks. To enhance its reasoning depth, the model introduces a new max reasoning effort, allowing it more time to process information deeply. Furthermore, a groundbreaking ultra mode is introduced, which moves beyond the limitations of a single AI agent by leveraging subagents. This subagent architecture is crucial for accelerating and managing highly complex workflows that require coordination across multiple specialized tasks, dramatically increasing the model's overall operational capacity.

Performance benchmarks demonstrate substantial improvements across several critical domains. In coding workflows, GPT-5.6 Sol sets a new state-of-the-art on the Terminal-Bench 2.1, a rigorous test designed for command-line workflows. This benchmark specifically evaluates the model's ability to handle planning, iterative execution, and complex tool coordination—skills vital for professional software development. Similarly, in the field of biology, the model shows marked improvements on GeneBench v1, which assesses long-horizon genomics and quantitative-biology analyses. Notably, it achieves stronger results than its predecessor, GPT-5.5, while simultaneously requiring fewer computational tokens, indicating a significant efficiency gain.

Perhaps the most detailed section concerns cybersecurity capabilities. GPT-5.6 Sol is touted as the most capable model yet for long-horizon security tasks, including vulnerability research and exploitation. On the ExploitBench², it performs competitively with Mythos Preview, but critically, it achieves this using only about one-third of the output tokens, highlighting a massive efficiency gain. The model also shows strong improvements across the ExploitGym benchmark, a collaborative effort by researchers from UC Berkeley and OpenAI. The developers emphasize that the goal is to shift the performance-efficiency frontier for defensive security work. The core philosophy is that GPT-5.6 Sol is significantly better at helping defenders find and fix vulnerabilities than it is at reliably executing end-to-end, malicious attacks.

Given the increased power and capability of GPT-5.6, the developers have implemented their most robust safety stack to date. This safety hardening involved weeks of intensive pressure-testing, specifically targeting weaknesses related to higher-risk activities, sensitive cyber requests, and patterns of repeated misuse. The safeguards are designed to be sophisticated: they aim to make prohibited offensive activity more difficult, uncertain, and detectable, all while meticulously preserving access for legitimate, beneficial uses. These legitimate uses include code review, vulnerability research, patch development, debugging, security education, and defensive testing. The developers explicitly state that while the model identified bugs and exploitation primitives in certain browser evaluations (like Chromium and Firefox), it did not autonomously produce a functional, full-chain exploit under the tested conditions. This careful balance of increased capability and enhanced safety is why the release is phased, starting with a limited preview for trusted partners, a step taken in coordination with the U.S. government. This limited access is framed as a short-term necessity to align with the development of a repeatable process for future model releases, rather than a long-term default, ensuring that the best tools remain accessible to developers, enterprises, and global partners.

Why it matters

  • It introduces a new, highly capable, and efficient flagship model (Sol) with significant performance leaps in complex domains.
  • The inclusion of 'ultra' mode and subagents represents a major architectural advancement for complex, multi-step reasoning.
  • The model sets new state-of-the-art benchmarks in critical areas like coding (Terminal-Bench) and cybersecurity (ExploitBench), while maintaining a strong focus on defensive use cases.

Key facts

  • GPT-5.6 Sol, Terra, and Luna are launched in a limited preview, with general availability expected soon.
  • The model features a new 'ultra' mode that utilizes subagents to accelerate complex work beyond single-agent capabilities.
  • GPT-5.6 Sol achieves state-of-the-art results on Terminal-Bench 2.1 for coding workflows.
  • The model demonstrates superior efficiency in cybersecurity, performing competitively with Mythos Preview using only ~1/3 of the output tokens.
  • Robust safeguards are implemented to restrict offensive use while preserving access for legitimate defensive activities.
Read the original

The full text is in the original source. Here we provide a brief summary and key facts.

/ related