Matthew Berman

GPT-5.6 SOL is HERE

Jul 9, 2026 1 h 9 min
openaigpt-5.6ai modelsbenchmarkingllm review
Watch on YouTube Follow Matthew Berman on Rundown — free

Summary

AI summaries can be incomplete or wrong. Verify anything important against the original video.

Matthew Berman reviews GPT-5.6 Sol, a new model from OpenAI, comparing its performance and cost against Fable 5. He finds GPT-5.6 Sol to be superior in benchmarks, reasoning, and general capabilities, particularly for coding tasks, while also being more affordable.

Matthew Berman provides an in-depth review of OpenAI's new GPT-5.6 Sol model, comparing it against Fable 5. He begins by detailing his personal experience using GPT-5.6 Sol over the past two months, highlighting its impressive performance on tasks like code generation, reasoning, and accuracy, especially noting its speed and ability to handle complex tasks.

He then dives into benchmarks, presenting data that shows GPT-5.6 Sol outperforming previous models like Fable 5 across various metrics. The review covers the Agents Last Exam, GDPval-AA v2, Management Consulting Tasks, Big Finance Bench, Artificial Analysis Intelligence Index, and others. Berman notes that while the numbers are included below, they are labeled rather than Fable-only results, and the OpenAI launch chart has no matching result. He also highlights that GPT-5.6 Sol is generally better at long-sustained work, completing complex projects with almost no intervention, and that the models kept adding levels, mods, and tiny details.

Berman also discusses the browser control capabilities, finding them the fastest and most accurate he's used. He mentions that while a data import ran, he was able to watch Supabase and resize the instances as the load changed, and it stayed on the dashboard and handled it. He also notes that he didn't set up an API or hand it keys, using Supabase the way he would.

The review concludes with a discussion of the model's capabilities, including its speed, cost, and performance on various benchmarks. He highlights that GPT-5.6 Sol is more capable than earlier models in both biology and cybersecurity, but does not cross the Critical threshold in either category. In cybersecurity, testing suggests GPT-5.6 is better at finding and fixing vulnerabilities than at reliably carrying out autonomous, end-to-end attacks against hardened targets—giving defenders an opportunity to strengthen systems before weaknesses are exploited. In biology, testing suggests GPT-5.6 can support legitimate research but does not provide the end-to-end capability needed to create, engineer, or synthesize a highly dangerous novel threat. He summarizes by saying that while he hasn't seen the whole picture, the model kept adding levels, mods, and tiny details. Days later, he killed the run himself. He never felt like he needed to jump back in and steer it onto the right path. He notes that the model kept adding levels, mods, and tiny details.

Verdict

GPT-5.6 Sol
ai model · $5 / million tokens

GPT-5.6 Sol is a significant improvement over previous models, offering better performance and cost-efficiency, though its raw capabilities remain within the established boundaries.

Buy

Pros

  • Significantly faster performance per token. 4:00
  • Better performance and cost-efficiency than Fable 5. 25:00
  • Better at finding and fixing vulnerabilities. 55:00
  • Faster and more capable. 19:10

Cons

  • Does not cross the Critical threshold in either category (biology/cybersecurity). 55:00

Specs

Tokens 25 billion 3:10
Output tokens $30 / million tokens 50:00
Input tokens $5 / million tokens 45:50
Cached input $0.50 / million cache hits 47:05
Luna $1 input / $6 output 44:10

Compared to

  • Fable 5

    GPT-5.6 generally performs better, especially on reasoning tasks, and is more affordable.

Best for

  • developers
  • researchers
  • teams needing advanced capabilities

Not for

  • projects requiring critical threshold crossing
  • creating highly dangerous novel threats

Claims & arguments

  • GPT-5.6 Sol Performance

    GPT-5.6 Sol is a significant step forward, outperforming previous models and competitors in key benchmarks.

    • 28:10 It beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost.
    • 34:10 GPT-5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost.
  • GPT-5.6 Sol Cost Efficiency

    GPT-5.6 Sol offers a lower cost model with competitive performance, making it more affordable.

    • 28:10 It beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost.
    • 34:10 GPT-5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost.
    • 45:50 GPT-5.6 Sol is priced per 1M tokens across three model sizes: Sol $5 input / $30 output, Terra $2.50 input / $15 output, and Luna is $1 input / $6 output.
  • GPT-5.6 Sol Capabilities

    GPT-5.6 models are more capable than earlier models in both biology and cybersecurity but do not cross the Critical threshold in either category.

    • 55:00 In cybersecurity, our testing suggests GPT-5.6 is better at finding and fixing vulnerabilities than at reliably carrying out autonomous, end-to-end attacks against hardened targets—giving defenders an opportunity to strengthen systems before weaknesses are exploited.
    • 57:50 In biology, our testing suggests GPT-5.6 can support legitimate research but does not provide the end-to-end capability needed to create, engineer, or synthesize a highly dangerous novel threat.
  • Codex Integration

    Codex is merging into the ChatGPT app.

    • 6:45 I switched back to GPT-5.5 to check. It went straight to the command line, even when the browser made more sense.

Key Points

  • 2:05 GPT-5.6 Sol is available now and is already making waves, costing less per token than previous models.
  • 2:10 The government approval held it up for two weeks, and it lands head-on against Claude Fable.
  • 3:00 Berman has used GPT-5.6 Sol internally for the past two months and has probably burned through more than 25 billion tokens.
  • 4:00 Codex has gotten much faster lately and is very noticeable.
  • 4:55 Cerebras is expected to soon serve full GPT-5.6, not a watered-down version.
  • 6:55 GPT-5.6 usually takes the short route to task completion, knows how to do things really well and directly.
  • 7:30 The Minecraft-style game is the best example, it started with /goal and walked away.
  • 8:20 The model kept adding levels, mobs, and tiny details.
  • 8:25 This model keeps adding levels, mods, and tiny details.
  • 10:20 The prompt stops sooner and is more likely to ask what comes next.
  • 10:40 The Codex browser is the first version of that idea I want to use all day.
  • 11:10 GPT-5.6 is better at long-sustained work.
  • 14:00 It keeps finding useful work when you give it a concrete finish line.
  • 14:10 It keeps finding useful work when you give it a concrete finish line.
  • 19:10 GPT-5.6 is faster than other models in two different ways: The raw generation speed is higher, something OpenAI has been putting effort into.
  • 20:10 It also takes a shorter path to solutions. It wonders less, changes less code, and generally knows how to get things done directly.
  • 22:10 GPT-5.6 Sol is capable of performing tasks faster than Fable, with roughly 3x fewer tokens per second.
  • 25:00 GPT-5.6 outperforms previous and competing frontier models with fewer tokens and at lower estimated cost.
  • 28:10 It beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost.
  • 34:10 GPT-5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost.
  • 45:50 GPT-5.6 is priced per 1M tokens across three model sizes: Sol $5 input / $30 output, Terra $2.50 input / $15 output, and Luna is $1 input / $6 output.
  • 48:30 GPT-5.6 Sol also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life.
  • 50:00 For GPT-5.6 and later models, cache writes are billed at 1.25x the model's uncached input rate, while cache reads continue to receive the 90% cached-input discount.
  • 55:00 GPT-5.6 is better at finding and fixing vulnerabilities than at reliably carrying out autonomous, end-to-end attacks against hardened targets

Worth watching if: If you're interested in the latest advancements in large language models, particularly from OpenAI, this detailed review of GPT-5.6 Sol by Matthew Berman offers valuable insights into its performance, cost, and capabilities compared to competitors like Fable 5. It's a deep dive for those curious about the frontier of AI.

Get every Matthew Berman video extracted like this

One daily email with structured extracts of every channel you follow. Free tier covers 15 videos a month.

Sign in with Google

No credit card. Free tier forever.

Watch on YouTube