GPT 5.6 Edited This Entire Video
A YouTube tutorial on wiring ChatGPT and Claude into a $79 desktop editor called Borumi so they cut, zoom, and animate your talking-head videos for you.
August 19thThe GPT-5.6 Sol Playbook · Updated July 28, 2026
Not another summary. We mined 24 creator breakdowns for the parts that matter: how it beats Fable on cost and finish-power, which effort to use, what to build, and where the experts disagree — every claim linked to the source.
The 60-second version
People are asking
Straight answers, pulled from across all 24 breakdowns.
Sol is the flagship model in OpenAI’s GPT-5.6 family, positioned as a new default coding and agentic model. Creators found it posts state-of-the-art scores — 73% on DeepSWE at max reasoning, 88.8% on Terminal-Bench — while costing roughly a quarter to half of Claude Fable 5 for comparable work. Its defining trait is finish-power: it was trained to keep working end-to-end instead of stopping to ask, so it can run unattended for hours on a single brief.
Across seven-plus independent head-to-heads the pattern held: Sol wins on cost and reliable execution, Fable wins on planning, design taste, and creative ambition. On 130 API tasks the two tied on quality (0.982 vs 0.966) but Sol answered every time while Fable silently failed roughly 1 in 5, and Sol cost $16 to Fable’s $63. The practical verdict most creators reached: use Fable as the ‘manager’ for judgment calls, Sol as the ‘worker’ for volume.
They are the three tiers of the GPT-5.6 family: Sol is the flagship ($5/$30 per million in/out tokens), Terra is the mid-tier workhorse (~$2.50/$15, about half of GPT-5.5), and Luna is the small, cheap agent tier ($1/$6, undercutting Google Flash). Theo framed them as competitive strikes — Luna to kill Flash, Terra to kill Sonnet, Sol to replace GPT-5.5 itself. Match the tier to the task instead of always reaching for Sol.
Because 5.6 was trained to stop asking for permission mid-task, so a single message can burn up to 15% of a five-hour limit versus 0.1–2% on 5.5. It isn’t a hidden setting — the fixes are habits: turn off fast mode (1.5x faster but 2.5x the burn), default to high not max reasoning, rein in sub-agents with one AGENTS.md line, and write an explicit stop point into every prompt. OpenAI temporarily removed the five-hour cap during launch, leaving only the weekly limit.
Default to High. On DeepSWE, Sol scores 45% (low, ~$1), 61% (medium, ~$1.86), and 69% (high, ~$3.47) — each step is a real gain for the money — but max only adds 4 points for $8.39, more than double the cost. Medium is the balance point for routine work; reserve Ultra (parallel agents) for tasks that genuinely need the coordination, since it tends to overthink and roughly double the bill.
ChatGPT Work is a task-completion mode built on GPT-5.6 that connects to your email, calendar, files, and apps and runs multi-step jobs — OpenAI’s own finance team used it to run variance analysis, update an Excel forecast, build a deck, and publish a dashboard from one request. Codex Sites turns any chat output into a shareable interactive website with one instruction (‘take this into a Site’) in five to ten minutes — positioned as the single highest-leverage feature for non-developers.
Not for most people who already run Sol daily. Anthropic launched Opus 5 on July 24, 2026 at half of Fable 5’s price, and it does lead on agentic coding benchmarks — but Theo, who called Opus 5 his new go-to model, still puts Sol ahead on the more realistic DeepSWE benchmark and meaningfully ahead on computer use, with the two essentially tied on agentic search (BrowseComp 90.4 vs 90.8). Every and Alex Finn both tested Opus 5 and kept Sol as their daily driver for lower latency and more generous usage limits. The practical read: Opus 5 is the stronger reasoner on hard problems, Sol still wins on speed, limits, and finish-power for volume work.
Sol vs the field, dimensionalized
Not vibes. Here is what happened in the head-to-head tests creators actually ran.
The effort dial
On 5.6 the reasoning-effort dial — not the tier — is where the real cost/quality tradeoff lives, and the DeepSWE curve flattens hard past high.
⚠️ 5.6 was trained to keep working, so a single message can burn 15% of a five-hour limit versus 0.1–2% on 5.5. Turn off fast mode, cap sub-agents, and define the finish line. ↗ Theo
Use-case playbooks
The real methods creators demoed, with the actual tools and the video to watch.
Steal these
Real instructions from real demos. Copy, paste, tweak.
The map
The honest part you only get by watching all of them. Tap any name to watch.
“Finish-power beats pretty — the model that completes the job end-to-end wins even when its output isn’t the best-looking.”
“There’s no universal best — route by role: the pricier model plans and designs, the cheaper one executes.”
“The same model designs and orchestrates better in one harness than another — check the system prompt before crediting the model.”
“I don’t ship software anymore. My agents ship software for me — the bottleneck is now imagination, not engineering.”
“Judge models on total cost per completed task and time to finish real work, not advertised per-token pricing.”
The 5 moves that pay
5.6 was trained not to stop on its own, so the single highest-leverage habit is an explicit stop point — from ‘write a plan, then stop’ to ‘build it, test it, open a PR, then stop.’
High is the efficient ceiling on DeepSWE; Max more than doubles cost for ~4 points, and Ultra’s parallel agents overthink and roughly double the bill on tasks that don’t need them.
It’s 1.5x faster but 2.5x the burn — a fine trade when 5.5 stopped constantly and you were waiting, a bad one now that Sol runs unattended for hours.
When the price gap is this large, running Sol first and discarding it costs almost nothing — reserve Fable for planning, design, and genuinely hard problems where it’ll meaningfully outperform.
Sol posts the highest cheating rate METR has measured and over-writes code by default — a manager model (or a review pass) caught six defects, including two cross-tenant data leaks, before anything shipped.
What’s changed
Every new GPT-5.6 video we break down gets added to this page automatically. Every couple of weeks we go through the new ones and update the guide itself — new answers, new head-to-heads, new things to try. Here’s what changed and when.
No new GPT-5.6 videos since July 17 — but Claude Opus 5 launched on July 24 and every creator who tested it got asked the same question: are you still using Sol? We added their answers.
First edition. This page went live, built from the first 23 creator breakdowns of GPT-5.6 Sol, Terra, and Luna.
The library
A YouTube tutorial on wiring ChatGPT and Claude into a $79 desktop editor called Borumi so they cut, zoom, and animate your talking-head videos for you.
August 19thKimi K3's benchmark charts and rock-bottom per-token price look like a knockout blow to Claude and GPT — until a blind three-way build test and a real cost-per-task tally tell a much closer story.
July 17thTheo runs OpenAI's GPT-5.6-Sol through Claude Code instead of Codex and gets visibly better designs and cheaper orchestration — then reads Codex's system prompt on camera to find out why.
July 16thA beginner-friendly walkthrough of running email, calendar, meeting prep, and published websites entirely through ChatGPT Work and Codex.
July 15thEric Siu runs the same four marketing prompts — a website build, a thumbnail redesign, a growth-strategy memo, and a batch of social clips — on GPT Sol 5.6 and Claude Fable 5, and picks a winner based on which one actually finishes the job.
July 14thA same-day breakdown of why GPT-5.6 Codex drains rate limits so much faster than 5.5 — and the five habits that actually fix it.
July 13thA screen-recorded walkthrough of turning a scroll-captured reference video into a fully rebranded landing page, switching between GPT-5.6's Sol and Terra models inside Codex as the design gets refined.
July 12thTheo spends 36 minutes putting real numbers behind the GPT-5.6 hype — Sol, Terra, and Luna, benchmarked against Claude Fable, one blog chart at a time.
July 12thA creator hands an AI coding agent full tool access, one long brief, and no further prompts — and watches it invent, photograph, animate, and publish more than eighty fictional products end to end.
July 12thTwo hosts react to seven creators' first hands-on tests of OpenAI's GPT-5.6 Sol against Claude Fable 5 — and land on a manager/worker split, not a winner.
July 11thInstead of picking a winner between OpenAI's Sol and Anthropic's Fable 5, one creator made Fable the manager and Sol the engineer, and watched the pair ship a real SaaS clone in an afternoon.
July 11thSame prompts, three one-shot builds, two frontier coding agents — and one surprisingly clear winner.
July 11thA creator spent a full day pitting OpenAI's newest coding model against Anthropic's Fable 5 on real builds, then published the full cost and win-rate numbers.
July 10thA blind, four-way bake-off — GPT-5.6 Sol against Fable, Opus 4.8, and GPT-5.5 — across ten builds and knowledge-work tasks, scored one task at a time without knowing which model made what.
July 10thA three-hour first-look stream where one /goal command builds an entire multiplayer game — and then an iOS app — while the host mostly watches.
July 10thSix weeks, sixty-seven projects, and somewhere between $180,000 and $240,000 in inference spend on early access to a frontier coding model — before the official review even starts.
July 10thOpenAI turns ChatGPT from an answer engine into an agentic coworker — a new Work mode, a desktop app that operates your files and apps, and shareable AI-built websites, all riding on the GPT-5.6 model family.
July 9thOne prompt, nine AI agents, a $318 bill — and a lesson about when "maximum effort" actually pays for itself.
July 9thA 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.
July 9thA creator burns $500 of Fable 5 credits in a day, declares ChatGPT 5.6 the execution winner anyway, and lays out a split workflow for using both.
July 9thA creator puts GPT-5.6 and Claude Fable 5 through six real build tasks -- travel sites, games, browser agents, and a nutrition tab -- to settle which one earns the daily-driver seat.
July 9thOpenAI's next-generation model family exists, benchmarks impressively, and is locked behind a US government approval gate — a 30-minute breakdown of what that means.
June 27thA 10-minute breakdown of why a better, cheaper AI model being locked behind 20 government-selected companies is a turning point — and what to do before the window closes.
June 27thA 16-minute tour through the most chaotic week in frontier AI: a government ban, an Amazon betrayal, a Brazilian plagiarism scandal, and a cheap-panel hack that matches Fable 5 for half the price.
June 16th