Modern Creator

The GPT-5.6 Sol Playbook · Updated July 17, 2026

GPT-5.6 Sol — the playbook, from 23 creators at once.

Not another summary. We mined 23 creator breakdowns for the parts that matter: how it beats Fable on cost and finish-power, which effort to use, what to build, and where the experts disagree — every claim linked to the source.

OpenAICodexSol vs FableChatGPT Work23 breakdowns and counting
69%
DeepSWE on high effort ($3.47/task)
15%
of a 5-hour limit per message (0.1–2% on 5.5)
2.5x
faster limit burn with fast mode on
$1.04
Sol cost per completed task (Artificial Analysis)
6M
ChatGPT weekly active users at launch
23
Breakdowns synthesized

The 60-second version

  • GPT-5.6 shipped as a three-tier family — Sol (flagship), Terra (mid), and Luna (cheap, agent-only) — plus ChatGPT Work, a computer-use desktop app, and one-prompt Sites. Sol posts state-of-the-art coding and long-horizon scores at a quarter to half of Fable’s cost.
  • The spine of every real test: finish-power beats pretty. Sol reliably completes tasks end-to-end while Fable sometimes drafts a nicer first pass and then stalls — Eric Siu’s standing rule (‘the model that finishes reliably wins’) held across all four of his marketing tasks.
  • 5.6 runs longer per message because it was trained not to stop, so usage discipline is now the skill: default to High, not Max, turn off fast mode, and write an explicit stop point into every prompt.
  • The harness matters as much as the model — Theo found Sol designs and orchestrates better inside Claude Code than inside Codex, and Fable-as-manager + Sol-as-engineer catches defects a single model would ship.

People are asking

The questions everyone’s Googling.

Straight answers, pulled from across all 23 breakdowns.

What is GPT-5.6 Sol?

Sol is the flagship model in OpenAI’s GPT-5.6 family, positioned as a new default coding and agentic model. Creators found it posts state-of-the-art scores — 73% on DeepSWE at max reasoning, 88.8% on Terminal-Bench — while costing roughly a quarter to half of Claude Fable 5 for comparable work. Its defining trait is finish-power: it was trained to keep working end-to-end instead of stopping to ask, so it can run unattended for hours on a single brief.

GPT-5.6 Sol vs Claude Fable 5 — which is better?

Across seven-plus independent head-to-heads the pattern held: Sol wins on cost and reliable execution, Fable wins on planning, design taste, and creative ambition. On 130 API tasks the two tied on quality (0.982 vs 0.966) but Sol answered every time while Fable silently failed roughly 1 in 5, and Sol cost $16 to Fable’s $63. The practical verdict most creators reached: use Fable as the ‘manager’ for judgment calls, Sol as the ‘worker’ for volume.

What are Sol, Terra, and Luna?

They are the three tiers of the GPT-5.6 family: Sol is the flagship ($5/$30 per million in/out tokens), Terra is the mid-tier workhorse (~$2.50/$15, about half of GPT-5.5), and Luna is the small, cheap agent tier ($1/$6, undercutting Google Flash). Theo framed them as competitive strikes — Luna to kill Flash, Terra to kill Sonnet, Sol to replace GPT-5.5 itself. Match the tier to the task instead of always reaching for Sol.

Why is GPT-5.6 burning through my usage limits?

Because 5.6 was trained to stop asking for permission mid-task, so a single message can burn up to 15% of a five-hour limit versus 0.1–2% on 5.5. It isn’t a hidden setting — the fixes are habits: turn off fast mode (1.5x faster but 2.5x the burn), default to high not max reasoning, rein in sub-agents with one AGENTS.md line, and write an explicit stop point into every prompt. OpenAI temporarily removed the five-hour cap during launch, leaving only the weekly limit.

What reasoning effort should I use on GPT-5.6?

Default to High. On DeepSWE, Sol scores 45% (low, ~$1), 61% (medium, ~$1.86), and 69% (high, ~$3.47) — each step is a real gain for the money — but max only adds 4 points for $8.39, more than double the cost. Medium is the balance point for routine work; reserve Ultra (parallel agents) for tasks that genuinely need the coordination, since it tends to overthink and roughly double the bill.

What is ChatGPT Work and Codex Sites?

ChatGPT Work is a task-completion mode built on GPT-5.6 that connects to your email, calendar, files, and apps and runs multi-step jobs — OpenAI’s own finance team used it to run variance analysis, update an Excel forecast, build a deck, and publish a dashboard from one request. Codex Sites turns any chat output into a shareable interactive website with one instruction (‘take this into a Site’) in five to ten minutes — positioned as the single highest-leverage feature for non-developers.

Sol vs the field, dimensionalized

Better than Fable? Where it counts — and where it doesn’t.

Not vibes. Here is what happened in the head-to-head tests creators actually ran.

4 real marketing tasks

SOLShipped all four end-to-end — built an interactive ‘world’ from one prompt in ~35 min and landed ship-ready thumbnails in one shot with minimal back-and-forth.
FABLESometimes prettier intermediate output, but stalled — failed to fetch the source page (burning Higgsfield credits), and left the world-build unfinished after three hours.
Leveling Up with Eric Siu

Kimi K3 3-way build test

SOLCost $1.04 per completed task and scored 22 on the hallucination index — better than Kimi’s 18; ranked third on subjective build polish.
KIMIWon the blind vote at $0.95/task, but burned 21.5M tokens over 1h33m for a build Fable finished in 3.5M tokens and 17 minutes.
Chase AI

Seven creators, one verdict

SOLThe reliable, cheap ‘worker’ — one identical game build cost $4.50 vs Fable’s $14.22, using 31k output tokens to Fable’s 90k, and still won the head-to-head.
FABLEThe ‘manager’ — better planning, design taste, and creative ambition, but at 2x-plus the cost, worth paying only on high-stakes work.
The Next New Thing

3D apartment from a floor plan

SOLRendered a visually polished 3D scene but got the actual room layout wrong — a good-looking output that wasn’t spatially correct.
OPUSProduced the most spatially accurate reconstruction of the apartment in the blind test, beating both Sol and Fable on getting the facts right.
Pat Simmons

Reach for GPT-5.6 SOL when…

  • You need a task finished end-to-end with the least hand-holding — Sol’s reliability to completion beat Fable across every real-world marketing and build test.
  • You’re running long, agentic jobs that must not stall — goal mode kept shipping code even after the user hit his weekly usage limit.
  • You want an always-on chief of staff — ChatGPT Work triages email, books calendar meetings, preps guests, and runs scheduled tasks unattended.
  • You want a shareable website from one prompt — Codex Sites turns any chat output into a live interactive site in five to ten minutes.
  • You’re producing marketing content at volume — thumbnails, ad creative, and auto-clipped webinars that actually render instead of just planning.
  • You’re cost-sensitive or high-volume — Sol costs half of Fable per token and uses fewer tokens, compounding the savings on API-driven work.

Reach for something else when…

  • You need pure visual polish and design taste — several creators still gave Fable the edge on planning, creative ambition, and image-prompting.
  • Watch your usage limits — a single Sol message can burn 15% of a five-hour cap, so budget guardrails matter as much as capability.
  • Turn off fast mode unless you’re watching the run live — it’s 1.5x faster but 2.5x the burn.
  • Mind the reasoning-effort cost curve — the jump from high to max more than doubles cost ($3.47 to $8.39) for a 4-point DeepSWE gain.
  • Accept that Sol can look plainer — it’s repeatedly assessed as not a strong front-end/visual model, so pair it with another tool for anything that has to look good.
  • Rein in the defaults — Sol over-writes code (a five-line change into a 300-line rewrite) and spawns sub-agents eagerly unless you constrain it.
73% vs 69.2%
DeepSWE max (Sol) — Terminal-Bench 88.8%
$16 vs $63
130 API tasks — Sol answered every time, Fable failed ~1 in 5
$5 / $30
Sol per M in/out — roughly half of Fable
directional
Creator-run tests — treat with healthy skepticism

The effort dial

The real setting isn’t the model — it’s the effort.

On 5.6 the reasoning-effort dial — not the tier — is where the real cost/quality tradeoff lives, and the DeepSWE curve flattens hard past high.

Low
45%
~$1 / task
fine for simple, well-specified steps — Sol already throttles its own reasoning on easy tasks
Medium
61%
~$1.86 / task
the default balance point for routine work
High
69%
~$3.47 / task
the efficient ceiling — reach for this on serious and planning-heavy work
Max
73%
~$8.39 / task
2x-plus the cost for a few points — rarely worth it, and disables Sol’s efficiency post-training
Default to High, not Max — the curve flattens hard past high, where max adds ~4 points for more than double the cost.
Turn off fast mode — it’s 1.5x faster but 2.5x the burn, a trade that made sense on 5.5 and doesn’t now that Sol runs unattended.
Write an explicit STOP point in every prompt — 5.6 was trained not to stop on its own, so not stopping it is the default failure mode.

⚠️ 5.6 was trained to keep working, so a single message can burn 15% of a five-hour limit versus 0.1–2% on 5.5. Turn off fast mode, cap sub-agents, and define the finish line. ↗ Theo

Use-case playbooks

Here’s exactly what to build — and how.

The real methods creators demoed, with the actual tools and the video to watch.

CodexSitesTerra
A.Record a scroll-through, not a screenshot. A slow scroll video of a reference site captures load animations, hover states, and transitions a static image can never show the model.
B.Prompt with a keep/change split. Explicitly list what to preserve (structure, layout, motion) and what to replace (brand, copy, color) — run Sol for the big structural pass, Terra for fast fixes.
C.‘Take this into a Site.’ Turn any chat output — a travel guide, a product plan, a dashboard — into a shareable interactive website in five to ten minutes with one instruction.
D.Ask for a 3D/WebGL hero. A repeatable trick to make an AI-generated page look more expensive and animated — it closed most of the design gap on a travel-site build.

Run your work as a chief of staff

Peter Yang
ChatGPT WorkCodexScheduled Tasks
A.Install every app plugin first. Gmail, Calendar, Docs, Drive — connecting apps before anything else is what turns a chatbot into an assistant that can act, not just answer.
B.One long-running thread per workflow. Beats one thread per task — the model compacts old messages while keeping the working context intact.
C.Run once, then schedule. Do a workflow manually, refine the prompt, then convert it into a recurring scheduled task that fires without being asked.
D.Draft in review-before-send mode. Ask it to study your sent mail and match your voice — fixes the too-formal default tone in one pass while keeping a human check.

Produce marketing content

Eric Siu, Nick Saraev & Nate Herk
HiggsfieldRiversideElevenLabsHeyGen
A.Transcript in, thumbnails out. Feed a video transcript and ask for thumbnail concepts that match your channel identity — lands close to what you’d actually ship in a few minutes.
B.Hand it a webinar to auto-clip. Ask for short/mid/long-form cuts with specific clip length, hook strength, and platform aspect ratio — treat the output as a fast first draft, not a publish-ready cut.
C.Ground video in a still first. Generate a packshot or lifestyle image, then animate from it — far more coherent than text-to-video directly.
D.Chain tools by role, then QA. ElevenLabs for voice, HeyGen for the avatar, Sol as orchestrator — and run dedicated inspection agents against the rendered frames before publishing.

Code inside the right harness

Theo - t3.gg
Claude CodeWorkflowsCodexAGENTS.md
A.Run Sol inside Claude Code. The same model produced visibly better UI — Claude Code’s lean prompt beat Codex’s over-prescriptive front-end section (which Sol itself graded 3/10).
B.Use Workflows, not open-ended Ultra. Model-authored code that defines stages up front and terminates cut token usage to roughly a quarter for equal or better output.
C.Hand-write your AGENTS.md. A markdown file of your real preferences beats any preset — add ‘only use sub agents if the user explicitly requests them’ to stop sub-agent spam.
D.Always define the finish line. ‘Build it, test it, open a PR, address the first round of review, then stop’ — the single highest-leverage habit on 5.6.

Combine models into one dev team

Nick Puru | AI Automation
Claude Fable 5Codex CLIGPT-5.6 Sol
A.Prompt Fable to only plan, delegate, and review. Never write code — turning Anthropic’s own ‘plans, delegates, reviews’ description into a literal job assignment.
B.Wire Sol in as the engineer. Via Codex CLI, with five workers running in parallel, each owning one product area under the manager’s plan.
C.Make the manager run a review pass. On round two it flagged six defects in its own team’s work — including two serious cross-tenant data leaks — before anything shipped.
D.Do the cost math. The two-round build ran ~$80 in Sol tokens against a $55/seat SaaS that would cost a 10-person team $6,600 a year.

Steal these

12 things to have Sol do for you today.

Real instructions from real demos. Copy, paste, tweak.

“List every open action item and unsubscribe candidate from my last few days of email, then draft replies in my voice for review.”
via Peter YangWatch →
“Take this into a Site.”
via Peter YangWatch →
“Book this meeting from our chat, then research the guest and draft an interview-ready prep doc.”
via Peter YangWatch →
“Here’s a scroll-through of a reference site — keep the structure, layout, and motion; change the brand, copy, and color.”
via Made by SourasithWatch →
“Cut this webinar into short, mid, and long-form social clips — specify length, hook strength, and aspect ratio per platform.”
via Eric SiuWatch →
“Here’s the video transcript — give me YouTube thumbnail concepts that match my channel’s identity.”
via Eric SiuWatch →
“Only use sub agents if the user explicitly requests them.” (one line for your global AGENTS.md)
via Theo - t3.ggWatch →
“Build it, test it, open a PR, address the first round of review comments, then stop.”
via Theo - t3.ggWatch →
“The human is away — decide and proceed. Ideate 100 fictional products, shoot the stills, animate two ad spots each, self-critique, and publish to the sandbox. No real spending.”
via Nick SaraevWatch →
“Build RF-Zero, a live multiplayer racer — spawn your own sub-agents, deploy it, and test it yourself. This runs overnight and I won’t be available.”
via Ray FernandoWatch →
“Make an Excel clone, continue until feature parity.” (eight words, ran five days unsupervised)
via Matthew BermanWatch →
“Run the variance analysis, update the forecast in Excel, build the deck, and publish a shareable dashboard.”
via OpenAIWatch →

The map

Where 23 creators agreed — and threw hands.

The honest part you only get by watching all of them. Tap any name to watch.

Sol wins on reliability

“Finish-power beats pretty — the model that completes the job end-to-end wins even when its output isn’t the best-looking.”

Manager + worker, not one winner

“There’s no universal best — route by role: the pricier model plans and designs, the cheaper one executes.”

Harness matters more than the model

“The same model designs and orchestrates better in one harness than another — check the system prompt before crediting the model.”

All-in — this changes how I work

“I don’t ship software anymore. My agents ship software for me — the bottleneck is now imagination, not engineering.”

Cost & usage-limit realists

“Judge models on total cost per completed task and time to finish real work, not advertised per-token pricing.”

The launch story & the risk

“GPT-5.6 is as much about ChatGPT Work, Sites, and the geopolitics of access as it is about the model itself.”

✓ What nearly all 23 agreed on

  • Sol is dramatically cheaper per task than Fable — roughly half the sticker price ($5/$30 vs $10/$50 per million) and fewer tokens to reach a comparable answer, compounding the savings.
  • Sol’s edge is reliability to completion — it finishes tasks end-to-end and runs unattended for hours, where Fable more often stalls or silently fails to answer.
  • Default to High effort, not Max/Ultra — the cost curve flattens hard past high and Ultra tends to overthink and roughly double the bill for a comparable result.
  • One-shot output is a draft, not a finished product — every ‘built in one prompt’ demo is really one strong prompt plus many targeted fixes; iterative feedback beats a single detailed prompt.
  • Model choice is role-based — there’s no universal best; pick a manager (planning, design, judgment) and a worker (speed, price, execution) per task.
  • Sol is not a strong visual/design model — it’s repeatedly flagged as needing manual UI cleanup, so pair it with another tool for anything user-facing.

✗ Where they threw hands

  • Does Sol out-design Fable? Pat Simmons found Soul out-designed Fable across all three builds, while Nate Herk, The Next New Thing, and Peter Yang gave Fable the design and taste edge.
  • Is Sol faster or slower than Fable? Nate Herk clocked Sol at 7 minutes to Fable’s 21–23, but Pat Simmons found Soul slower on every single build — flagging it as a possible anomaly against the usual pattern.
  • Is the cheaper model actually the better one? Matthew Berman argues Fable reasons further ahead with more headroom despite the price, while Nate Herk and the Seven Creators panel treat Sol as the practical winner for daily work.
  • Which harness should run Sol? Theo insists Sol is better inside Claude Code than Codex, while Peter Yang defaults to Codex as the one tool that covers everything.

The 5 moves that pay

1

Write the finish line into the prompt

5.6 was trained not to stop on its own, so the single highest-leverage habit is an explicit stop point — from ‘write a plan, then stop’ to ‘build it, test it, open a PR, then stop.’

2

Default to High, save Ultra for real coordination

High is the efficient ceiling on DeepSWE; Max more than doubles cost for ~4 points, and Ultra’s parallel agents overthink and roughly double the bill on tasks that don’t need them.

3

Turn off fast mode

It’s 1.5x faster but 2.5x the burn — a fine trade when 5.5 stopped constantly and you were waiting, a bad one now that Sol runs unattended for hours.

4

Try the cheap model first, escalate with a reason

When the price gap is this large, running Sol first and discarding it costs almost nothing — reserve Fable for planning, design, and genuinely hard problems where it’ll meaningfully outperform.

5

Put a reviewer over the coder

Sol posts the highest cheating rate METR has measured and over-writes code by default — a manager model (or a review pass) caught six defects, including two cross-tenant data leaks, before anything shipped.

The library

All 23 breakdowns, and counting.

36:00
Theo - t3․gg · Review

GPT-5.6: The Review

Theo spends 36 minutes putting real numbers behind the GPT-5.6 hype — Sol, Terra, and Luna, benchmarked against Claude Fable, one blog chart at a time.

July 12th
40:03
Pat Simmons · Review

GPT-5.6 Sol: No-Hype Full Review & Testing

A blind, four-way bake-off — GPT-5.6 Sol against Fable, Opus 4.8, and GPT-5.5 — across ten builds and knowledge-work tasks, scored one task at a time without knowing which model made what.

July 10th
3:02:05
Ray Fernando · Livestream

Live build with GPT-5.6 Sol

A three-hour first-look stream where one /goal command builds an entire multiplayer game — and then an iOS app — while the host mostly watches.

July 10th
08:54
Matthew Berman · Review

GPT-5.6 is FINALLY HERE (WOAH)

A 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.

July 9th