Same one-shot prompts, two models, three real builds — a website, an app clone, and a 3D game — scored on intelligence, cost, and time, not just vibes.
Kimi K3, a brand-new open-weight model, produces results close enough to Claude Fable 5's quality across three real one-shot builds that the deciding factor becomes ROI — cost and time — rather than raw intelligence.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You use Claude Code or a similar coding-agent harness and want to know if a cheaper model is worth routing tasks to.
You're deciding whether to keep paying for a Claude subscription versus switching to usage-based API pricing on a different model.
You want a real cost-and-time benchmark for one-shot AI app builds, not just an intelligence leaderboard screenshot.
You're evaluating whether an existing app idea (like a stretching or habit app) is now trivial to clone with AI.
SKIP IF…
You want a deep technical explainer of Kimi K3's architecture — this is a hands-on build comparison, not a model-internals breakdown.
You're not using AI coding agents at all and just want general AI news.
TL;DR
The full version, fast.
The creator ran identical one-shot prompts through Claude Fable 5 and the newly launched open-weight Kimi K3 across three builds: a Starbucks x FIFA World Cup campaign site, a clone of a five-million-download stretching app, and a playable 3D Portal clone. Kimi K3 matched Fable 5 closely on Artificial Analysis's intelligence benchmark and even beat it on Arena.ai's blind-vote design benchmark, while costing a fraction as much per task — $4.70-$9.37 versus $20-$29 — at roughly double the build time. The conclusion: a flat-rate Claude subscription still beats usage-based pricing for regular use, but Kimi K3 becomes the obvious fallback once subscription credits run dry, and its scheduled open-weight release means its capability will propagate into every other model over time.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Same prompts sent to Fable 5 and Kimi K3 across three builds — this is the head-to-head reveal.
00:23 – 01:48
02 · Why Kimi K3 is making waves
Kimi K3 launched days earlier and is already placing top-3 on independent coding benchmarks as an open-weight, lower-cost model.
01:48 – 03:42
03 · Model ROI: intelligence, cost, speed
ROI framing: intelligence is the return, cost and speed are the investment. Kimi K3 nearly matches Fable 5's intelligence at less than half the cost, but runs about twice as slow.
03:42 – 04:38
04 · The Arena.ai frontend design upset
Arena.ai's blind-vote benchmark shows Kimi K3 beating Fable 5 on front-end code and design.
04:38 – 06:15
05 · Test setup: Kimi K3 inside Claude Code
Kimi K3 is wired into Claude Code as its harness so both models have access to the same skills and tools for a fair comparison.
06:15 – 12:58
06 · Test 1: Starbucks x FIFA World Cup site
Both models one-shot a branded campaign site; Fable 5 leans on a recognizable default aesthetic while Kimi K3 improvises a 3D-wrapped cup render and stronger copy.
12:58 – 18:16
07 · Test 2: Stretch app clone
Both models clone the core functionality of a five-million-download stretching app in Flutter; results are functionally near-identical but Kimi K3 costs a quarter as much.
18:16 – 23:30
08 · Test 3: Portal clone
Both models one-shot a playable 3D physics puzzle game with two difficulty levels; Kimi K3 looks better, Fable 5 runs smoother.
23:30 – 24:56
09 · Verdict: keep your Claude subscription?
A flat-rate subscription still beats usage-based pricing for regular use; Kimi K3 is the fallback once subscription credits are exhausted.
24:56 – 26:07
10 · Open weights and final takeaways
Kimi K3's scheduled open-weight release means every other provider can absorb its capability, lifting the baseline for everyone.
Atomic Insights
Lines worth screenshotting.
Kimi K3 scores nearly the same on Artificial Analysis's Intelligence Index as Claude Fable 5 while costing less than half as much per task.
Kimi K3 runs at roughly half the speed of Fable 5 and GPT-5.6, its clearest weakness against faster frontier models.
On Arena.ai's blind-vote benchmark for front-end code and design, Kimi K3 has already surpassed Fable 5 by a real margin.
Kimi K3 ships as an open-weight model, meaning any provider can download its published weights and fold its capability into their own models.
A branded campaign site cost $20 and 40 minutes to build with Fable 5 versus $4.70 and 1 hour 17 minutes with Kimi K3 for a comparable one-shot result.
Cloning the core functionality of a five-million-download, roughly $600K-MRR stretching app took Fable 5 $20.70/31 minutes versus Kimi K3's $5.39/1 hour 10 minutes — a 75% cost saving for about double the time.
A playable 3D Portal clone with real physics cost $29.44 (53 minutes) with Fable 5 versus $9.37 (2 hours 31 minutes) with Kimi K3, both one-shot with self-verification loops.
Kimi K3 invented a detail the prompt never requested — wrapping a generated product image around a 3D cup render — showing initiative beyond the literal brief.
Coding execution is no longer the bottleneck for building a viable app: a single detailed one-shot prompt can now clone the core functionality of a five-million-download, six-figure-MRR app.
Even with Kimi K3 priced far below Fable 5, a flat-rate Claude subscription ($100-200/month) still beats usage-based API pricing for regular heavy use, because subscriptions are heavily subsidized.
Kimi K3's provider closed new subscription signups within days of launch because demand outstripped their serving capacity, leaving usage-based API access (e.g. via OpenRouter) as the only path in.
The weakest point across both models is copywriting — AI-generated site copy still reads noticeably 'AI' and needs a human pass to sound on-brand.
A good game idea doesn't require advanced graphics: both AI-built Portal clones proved playable puzzle mechanics matter more than visual fidelity, echoing why simple-mechanic games succeed.
Takeaway
The cheap model matches the expensive one
COST VS QUALITY
Across three identical one-shot builds, Kimi K3 delivered results close to Claude Fable 5's quality at a quarter to a fifth of the cost, in exchange for roughly double the build time.
02Why Kimi K3 is making waves
Kimi K3 launched five days before this comparison and immediately placed second or third on several independent coding benchmarks despite being an open-weight model.
Artificial Analysis's blended Intelligence Index — backed by Andrew Ng and former GitHub CEO Nat Friedman — is treated as a trustworthy quick reference because it scores models across a wide range of hard tasks rather than one narrow test.
03Model ROI: intelligence, cost, speed
Judging a model by intelligence alone is incomplete: real ROI comes from dividing the output (intelligence score) by the investment (cost per task plus time to complete it).
Kimi K3 scores nearly identical to GPT-5.6 and close to Fable 5 on intelligence while costing less than half of Fable 5's price per task.
Kimi K3's clearest weakness is speed — it runs roughly twice as slow as Fable 5 and GPT-5.6, so the savings come with a real time cost.
04The Arena.ai frontend design upset
Arena.ai's blind-vote benchmark, where users pick a preferred design without knowing which model made it, showed Kimi K3 beating Fable 5 on front-end code and design by a real margin.
Design taste is inherently subjective, so a blind-vote benchmark carries different weight than a scored intelligence test — it measures what people actually prefer, not what a rubric rewards.
05Test setup: Kimi K3 inside Claude Code
Wiring an open-weight model into an existing agent harness lets it use the same skills and tools already built for the primary model, keeping a comparison apples-to-apples.
Because the harness stayed identical across both models, the comparison isolated model quality from tooling differences — a useful pattern for anyone benchmarking two models fairly.
06Test 1: Starbucks x FIFA World Cup site
On a one-shot 'showcase of extreme capability' prompt for a branded campaign site, Fable 5 defaulted to a font and layout style that's become recognizably overused by heavy Claude users.
Kimi K3 invented a detail the prompt never asked for — wrapping a generated product image around a 3D cup render — showing initiative beyond literal instruction-following.
AI-generated marketing copy is still the weak link: even a good one-shot build produces copy that reads distinctly 'AI' and needs a human copywriting pass to sound on-brand.
Fable 5 finished the campaign site in about 40 minutes for an API-equivalent $20; Kimi K3 took 1 hour 17 minutes for $4.70 — under a quarter of the cost for a comparable result.
07Test 2: Stretch app clone
Replicating the core functionality of an existing five-million-download, roughly $600K-MRR app took only one detailed one-shot prompt with each model.
Both models built functionally near-identical Flutter apps from the same brief, meaning coding execution itself is no longer the differentiator between competitors copying an app idea.
Kimi K3 completed the same build for $5.39 versus Fable 5's $20.70 — a 75% cost reduction — but took roughly twice as long.
08Test 3: Portal clone
A playable 3D puzzle game with real physics was buildable in a single one-shot prompt by both models, each self-verifying by taking screenshots and looping to fix issues.
A good game idea doesn't require advanced graphics — the core mechanic and puzzle logic carried both clones, echoing why simple-mechanic games like Flappy Bird succeeded.
Kimi K3's build had noticeably better visual polish, but Fable 5's version ran with less input lag — the two models traded off on different axes of quality within the same task.
The Portal clone cost $29.44 (53 minutes) on Fable 5 versus $9.37 (2 hours 31 minutes) on Kimi K3 — the widest time gap of the three tests.
09Verdict: keep your Claude subscription?
A flat-rate AI subscription is still cheaper than switching to usage-based pricing on a lower-cost model for regular heavy use, because subscription tiers are heavily subsidized relative to raw API pricing.
The right time to reach for a cheaper model like Kimi K3 is specifically when subscription credits or usage caps are already exhausted, not as a wholesale replacement.
Access to Kimi K3 was usage-based only at the time of testing because its provider closed new subscription signups within days of launch — demand had outstripped serving capacity.
10Open weights and final takeaways
An open-weight model publishing its weights lets every other AI provider download and study them, which tends to lift the baseline intelligence of the entire field over time.
For consumers, open-weight releases from frontier-adjacent models are a net win regardless of which specific model you use, because competitive pressure pushes free and low-cost tiers upward across the board.
Glossary
Terms worth knowing.
Artificial Analysis Intelligence Index
A blended benchmark score from artificialanalysis.ai that runs models through a range of hard tasks and averages the results into a single comparable number.
Open-weight model
A model whose trained parameters are published publicly, letting anyone download and run it themselves, or fold its capability into their own systems.
Arena.ai
A benchmark site where users are shown two anonymous model outputs side by side (e.g. two front-end designs) and blind-vote for the one they prefer, aggregating the votes into a ranking.
OpenRouter
A pay-per-use API gateway that lets developers access many different AI models through one billing account, priced by token usage rather than a flat subscription.
One-shot prompt
A single prompt sent to a model with no follow-up corrections or revisions — the model's first output is the final result being judged.
Goal prompt framework
A structured prompting format used in this video that spells out the ask, the goal/definition of done, examples, things to avoid, and available tools in one prompt.
“I think for this one, I would probably vote for Kimi K3.”
concise, opinionated verdict with a clear stance→ TikTok hook↗ Tweet quote
13:26
“The coding part isn't really the roadblock anymore. It is mostly around getting an audience as well as probably marketing this app.”
a broader, quotable claim about the state of AI coding→ newsletter pull-quote↗ Tweet quote
23:55
“Should you now cancel that and just go all in on something like Kimi K3 because of a lower cost? Probably not.”
sets up and answers the video's central practical question in one line→ IG reel cold open↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
I just gave the exact same prompts to CloudFable five and China's brand new Kimi k three across a handful of builds. A Starbucks and FIFA World Cup campaign website, a mobile app that clones a six figure MRR Android application, and finally, a playable three d copy of the video game portal.
Same rules for both, only one prompt each without any revisions. This is Fable five versus Kimi k three head to head. Let's dive into it.
So if you've been around X or YouTube lately, for sure you have seen Kimi k three. They've launched something like five days ago and they did so with this really good product launch video, I must say. But the big reason why it's making so much waves in the AI space is because it is placing at the top of a lot of different benchmarks.
So this chart, you probably have seen it where you can see Kimik A3 placing at number two or number three in these different coding benchmarks despite being an open weight model and also at a lower cost. Now there's a lot of independent benchmarks out there, but the one that I always come back to is this one by artificialanalysis.ai.
Because if you go to the homepage of their website, they actually have this really nice blended intelligence benchmark where essentially what they do is they throw these models through a variety of really hard tasks and they just measure the score of each of these models and give it a blended score. Apart from that, this firm artificialanalysis.ai is also backed by people like Andrew Ng who founded Coursera and also Nat Friedman who was ex CEO of GitHub.
GitHub. And so apart from just coming up with really good benchmarks including the speed and cost, which I'll show later, I think if you need a quick reference of all these new models as they come out, like for example, Quen three dot eight is I think coming out soon as well, then I'll probably come back to this and see where it will place in this leader board.
Any case, you can see here the intelligence index on a score out of a 100. You can see that Cloud Table five is still topping the charts, while g p d five dot six and q m k three are trailing pretty closely as number two and number three.
But apart from the intelligence, which is sort of the output or the return that you're getting, there's two things that is important for you to understand when it comes to assessing a model. Because really when you talk to a model, what you want to have is a higher return on investment.
And when you think about ROI, a return on investment, there's always two components to it. The return, which is the output that you get, which right now the best benchmark summaries that we can get is this intelligence index. But the investment part of the ROI equation, you can pretty much separate into two.
One is the cost it takes to accomplish a task, and number two is the speed by which it does that task. And so if you look at the cost per task, this is a big reason why Kimi k three is making a lot of waves. It is even cheaper than g p d five dot six sol, much cheaper, less than half the price of Cloud Table five.
But if you look at the intelligence score here, it is giving us almost the same output in terms of this index. Add to that the fact that Kimi k three is an open weight model. The reason why open weight models are really important is because once they publish their weights, essentially these other providers or even yourself if you have the hard way to run it, you can download those weights and you can run it yourself and build your own models.
So you can imagine once you have a frontier model intelligence like this that is going to release its weights that all of these other players can download, that will essentially lift the tide of each of these players so that the baseline gets higher. Which if you're thinking about access of us as users of these AI tools, then that is a massive win.
However, there is one glaring drawback with GimmickA3, which is the fact that right now, it is one of the slowest models out there. So when we talk about speed, you can see Fable five and GPT five dot six are around almost two times faster versus Kimi K three.
So it does have its drawback, but because of its performance as well as its cost and the fact that it's open weight, then as of right now and as per these benchmarks, then I think kimikatree really deserves your attention. Also, by the way, the other benchmark that is making the waves around x recently is this one by arena.ai, where they are showing how kimikatree pretty much surpassed CloudFable five when it comes to front end code and design.
And in case you don't know what arena.ai is, basically, you go to that website, you can actually try out these models for free. But essentially, when you use their tool, like, example, you want to create a front end design for a website, what you'll actually get is two results from two different random models, and you will essentially blind vote on which one you prefer.
And so this benchmark, what it does is it essentially aggregates people's votes on these front end website designs. And so this is a big deal because when it comes to matters of design and taste, everything is subjective.
And so the fact that Kimi three is eclipsing Cloud Table five in this sort of blind voting test and by quite a margin is actually a significant achievement for them. And so one of the tests that we'll be doing for sure is to compare Kimi k three's front end design versus that of CloudFable five. But alright.
Enough of the benchmarks. What I'll now be doing is I'll be testing out CloudFable five, which right now is the most frontier amongst the frontier models in terms of intelligence, but it is expensive. And I'll also be testing Kimi k three here on the right where we'll be sending the same prompt, one shot without any revisions.
So here on the left, we'll have Fable five at max effort. And here on the right, we'll have Kimi k three with max effort as well, which I've just wired up to Cloud Code as the harness. Now I've wired up Kimi k three to Cloud Code as the harness because of one simple reason.
I actually want k three to have access to the same skills that I have in my workspace as Cloud Fable five. So for example, in the website design test later, I'm going to give the freedom to these models to generate images or videos using generative AI models, which I have already linked up to some skills in ClaudeCode. So that's quite important.
Now if you want to set up Kimiko three in ClaudeCode yourself, I will link down below a guide of how I did it, which you can just access for free. And you can just either read through that or give it to your Cloud Code in order to set k m k three up in your instance. And by the way, if you want to learn how to build and sell AI systems that businesses actually pay for, then that's pretty much all we do over at the Robo Nuggets community, where not only do you get access to the Cloud Living master class, which we update every week and takes you from zero to mastery with the latest on AI, but you also get access to our agents as a service course, which walks you through how to actually get paid for all these AI skills that you are learning.
You also get to be part of a genuinely great community of AI builders. In fact, you can see just some of the recent wins our members are getting from the program right here. So if you want to start earning from AI, then check that just in the pinned comment below.
Now back to the video. Alright. Now let's go to the first test.
So I'll just send over these prompts and then I'll show you. Basically, what this first test does is it's going to be a website design for a campaign for a Starbucks x FIFA World Cup twenty twenty six partnership where we're asking them to build one self contained HTML file. And, specifically, we want this to be treated as a showcase of extreme capability in terms of visuals.
The core idea being is that Starbucks is offering limited edition drinks for of the top eight countries of the FIFA World Cup, and this is a goal prompt as well. So we're just using the agent framework for goal prompts that we have mentioned before where we have the ask, we have the goal, we have some examples, some negations of what not to do, as well as tools.
But pretty much if you read through this, which I'll probably also link below, you can see that this is quite general. The only other thing that I wanted to mention here is I actually give them a $10 budget to generate images or videos as they see see fit using GPT image two for stills and Clang three dot zero for motion.
So we'll see what that looks like once it's done. Alright. So now both of these sessions have completed their task.
Alright. So this is from Cloud Fable five. So there, you can see eight nations, eight cups, one final pour.
So the thing with Fable is it really likes this font with the fancy f's. I think it's called Francis. And although not a lot of people looking at this will consider it vibe coded because not a lot of people probably know what Fable's designs default looks like.
But if you were building out and using Claude quite a lot, then you might notice this font as well as this design being overused, I would say. But I think overall for a one shot prompt, this is quite good. Like, the fact that you have animations for attacks there is quite nice.
And here we go. So it was able to generate these using GPT image two, and it decided that Starbucks is going to launch these cups for each of these countries in this sort of dome shaped format.
And as I scroll through them, you'll be able to see each of these countries here, the top eight participating in the World Cup. And you have all of the portfolio of those cups available to you.
So it did it pretty well again coming from just one shot prompt. The only thing that I think is still missing with a lot of these models when it comes to the taste component is when it comes to the copy, meaning what it actually says. So if I just scroll back up here, like this one, the torment last month, the queue last minutes, the cup in your hand remembers.
That whole paragraph reads very AI to me. But the good news is if you have some level of copywriting skill, you can probably tweak that so that it reads more like a human and it reads more like in Starbucks' tone of voice. And that's much easier to tweak versus the overall vibe and the design of this site.
So that is what Fablefire created for us with that one prompt. Alright. So now this is from Kimi k three.
Here we go. So it also generated a video in the background. Eight countries, eight drinks.
It has that nice sort of negative space there for the eight drinks, which is looking pretty good, and it even has the 26 at the bottom. So if I scroll down, the world's game, poured by the world's coffee house, taste the final eight. Now I think that copy is much better versus what Fable five gave us.
A bit more catchy, I would say. And then it has some subtext here. And then if you go down interesting.
Alright. So if I scroll down here, what it actually did, if you can notice, is it created this three d rendering of a cup, and it created also this is probably an image that it created using g p d image too, and it decided to put the image wrapping this three d object as its rendering of the cup. So I did not provide that prompt to it.
It came up with that idea all by itself. Now, obviously, if you were designing this for real, like, is a really good idea. I'll probably run with this if I were to create this for real.
But, uh, obviously, the aspect ratio of of the logo there shouldn't be oval shaped. And so you can probably optimize this image more and the three d renderings so that it is representing the cup more accurately if we were doing this campaign for real. But, like, this idea of just scrolling down and, uh, it transitions into this new country cup is actually really, really good, and it's a good idea to start off with.
And I really like how it gave us these flavors as well. So that's a very novel idea. And I think Kimi k three, like, if I were to choose which one won this round, it's probably k three, to be honest.
Here for the tournament, gone with the trophy. Alright. So this is what Kimi k three gave us with a one shot prompt with access to g p d image two as well as cling k three.
So this is the one from Fable five, and this is the one from Kimi k three. Very, very close. And I think for this one, I would probably vote for Kimi k three.
So let me know below which one you prefer. Because, again, front end design, very subjective, but I think the level of what we have with these AI models now, it is really, really easy to create websites that are a bit more eye catchy than your usual vibe coated purple AI websites. But like what I said earlier, the important thing to consider when it comes to these AI models is to measure your ROI from them.
So your return on your investment. Right? Because at least for this use case, this is the output that you're getting, but how much did it actually cost and how much time did it actually take to produce these results.
And so what I just did is have each of those models self report on these use cases on what the token cost is as well as the time that it took. Now the cost here, just for clarity and avoidance of doubt, is the cost that would have been if we used API usage based pricing. So for CloudFable five, because I'm in subscription, I actually didn't pay $35.
For Kimikay three, since I'm using it via a service called OpenRouter, I actually did pay $4.7 for that build. Now do note the reason why this is probably more expensive than usual is because I actually asked if you go back to that prompt, I actually asked these models to self check their work, which is why they took screenshots as well as repeated a few loops in order to just refine their work.
Now when you're constrained for budget, you probably won't be doing that. But just know that if you have these models loop to check their work before giving it to you, then obviously, it would cost much more. Now this is what we're saying earlier.
Right? Because as we know, Cloud Fable five, terribly expensive, but it does finish a task much faster. So end to end, it took roughly forty minutes to build that site.
And remember, this is at max effort, so that's why it probably took longer. While with Kimikay three, it took around one hour and seventeen minutes. So if you're doing front end website design, what is the key takeaway?
Well, I think the key takeaway is the fact that Kimi k three is really good at it. It depends with your prompt, obviously, but I think that arita.a I benchmark that we saw earlier is quite real.
Like, I can imagine how k three can surpass Fable five with, uh, different prompts, with different brands, with different designs. And at this level of cost, especially if, for example, you've already maxed out your Fable five credits in your cloud subscription and you need to pay usage based pricing anyway, I would probably see if I can just offload that task with Qumik three even despite the fact that it is taking a bit longer.
Alright. Let's move on to the second prompt. So I'm going to send this two prompts, and then I'll just take you through it.
Because what we're now going to do is to build an Android app. And if we expand this prompt, you can see we're building this Android app called Stretch, which offers guided stretch routines with beautiful illustrations and timers. Now if you read through that prompt, which I'll also provide below, you can see that it is pretty much emulating the features of this app, this existing app called Bend, and it has already a 157,000 reviews with 5,000,000 plus downloads.
And when I did a quick Google search, they were reportedly making something along the lines of 600 k in terms of MRR, which is pretty insane because if you actually download this app, all it really does is give you a set of stretching routines with this really nice set of images to go along with it. So it's sort of like a timer app, but you can definitely see, like, from just the reviews here and the amount of usership that it has that it is quite valuable.
So if Fable and Kimikatree are successful in building this out, then that just alludes to the point that the coding part isn't really the roadblock anymore. It is mostly around getting an audience as well as probably marketing this app. So we'll see how both of those models do once they finish that prompt.
And just to go through that prompt once again, you can see I am using the agent framework for this goal prompt, which I declared a clear ask. I am declaring a goal, so stop only with the app builds and launches on a standard Android emulator. And at least for this build so that it won't take so much time, I just ask it to include eight distinct structures, each with its own illustration, every stretch illustration shown, again, I'm using or having these models use GPT image two to render images, and I also want at least four routines built from those stretches.
I gave it some guidance on some examples, but not too specific, and some guidance on what not to do as well as the tools that it can use. And then once again, I gave it a hard cap of how much to spend so that we have a fair playing field between these two models.
Alright. So now with those two tasks done, you can see they used Flutter as well as Dart or at least, uh, k m k three did, and Fable five probably also did.
So, yeah, you can see the stack is Flutter, and everything got built in their individual emulators, which I have on my other screen. So this is the one from CloudFable five, and there you go.
It has four routines, one for waking up, one for anytime, one for a desk break, post workout, and before sleep. So let's say we wanna try out the waking up routine.
You can see there's a few stretches in here, and all these images are generated via g b d image too. And I like how consistent it all is, and it's just that one design. And, again, it decided the prompts for these icons by itself.
Because this is just a simple timer, you can see it has this neck release stretch. It has instructions on how to do it, a timer on doing it. And then if we go to the next, we can actually do that to do the cat curl.
So you have instructions there as well. And, uh, yeah, it's pretty basic, but if you try the original inspiration for this, the Bend app, this pretty much what it does. And it has what, like, more than a million users.
Right? And so something as simple as this actually does add a lot of value for its users. Right now, we just did four routines in here, but you can just imagine how you can probably have CloudFable five ideate, like, 10 more routines for the different exercises, maybe like a nighttime routine, maybe one that's more focused on yoga.
And at least right now in its most basic form, it does function well. And I think the look of it, especially these images, are also as per the brief, as per the prompt that we gave it. So props to CloudTable five for that, especially given that it was only done with just one prompt.
Now if we go and see what, um, Kimi k three did, it is actually quite similar. Like, this is the homepage of it. If I go back to the homepage of Club Fable five, I think the only difference is, number one, it doesn't really have the core images of these workout routines here at the very top, but it did include all of the stretch es here at the bottom if in case you want to try them out.
So if we try one of these, it's pretty much the same, but the difference is the images are slightly different, but it functions similarly, and the front end design of it is also the same. So, really, if you're after an Android app, you can pretty much one shot this whole emulator experience running on your own desktop and just create an APK file from this to run it on your own phone and then just submit it to Google Player.
Right? And so the point being the barrier to entry to creating applications like these have never been lower. Now, obviously, because these two are so similar, probably want to just look at the cost and the time that it took for these two models to have created these builds.
And again, it's the same pattern as earlier. So CloudFable five costed around $20 if you were using API usage based pricing at around half the time that it took Kimi k three to do it. But that is four times the cost of Kimi k three at only $5 to create that whole experience.
Now time wise, it did take an hour and ten minutes, so roughly twice as long as Cloud Fable five. But this is just one prompt, and you wanted it to run-in the background and realize, like, 75% savings in the process, then that's a really big deal, especially since the output, as you saw earlier, were pretty much the same in terms of the design aesthetic as well as the functionality.
Now onward to our third test. And this prompt is quite interesting because what we'll be doing is we'll be replicating this video game called Portal. And in case you haven't played Portal before, it is quite an interesting test because the core mechanic of it is you actually can shoot out these blue and orange portals so that you can teleport to wherever the other portal is leading to.
And I think Portal would be a really good test because it's essentially a puzzle game. Right? And it has a lot of these physics based movements that I think would be a sort of a good evaluation of how good these models really are.
The other reason is with games like these, you actually don't need a lot of really good graphics in order to have a good idea for a game. As long as the core idea and mechanic is good, then you'll actually be able to have a pretty successful game. Like, for example, Minecraft or even Flappy Bird, if you remember that.
Those have been pretty successful just because of the core idea and the core mechanic. That said, we are going to ask these models to do this in just one shot.
And what we're asking them to do is to build a playable three d clone of Portal as a single self contained HTML file. So we'll be opening it in our browser. And so we just enumerated a few of the core mechanics that must work and they must verify.
And the other thing is we wanted two levels for now. So one that is quite easy and the second one which is more difficult. And then same with the framework that we have for this goal prompt, we gave it some written instructions on a few examples, a few negations, as well as tools.
Other thing is, again, because we want it to be visually appealing, we gave it the chance to use GPD image two with a budget if you need to render images, let's say textures for in game assets. Alright. So both of those sessions are now done.
And this is the one from Clawtable five called Slingshot, two chamber portal physics test. And this is wild. So you can see that the goal, it seems, is we need to stand on this button, but we need to get that box over at the top.
So if I click on the left mouse button that launches the blue portal, so it has that capability. So Fable was able to do that pretty well. And I think it says here I can press e to grab this, jump, and just put that there.
There you go. We solved level one pretty quickly. K.
So chamber two supposedly needs to be harder. So it seems that we'll probably need to have a portal here, then launch ourselves here. And then if I can fling myself there, I'll be able to go to this side.
So that probably takes a bit more precision. Alright. Level two, because our brief is to make it super difficult.
It is actually really difficult. But the core idea is if I place the portals here, I'll be able to fling myself to this platform in the center with this box while launching myself like that.
There you go. Okay. And then I'll just place it here.
K. And then we'll finish the level that way. Perfect.
So that is a pretty good run, actually. That's a pretty good one shot prompt. Basically, just being able to have that pretty complex physics built in and having, like, a full portal game that obviously the whole visuals of this you can improve.
It's not like the actual portal game. But the fact that you have everything in order to make, like, a proper puzzle game, I think is really powerful, especially coming from just a one shot prop. So kudos to Cloud Table five.
Alright. Now we go to what Kimi k three gave us. So it's called aperture Kimi.
So pretty much the same controls. Let's click to begin. K.
So basically, we need to get that. Right? And if we go and oh.
Alright. Kinda works. I do have to say though that there is a bit of lag.
So if we are probably going for speediness, we can give it a second prompt to improve that. But that first level is pretty easy as instructed. Okay.
So now his second level, it's kinda hard. I I don't think I can solve this, but let me try. I think the graphics are a bit better with QEMI.
So again, alluding to that front end design benchmark that we saw earlier. But in terms of speediness, Fable five was able to give us a much better experience because there is a bit of lag here.
But the fact that both of these models are giving us, especially Kimik three with this pretty complex puzzle that can probably pass as an actual stage in a game like Portal, is pretty insane, honestly. So, yeah, there is that exit.
It'll probably take me a few minutes to figure this out, guys. Especially the point being that if you have, like, an idea, especially that of a game and even if it has, like, complex physics like these, then you can probably do it now with the smartness of these models. Now to build that portal clone, again, it's the same pattern as we saw earlier.
Cloud Fable five, at least for this run, it was three times more expensive versus Kimi. Like, imagine that two stage thing that was built with just one shot. It did take two hours and thirty minutes though because the prompt that I gave, I always ask them to verify the stages.
And at least with the Kimi k three one, the second stage was a bit more complex, so it probably took some turns in order to make sure that that one is solvable. But that took two point five hours to finish. But at a cost that is less than $10 to build out that whole thing, that is pretty incredible.
Right? Now with CloudFable five, again, that takes less time but at a much higher cost. So what do we now take away from this?
Because if you have a subscription to Claude, should you now cancel that and just go all in on something like Kimikatree because of a lower cost? Probably not. Right?
Because at the end of the day, the subscriptions that Entropic as well as OpenAI is offering us is still much cheaper and highly subsidized to access these models. So take advantage of that, especially since there's still some question on whether that will last years into the future.
But at least right now, having, let's say, a max plan of a $100 a month or $200 a month to give you a lot of tokens and access to models like Fable five is still the cheapest ways by which you can access intelligence levels of this skill. Now that said, if you do run out of table five credits or usage, then that is when you should check out Kimik A Tree, especially since right now you can only access Kimik A Tree through usage based pricing.
And the reason why that is is because if you go to the Kimi K three membership plans, they actually closed this recently. So you can see it says join waitlist here right now. So you can't really get Kimi K three on a subscription plan as of the moment.
And the reason why they did that, they mentioned it in this x post where basically over the past couple of hours, the demand for Kimi K three was so much that they couldn't support it at current capacity, so they had to close the subscriptions as of right now. But look, Kimikitri, it's an open weight model.
I think five days from now, July 27, they'll release the weights. And one of the big implications of that, like I mentioned in the beginning, is the fact that every other model company will definitely download that file and implement it in order to improve their own models. So all of these ones will probably get to something like almost Fable level intelligence at some point because Kimikatree is going to open weight and provide basically access to their secret sauce in order to make this intelligence level possible.
So overall, it's a win for the consumer. It's a win for us as users. I think definitely try Kimi K Tree out.
It is almost as good as Fable as you can see in the tree tests that we did and at a fraction of the cost. So if you want to test it out over a cloud code as your harness, then the guide for that, I just ship that from my system, from my workspace so that you can try it out as well.
But there, I hope that was useful. And as always, thank you for watching until the end, and I'll see you all next time. Cheers.
The Hook
The bait, then the rug-pull.
Same prompts, same rules, zero revisions — the creator pits Claude Fable 5 against the just-launched open-weight Kimi K3 on three real builds: a campaign site, an app clone, and a playable 3D game. What comes back isn't a benchmark screenshot, it's a receipt.
Frameworks
Named ideas worth stealing.
01:50model
Model ROI (Return / Investment)
Return = intelligence score
Investment = cost per task
Investment = time/speed to complete task
A model's raw intelligence score is only half the picture — real ROI divides that return by what it costs in money and time to get it.
Steal forframing any AI-model or tool comparison beyond a single leaderboard number
06:15list
Goal prompt agent framework
Ask
Goal
Examples
Negations (what not to do)
Tools
A structured one-shot prompt template used for all three tests: state the ask, define the goal/done-condition, give examples, list what to avoid, and specify available tools.
Steal forwriting repeatable one-shot prompts for coding agents
CTA Breakdown
How they asked for the click.
VERBAL ASK
05:30product
“check that just in the pinned comment below”
soft mid-video pitch for the free Kimi K3 + Claude Code setup guide and the paid RoboNuggets community, woven into the test-setup segment rather than a hard break
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Anthropic shipped a Figma-style artboard editor into Claude Code itself, and one identical prompt with real context beat Claude Design's default template every time.
Two five-minute configuration changes, a custom output style and an on-demand skill, turn Opus 5's dense jargon and wall-of-text replies into plain, scannable answers.
A three-line prompt that fans Claude out into paired builder and critic sub-agents until every piece clears a stated quality bar — and the one condition that decides whether it helps or hurts.
A walkthrough of routing Claude Code through a ChatGPT subscription's GPT-5.6 Sol model, plus a two-model skill that has Claude plan while Sol builds, for roughly a quarter to a half less per task.