Modern Creator
Jay E | RoboNuggets · YouTube

I Made Fable 5 and Kimi K3 Build the Same App

Same one-shot prompts, two models, three real builds — a website, an app clone, and a 3D game — scored on intelligence, cost, and time, not just vibes.

Posted
yesterday
Duration
Format
Review
educational
Views
9K
206 likes
Part of the collectionThe Fable 5 PlaybookAll 45 Fable 5 breakdowns, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Kimi K3, a brand-new open-weight model, produces results close enough to Claude Fable 5's quality across three real one-shot builds that the deciding factor becomes ROI — cost and time — rather than raw intelligence.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use Claude Code or a similar coding-agent harness and want to know if a cheaper model is worth routing tasks to.
  • You're deciding whether to keep paying for a Claude subscription versus switching to usage-based API pricing on a different model.
  • You want a real cost-and-time benchmark for one-shot AI app builds, not just an intelligence leaderboard screenshot.
  • You're evaluating whether an existing app idea (like a stretching or habit app) is now trivial to clone with AI.
SKIP IF…
  • You want a deep technical explainer of Kimi K3's architecture — this is a hands-on build comparison, not a model-internals breakdown.
  • You're not using AI coding agents at all and just want general AI news.
TL;DR

The full version, fast.

The creator ran identical one-shot prompts through Claude Fable 5 and the newly launched open-weight Kimi K3 across three builds: a Starbucks x FIFA World Cup campaign site, a clone of a five-million-download stretching app, and a playable 3D Portal clone. Kimi K3 matched Fable 5 closely on Artificial Analysis's intelligence benchmark and even beat it on Arena.ai's blind-vote design benchmark, while costing a fraction as much per task — $4.70-$9.37 versus $20-$29 — at roughly double the build time. The conclusion: a flat-rate Claude subscription still beats usage-based pricing for regular use, but Kimi K3 becomes the obvious fallback once subscription credits run dry, and its scheduled open-weight release means its capability will propagate into every other model over time.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:23

01 · Cold open

Same prompts sent to Fable 5 and Kimi K3 across three builds — this is the head-to-head reveal.

00:2301:48

02 · Why Kimi K3 is making waves

Kimi K3 launched days earlier and is already placing top-3 on independent coding benchmarks as an open-weight, lower-cost model.

01:4803:42

03 · Model ROI: intelligence, cost, speed

ROI framing: intelligence is the return, cost and speed are the investment. Kimi K3 nearly matches Fable 5's intelligence at less than half the cost, but runs about twice as slow.

03:4204:38

04 · The Arena.ai frontend design upset

Arena.ai's blind-vote benchmark shows Kimi K3 beating Fable 5 on front-end code and design.

04:3806:15

05 · Test setup: Kimi K3 inside Claude Code

Kimi K3 is wired into Claude Code as its harness so both models have access to the same skills and tools for a fair comparison.

06:1512:58

06 · Test 1: Starbucks x FIFA World Cup site

Both models one-shot a branded campaign site; Fable 5 leans on a recognizable default aesthetic while Kimi K3 improvises a 3D-wrapped cup render and stronger copy.

12:5818:16

07 · Test 2: Stretch app clone

Both models clone the core functionality of a five-million-download stretching app in Flutter; results are functionally near-identical but Kimi K3 costs a quarter as much.

18:1623:30

08 · Test 3: Portal clone

Both models one-shot a playable 3D physics puzzle game with two difficulty levels; Kimi K3 looks better, Fable 5 runs smoother.

23:3024:56

09 · Verdict: keep your Claude subscription?

A flat-rate subscription still beats usage-based pricing for regular use; Kimi K3 is the fallback once subscription credits are exhausted.

24:5626:07

10 · Open weights and final takeaways

Kimi K3's scheduled open-weight release means every other provider can absorb its capability, lifting the baseline for everyone.

Atomic Insights

Lines worth screenshotting.

  • Kimi K3 scores nearly the same on Artificial Analysis's Intelligence Index as Claude Fable 5 while costing less than half as much per task.
  • Kimi K3 runs at roughly half the speed of Fable 5 and GPT-5.6, its clearest weakness against faster frontier models.
  • On Arena.ai's blind-vote benchmark for front-end code and design, Kimi K3 has already surpassed Fable 5 by a real margin.
  • Kimi K3 ships as an open-weight model, meaning any provider can download its published weights and fold its capability into their own models.
  • A branded campaign site cost $20 and 40 minutes to build with Fable 5 versus $4.70 and 1 hour 17 minutes with Kimi K3 for a comparable one-shot result.
  • Cloning the core functionality of a five-million-download, roughly $600K-MRR stretching app took Fable 5 $20.70/31 minutes versus Kimi K3's $5.39/1 hour 10 minutes — a 75% cost saving for about double the time.
  • A playable 3D Portal clone with real physics cost $29.44 (53 minutes) with Fable 5 versus $9.37 (2 hours 31 minutes) with Kimi K3, both one-shot with self-verification loops.
  • Kimi K3 invented a detail the prompt never requested — wrapping a generated product image around a 3D cup render — showing initiative beyond the literal brief.
  • Coding execution is no longer the bottleneck for building a viable app: a single detailed one-shot prompt can now clone the core functionality of a five-million-download, six-figure-MRR app.
  • Even with Kimi K3 priced far below Fable 5, a flat-rate Claude subscription ($100-200/month) still beats usage-based API pricing for regular heavy use, because subscriptions are heavily subsidized.
  • Kimi K3's provider closed new subscription signups within days of launch because demand outstripped their serving capacity, leaving usage-based API access (e.g. via OpenRouter) as the only path in.
  • The weakest point across both models is copywriting — AI-generated site copy still reads noticeably 'AI' and needs a human pass to sound on-brand.
  • A good game idea doesn't require advanced graphics: both AI-built Portal clones proved playable puzzle mechanics matter more than visual fidelity, echoing why simple-mechanic games succeed.
Takeaway

The cheap model matches the expensive one

COST VS QUALITY

Across three identical one-shot builds, Kimi K3 delivered results close to Claude Fable 5's quality at a quarter to a fifth of the cost, in exchange for roughly double the build time.

02Why Kimi K3 is making waves
  • Kimi K3 launched five days before this comparison and immediately placed second or third on several independent coding benchmarks despite being an open-weight model.
  • Artificial Analysis's blended Intelligence Index — backed by Andrew Ng and former GitHub CEO Nat Friedman — is treated as a trustworthy quick reference because it scores models across a wide range of hard tasks rather than one narrow test.
03Model ROI: intelligence, cost, speed
  • Judging a model by intelligence alone is incomplete: real ROI comes from dividing the output (intelligence score) by the investment (cost per task plus time to complete it).
  • Kimi K3 scores nearly identical to GPT-5.6 and close to Fable 5 on intelligence while costing less than half of Fable 5's price per task.
  • Kimi K3's clearest weakness is speed — it runs roughly twice as slow as Fable 5 and GPT-5.6, so the savings come with a real time cost.
04The Arena.ai frontend design upset
  • Arena.ai's blind-vote benchmark, where users pick a preferred design without knowing which model made it, showed Kimi K3 beating Fable 5 on front-end code and design by a real margin.
  • Design taste is inherently subjective, so a blind-vote benchmark carries different weight than a scored intelligence test — it measures what people actually prefer, not what a rubric rewards.
05Test setup: Kimi K3 inside Claude Code
  • Wiring an open-weight model into an existing agent harness lets it use the same skills and tools already built for the primary model, keeping a comparison apples-to-apples.
  • Because the harness stayed identical across both models, the comparison isolated model quality from tooling differences — a useful pattern for anyone benchmarking two models fairly.
06Test 1: Starbucks x FIFA World Cup site
  • On a one-shot 'showcase of extreme capability' prompt for a branded campaign site, Fable 5 defaulted to a font and layout style that's become recognizably overused by heavy Claude users.
  • Kimi K3 invented a detail the prompt never asked for — wrapping a generated product image around a 3D cup render — showing initiative beyond literal instruction-following.
  • AI-generated marketing copy is still the weak link: even a good one-shot build produces copy that reads distinctly 'AI' and needs a human copywriting pass to sound on-brand.
  • Fable 5 finished the campaign site in about 40 minutes for an API-equivalent $20; Kimi K3 took 1 hour 17 minutes for $4.70 — under a quarter of the cost for a comparable result.
07Test 2: Stretch app clone
  • Replicating the core functionality of an existing five-million-download, roughly $600K-MRR app took only one detailed one-shot prompt with each model.
  • Both models built functionally near-identical Flutter apps from the same brief, meaning coding execution itself is no longer the differentiator between competitors copying an app idea.
  • Kimi K3 completed the same build for $5.39 versus Fable 5's $20.70 — a 75% cost reduction — but took roughly twice as long.
08Test 3: Portal clone
  • A playable 3D puzzle game with real physics was buildable in a single one-shot prompt by both models, each self-verifying by taking screenshots and looping to fix issues.
  • A good game idea doesn't require advanced graphics — the core mechanic and puzzle logic carried both clones, echoing why simple-mechanic games like Flappy Bird succeeded.
  • Kimi K3's build had noticeably better visual polish, but Fable 5's version ran with less input lag — the two models traded off on different axes of quality within the same task.
  • The Portal clone cost $29.44 (53 minutes) on Fable 5 versus $9.37 (2 hours 31 minutes) on Kimi K3 — the widest time gap of the three tests.
09Verdict: keep your Claude subscription?
  • A flat-rate AI subscription is still cheaper than switching to usage-based pricing on a lower-cost model for regular heavy use, because subscription tiers are heavily subsidized relative to raw API pricing.
  • The right time to reach for a cheaper model like Kimi K3 is specifically when subscription credits or usage caps are already exhausted, not as a wholesale replacement.
  • Access to Kimi K3 was usage-based only at the time of testing because its provider closed new subscription signups within days of launch — demand had outstripped serving capacity.
10Open weights and final takeaways
  • An open-weight model publishing its weights lets every other AI provider download and study them, which tends to lift the baseline intelligence of the entire field over time.
  • For consumers, open-weight releases from frontier-adjacent models are a net win regardless of which specific model you use, because competitive pressure pushes free and low-cost tiers upward across the board.
Glossary

Terms worth knowing.

Artificial Analysis Intelligence Index
A blended benchmark score from artificialanalysis.ai that runs models through a range of hard tasks and averages the results into a single comparable number.
Open-weight model
A model whose trained parameters are published publicly, letting anyone download and run it themselves, or fold its capability into their own systems.
Arena.ai
A benchmark site where users are shown two anonymous model outputs side by side (e.g. two front-end designs) and blind-vote for the one they prefer, aggregating the votes into a ranking.
OpenRouter
A pay-per-use API gateway that lets developers access many different AI models through one billing account, priced by token usage rather than a flat subscription.
One-shot prompt
A single prompt sent to a model with no follow-up corrections or revisions — the model's first output is the final result being judged.
Goal prompt framework
A structured prompting format used in this video that spells out the ask, the goal/definition of done, examples, things to avoid, and available tools in one prompt.
Resources

Things they pointed at.

03:48toolarena.ai
13:23productBend: Stretching & Flexibility (Google Play)
Quotables

Lines you could clip.

11:00
I think for this one, I would probably vote for Kimi K3.
concise, opinionated verdict with a clear stanceTikTok hook↗ Tweet quote
13:26
The coding part isn't really the roadblock anymore. It is mostly around getting an audience as well as probably marketing this app.
a broader, quotable claim about the state of AI codingnewsletter pull-quote↗ Tweet quote
23:55
Should you now cancel that and just go all in on something like Kimi K3 because of a lower cost? Probably not.
sets up and answers the video's central practical question in one lineIG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
00:00I just gave the exact same prompts to CloudFable five and China's brand new Kimi k three across a handful of builds. A Starbucks and FIFA World Cup campaign website, a mobile app that clones a six figure MRR Android application, and finally, a playable three d copy of the video game portal.
00:16Same rules for both, only one prompt each without any revisions. This is Fable five versus Kimi k three head to head. Let's dive into it.
00:23So if you've been around X or YouTube lately, for sure you have seen Kimi k three. They've launched something like five days ago and they did so with this really good product launch video, I must say. But the big reason why it's making so much waves in the AI space is because it is placing at the top of a lot of different benchmarks.
00:39So this chart, you probably have seen it where you can see Kimik A3 placing at number two or number three in these different coding benchmarks despite being an open weight model and also at a lower cost. Now there's a lot of independent benchmarks out there, but the one that I always come back to is this one by artificialanalysis.ai.
00:57Because if you go to the homepage of their website, they actually have this really nice blended intelligence benchmark where essentially what they do is they throw these models through a variety of really hard tasks and they just measure the score of each of these models and give it a blended score. Apart from that, this firm artificialanalysis.ai is also backed by people like Andrew Ng who founded Coursera and also Nat Friedman who was ex CEO of GitHub.
01:20GitHub. And so apart from just coming up with really good benchmarks including the speed and cost, which I'll show later, I think if you need a quick reference of all these new models as they come out, like for example, Quen three dot eight is I think coming out soon as well, then I'll probably come back to this and see where it will place in this leader board.
01:35Any case, you can see here the intelligence index on a score out of a 100. You can see that Cloud Table five is still topping the charts, while g p d five dot six and q m k three are trailing pretty closely as number two and number three.
01:48But apart from the intelligence, which is sort of the output or the return that you're getting, there's two things that is important for you to understand when it comes to assessing a model. Because really when you talk to a model, what you want to have is a higher return on investment.
02:03And when you think about ROI, a return on investment, there's always two components to it. The return, which is the output that you get, which right now the best benchmark summaries that we can get is this intelligence index. But the investment part of the ROI equation, you can pretty much separate into two.
02:17One is the cost it takes to accomplish a task, and number two is the speed by which it does that task. And so if you look at the cost per task, this is a big reason why Kimi k three is making a lot of waves. It is even cheaper than g p d five dot six sol, much cheaper, less than half the price of Cloud Table five.
02:37But if you look at the intelligence score here, it is giving us almost the same output in terms of this index. Add to that the fact that Kimi k three is an open weight model. The reason why open weight models are really important is because once they publish their weights, essentially these other providers or even yourself if you have the hard way to run it, you can download those weights and you can run it yourself and build your own models.
02:58So you can imagine once you have a frontier model intelligence like this that is going to release its weights that all of these other players can download, that will essentially lift the tide of each of these players so that the baseline gets higher. Which if you're thinking about access of us as users of these AI tools, then that is a massive win.
03:17However, there is one glaring drawback with GimmickA3, which is the fact that right now, it is one of the slowest models out there. So when we talk about speed, you can see Fable five and GPT five dot six are around almost two times faster versus Kimi K three.
03:32So it does have its drawback, but because of its performance as well as its cost and the fact that it's open weight, then as of right now and as per these benchmarks, then I think kimikatree really deserves your attention. Also, by the way, the other benchmark that is making the waves around x recently is this one by arena.ai, where they are showing how kimikatree pretty much surpassed CloudFable five when it comes to front end code and design.
03:55And in case you don't know what arena.ai is, basically, you go to that website, you can actually try out these models for free. But essentially, when you use their tool, like, example, you want to create a front end design for a website, what you'll actually get is two results from two different random models, and you will essentially blind vote on which one you prefer.
04:12And so this benchmark, what it does is it essentially aggregates people's votes on these front end website designs. And so this is a big deal because when it comes to matters of design and taste, everything is subjective.
04:23And so the fact that Kimi three is eclipsing Cloud Table five in this sort of blind voting test and by quite a margin is actually a significant achievement for them. And so one of the tests that we'll be doing for sure is to compare Kimi k three's front end design versus that of CloudFable five. But alright.
04:39Enough of the benchmarks. What I'll now be doing is I'll be testing out CloudFable five, which right now is the most frontier amongst the frontier models in terms of intelligence, but it is expensive. And I'll also be testing Kimi k three here on the right where we'll be sending the same prompt, one shot without any revisions.
04:55So here on the left, we'll have Fable five at max effort. And here on the right, we'll have Kimi k three with max effort as well, which I've just wired up to Cloud Code as the harness. Now I've wired up Kimi k three to Cloud Code as the harness because of one simple reason.
05:08I actually want k three to have access to the same skills that I have in my workspace as Cloud Fable five. So for example, in the website design test later, I'm going to give the freedom to these models to generate images or videos using generative AI models, which I have already linked up to some skills in ClaudeCode. So that's quite important.
05:27Now if you want to set up Kimiko three in ClaudeCode yourself, I will link down below a guide of how I did it, which you can just access for free. And you can just either read through that or give it to your Cloud Code in order to set k m k three up in your instance. And by the way, if you want to learn how to build and sell AI systems that businesses actually pay for, then that's pretty much all we do over at the Robo Nuggets community, where not only do you get access to the Cloud Living master class, which we update every week and takes you from zero to mastery with the latest on AI, but you also get access to our agents as a service course, which walks you through how to actually get paid for all these AI skills that you are learning.
06:02You also get to be part of a genuinely great community of AI builders. In fact, you can see just some of the recent wins our members are getting from the program right here. So if you want to start earning from AI, then check that just in the pinned comment below.
06:13Now back to the video. Alright. Now let's go to the first test.
06:16So I'll just send over these prompts and then I'll show you. Basically, what this first test does is it's going to be a website design for a campaign for a Starbucks x FIFA World Cup twenty twenty six partnership where we're asking them to build one self contained HTML file. And, specifically, we want this to be treated as a showcase of extreme capability in terms of visuals.
06:37The core idea being is that Starbucks is offering limited edition drinks for of the top eight countries of the FIFA World Cup, and this is a goal prompt as well. So we're just using the agent framework for goal prompts that we have mentioned before where we have the ask, we have the goal, we have some examples, some negations of what not to do, as well as tools.
06:56But pretty much if you read through this, which I'll probably also link below, you can see that this is quite general. The only other thing that I wanted to mention here is I actually give them a $10 budget to generate images or videos as they see see fit using GPT image two for stills and Clang three dot zero for motion.
07:14So we'll see what that looks like once it's done. Alright. So now both of these sessions have completed their task.
07:19Alright. So this is from Cloud Fable five. So there, you can see eight nations, eight cups, one final pour.
07:25So the thing with Fable is it really likes this font with the fancy f's. I think it's called Francis. And although not a lot of people looking at this will consider it vibe coded because not a lot of people probably know what Fable's designs default looks like.
07:39But if you were building out and using Claude quite a lot, then you might notice this font as well as this design being overused, I would say. But I think overall for a one shot prompt, this is quite good. Like, the fact that you have animations for attacks there is quite nice.
07:53And here we go. So it was able to generate these using GPT image two, and it decided that Starbucks is going to launch these cups for each of these countries in this sort of dome shaped format.
08:05And as I scroll through them, you'll be able to see each of these countries here, the top eight participating in the World Cup. And you have all of the portfolio of those cups available to you.
08:15So it did it pretty well again coming from just one shot prompt. The only thing that I think is still missing with a lot of these models when it comes to the taste component is when it comes to the copy, meaning what it actually says. So if I just scroll back up here, like this one, the torment last month, the queue last minutes, the cup in your hand remembers.
08:32That whole paragraph reads very AI to me. But the good news is if you have some level of copywriting skill, you can probably tweak that so that it reads more like a human and it reads more like in Starbucks' tone of voice. And that's much easier to tweak versus the overall vibe and the design of this site.
08:51So that is what Fablefire created for us with that one prompt. Alright. So now this is from Kimi k three.
08:56Here we go. So it also generated a video in the background. Eight countries, eight drinks.
09:00It has that nice sort of negative space there for the eight drinks, which is looking pretty good, and it even has the 26 at the bottom. So if I scroll down, the world's game, poured by the world's coffee house, taste the final eight. Now I think that copy is much better versus what Fable five gave us.
09:15A bit more catchy, I would say. And then it has some subtext here. And then if you go down interesting.
09:21Alright. So if I scroll down here, what it actually did, if you can notice, is it created this three d rendering of a cup, and it created also this is probably an image that it created using g p d image too, and it decided to put the image wrapping this three d object as its rendering of the cup. So I did not provide that prompt to it.
09:42It came up with that idea all by itself. Now, obviously, if you were designing this for real, like, is a really good idea. I'll probably run with this if I were to create this for real.
09:50But, uh, obviously, the aspect ratio of of the logo there shouldn't be oval shaped. And so you can probably optimize this image more and the three d renderings so that it is representing the cup more accurately if we were doing this campaign for real. But, like, this idea of just scrolling down and, uh, it transitions into this new country cup is actually really, really good, and it's a good idea to start off with.
10:13And I really like how it gave us these flavors as well. So that's a very novel idea. And I think Kimi k three, like, if I were to choose which one won this round, it's probably k three, to be honest.
10:22Here for the tournament, gone with the trophy. Alright. So this is what Kimi k three gave us with a one shot prompt with access to g p d image two as well as cling k three.
10:31So this is the one from Fable five, and this is the one from Kimi k three. Very, very close. And I think for this one, I would probably vote for Kimi k three.
10:39So let me know below which one you prefer. Because, again, front end design, very subjective, but I think the level of what we have with these AI models now, it is really, really easy to create websites that are a bit more eye catchy than your usual vibe coated purple AI websites. But like what I said earlier, the important thing to consider when it comes to these AI models is to measure your ROI from them.
11:02So your return on your investment. Right? Because at least for this use case, this is the output that you're getting, but how much did it actually cost and how much time did it actually take to produce these results.
11:13And so what I just did is have each of those models self report on these use cases on what the token cost is as well as the time that it took. Now the cost here, just for clarity and avoidance of doubt, is the cost that would have been if we used API usage based pricing. So for CloudFable five, because I'm in subscription, I actually didn't pay $35.
11:33For Kimikay three, since I'm using it via a service called OpenRouter, I actually did pay $4.7 for that build. Now do note the reason why this is probably more expensive than usual is because I actually asked if you go back to that prompt, I actually asked these models to self check their work, which is why they took screenshots as well as repeated a few loops in order to just refine their work.
11:55Now when you're constrained for budget, you probably won't be doing that. But just know that if you have these models loop to check their work before giving it to you, then obviously, it would cost much more. Now this is what we're saying earlier.
12:06Right? Because as we know, Cloud Fable five, terribly expensive, but it does finish a task much faster. So end to end, it took roughly forty minutes to build that site.
12:14And remember, this is at max effort, so that's why it probably took longer. While with Kimikay three, it took around one hour and seventeen minutes. So if you're doing front end website design, what is the key takeaway?
12:25Well, I think the key takeaway is the fact that Kimi k three is really good at it. It depends with your prompt, obviously, but I think that arita.a I benchmark that we saw earlier is quite real.
12:34Like, I can imagine how k three can surpass Fable five with, uh, different prompts, with different brands, with different designs. And at this level of cost, especially if, for example, you've already maxed out your Fable five credits in your cloud subscription and you need to pay usage based pricing anyway, I would probably see if I can just offload that task with Qumik three even despite the fact that it is taking a bit longer.
12:58Alright. Let's move on to the second prompt. So I'm going to send this two prompts, and then I'll just take you through it.
13:05Because what we're now going to do is to build an Android app. And if we expand this prompt, you can see we're building this Android app called Stretch, which offers guided stretch routines with beautiful illustrations and timers. Now if you read through that prompt, which I'll also provide below, you can see that it is pretty much emulating the features of this app, this existing app called Bend, and it has already a 157,000 reviews with 5,000,000 plus downloads.
13:27And when I did a quick Google search, they were reportedly making something along the lines of 600 k in terms of MRR, which is pretty insane because if you actually download this app, all it really does is give you a set of stretching routines with this really nice set of images to go along with it. So it's sort of like a timer app, but you can definitely see, like, from just the reviews here and the amount of usership that it has that it is quite valuable.
13:51So if Fable and Kimikatree are successful in building this out, then that just alludes to the point that the coding part isn't really the roadblock anymore. It is mostly around getting an audience as well as probably marketing this app. So we'll see how both of those models do once they finish that prompt.
14:08And just to go through that prompt once again, you can see I am using the agent framework for this goal prompt, which I declared a clear ask. I am declaring a goal, so stop only with the app builds and launches on a standard Android emulator. And at least for this build so that it won't take so much time, I just ask it to include eight distinct structures, each with its own illustration, every stretch illustration shown, again, I'm using or having these models use GPT image two to render images, and I also want at least four routines built from those stretches.
14:39I gave it some guidance on some examples, but not too specific, and some guidance on what not to do as well as the tools that it can use. And then once again, I gave it a hard cap of how much to spend so that we have a fair playing field between these two models.
14:54Alright. So now with those two tasks done, you can see they used Flutter as well as Dart or at least, uh, k m k three did, and Fable five probably also did.
15:04So, yeah, you can see the stack is Flutter, and everything got built in their individual emulators, which I have on my other screen. So this is the one from CloudFable five, and there you go.
15:16It has four routines, one for waking up, one for anytime, one for a desk break, post workout, and before sleep. So let's say we wanna try out the waking up routine.
15:25You can see there's a few stretches in here, and all these images are generated via g b d image too. And I like how consistent it all is, and it's just that one design. And, again, it decided the prompts for these icons by itself.
15:37Because this is just a simple timer, you can see it has this neck release stretch. It has instructions on how to do it, a timer on doing it. And then if we go to the next, we can actually do that to do the cat curl.
15:49So you have instructions there as well. And, uh, yeah, it's pretty basic, but if you try the original inspiration for this, the Bend app, this pretty much what it does. And it has what, like, more than a million users.
16:01Right? And so something as simple as this actually does add a lot of value for its users. Right now, we just did four routines in here, but you can just imagine how you can probably have CloudFable five ideate, like, 10 more routines for the different exercises, maybe like a nighttime routine, maybe one that's more focused on yoga.
16:21And at least right now in its most basic form, it does function well. And I think the look of it, especially these images, are also as per the brief, as per the prompt that we gave it. So props to CloudTable five for that, especially given that it was only done with just one prompt.
16:36Now if we go and see what, um, Kimi k three did, it is actually quite similar. Like, this is the homepage of it. If I go back to the homepage of Club Fable five, I think the only difference is, number one, it doesn't really have the core images of these workout routines here at the very top, but it did include all of the stretch es here at the bottom if in case you want to try them out.
16:59So if we try one of these, it's pretty much the same, but the difference is the images are slightly different, but it functions similarly, and the front end design of it is also the same. So, really, if you're after an Android app, you can pretty much one shot this whole emulator experience running on your own desktop and just create an APK file from this to run it on your own phone and then just submit it to Google Player.
17:24Right? And so the point being the barrier to entry to creating applications like these have never been lower. Now, obviously, because these two are so similar, probably want to just look at the cost and the time that it took for these two models to have created these builds.
17:38And again, it's the same pattern as earlier. So CloudFable five costed around $20 if you were using API usage based pricing at around half the time that it took Kimi k three to do it. But that is four times the cost of Kimi k three at only $5 to create that whole experience.
17:56Now time wise, it did take an hour and ten minutes, so roughly twice as long as Cloud Fable five. But this is just one prompt, and you wanted it to run-in the background and realize, like, 75% savings in the process, then that's a really big deal, especially since the output, as you saw earlier, were pretty much the same in terms of the design aesthetic as well as the functionality.
18:17Now onward to our third test. And this prompt is quite interesting because what we'll be doing is we'll be replicating this video game called Portal. And in case you haven't played Portal before, it is quite an interesting test because the core mechanic of it is you actually can shoot out these blue and orange portals so that you can teleport to wherever the other portal is leading to.
18:40And I think Portal would be a really good test because it's essentially a puzzle game. Right? And it has a lot of these physics based movements that I think would be a sort of a good evaluation of how good these models really are.
18:51The other reason is with games like these, you actually don't need a lot of really good graphics in order to have a good idea for a game. As long as the core idea and mechanic is good, then you'll actually be able to have a pretty successful game. Like, for example, Minecraft or even Flappy Bird, if you remember that.
19:09Those have been pretty successful just because of the core idea and the core mechanic. That said, we are going to ask these models to do this in just one shot.
19:17And what we're asking them to do is to build a playable three d clone of Portal as a single self contained HTML file. So we'll be opening it in our browser. And so we just enumerated a few of the core mechanics that must work and they must verify.
19:30And the other thing is we wanted two levels for now. So one that is quite easy and the second one which is more difficult. And then same with the framework that we have for this goal prompt, we gave it some written instructions on a few examples, a few negations, as well as tools.
19:47Other thing is, again, because we want it to be visually appealing, we gave it the chance to use GPD image two with a budget if you need to render images, let's say textures for in game assets. Alright. So both of those sessions are now done.
20:00And this is the one from Clawtable five called Slingshot, two chamber portal physics test. And this is wild. So you can see that the goal, it seems, is we need to stand on this button, but we need to get that box over at the top.
20:16So if I click on the left mouse button that launches the blue portal, so it has that capability. So Fable was able to do that pretty well. And I think it says here I can press e to grab this, jump, and just put that there.
20:31There you go. We solved level one pretty quickly. K.
20:34So chamber two supposedly needs to be harder. So it seems that we'll probably need to have a portal here, then launch ourselves here. And then if I can fling myself there, I'll be able to go to this side.
20:46So that probably takes a bit more precision. Alright. Level two, because our brief is to make it super difficult.
20:52It is actually really difficult. But the core idea is if I place the portals here, I'll be able to fling myself to this platform in the center with this box while launching myself like that.
21:04There you go. Okay. And then I'll just place it here.
21:08K. And then we'll finish the level that way. Perfect.
21:11So that is a pretty good run, actually. That's a pretty good one shot prompt. Basically, just being able to have that pretty complex physics built in and having, like, a full portal game that obviously the whole visuals of this you can improve.
21:26It's not like the actual portal game. But the fact that you have everything in order to make, like, a proper puzzle game, I think is really powerful, especially coming from just a one shot prop. So kudos to Cloud Table five.
21:37Alright. Now we go to what Kimi k three gave us. So it's called aperture Kimi.
21:42So pretty much the same controls. Let's click to begin. K.
21:44So basically, we need to get that. Right? And if we go and oh.
21:51Alright. Kinda works. I do have to say though that there is a bit of lag.
21:55So if we are probably going for speediness, we can give it a second prompt to improve that. But that first level is pretty easy as instructed. Okay.
22:05So now his second level, it's kinda hard. I I don't think I can solve this, but let me try. I think the graphics are a bit better with QEMI.
22:12So again, alluding to that front end design benchmark that we saw earlier. But in terms of speediness, Fable five was able to give us a much better experience because there is a bit of lag here.
22:24But the fact that both of these models are giving us, especially Kimik three with this pretty complex puzzle that can probably pass as an actual stage in a game like Portal, is pretty insane, honestly. So, yeah, there is that exit.
22:38It'll probably take me a few minutes to figure this out, guys. Especially the point being that if you have, like, an idea, especially that of a game and even if it has, like, complex physics like these, then you can probably do it now with the smartness of these models. Now to build that portal clone, again, it's the same pattern as we saw earlier.
22:56Cloud Fable five, at least for this run, it was three times more expensive versus Kimi. Like, imagine that two stage thing that was built with just one shot. It did take two hours and thirty minutes though because the prompt that I gave, I always ask them to verify the stages.
23:10And at least with the Kimi k three one, the second stage was a bit more complex, so it probably took some turns in order to make sure that that one is solvable. But that took two point five hours to finish. But at a cost that is less than $10 to build out that whole thing, that is pretty incredible.
23:25Right? Now with CloudFable five, again, that takes less time but at a much higher cost. So what do we now take away from this?
23:33Because if you have a subscription to Claude, should you now cancel that and just go all in on something like Kimikatree because of a lower cost? Probably not. Right?
23:41Because at the end of the day, the subscriptions that Entropic as well as OpenAI is offering us is still much cheaper and highly subsidized to access these models. So take advantage of that, especially since there's still some question on whether that will last years into the future.
23:57But at least right now, having, let's say, a max plan of a $100 a month or $200 a month to give you a lot of tokens and access to models like Fable five is still the cheapest ways by which you can access intelligence levels of this skill. Now that said, if you do run out of table five credits or usage, then that is when you should check out Kimik A Tree, especially since right now you can only access Kimik A Tree through usage based pricing.
24:26And the reason why that is is because if you go to the Kimi K three membership plans, they actually closed this recently. So you can see it says join waitlist here right now. So you can't really get Kimi K three on a subscription plan as of the moment.
24:40And the reason why they did that, they mentioned it in this x post where basically over the past couple of hours, the demand for Kimi K three was so much that they couldn't support it at current capacity, so they had to close the subscriptions as of right now. But look, Kimikitri, it's an open weight model.
24:59I think five days from now, July 27, they'll release the weights. And one of the big implications of that, like I mentioned in the beginning, is the fact that every other model company will definitely download that file and implement it in order to improve their own models. So all of these ones will probably get to something like almost Fable level intelligence at some point because Kimikatree is going to open weight and provide basically access to their secret sauce in order to make this intelligence level possible.
25:30So overall, it's a win for the consumer. It's a win for us as users. I think definitely try Kimi K Tree out.
25:35It is almost as good as Fable as you can see in the tree tests that we did and at a fraction of the cost. So if you want to test it out over a cloud code as your harness, then the guide for that, I just ship that from my system, from my workspace so that you can try it out as well.
25:51But there, I hope that was useful. And as always, thank you for watching until the end, and I'll see you all next time. Cheers.
The Hook

The bait, then the rug-pull.

Same prompts, same rules, zero revisions — the creator pits Claude Fable 5 against the just-launched open-weight Kimi K3 on three real builds: a campaign site, an app clone, and a playable 3D game. What comes back isn't a benchmark screenshot, it's a receipt.

Frameworks

Named ideas worth stealing.

01:50model

Model ROI (Return / Investment)

  1. Return = intelligence score
  2. Investment = cost per task
  3. Investment = time/speed to complete task

A model's raw intelligence score is only half the picture — real ROI divides that return by what it costs in money and time to get it.

Steal forframing any AI-model or tool comparison beyond a single leaderboard number
06:15list

Goal prompt agent framework

  1. Ask
  2. Goal
  3. Examples
  4. Negations (what not to do)
  5. Tools

A structured one-shot prompt template used for all three tests: state the ask, define the goal/done-condition, give examples, list what to avoid, and specify available tools.

Steal forwriting repeatable one-shot prompts for coding agents
CTA Breakdown

How they asked for the click.

VERBAL ASK
05:30product
check that just in the pinned comment below

soft mid-video pitch for the free Kimi K3 + Claude Code setup guide and the paid RoboNuggets community, woven into the test-setup segment rather than a hard break

MENTIONED ON CAMERA
Storyboard

Visual structure at a glance.

open
hookopen00:00
Starbucks build
valueStarbucks build07:29
Stretch app costs
valueStretch app costs18:09
Portal clone costs
ctaPortal clone costs23:20
Frame Gallery

Visual moments.

Watch next

More from this channel + related breakdowns.

14:18
Jay E | RoboNuggets · Tutorial

STOP Prompting Claude

A 14-minute tutorial on the three tiers of self-running Claude Code workflows — and why the creator of Claude Code stopped prompting it manually.

June 12th
Chat about this