Modern Creator
Matthew Berman · YouTube

Cancel Your Subscriptions, Ox-Alpha Is Here (GLM-5.3-Flash)

The mystery model that took over OpenRouter turns out to be a Chinese open-weights release that nearly matches frontier intelligence for a few cents a task.

Posted
1 weeks ago
Duration
Format
Review
educational
Views
145.6K
2K likes
Big Idea

The argument in one line.

GLM-5.3-Flash, the model that anonymously topped OpenRouter as 'Ox-Alpha,' delivers benchmark scores close to Claude Opus and GPT-5.6 while costing about 9 cents per task instead of several dollars, and it does it running entirely on Chinese-made AI chips instead of Nvidia hardware.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You build with AI coding agents and want to know if a cheaper open-weights model can replace part of your Claude or GPT spend.
  • You're tracking whether Chinese AI labs can match frontier model quality without Nvidia chips.
  • You want a real cost-per-task comparison across Claude Opus, GPT-5.6, and open-weights alternatives before picking a default model.
  • You're curious what open weights actually buys you beyond price: self-hosting, fine-tuning, no vendor lock-in.
SKIP IF…
  • You need the single highest intelligence score available regardless of price and already default to Claude Opus or GPT-5.6 max effort.
  • You have no interest in model benchmarks, pricing tables, or the US-China AI chip race.
TL;DR

The full version, fast.

A mystery model called Ox-Alpha went viral on OpenRouter, and it turned out to be GLM-5.3-Flash, a new open-weights release from the Chinese lab Z.ai. It's a 320-billion-parameter mixture-of-experts model with only 18 billion active parameters, priced at roughly 9 cents per completed task versus $3.14 for Claude Opus 4.8 and 95 cents for GPT-5.6 Soul, while scoring within about 7% of frontier intelligence on the Artificial Analysis index. It uses more output tokens per task than cheaper rivals like GPT-5.6 Luna, so it isn't the single best value model, but it's close to the top of the intelligence-per-dollar frontier. The bigger story: Z.ai says it served the model at 100 trillion tokens a day entirely on domestic Chinese AI chips, no Nvidia hardware involved, which is evidence China's chip-plus-model co-design strategy is starting to close the gap. Live demos (a physics-accurate Rubik's cube, five website builds against GPT-5.6 Sol, and an on-brand PowerPoint deck) show real but uneven quality: it beat Sol on some builds, lost on others, and nailed pulling live brand colors off a real website for the deck.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0001:17

01 · The Ox-Alpha mystery

A hyped, unlabeled model on OpenRouter turns out to be Z.ai's new GLM-5.3-Flash.

01:1702:35

02 · Z.ai unveils GLM-5.3-Flash

320B total / 18B active mixture-of-experts model, MIT licensed, one-tenth the price of GLM-5.2, approaching Claude Opus 4.8 on coding and agentic benchmarks.

02:3505:02

03 · Benchmarks: closing in on the frontier

Terminal Bench, DeepSWE, Agents' Last Exam, AutomationBench, HLE, and GDPval scores compared against GLM-5.2, DeepSeek, Opus, GPT-5.6 Terra, and Gemini.

05:0206:20

04 · Sponsor break: Higgsfield

Ad read for Higgsfield's unified image/video generation API (Veo 3, Kling, Nano Banana, Higgsfield Soul).

06:2009:51

05 · The real story: cost per task

GLM-5.3-Flash costs about 9 cents per Intelligence Index task versus $3.14 for Claude Opus and 95 cents for GPT-5.6 Soul, though it uses more output tokens per task than cheaper rivals like GPT-5.6 Luna Max.

09:5111:43

06 · 100 trillion tokens a day, zero Nvidia chips

1M-token context, launch pricing details, and a SemiAnalysis report that Z.ai served the model at massive scale entirely on Chinese-made AI chips.

11:4312:38

07 · Getting access: API keys and OpenCode

How to spin up a z.ai API key and plug the model into any OpenAI-compatible tool, demonstrated inside OpenCode.

12:3813:53

08 · Demo: a physics-accurate Rubik's cube

GLM-5.3-Flash one-shots an interactive 3D Rubik's cube simulation with working scramble/solve and a customizable 5x5 variant.

13:5317:59

09 · Demo: GLM vs GPT-5.6 Sol, five websites

Head-to-head website builds (apples, DGX Spark, rubber ducks, Galaxy Z Fold, Tesla Model Y) judged on design taste, with mixed results both ways.

17:5918:47

10 · Demo: an on-brand PowerPoint deck

GLM pulls a real company's live brand colors and logo to build an accurate data-center presentation.

18:4718:57

11 · Sign-off

Wrap-up and pointer to a deeper video on Chinese open-source AI.

Atomic Insights

Lines worth screenshotting.

  • GLM-5.3-Flash costs about 9 cents per completed task on the Artificial Analysis Intelligence Index, versus $3.14 for Claude Opus 4.8, roughly 3% of the price.
  • It's a 320-billion-parameter model with only 18 billion active parameters at any time, the mixture-of-experts design that lets it run cheap while staying capable.
  • GLM-5.3-Flash is one-tenth the price of the previous GLM-5.2 release, which was already considered cheap.
  • On the Artificial Analysis Intelligence Index, GLM-5.3-Flash scores 57 against Claude Fable 5's 62, a 7-8% intelligence gap for roughly 97% less cost.
  • GPT-5.6 Luna Max is still the single best cost-to-intelligence tradeoff at about half GLM-5.3-Flash's price, but GLM scores about 5 points higher on the intelligence index.
  • GLM-5.3-Flash uses roughly 47,000 output tokens to complete a task versus about 20,000 for GPT-5.6 Luna Max, meaning it's less token-efficient even though its per-task dollar cost is close.
  • Z.ai reports serving GLM-5.3-Flash at 100 trillion tokens per day entirely on Chinese-made AI chips, with zero Nvidia hardware in the serving stack.
  • Z.ai claims a 3x improvement in end-to-end serving performance on the same Chinese chip hardware compared to their own prior baseline, reaching efficiency comparable to mainstream Nvidia GPUs.
  • The model ships with a 1-million-token context window and 131,000 max output tokens, with max reasoning effort enabled by default.
  • Launch pricing via OpenRouter is $0.075/M input, $0.25/M output, and $0.015/M cached input tokens, at 50% off through September 9, after which it moves to $0.15/$0.50/$0.03 per million.
  • The model is released under the MIT license with open weights, meaning anyone can self-host, fine-tune, or customize it rather than depending on a single API provider.
  • In a live head-to-head against GPT-5.6 Sol, GLM won the apple-orchard website and the DGX Spark website comparisons on design taste, lost clearly on the seven-biome-diorama 3D scene test, and split evenly on the rubber duck, Galaxy Z Fold, and Tesla Model Y builds.
  • GLM-5.3-Flash successfully pulled a real company's actual logo and brand colors from its live website to build an on-brand PowerPoint deck, rated 'phenomenal' in the demo.
Takeaway

A cheap open-weights model just closed most of the gap to frontier AI.

WHAT TO LEARN

GLM-5.3-Flash proves that mixture-of-experts open-weights models can land within single-digit percentage points of frontier intelligence at roughly 3% of the cost, which should change how you pick a default model for agentic coding work.

02Z.ai unveils GLM-5.3-Flash
  • Mixture-of-experts models activate only a fraction of their total parameters per task, which is how a 320B-parameter model can run at flash-tier cost and speed.
  • A model can outperform its own prior generation by 10x on price while still improving benchmark scores, which is the pace open-weights competition is moving at.
03Benchmarks: closing in on the frontier
  • Benchmark comparisons only make sense within a size class; comparing a 320B MoE model directly to a multi-trillion-parameter frontier model on raw score ignores the compute gap.
  • GDPval (real-world knowledge work) is a more predictive benchmark for day-to-day usefulness than academic reasoning tests.
05The real story: cost per task
  • Cost per completed task (which factors in both token price and token volume) is a more honest comparison metric than per-million-token pricing alone.
  • A model can be cheaper per token but more expensive per task if it needs significantly more tokens to reach the same answer.
  • The 'most attractive quadrant' plot (high intelligence, low cost) is a reusable framework for evaluating any competing tools, not just language models.
06100 trillion tokens a day, zero Nvidia chips
  • Serving scale claims (tokens per day, hardware used) are as important a signal as benchmark scores when judging whether a lab can actually deliver a model in production.
  • Chip-and-model co-design, rather than adapting a model to existing hardware, is the strategy closing the gap between Chinese and Western AI infrastructure.
07Getting access: API keys and OpenCode
  • Open-weights models being available through many inference providers (not just the originating lab) creates price competition that benefits the end user directly.
08Demo: a physics-accurate Rubik's cube
  • A working interactive 3D physics simulation from a single prompt is now table stakes for a capable model, not a differentiator.
09Demo: GLM vs GPT-5.6 Sol, five websites
  • No model wins every task category; test candidate models on your specific use case rather than trusting one benchmark or one demo.
  • Internet access (or lack of it) inside a coding tool materially changes output quality when a task depends on real images or live data.
  • A model inventing a plausible but fake detail (like a real street address it didn't verify) is a concrete reminder to fact-check anything AI-generated before publishing it.
10Demo: an on-brand PowerPoint deck
  • A model that can pull live brand assets (logo, colors) from a real website and apply them accurately is a genuinely useful, testable capability for brand-sensitive deliverables.
Glossary

Terms worth knowing.

Mixture of experts (MoE)
A model architecture where only a subset of the model's total parameters activate for any given task, letting a very large model run at a fraction of the compute cost of using all parameters at once.
Open weights
A model release where the trained parameters are published publicly, letting anyone download, run, customize, or fine-tune the model instead of only accessing it through a paid API.
Intelligence Index
Artificial Analysis's composite benchmark score combining multiple evaluations (coding, reasoning, real-world knowledge work) into one number used to rank and compare models.
GDPval
An OpenAI benchmark that tests a model's performance on real-world, economically valuable knowledge work tasks rather than academic puzzles.
Cost per Intelligence Index task
A metric combining how many tokens a model needs to complete a task and the price per token, producing a single dollar figure for 'cost to get one unit of benchmark-measured work done.'
Diarization / speaker ID
Not used here, but referenced implicitly in model context handling: the process of identifying which parts of input correspond to different sources.
Resources

Things they pointed at.

01:29toolZ.ai
12:22toolOpenCode
10:18channelSemiAnalysis
05:22productHiggsfield API
Quotables

Lines you could clip.

1:56:36
GLM 5.3 Flash coming in at 9 cents per task completed, and that is just a smidge off of the absolute frontier of Fable five intelligence.
The single number that carries the whole video's thesis.TikTok hook↗ Tweet quote
42:03
You can't really compare a seven or eight-trillion-parameter model, which is what Fable is, to a 320-billion-parameter model.
Frames the size-vs-performance gap in one line.IG reel cold open↗ Tweet quote
3:09:11
This demonstrates that Chinese chips can support frontier model inference efficiently and economically at scale.
The geopolitical punchline of the video, straight from Z.ai's own blog post.newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

A few days ago, rumors started swirling around this mystery model that showed up on OpenRouter. It was called AUX Alpha, and people were speculating as to what it could be.
Maybe it's the new Gemini model. Maybe it's some new continual learning model. Here's the thing.
It was fantastic. People were loving it, and it was seemingly comparable to the top frontier models out there. And then we found out what it was.
This is the newest open source open weights model from z AI. This is GLM 5.3 flash, and you know if it says flash, that means fast, cheap, and smaller than a usual model, and that's what makes it super special.
This model performed incredibly well despite its small size. And this might be one of the most exciting model releases to date because you actually might be able to download this model and run it locally, but certainly, you're gonna be paying a fraction of the price that you're paying for Fable or GPT 5.6 Soul for a model that is insanely capable.
So I'm gonna tell you about the model. I'm gonna show you the benchmarks, the pricing, which is crazy. I'll show you how to use it and some of the demos that I created.
But there's something more to this story than just how the model performs, and it's going to have a massive impact on the world. And this video is brought to you by Higgs Field.
More on them later. So the first thing you need to know is ZAI is a Chinese AI lab. And like usual, they're putting out incredible open source and open weights models.
They're telling the world how they're baking the model, and they're giving the weights. So you can actually use it. You can customize it.
You can fine tune it. You can host it yourself. It is just such an incredible value to the world, these open source models.
So this is a relatively small model. It is 320,000,000,000 parameters.
18,000,000,000 are active. That's called mixture of experts.
It just means it can run really efficiently while also being extremely capable. It outperforms GLM 5.2, which is the last version of the GLM series of models, but also the full version of it.
And this is a flash version, and it is one tenth of the price of the previous GLM 5.2, which was already cheap. And it is approaching Claude Opus 4.8 on coding and agentic benchmarks.
Now if you were expecting it to be on the same level as a Fable, it's not.
It's not that far from it, but you can't really compare a seven or 8,000,000,000,000 parameter model, which is what Fable is, to a 320,000,000,000 parameter model. I mean, they're just completely different class sizes, and the fact that it's even close at all is what is so mind blowing here.
So here are some benchmarks, and keep in mind, they're comparing this new GLM 5.3 flash to other models that are comparable in size. So here we go. Here's terminal bench coming in at 84.3.
A nice bump over the last full size 5.2 model. Here's DeepSeq. Here's Opus 4.8.
Here's GPT 5.6 Terra, which remember Terra is kind of that middle child between Sol and Luna on the GPT 5.6 family. And then here's Gemini 3.7 Flash.
Here's DeepSweet, which is probably the most accurate benchmark for how people are actually feeling about a model, how good it is when it's really being used in the real world. So here it is at 63.4, a massive jump from 5.2, and we can see quite competitive with 5.6 tera.
Here is Opus. Here's DeepSeek. Here's Gemini.
Here's GDPVal, the OpenAI benchmark testing real world knowledge work, and it is number one by a large margin.
So now let's talk about the benchmarks that I think matter more, which is the total cost per task completed. And we're gonna start adjacent to that with this right here, which is a JNC coding performance by effort level. Now what you're seeing right here is the output tokens per task, and then on the y axis over here, we have accuracy.
So higher up on the y axis is better, and to the left of the x axis means less tokens used, more density, more intelligence per token, which is also better. And what we're seeing right here is GLM 5.3.
So here's Fable, which is crazy. GLM 5.3 max sits nearly as high as Fable high, and it does use a little bit more tokens.
But the fact that it's even getting close and that it is a fraction of the price, and I'm gonna get to price in a few minutes, is kinda just insane to look at. Now when you go to max effort on Fable, you do get a massive jump, but you're using a lot more tokens, and each one of those tokens is much more expensive.
So here's GLM 5.3 Flash on the artificial analysis intelligence index coming in at 57. Claude Fable five is 62. 62 compared to 57, and we are talking about a model that is literally a fraction of the size and price.
And I'm able to make videos just like this because of the sponsor of today's video, Higgs Field. Every week, there's a new AI image or video model that you might wanna use inside of your product. Adding each one usually means a separate API, separate billing, and more code to maintain.
Higgs Field API solves it. It gives you models like cling, nano banana, v o three, and Higgs Field Soul all with a single API key.
You pay per generation rather than paying a monthly fee, so you're literally only paying for what you use. And they tell you the cost before the generation. Failed generations are not charged.
And as they add new video models, new image models, you don't need to change anything in your code. Hicksfield also supports webhooks and concurrent jobs, which makes it easier to run real applications for a team or your customers.
Higgs Field is one of the fastest growing companies of all time. There are so many people out there who are using it and loving it right now. And during your first week, you can choose two image models and two video models to get at a discounted price.
So go click it down below. Thanks to Higgs Field. Now back to the video.
Now here's where it gets wild. Look at this. So this is the cost per intelligence index task.
This is how much does it cost to actually get things done, and that factors in multiple things. That factors in how many tokens does it need to complete a task? How expensive are those tokens?
So we see Claude Fable five at $3.14, incredibly expensive. We have GPT 5.6 Soul coming in at 95¢, a third of the price of Claude Fable five.
Here's Kimmy k three at 84¢, which when we first saw it, that's amazing. Now look at this.
GLM 5.3 Flash coming in at 9¢. 9¢ per task completed, and that is just a smidge off of the absolute frontier of Fable five intelligence.
Not that much off. It's like maybe, you know, 7% off of the absolute frontier of intelligence, and it's about two or 3% of the price.
That is so crazy to see. Now this is really the important chart, intelligence versus cost.
And where you wanna be is in this quadrant right here, that's why it's green. You wanna be as high up and to the left as possible. So what's really cool is GPT 5.6 Luna max is incredibly cheap.
Actually about half the price of GLM 5.3 flash. But for double the price, you get about five points higher on the intelligence index. So that's a nice trade off, but you also get open weights.
You also get full control over the model. You can customize the model. So, look, GPT 5.6 Luna is fantastic because it is so cheap, especially after that 80% discount a couple weeks ago.
But GLM flash is a better model. It has a higher intelligence score, and it's about 40% more cost. But we're talking about 5¢ versus, like, 9¢.
So it's really, really inexpensive. So when I say 40% more, both of them are still incredibly cheap.
So here's DeepSeek v four Pro coming in at 27¢, much more expensive, but also not as good.
So GLM flash is sitting in a great place. If you really care about cost, GPT 5.6 Luna still is the number one model in terms of the best trade off of cost and quality. But GLM 5.3 flash, man, that offers a lot of benefits.
Now here's something really interesting. If we look at GPT 5.6 Luna max, this is the output tokens per intelligence task. Basically, how many tokens does it take to arrive at the same answer as another model?
And, again, here's Luna at 20,000 tokens on average. Now if we go all the way to the top, GLM 5.3 Flash is actually one of the most token intensive models out there, which is kinda crazy at 47,000 tokens.
So when you go back to this chart and you see LunaMax here, you see GLM 5.3 Flash here, LunaMax is less expensive because it uses far fewer tokens, less than half the amount of tokens to arrive at the same solution as GLM 5.3 flash.
So, again, these are all the different factors you need to keep in mind when evaluating a given model. But the nice thing is with open weights, open source, they will continue to iterate and everybody will have their eyes on it, and they will improve it and potentially improve the number of tokens used per task, theoretically more quickly than what an OpenAI can do.
Now I'm saying this all in a vacuum. This is all speculation, but I really am a big proponent of open source open weights for that reason.
Now a couple last things before I reveal the most shocking part of this entire story. So number one, it's a million token context. Wonderful.
A 131,000 max output tokens and max reasoning by default.
So this is all really good stuff. You know, 1¢ per million cashed input. It's it's incredibly cheap, incredibly great.
But here's the crazy part. Reported by Semi Analysis, it was serving a 100,000,000,000,000 per day on purely Chinese chips.
That is crazy. That type of capacity without using a single NVIDIA chip is wild, and I've been talking about this on the channel for a little while now.
China is developing their own chips. They are codesigning it with the models that are incredible.
Plus, of course, the model's small. It's efficient. It's a flash version.
So when you pair these things together, the fact that they can have near frontier intelligence at a fraction of the price served entirely on their own infrastructure is quite surprising.
I was not expecting this. So over the past week, we have served g l m five three flash on a large scale cluster of Chinese AI chips supported by high bandwidth interconnect and a serving stack optimized for the underlying hardware.
All of the parts of the AI stack are being codesigned for Chinese models, Chinese hardware, Chinese interconnects.
All of these things are being codesigned together, and they have capacity now. Hardware efficiency and per token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier model inference efficiently and economically at scale.
This is all great. This is all really cool news. So the way that I tested it out, the way that you can test it out is by going to z.ai, which is being served from China, so keep that in mind.
And I just spun up an API key. I'm not doing anything sensitive. I'm not sending sensitive information.
But if you just wanted to try it out, this is probably the most straightforward way. You can go to OpenRouter as well. You can go to a bunch of different inference providers because it is an open weights model, and so anybody can serve it.
And they're all gonna compete on price. They're all going to look for their own optimizations to eke out every penny they possibly can out of that price, which, of course, benefits us. This is the promise of open source.
And today, I'm plugging it into OpenCode. You can plug your API key into anything, your project. You can plug it into t three.
You can plug it into anything you want as long as it supports an OpenAI compatible API endpoint and the API keys. And, of course, the first thing I wanted to test is a Rubik's cube simulation. And here it is.
It looks hyper realistic. It's smooth. I can grab one of the sides and turn it easily.
I asked it to create a bunch of different sliders. We can see right here you can change all of these.
Now click scramble right there and it scrambles. Everything looks accurate. It scrambled correctly.
The physics are all correct. And then, of course, we can solve it just like this. Now all recent models can pretty much create the Rubik's cube simulation, but it's a good baseline to just see how it does.
And by the way, we can also change the size of it. We can change the colors. So these are all settings that it decided to create.
So here's this. I'll do scramble. Now you can see this is a five by five cube, and let's solve it.
Yeah. Look at that. We can change the turn speed, the scramble length.
Let's turn that up. We could do auto spin, so you can see right there. Field of view, zoom in and out, the glow, the exposure, the key light.
Here's some shadow. Material, we can change the reflection amount, the metalness, which I don't even know what that is, and then clear coat so you can make it shinier or less shiny.
So very cool. Worked very well. Now I also wanted to test it against GPT 5.6 sol.
So I gave Soul and GLM 5.3 Flash the same prompts to create a few different demos. Let me show you those. Alright.
So here's the first experiment. Build a beautiful interactive three d scene featuring seven miniature biome dioramas floating against the deep navy background. So I'm basically trying to create these little low poly diorama type things, and let me show you the results.
Alright. So the final comparison on the left is Sol, on the right is GLM. Now GLM looks good, but you can just clearly see the winner by far is Seoul.
Just the amount of detail, the coherency of it, it just looks a lot better. Okay.
Next, I had Seoul and GLM five three create five different websites just to see how it does on website design. So I had it create one about apples, a DGX Spark, Rubber Ducks, Galaxy Z Fold, and the Tesla Model Y.
Now here's the website about apples. It just, like, kinda plopped a photo of apples right in the middle.
It really does not look good at all. And, actually, one other thing I want you to keep in mind is that GPT 5.6 Sol and within Codex has access to the Internet, and open code right now kinda doesn't.
So just keep that in mind. But, you know, the website overall is okay. It's pretty sparse.
The colors are okay, but just plopping this image of three apples right in the middle does not look good at all because it overlaps with the text in the background. And here's GLM five point three. I think this is overall a much better website.
Obviously, it does not have access to the Internet because it did not grab an image of a real Apple, but overall, I think it looks better.
Here's something kinda interesting. It actually put an address here. 42 Frosthill Road in Vermont.
Let's see if that's even a real place. Yeah. Okay.
So here it is. It is a real place. I can't tell if it's actually an apple orchard, but yeah.
There it is. I wonder why it chose this address. That's kinda wild.
Alright. So here it is side by side. I think G l m five three won by a pretty big margin on this one.
Now here's the comparison of a website about the DGX Spark. On the right, we have GLM53. On the left, we have Soul.
And once again, I actually think GLM one. It has better design taste than Soul.
Soul made up this image of what the DGX looks like. GLM five three didn't have an image, so it didn't try to make it up. I really like that it has this kinda terminal looking UI here.
But, yeah, overall, soul is not as good. This is much better. Next, rubber ducks.
Here's a website about a rubber duck. So this is on the left soul, on the right, GLM, and I actually think they're both really good. This is okay.
I don't know what these are. They kinda look like parts of a duck. Uh, all of them tend to be very, very simplistic.
I would still give the overall win to GLM. I mean, the duck actually looks like a duck here. Obviously, Sol was able to pull a real image.
But, like, look at these recreations of ducks on the left, which is Sol on the right, which is GLM. Here's a website about a Galaxy Z Fold. And, again, it was actually able to pull up an image.
I actually think they both are pretty bad. This is, a very simplistic website. It did pull some information about it, although I don't think these are accurate.
This just has completely blank space that doesn't look good at all. Yeah. These are both really bad.
And then a website about the Tesla Model y. On the right, GLM. This is embarrassing.
They basically just copied the Tesla website. It looks identical to this and then somehow created this SVG, uh, of I don't even know what it is.
A limousine. If you scroll down I mean, the website itself looks good. It's just, you know, more or less a copy of the Tesla website, whereas Seoul actually built a website.
Now this one has an image. These two do not. Yeah.
They're both quite bad. I don't know if either of them wins. Okay.
Next, I had to put together a PowerPoint presentation about data centers using Forward Futures brand guidelines, and it actually did go to the website forwardfuture.com, downloaded our branding.
This is the Accurate logo. These are the actual colors, and it's quite good, surprisingly good.
Look at this. And this, by the way, is what would show up in the GDP val benchmark is, like, creating PowerPoints and real knowledge work. So here we go.
All the colors are accurate. The text looks good, and we see a little m dash right here. Yeah.
I mean, this is a great deck. Yeah. This is phenomenal.
I'm very impressed with this. And once again, thank you to Higgs Field for sponsoring this video. I'm gonna drop a link down below so you can go check them out.
Click that link. Let them know I sent you. Open source is just so important.
The fact that we're getting these open source, open weights models out of China is incredible, and I break it down in full. Go check out that video right here.
The Hook

The bait, then the rug-pull.

A mystery model called Ox-Alpha quietly topped OpenRouter's leaderboard with no one sure what it was. It turned out to be GLM-5.3-Flash, a new open-weights release out of China, and the price tag is what makes it a story.

Frameworks

Named ideas worth stealing.

07:13model

Intelligence vs. Cost quadrant

  1. Artificial Analysis Intelligence Index (y-axis)
  2. Cost per Intelligence Index task, log scale (x-axis)
  3. Pareto frontier line

A scatter chart plotting model intelligence against cost per task, with a green 'most attractive' quadrant (high intelligence, low cost) and a dotted Pareto line showing which models are actually worth their price.

Steal forAny internal model-selection decision: plot candidate models on intelligence vs. cost before defaulting to the most familiar API.
CTA Breakdown

How they asked for the click.

VERBAL ASK
05:22product
go click it down below... thanks to Higgs Field

Mid-roll sponsor read for Higgsfield, delivered as a natural break between the benchmark section and the cost-analysis section rather than interrupting a demo.

MENTIONED ON CAMERA
Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
the reveal
promisethe reveal01:29
cost vs intelligence chart
valuecost vs intelligence chart06:36
Chinese chip reveal
valueChinese chip reveal10:18
SOL vs GLM demo battle
valueSOL vs GLM demo battle18:14
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

15:48
Matthew Berman · Tutorial

You Aren't Using Codex Like Me

Eleven power-user habits from someone who has logged over a thousand hours in OpenAI's Codex CLI — model tiers, thread delegation, safety hooks, and remote control from a phone.

July 14th
08:54
Matthew Berman · Review

GPT-5.6 is FINALLY HERE (WOAH)

A 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.

July 9th