Modern Creator
Pat Simmons · YouTube

OpenAI's Decisions API, Tested in 7 Real Builds

A new OpenAI endpoint skips writing an answer and just hands back odds. Pat Simmons races it against Jev and a regular chat model, then builds seven real tools with it live.

Posted
2 days ago
Duration
Format
Tutorial
hype
Views
21.9K
364 likes
Big Idea

The argument in one line.

OpenAI's Decisions API swaps token-by-token generation for a fixed multiple-choice answer, landing in about 150 milliseconds for a fraction of a cent, which makes real-time AI decisions viable inside games, screen-watchers, and editing tools.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You build with a coding agent like Claude Code and want a fast, cheap way to add real-time classification instead of a full LLM call.
  • You're deciding between OpenAI's Decisions API and Jev (Typesafe AI) and want real cost and speed numbers from working builds, not marketing claims.
  • You want concrete project ideas for screen-watching agents, feed filters, or editing tools that need a yes/no or pick-one decision many times per second.
  • You're curious what becomes possible once vision-capable AI decisions cost fractions of a cent and land in milliseconds.
SKIP IF…
  • You need the model to reason, write, or explain its answer. The Decisions API only returns a probability-ranked pick from a list you already wrote.
  • You want API documentation or a request/response walkthrough. This video is build demos, not reference docs.
  • You don't use a coding agent in your workflow, since every build here leans on one to write the integration from a one-paragraph prompt.
TL;DR

The full version, fast.

OpenAI's Decisions API is a classification endpoint: hand it text or an image plus a short list of answers, and in about 150 milliseconds it returns the best pick with a confidence score, for a fraction of a cent. It works like the startup model Jev, except it also reads images and costs $0.10 per million tokens versus Jev's $0.042. Pat Simmons builds seven tools around it with a coding agent: a Pomodoro timer that scolds him for scrolling X, hands-free Mac voice control, an X feed ad-hider, a privacy auto-blur tool, a Wikipedia race against Jev, a calorie tracker, and a bad-take finder for raw footage.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 00:44

01 · Intro

A salt cube FPS bot plays itself in real time, making movement and combat decisions several times a second via the Decisions API. States the premise: seven real use cases, by the end you'll know what the API is and how to start using it.

00:44 – 01:47

02 · Kicking off the Pomodoro timer build

Pastes a build prompt into Claude Code (Opus 5.5) for a next-gen Pomodoro timer: an animated orb that screenshots three displays every few seconds and asks the Decisions API whether the user is on-task.

01:47 – 03:43

03 · How the Decisions API works

Explains the pattern: hand it text or an image plus a short list of answers, it picks one with a confidence score. Demonstrates with a customer-service routing example and the Assault Cube bot's per-frame decisions, then contrasts the 150ms decision speed against a 1.6-second regular chat call.

03:43 – 04:59

04 · Decisions API vs Jev

Covers Jev's backstory (Typesafe AI, founded by an ex-OpenAI researcher, sparked a wave of fast-classification builds) and compares it to the Decisions API: Jev is cheaper per token but text-only, while the Decisions API costs more but reads images.

04:59 – 05:30

05 · How to start using it

Three on-ramps in order of effort: get an API key, try it with no code in the Decisions Playground, or paste a ready-made build prompt into a coding agent.

05:30 – 06:54

06 · Use case 1: Pomodoro focus orb

Tests the finished orb live. It catches the creator scrolling X twice during a 25-minute session and calls him back to work each time.

06:54 – 10:47

07 · Use case 2: Voice control for my Mac

Builds a hold-to-talk voice assistant that transcribes speech, then uses the Decisions API to pick an action from a fixed list (open app, switch app, browser navigation). Works almost instantly but only executes actions explicitly written into the spec.

10:47 – 12:37

08 · Use case 3: X feed cleaner

A Chrome extension sends each X post's text and image to the Decisions API and hides anything classified as an ad, sponsored post, or ragebait before it fully renders.

12:37 – 14:00

09 · Use case 4: Auto-blur private info

A tool scans a screen recording for sensitive on-screen information and blurs it automatically, finishing a one-minute test video in under 10 seconds across extraction, scanning, locating, and rendering.

14:00 – 17:13

10 · Use case 5: Wikipedia race vs Jev

A live three-way race between the Decisions API, Jev, and a regular chat model, each picking links to navigate from a start article to a target article. The Decisions API and Jev consistently beat the chat model on speed and cost across several rounds.

17:13 – 19:35

11 · Use case 6: iPhone calorie tracker

An instant calorie tracker: photograph food, the Decisions API classifies the food category and portion size, and plain code multiplies that against a USDA FoodData Central database for the calorie count.

19:35 – 24:43

12 · Use case 7: Editing out my bad takes

Builds a tool that transcribes a raw recording, classifies each line (clean take, false start, restarted line, dead air) with the Decisions API, and outputs a cut list. Catches all seven planted bad takes in 1.3 seconds for half a cent.

24:43 – 25:31

13 · Why it's a game changer

Closing thoughts: the Decisions API's vision capability points toward robotics and real-time object identification as the next frontier, and every build in the video came from a one-paragraph prompt handed to a coding agent.

Atomic Insights

Lines worth screenshotting.

  • The Decisions API skips writing out an answer entirely. It returns probability scores across a list of options you already gave it, which is why it runs in about 150 milliseconds instead of the 1.6 seconds a normal AI call takes.
  • A regular chat model call costs money for every word it generates. The Decisions API never generates words, so a single decision costs roughly a tenth of a cent.
  • OpenAI's Decisions API costs $0.10 per million input tokens with no output cost. The rival model Jev costs about $0.042 per million but can't read images.
  • Vision is the real gap between the two models: because the Decisions API can look at a screenshot, it can run things Jev can't, like a bot that plays Assault Cube by reading the screen several times a second.
  • A screen-watching Pomodoro timer that takes a screenshot every few seconds and asks the Decisions API whether the user is on-task caught the creator scrolling X twice in one short test run.
  • A voice-controlled Mac assistant built on the Decisions API switched apps before the creator finished speaking the command, because the decision itself takes a fraction of a second.
  • The voice control build only worked for actions explicitly listed in its spec. It couldn't free-type text into a field because that wasn't one of the fixed answers it had been given.
  • A Chrome extension fed each X post's text and image to the Decisions API and hid anything scored as an ad or ragebait before the post finished rendering.
  • An auto-blur tool found and masked sensitive on-screen information in a one-minute screen recording in under ten seconds total: extract frames, scan, locate, and render the blur.
  • In a head-to-head Wikipedia race, the Decisions API linked Banana to Moon Landing in 6 seconds for $0.0003. Jev took 7 seconds, and a regular chat model took 9 seconds and cost about 9 cents.
  • An iPhone calorie tracker built on the Decisions API never calculates calories itself. It classifies the food and portion size from a photo, then plain code multiplies that against a USDA food database for the actual number.
  • A bad-take finder built a classification system with categories like 'complete line fail,' 'false start,' and 'small slip worth keeping,' then caught all seven planted bad takes in 1.3 seconds for half a cent.
  • Treating an editing decision as a fixed multiple-choice question instead of an open-ended judgment call is what makes the Decisions API usable for tasks a reasoning model would otherwise grind through slowly.
  • The creator estimates his existing reasoning-model video-editing agent costs hundreds of dollars per video in API calls. Handing first-pass classification to the Decisions API is meant to cut that bill, not replace the reasoning pass.
Takeaway

Why Fixed-Answer AI Calls Are So Fast

WHAT TO LEARN

OpenAI's Decisions API turns a slow, token-by-token AI call into a 150-millisecond multiple-choice pick, and handing a coding agent a one-paragraph spec is now enough to wire it into seven working tools.

02Kicking off the Pomodoro timer build
  • Starting the build before explaining the tool let the viewer see a working screen-watcher before getting the theory, which is a good way to sell a new API: show the output first.
  • The Pomodoro build spec was literal: screenshot three displays every few seconds and ask one yes/no question, 'is this what I said I'd work on,' rather than asking the model to reason generally about focus.
03How the Decisions API works
  • The entire interface is multiple-choice: hand the model text or an image plus a short list of answers you already wrote, and it returns the most likely one with a confidence score.
  • A routing example, an angry customer message sorted into billing, technical, shipping, or other, shows the pattern: obvious cases get near-100% confidence, which is exactly the signal a pipeline needs to auto-act without a human check.
  • A regular AI call writes its answer one token at a time, and every token costs time and money. The Decisions API skips writing entirely and just returns odds, so it's faster and cheaper by design, not by optimization.
04Decisions API vs Jev
  • Jev, from the startup Typesafe AI, pioneered this same fast-classification pattern and got people building things like automatic Zillow listing tags and mass email sorting.
  • OpenAI's version costs more per token, $0.10 vs Jev's $0.042 per million, but is the only one of the two that reads images, which is the whole reason the Assault Cube bot demo was possible.
05How to start using it
  • Getting started is three options in order of effort: grab an API key, test it with no code in OpenAI's Decisions Playground, or paste a ready-made build prompt into a coding agent.
  • Handing the build prompt straight to a coding agent, rather than writing the integration by hand, is the real on-ramp this video sells: you don't need to know the API to ship something with it.
06Use case 1: Pomodoro focus orb
  • The screen-watcher caught distraction in both directions: it noticed him switch to X and called him out by name, then noticed him switch back and said so.
  • This is the simplest viable shape for this API: one repeating yes/no question running on a timer, cheap enough to ask every few seconds without worrying about the bill.
07Use case 2: Voice control for my Mac
  • The 150ms decision speed is what makes hands-free voice control feel instant instead of laggy, something past attempts with regular LLMs couldn't pull off cost-effectively.
  • The build only executes actions explicitly listed in the spec, so 'type lovable in the browser' failed outright; a fixed-answer API can't improvise an action nobody pre-wrote.
  • A low-confidence result, 67% for 'quit Spotify,' is itself useful information: it's where the spec was ambiguous and needs another explicit answer added, not a smarter model.
08Use case 3: X feed cleaner
  • Feeding each post's text and image to the Decisions API and hiding it before it fully renders only works because the decision lands in well under the time it takes to scroll past.
  • The extension's false positives, flagging a brand's own post as an ad, show the limit of a fixed-answer classifier: it can only be as precise as the categories and examples it was given.
09Use case 4: Auto-blur private info
  • Blurring sensitive on-screen info in a one-minute recording broke into four fast sub-steps, extract frames, scan, locate, render, each finishing in single-digit seconds.
  • This replaces a workflow that used to require handing whole videos to a general vision model just to find what needs blurring, at a fraction of the cost and wait.
10Use case 5: Wikipedia race vs Jev
  • In a live three-way race, Decisions API vs Jev vs a regular chat model, the Decisions API won most rounds on speed, often finishing a multi-hop path in 1-6 seconds for under a cent.
  • The regular chat model wasn't just slower, it was an order of magnitude more expensive per decision, about 9 cents vs fractions of a cent, because it was still writing out reasoning instead of just picking.
  • Jev beat the Decisions API outright in one round, a reminder that 'faster and cheaper' between the two varies by task rather than having one fixed winner.
11Use case 6: iPhone calorie tracker
  • The tracker never asks the API to compute a calorie number. It asks two classification questions, food category and portion size, then plain code does the math against a USDA database, narrowing the API's job to what it's actually good at.
  • Treating a reasoning-adjacent task like estimating calories as two small classification questions instead of one open-ended one is the reusable move: decompose until each piece is a pick-from-a-list.
12Use case 7: Editing out my bad takes
  • The classification system worked because the categories were defined ahead of time, complete line fail, false start, restarted line, small slip worth keeping, turning 'is this a good take' into a multiple-choice question instead of an open judgment call.
  • The tool caught all seven planted bad takes in 1.3 seconds for half a cent, a result framed as a first-pass filter to run before, not instead of, a slower reasoning-based editor.
  • Confidence scores double as a triage tool: a 98%-confidence dead-air cut can be trusted automatically, while anything uncertain gets routed to a slower reasoning model for a second look.
13Why it's a game changer
  • The closing bet is that vision is the durable edge: anything that needs to read a screen or a camera a few times a second, point-of-sale, robotics, security, becomes viable once decisions cost fractions of a cent and land in milliseconds.
  • None of the seven builds required writing the API integration by hand; a coding agent did it from a one-paragraph prompt each time, which is the bigger claim: the barrier to building with a new API is now a clear English spec, not code.
Glossary

Terms worth knowing.

Decisions API
An OpenAI endpoint that takes text or an image plus a short list of possible answers and returns the most likely one with a confidence score, instead of generating a written response.
Jev
A fast, cheap classification model from the startup Typesafe AI, founded by a former OpenAI researcher, that works the same way as the Decisions API but is text-only and does not read images.
Classification call
An AI request that picks one answer from a fixed list you provide rather than writing free-form text, which lets it skip token-by-token generation and return almost instantly.
Confidence score
The probability the model assigns to its top answer, used to flag low-certainty decisions, like an ambiguous voice command or a borderline editing cut, for a human or a reasoning model to double-check.
Resources

Things they pointed at.

03:52toolJev (Typesafe AI)
02:55toolGPT-6 Luna
00:00toolClaude Code (Opus 5.5)
Quotables

Lines you could clip.

02:55
“OpenAI says the decisions API makes a decision at about 150 milliseconds, while a regular call... takes about 1.6 seconds.”
concrete, quotable speed stat with a clean comparison→ TikTok hook↗ Tweet quote
02:26
“I'm 99 sure that we should advance, 100 sure we should fire.”
plain-English readout of an AI decision, easy to visualize on screen→ IG reel cold open↗ Tweet quote
22:56
“Classification correctly caught all seven bad takes with Decisions running in 1.3 seconds and half a cent of cost.”
the single best proof-of-value line in the video→ newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

So check this out. This is a salt cube. It is a first person shooter, but I'm not touching anything.
It's playing completely on its own. Several times a second, it's looking at what's happening, deciding where to go, who to aim at, and when to shoot. And in literally milliseconds, it is making each decision for a fraction of a penny.
The brain behind it is the just dropped decisions API from OpenAI. It is their answer to Jev, the startup model that everyone in AI was talking about last month. In this video, we're going to be putting it through seven real use cases from controlling my entire Mac with my voice and...
building a Pomodoro timer that watches my screen and calls me out for getting distracted to editing the bad takes out of my own video. By the end, you'll know what the decisions API is, why I think it's an absolute game changer for AI builders and how to start using it yourself, even if you've never touched an API. So let's get right into it.
And what we're going to do is we're going to kick off a build just to show you how easy this is. And then I'm going to explain exactly what the decisions API is. So here's what we're going to do first.
I'll just paste this in, let Opus 5 .5 running on medium effort, build this out. What I'm asking it to do, and I'll put the prompt on screen now, build a next generation Pomodoro timer, a little animated orb that watches my screen and keeps me focused while the timer is running. When I start a Pomodoro, I type what I'm working on.
Then the orb starts the 25 minute timer, shows as a ring around itself. And while the timer is running every few seconds, it takes a screenshot of each of my three displays and asks OpenAI's decision API whether what I'm doing right now is working on what I said I'd work on. So this is kind of like a next gen Pomodoro timer.
Of course, the Pomodoro timer is one. the first things us AI builders started to tinker with. So now we're giving it reasoning and vision ability, and it's just going to feed it right into the decisions API and determine if I'm getting distracted.
So it's going to see on my screen, you know, if I'm scrolling Twitter and should be working on writing an email, it's going to know that. And while Claude is building that, let me show you what the decisions API actually is. So the way the decisions API works is actually pretty simple.
All you're doing is you're handing it a piece of text or an image along with a short list of possible answers, and it's going to pick the right one. tell you how sure it is so here i'm handing it a customer service message that says i was charged twice for my order and if the task is to route this to the correct department we're going to give it a set of possible answers in this case is going to be billing technical shipping or other and because this one's pretty obvious it's going to choose billing with 100 assurance that it should go to the billing department and that's exactly how our assault cube bot that we saw earlier that's exactly how it was running so several times a second it was getting handed a screenshot of the game plus where the players are the enemies where the items are we gave it a set of possible answers and really for a game like this there's only a small amount of possible answers it's really should i advance back off straight left straight right stay put and then on top of that should i be firing should i be reloading should i jump should i switch a gun and so we're handing those decisions to the decisions api and it's assigning probabilities so in this case it's saying i'm 99 sure that we should advance 100 sure we should fire and it's fed back to the bot and that's exactly what the bot does and all of
this is happening in real time, in fractions of a second. Pretty freaking cool. And that's also why it's so fast and cheap.
So OpenAI says the decisions API makes a decision at about 150 milliseconds, while a regular call to GPT -6 Luna, which is in itself incredibly fast and cost efficient, takes about 1 .6 seconds. And that is because what we talked about a second ago, where a regular AI call, it's going to write out its answer one word at a time.
And every one of those words, however fast it may be, it still takes a little bit of time and therefore cost money. Meanwhile, the Decisions API, it's skipping that writing entirely. And all it's doing is it's just handing back the odds, exactly like what we talked about, for a set of answers we've already given it.
So it's already set up to be much faster, much cheaper, and that's exactly what it is. You know, it's making those real -time decisions in 150 milliseconds. Not to mention it costs 10 cents per 1 million input tokens.
Now, if this sounds familiar, that's because OpenAI wasn't the first one to do this, of course. Last month, a startup called Typesafe AI, started by an ex -OpenAI researcher, launched a model called Jiv that does exactly that. and it kind of shook up the ai world because people immediately started building really cool stuff with it it's so fast and so cheap it works great for this kind of like mass amount of data that needs to be classified really quickly for example sorting through thousands of emails and filtering them by anything you know technology related or another guy he sorted through thousands of zillow listings in seconds and he was tagging different things that you normally can't filter for like the style of the house or whether it was close to a freeway stuff like that so the implications of jev really got people building some cool stuff.
And then OpenAI came out with their Jev killer, but we'll actually see if that's the case. We're going to put Jev and OpenAI to the test for a couple of these examples to see if that's really true. But they're essentially the same with a couple minor differences.
First is just cost. So Jev is quite a bit cheaper. It's about half the cost at $0 .04 for 1 million input tokens, whereas OpenAI is $0 .10.
Of course, no output tokens. And then the big thing that the Decisions API has that Jev does not is images. So the Decisions API, it unlocks a ton more things because it can look at screenshots it can use its vision which is exactly what made that assault cube first person shooter demo possible so with that brief rundown of the decisions api and the difference between jev let's get this thing set up it's very simple all you need to do is just get an api key from platform .openai .com generate the api key like you normally would you can also mess around with it in their decision playground and then you can just go to your coding agent and say i want to build an ai that plays assault cube which by the way you'll have this full prompt if you do want to do this with a game like Assault Cube.
It'll be in the blog post below, but it really is that straightforward. So once you have the OpenAI API key stored in .env, you can just start cooking with Claude. And that's exactly what we were doing with our Pomodoro timer.
So let's check in on how that thing's done. So it looks like it's all ready to go. And I'm just going to ask Claude to start this locally so we can test it.
Please start this locally, Claude. And boom, look at this. We've got our little orb here.
Of course, it couldn't resist doing a purple gradient, but here it is. We've got a what are you working on here. gonna say coding with Claude start 25 all right we've got a nice little 25 minute timer and then I'm going to test it so it's just taking screenshots every second here monitoring my screen and let's just test this by going to X see if it picks up on it so I'm scrolling oh look at this it's changing you're on X reading agent mail ads Claude doesn't need a fan club back to coding nice Nice.
Okay, what about... Is it going to pick up on this again? Whoa, okay.
All right, nice. We got it. Scrolling X again, huh?
Back to coding with Claude before the timeline steals your afternoon. And then I can go back, and I'm assuming if I go back to Claude... Jeez, this thing really locks me out.
Oh, look at that. And there it goes. Went away.
Pretty freaking cool. All right, there we go. There's one use case.
This actually, you know, that's actually something you could probably use. I might start doing this myself instead of just constantly jumping between 100 different tasks. You'll love to see that.
So for our second use case, we've got the Pomodoro focus orb done. Now we're going to do a voice control. So here is the prompt, build a voice control for my whole Mac.
I hold a key and talk, turn my speech into text, then use OpenAI's decision API to pick which action I meant for my fixed list and run it instantly. So things like open an application, switch to another application. In the browser, go back, go forward, switch tabs, all that kind of stuff.
So I'm just going to give this to Claude and I'll just put it in the same chat. Why not? I'm just going to say, organize this code, separate it and do different builds.
And this example, this holding a key, talking to an agent and then having the agent control something on your computer isn't exactly new. You can use this with any sort of LLM here. I've done this with GPT Live.
You could even use an open source model to do this. However, there's a ton of latency with this and it's not exactly cost effective. Like we're not at the point where we can have like you know jarvis agents actually doing this kind of stuff but with the decisions api that becomes a lot more possible so that's we're going to test right now and see how well it does and this of course is why you hand this off to a coding agent too because claude is mapping out all these different probabilities it's looking at things like the api refuses when a question doesn't apply it gave only 67 confidence for quit spotify since action and app were split into several questions so claude's just going to map out all the logic for all of these different applications there's all sorts of nuances we'll probably have to go through some tests I may even just do that off camera until it's somewhat working because there is all of these nuances with opening up applications and stuff.
When you have a reasoning model that's doing this, it can do a lot more thinking than something like the decisions API where it's assigning probabilities and making actions that it thinks it should do. And Claude's humming on this. We can see it's testing.
Scroll down. I don't know what it's testing, but it's looking somewhere. Open the notifications.
Click notifications. It's even getting into social media, like liking a person's post. That would be incredible.
And Claude's got it live for us. That took about 15 minutes, I'd say. So I've built it in voice decisions, but haven't launched it yet.
First start, macOS will show several permission prompts. I'm just going to say start it up. And here are our permission prompts.
Okay. All right, we've given permission to this Carmen Electron app. Try it now, Claude.
All right, let's see if this thing's actually working. Right option. Oh, it's listening.
Hey, oh wait, you can't talk back to me. Open, open Arc. Whoa, that, what?
Open Claude Code. Oh, that's so fast. My goodness.
Open Google Chrome. Jeez, that is insane. Before I can even get the words out, it's switching.
Type lovable in the browser. i don't know how to do that open a new tab oh look at that okay so i guess i can't type that's really freaking cool claude what what exactly can it do can it type into things or can it just like open different apps again while claude's doing this look how cool this is open arc cheese open a new tab go to x doesn't work But I mean, that's still awesome.
Okay. What did Claude say? It only does what your spec listed.
Okay. So yeah, we put in this whole spec. So you could, in theory, just have Claude build out this whole spec.
We're just going to test this quickly. So we would have to have every single permutation of, you know, type whatever in the browser, go to google .com, that kind of stuff, because it's not actually like using its reasoning. it's still pretty amazing just cycling in between apps and how fast that is that's like just so much faster than typing a keyboard and for example like if i'm working in cloud and i'm working in a bunch of different sessions i can just give it the names of you know these different repos that i'm in and say switch to this switch to this repo switch to this repo it'll be smart enough to know that and then i can say now start recording and i can just record via a speech to text model right into the chat box that's pretty amazing so you can really start to see the potential with that one as well it's just so fast at it one more time open google chrome that's crazy open arc new tab very very cool okay there we go Okay, so next one is our X feed cleaner.
We're going to build a Chrome extension that's going to clean up my X feed while I scroll. This was in a Jev demo from someone, but I did love this one. So as each post loads, send its text and any image to OpenAI's decision API.
If it's some sort of ad or sponsored post, just hide it. If it's engagement bait, rage bait, or a low effort post. also hide it so we're gonna do that we're gonna copy this in and we're gonna go back to claude a lot is so incredibly fast so it's mocking up a fake x feed right now to test the model in freaking four minutes opus five five goaded and one minute later claude has our Chrome extension built.
So we're just going to go over to arc here and load in this Chrome extension unpacked X feed cleaner extension select and let's go to X and see if this actually works. Open up X in a new tab and we're scrolling over there. Look at this.
It's assigning something hidden at 100 % because it already did it. That's crazy. What that is crazy.
Before we can even see it, it does it Cool. Bookmark that.
Let's see. And we're scrolling. I don't see any ads yet.
Oh, look at those. Look at those ads hidden. That's insane.
Boom. Okay. Actually, it's not an ad.
That's just a promotion from Piper Friends. That's not an ad either. Okay.
Not totally right. But I mean, if I'm being honest, that's kind of cool. Sure.
It's hiding like company promotion, which is good. Let's keep scrolling. See if this continues to work.
Ad. Uh, no, just company promotion. I mean, this is something that you can just feedback to Claude and say, hey, it's seeing company social posts and thinking they're ads.
I didn't find these ads now, but it literally is. says add in the top right. So we can just look at that.
But I think you get the idea. So that's just an easy fix in Cloudback. You can see just how fast it is.
Before I even get there, it just shows up. Where is it? There it is.
Nope, that's not an ad. But introducing open source clay killer. It sounds like an ad.
Tebow post. I guess that's an ad. But yeah, all it needs to do is just look for the word ad in the top right.
And you've got a real -time ad blocker. So next up, we have auto blur. So this is something I actually deal with quite a bit.
And I've had an agent in the past. I was actually using Gemini because I could just feed it videos to have Gemini go through all the frames and find anything that I reveal that's sensitive information in my video. But now we can just get the decisions API to do this.
And it will be much faster and much cheaper. So I'm just going to paste this in to Claude yet again, we'll have Claude build this out. And then I had Opus just in this test put together this recording here.
So we can test how well the decisions API doesn't this a separate Opus agent just doing a great job scrolling a fake email UI like I love when it was five five just goes ham bony on these types of tasks like it really takes his job seriously and making sure that we're testing all this properly. So yeah, we have this one minute long video.
We'll see how well it does in actually blurring out that information. All right, Claude's got a blur uploader ready for us with the decisions API. So I'm gonna click blur another and then just drag in this test recording here.
Same one we just looked at. Look how fast this is. Look at it just finding these.
It's crazy, insane. Locate and then find grid and blur render. So it just went upload and then extract the frames 0 .5 seconds.
It scanned in four seconds. It located any sensitive information in three seconds. It did another further check.
in one second and then render those blurs in 5 .9 seconds. I mean, that is just insane. I hope this has given you ideas on how you can use this.
So there we go for our fourth example. Now let's do a little Wikipedia race. Everyone loves a good wiki race.
It was a classic in college for me. So let's see how the AIs do now. So here's the prompt.
Build a Wikipedia race between OpenAI's decision API and Jev. Pick a start article and a target article at each step. Give the model the current page's links as the options and ask which link gets closest to the target.
Follow it and repeat. So we're going to do that. I'm going to go to Claude.
By the way, I'm giving Claude access. You should have access. access to debt to jev via my open router API key just hit enter never mind it looks like I need API keys from typesafe for some reason it's not on open router so I'm going to get those real quick I've saved console by the way makes it very easy to get API keys you just go to balance you hit add card and then you can create the key so we're going to create this key now geez Claude already built this application and ran a full test before I even had a chance to check in on its progress but here we go we've got a Wikipedia race you can open it up here and you just type in any Albert Einstein and pizza.
And then we're also comparing this to a chat model as well. In this case, we're just going to do GPT 55. Let's just choose Yeah, banana and moon landing.
So we're gonna have the decision API, Jeff and a regular chat model. And let's click race and see what happens. And they're off.
Jeez. Oh my goodness. Look how fast the decision is.
Six seconds. It went banana, domesticated plants and animals. James Cook, 1769 transit of Venus observed from Tahiti.
Transit of Venus, moon landing. That is crazy. Oh my God.
That's amazing. And then, okay. So the costs for OpenAI decisions, zero, zero, three cents.
And then Jeb, a little bit slower, took seven seconds. OpenAI. Two seconds.
This is insane. And then regular chat model took nine seconds and costs about nine cents. Wow, this is actually really fun.
Okay, let's go Taylor Swift and photosynthesis. Let's see what happens. And they're off.
Look, oh, I'm having too much fun. Taylor Swift to carbon emissions, to carbon cycle, to photosynthesis. What is in that Taylor Swift article that's talking about carbon emissions?
Okay, this one went. Taylor Swift, environmentalism, ecology, photosynthesis, and the regular chat model still working. Okay, it takes 16 seconds for the regular chat model, two seconds for Jev, and 1 .2 seconds for the opening decisions API.
Look how fast the decisions API is. All right, I'm having so much fun. Let's go Pokemon Roman Empire.
And let's do it. Okay, three seconds. Oh, two seconds for Jev.
Look at Jev, much, much faster. Look at this chat model, still running, still running on 5 .5. Pokemon, Pokemon Conquest.
Nabunga's Ambition, History of Japan, Book of Han, Sina Roman, Relations, Roman Empire, and then Sojevwent, Pokemon Conquest, Tecmo Koei, sorry if I'm butchering that, Warriors Legends of Troy, Ayanide, I don't know how to pronounce that poem by Virgil, and then Roman Empire, and then look at this, this chat model. Hey, more efficient.
I went Pokemon globalization to Roman Empire, but it took 14 seconds. Wow, this is so fun. Reminder that all the prompts will be available to you in the link below.
I'm not gonna include this build only because it's using my API here, but that is pretty freaking fun. Okay, fifth real use case, done and dusted. Next, we're going to build a calorie tracker.
Yes, we are going to clone CalAI, not actually the UI or anything, just in the ability to track calories from a photo. So here is the prompt. build an instant calorie tracker that i can use on my iphone i tap a camera button take a photo of my food and then it will using the decisions api sign a calorie count and the way we're going to do that is we're going to pull from the usda food data central food list this is a free download So I guess Claude's going to have to download that, put that into some kind of database too, because remember, it's not going to use any reasoning.
It's just going to be categorizing. So we're going to have two questions in the call. The first is what category is the food?
And then what size is the portion? And then we just have some sort of math that's going to output the calories from the size. You know, it's not going to be that totally accurate because the options are small, medium, large, extra large, but it might be kind of close.
So size of the food and then the actual food. So let's give that to Claude now. Paste in the prompt and get going.
Let's go, Claude. Oh my God, Claude, I cannot keep up. you literally in five minutes it's already testing this full application built the database of 5000 food items and it was testing the decisions api already look at this all right two minutes later claude has completed the application so let's see what we got here calorie trackers running on your mac on your iphone same wi -fi open safari and go to this link copy that got safari open here we're going to paste this in bottom boom all right tap to photograph your food So I'm just going to pull up a photo here.
I'm going to go massive burrito. Eight pounds of bacon wrapped burrito. Let's see how well this does.
So on my phone here, we're just going to hit this. Zoom in on those, baby. Boom.
Use photo. Identify. Go fast.
That's crazy. 964 calories. Okay, yeah, we need a little bit more.
Maybe hand it over to a small model that does some reasoning here. Or it's probably because we only have extra large. I mean, it did identify it as extra large.
913 calories seems a little low for something like that let's try it again let's just do cheeseburger here but tap to photograph your feed zoom in on that burger use photo 488 calories okay yeah not bad i mean just look you can just see how fast it is it's just insane 1000 milliseconds and it costs zero zero zero zero five four three zeros five four To identify that.
Pretty freaking cool. And so for the seventh and final use case of the OpenAI Decisions API, we're going to have it be editing out my bad takes. So I've been working on this skill with Opus 5 .5 in editing my videos in Premiere Pro.
And we've gotten to about 98 % there with just Opus 5 .5, where it's editing everything. It's actually freaking incredible. I can't wait to debut it.
I'm going to put it in an upcoming video. So now would be a good time to subscribe. But the one problem is it's incredibly slow and it's very token intensive.
Like, I don't know. what the API cost would be, but it's definitely in the hundreds of dollars per video. So we're going to see if we can hand off some of that legwork over to the decisions API and just edit out some bad takes, maybe even just first pass of editing before we hand it over to Opus 5 .5.
So I'm going to give this prompt to Claude, build a tool that finds the bad takes in a raw video recording so I can cut them, give it a raw recording. I'm not going to have it in Premiere. We're just going to do it in some kind of interface.
I can have Claude just put together a timeline with a waveform and just show it that way. and then we can see how good the Decisions API actually does in finding these bad takes. And we'll probably have some sort of audio transcription as well.
I don't know how Claude's going to classify this. We'll just leave it up to Opus to do all that. Paste that into Claude.
I'm just going to give it a little bit more direction. I'm going to say, come up with some sort of classification system. I'm not sure.
You can look at my Premiere human editing to understand how we do this right now. And I probably want some sort of Premiere type timeline just in a local host where I can just upload a video and we can test this. this, but I want to see how well the decisions API does.
I'll leave it up to you. Like I said, to actually come up with the classification rules here, but really think this through. I just give cloud this random assortment of whereas grok bod and muse while they can code, wait for a bad take.
Here we go. There's one. Plenty of those.
We're going to feed that into the decisions API and cut out that stuff. Oh, also I should say also cut out pauses, have it do that. So Claude's putting all the logic together as always being more thorough than it has to be.
It's testing OpenAI's Whisper 1 and locally installed MLX Whisper to see which video 2 transcription model is going to be the most accurate and cost effective. And it looks like Whisper 1 outperformed the local model here. No surprise.
That means I'm going to have to pay for the API, but that's fine. And then here is the classification system it's building, splitting into lines. So breaking the transcript at sentence ends, at Whisper's cutoff markers and its segment.
breaks, and then finding any repeated attempts. So if I do a bad take, then do the bad take, then do it over again, I'm also going to find that, and then judge every line. Is it a bad take?
Yes or no? What's the reason for that take? Is it a complete line fail?
Is it a false start? Is it a restarted line? Stumble editor queue, small slip worth keeping, interesting, and then aside.
So you can start to see when it's mapping out this classification system where there really is, for a lot of these types of tasks, only so many decisions you're making. And so if you do the work of assigning, these ahead of time, or I guess Claude does that work, then you can have something like the Decisions API just moving much faster for much cheaper.
Now you'll certainly find edge cases that Reasoning and Model will be able to identify that this won't, but with a couple iterations of this I can guarantee you that it's going to take a lot away from what Opus 5 .5 is doing when it's editing my videos in Premiere. Dang it, it looks like this classification actually is working too.
So Claude just tested this. Classification correctly caught all seven bad takes with Decisions running in 1 .3 seconds and Half a cent of costs, just amazing.
Opus in 21 minutes, clone Premiere, not actually even the UI looks bad, but still amazing. And we all be able to see exactly where like dead air, all of these different things that are being detected by the decision API. So we've got our bad take detector slash Premiere Pro clone ready to go.
Confusing on how to actually work this. Claude clearly didn't spend any time on the UI. Okay, so I need to upload a new file somehow.
I just refresh. this there we go bad take finder we're gonna drop in this bad take then it's going to upload okay and then it's pulling out the audio okay and then we need to transcribe it this isn't the decisions api yet hey judge every line jeez 1 .7 seconds 40 calls look at that and then it cut it out and so we can see it's already categorized all of these dead air rejected take so i wonder if i just go yeah okay so it can i can see exactly what takes were kept with confidence render cut video is that going to render i want to just actually collapse all the cuts first but i don't know how to do that cloud made this incredibly confusing but looks like we're rendering it now just look at one of these, some admin work, then we'll get progressively more complex.
Oh, look at that. Yeah. So it's got the dead air there.
So I can just click another dead air directly to this agent on the website. So yeah, you can see it got paused there. If I show this directly to this agent.
Yes, look at that. I mean, I would have to take another look at this. And that's even where you would just have Opus take a second look, make sure the decisions API was correct.
It could even look at, because it's assigning a confidence score, it could look at the ones where it has low confidence. Like it can just skip this. you know, 98 % confidence and probably the dead air because that's easy to just cut out.
But all the others, you can just feed it to a reasoning model to then just finish this up. And it costs pennies on the dollar like this, three cents from the decisions API and two cents to transcribe it with OpenAI Whisper. So there we go.
There are seven real use cases with OpenAI's new decision API. If you can't tell, I'm incredibly excited about this. I think this is going to be a game changer for AI builders or tankers.
I mean, the possibilities, especially with OpenAI's vision capabilities, which I'm sure Sure, Jev is going to roll out pretty soon, but because OpenAI's decision API has that vision capability, you can really start to think about all of the endless possibilities. I mean, think about robotics.
You could hook this up to like a Raspberry Pi with a camera and you could build a little robot that's just identifying objects in real time. And it's only going to get faster. It's only going to get cheaper and just better overall too.
So I highly recommend you just having fun with this, putting it to the test. If there's anything that you've experimented with, make sure to drop in the comments with OpenAI or Jev. And let me know if you come up with any other...
cases too and maybe i'll make a future video about it as well alright with that i will see in the next one
The Hook

The bait, then the rug-pull.

A first-person-shooter bot plays itself, several times a second deciding where to move, who to aim at, and when to fire, for a fraction of a penny per decision. That's the opening demo for OpenAI's new Decisions API, an endpoint that skips writing an answer entirely and just hands back the odds on a list you already gave it.

Frameworks

Named ideas worth stealing.

01:47concept

Decisions API request pattern

  1. Hand it text or an image
  2. Give it a short, fixed list of possible answers
  3. It returns the most likely answer plus a confidence score

The entire interface is multiple-choice: you write the options ahead of time and the model picks one instead of composing a response.

Steal forAny workflow currently burning a full LLM call on a yes/no, route-to-X, or pick-one-of-N decision
CTA Breakdown

How they asked for the click.

VERBAL ASK
19:03subscribe
“So now would be a good time to subscribe.”

Dropped casually mid-build during the bad-takes editor segment, tied to a teased upcoming video about his Opus 5.5 editing workflow, rather than staged as a dedicated ask.

MENTIONED ON CAMERA
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
OTHER LINKSAlso linked in the description.
Storyboard

Visual structure at a glance.

open
hookopen00:00
mechanism explained
promisemechanism explained02:04
first build live
valuefirst build live06:02
X feed cleaner
valueX feed cleaner11:00
Wikipedia race
valueWikipedia race15:28
bad takes editor
valuebad takes editor20:15
close
ctaclose25:02
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

36:06
Pat Simmons · Review

Kimi K3 Is Here! (Better Than Opus 4.8?)

Ten identical builds, five models, blind-ranked before the reveal — a real-world stress test of Moonshot AI's new open-source model against GPT-5.6 Sol, Opus 4.8, GLM 5.2, and its own predecessor.

July 17th