
Get every episode summarized
Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“We're laughing because I literally came in and Andy pointed and I just I was about to say is it you is it me? This was one of those days where I had 15 minutes to the show and then I looked up and literally said it was 10 a.m.”From the transcript
The episode opened with Gemini 4 Argon, Google’s new frontier model currently limited to cybersecurity researchers. The hosts compared its early Artificial Analysis results with Astra, Fable, Opus 5.5 and Sol 6.1, then noticed an unexpected coding result: Sonnet 5.5 ranked above Opus 5.5 and Gemini 4 on the coding-agent index they reviewed.
That led to a deeper discussion about multimodal AI and what it would take for a model to truly understand video. Brian described how his current thumbnail system samples individual frames, while the next step requires understanding expressions, audio, movement and events across time rather than treating each image independently. The conversation also covered Figure’s unusual decision to train its Figure 02 robots to autonomously jump into molten steel during decommissioning.
The second half shifted toward agents. OpenAI’s Decisions API was compared with JEV, while Gareth described Dot interrupting his work to surface an urgent school security email and later notifying him when the situation was resolved. Brian shared how Muse helped surface the used Kia Niro he ultimately purchased. Those examples pushed the hosts into a larger question about AI education: as agents handle more prompting, research and orchestration themselves, should new users still start with traditional prompting skills or learn how to define goals, judge outputs and work with agents instead?
The hosts also discussed the voluntary White House AI safety accord signed by major AI companies and the FTC’s investigation into potential consumer risks from AI systems. Both developments were reported this week. AP News
Key Points Discussed
00:02:01 Gemini 4 Argon Enters The Frontier Model Race
00:04:04 Gemini 4’s Artificial Analysis Results
00:05:34 Gemini 4 Versus Sol On Coding
00:06:15 Sonnet 5.5 Surprisingly Leads The Coding Index
00:08:16 Figure 02 Robots Jump Into Molten Steel
00:15:34 The White House AI Safety Accord
00:21:40 Has Opus 5.5 Already Been Dialed Back?
00:23:39 Gemini 4 And The Future Of Video Understanding
00:29:24 How AI Chooses The Best Video Frame
00:31:47 Why Understanding Video Requires Context Over Time
00:34:33 FTC Investigates AI Risks To Consumers
00:36:15 Chinese Model Distillation And Cybersecurity
00:38:50 OpenAI’s Decisions API Versus JEV
00:41:30 Why Codex Was Slowing Down
00:42:51 Gareth’s Dot Surfaces An Urgent School Alert
00:44:55 Muse Helps Brian Find His Next Car
00:46:49 Should AI Training Still Start With Prompting?
00:49:05 Ethan Mollick And The “Bitter Lesson”
00:52:38 Teaching People To Define Success Instead
00:54:42 Should Skills And Agents Become The New Basics?
00:56:02 Meta Hires MongoDB CEO CJ Desai
00:57:37 Meta’s Reported $4 Billion Data Center Tax Credits
01:02:08 Episode Wrap-Up
The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Gareth Hood, Beth Lyons, Karl Yeh
Get every episode summarized
Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
569 searchable segments. Every word is indexed and playable.
Full transcript
The Daily AI Show — Is Gemini Back At the Frontier with Argon 4?. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hey, it was fun on everybody. Welcome to the Daily AI show. Today is October 1st, 2026. We're laughing because I literally came in and Andy pointed and I just I was about to say is it you is it me? This was one of those days where I had 15 minutes to the show and then I looked up and literally said it was 10 a.m. So my apologies for being a hair second late there. Anyway, Andy is here, Garrett is here, thanks Garrett's for being here right on time today to help us out. I'm Brian. Welcome to the show. I wasn't here yesterday, but as these were my exact words to bet on Slack yesterday, y'all hate me. And my reason for saying that back to everybody is because she I had just a crazy day yesterday and I was all it was a crazy day. And so I couldn't be on the show, but I do the post show editing as you guys know, but I'm telling the audience. And Beth just
sends me a Slack message that says, Hey, the show was an hour and a half today. They might be, you know, blah, blah, blah. And I always joke about that because when Jimmy was able to be on here more and of course we'll get Jimmy back when he can, that used to happen whenever I wasn't on a show, which you would think with my mouth and my ability to talk. I would be the one on the hour and a half shows, but it's usually when I'm gone, these shows go super long. And then I'm like doing this super editing and it just strains the systems and everything else. So I know you all, I know Andy, you weren't here for the whole thing. Because I saw it in the edit, but I guess I knew you were probably here for the vast majority of you as well. So I'm sure you guys already talked about all the super cool stuff that came out of Dev Day. But did you talk about Gemini for our gone at all? Did that happen? And not yet. So let's take a look at it. Yeah, let's take a look at it. Because I don't have a ton at the moment and it's not out yet for everybody. Yeah, let's see what we do know about it. Right. We haven't been able to try it ourselves, but artificial analysis has done its sweep of it. I'll show you where Gemini 4 fits in the world. But here
you can see a comparison against Astra and Fable and Opus 5.5 and pretty much across the board with a very few exceptions. It beats them all. So it's state of the art at the frontier level on many of the benchmarks that matter. And it's right now on sale. It's going to end up being about you know, about the same cost as Astra when it goes off promotion. But right now it's 50% off. But that doesn't matter to almost anyone because it's only available to cyber security firms at the moment. So it's not publicly available. Let's take that off the screen and I will show you another view of how it performs that you're more familiar with, which is the artificial analysis index. That which I was just showing you is their own comparison. Right. So that's
using standard benchmarks, but that was built by them. But the same story is available from, let me share my screen now from artificial analysis, which is this one. Hang on. You do what I just said instead of sharing your remute date. Yeah. I do that all the time. The buttons right where it shouldn't be. While you're bringing that back up, the pricing, the introductory price that you were just talking about is $2 per million input tokens and $10 per million output tokens, which is affordable, you know, comparatively to other models out there. So that's that's yeah. So here we are now with the artificial analysis index. And you see on this top left one, the intelligence index, that Opus 5.5 still beats the field by a significant margin at 58 while side by side. Now you have Fable GPT 6 Astra and Gemini 4 at 53. And GPT 6.1 sold
just barely behind at 52. So here you see the discounted price $1.99 for a cost per task. That's this is artificial analysis doing something akin to what Brian has instrumented on his application development where every every single run through the API is it measures what it's costing. And that's what artificial analysis is doing when it runs all of these models against the tasks in the benchmarks. It counts with a meter exactly how much that's costing in their models API. So it's there at $2. Imagine it going up to about $4, which will put it kind of a kin to GROC 4.7 in terms of cost. But you see over here that GROC 4.7 only performs at 46 compared to the 53 that you see in this sort of top tier of the intelligence. Now I wanted to show you something that's really kind of interesting down here. And this is about coding. I was surprised to see,
okay, so here you see anti-gravity, the harness anti-gravity CLI with Gemini 4 argon performing at 64, which is this just here above codex running GPT 6.1 sold. So it came out with something that open AI hadn't been able to best, you know, even with its most recent last week release of 6.1 sold. Or was it this week? I can't remember. Like a few days ago. But now look at this. So Opus 5.5 at 66. But Sonnet 5.5 at 68. Now I never seen any news or mention that on a coding agent index that covers all these various things that Sonnet 5.5 was beating the field. So that's a good takeaway because Sonnet is going to be less expensive for us vibe coders
than Opus 5.5 and it's doing better work apparently. Yeah, that's weird because I mean, I think it's moved so fast. I had been using Sonnet full time and then yesterday I was like, why am I using Sonnet? Why am I not using the other ones? Opus 5.5. Yeah, or Opus, I knew why I was doing over Opus 5 because it was cheaper and it was better I guess. But I didn't realize it was above the wow. Coding X 5.6. Yeah, yeah, yeah, significantly. Five points on the artificial analysis index. So it gives, it depends what you're doing in your coding, right? Because if you are incredibly explicit and something could just run the gamut on that, it's helpful to use a model that defaults to less reasoning, right? I'm just going to follow the instructions. I'm going to do the thing. If you need a bunch of reasoning, then that coding
job is not going to be as good. It's this, I feel like it's a similar concept that we get to when we're saying, we'll use Opus 5.5, but use it on medium. Don't ever go above medium if you aren't doing some very high reasoning. Hi, everybody. Best year. I think that's all right. So you do it. I like, I totally, I'm looking out to the other screen because you bet and whoever else, Jeff, you guys see that figure O2 decommissioned video? I am totally absorbed with that on my other screen. It's got all my adventure. Is this where they show robots diving into metal, molten metal? Yes. Wow. That's, that's perverse. Arnold Schwarzenegger, which is even cooler if you're geeking out in my, you know, of a certain age and saw the T2 movies and stuff like that. Yeah, that's, I haven't even watched the whole thing. Now, is that, I mean, do you really have to go to a steel mill to get
molten metal or did they build their own smelter? I just assumed it was all AI. Honestly, I assume none of this was real. Yeah, I don't know. It's where my brain went first. I'm like, I'm sure I thought it was real. Yeah. Well, it looks really good. It looks really good. So if people want to know what I'm talking about there, figure, we love figure O2. It's an incredible robot figure, F O2 or figure O2. Marked many first for the company, our first deployment at BMW, the birth of Helix, our first robot doing house chores, and our first logistics development as our figure O3 fleet grows, maintaining O2 fleet no longer makes sense. Decommissioning the fleet of robots for a full full of custom built actuators. Another intellectual intellectual property is no easy task. We needed to dispose of the robots in a way that protected our proprietary hardware. This assembling each robot would take our technical staff so much time it would delay the launch of figure four. So we asked the internet, what should we do? The resounding answer came from an unlikely source, Arnold Schwarzenegger, who told us to melt them. We reached out to Arnold and when he
said, and when he said he wanted to be involved, we put everything in motion. So then they show a screenshot of Arnold Schwarzenegger responding back to Brett Akok saying you should melt them. Turing Post says it's real. Here's what Turing Post says. They trained them in a simulation first and they did that with stunt performers to show the jumping into the bucket, which they're going to go to a foundry, so a steel foundry that these robots have never seen before. So they have to kind familiarize themselves with the environment and then they jump into the bucket of molten steel, one of those big, big buckets that tilts and it took about 20 minutes for each one of them to melt. And during that, the heat and the electromagnetic fields created were knocking out the cameras.
So figures says the footage is real. You know what? I'm going to bring this up on stage. I should have kept reading in the article because I guess this is exactly what you're talking about, Andy. They're showing here. This is the simulation of them jumping in and I guess them actually doing tests on the right hand side. So 3D simulation in the left, doing some results into a pad. And then if we go here, I guess this is like you said, the melt was in Finland. So you can see, like you said here, the melting bars, as machine limited, let's see here. Most of the F2 fleet is gone, just a free remain in storage in HQ. With the help of Arnold Schwarzenegger and our film crew, we pulled off something pretty wild. In an age of AI generated footage, it's worth being clear. We really trained robots to jump autonomously, ship them to Finland and had them leak into a bat emulten steel. We hope you enjoy the film. So props to, props to figure.
Figure that I wanted great publicity stunts. What a great publicity stunt. Really great. I hope this makes all the news, hits all the news rounds because like what a different way to kind of go about this. And yeah, very, very cool. And then it says something about limited edition artifact purchase here. Okay, well, I have to know what that is. Let's just look. Oh, yeah. I suspect that they're taking the bucket of steel now, which contains many robots parts. That's cool. And now you can get a little medallion, you know, with figure 02 on it or something. I guess so. That's yeah. So you can get a small one-third scale pre-order, five hundred dollars. And there you go. There's a look at it, I guess. The small artifact is one-third scale of an actual F2 head unit. There's 290 of those remaining. And then the $1,900 once says there's zero remaining. And that's the larger artifact is two-third scale of an actual F2. And it says none remaining. Now whether that's live or literally that many people bought it.
There was none remaining in the first place. They just did it. Maybe. Yeah, I do like that's the part I don't know. That's the manufacturing scarcity on that one. I know, but no kidding. Yeah, listen, I'm sorry, this is cool. Yeah, it is cool. Okay. Go ahead, Beth. I just want to say that this has the potential to play badly with people who think everything's going too fast and things are taking over. And why are we doing this? Like once this is real footage, once this real footage gets out, I feel like I feel like this was one of those moments like this is cool. This is great. We should totally do this. And day one think about metal consequences for the people who seem to be just on the edge right now. That's and wait until AI finds this video and realizes that we're teaching them how to jump
in temples and love. Right. Because if you go as part of the training data, like, yes, if you feel like you're being tested and about to be destroyed, maybe you are, right? Like all that said, very, very cool video. Absolutely cool idea. Just just because you can't. Doesn't mean you should. I was like, what was the last part of this going to be? I was like, wait for like the bet. But then yeah, you can. Right. I mean, all of us thought it was AI in the first place. You could have just said no worries. This is AI. We didn't really train the robots. But here's kind of the cool part about this bet. Me and Gareth were like, oh, that's probably just AI. Then Andy's like, no, and I literally just kept scrolling in the same damn article I was reading. I was like, no, you're right. They're making a point of saying this is not AI. That would totally, I mean, nobody would have even, I don't even think what a carrot
would have been like, that's a great video. And they could have said like, we did one. You know, or something like that. But the idea that they're like, doubling gamut saying, oh, this happened in Finland. We really did this. We hope you enjoy it, whatever. I think just put this figure at this sort of like next level because the easy answer would have been like, we created an AI video and it looks totally real. And I think everybody would have said, oh, it's still pretty cool. You know, good idea, whatever. You know, so anyway, I don't know. Okay. Well, I want to move on to another thing that happened yesterday. There was another smelting happening in Washington, DC as major tech executives gathered to sign a voluntary super intelligence accord at the White House. And so that was Musk and Zuckerberg and Amade and Huang and Sundar Peshai and open AI's not Sam Altman, but Greg Brockman, who is the president. And they all signed a one page pledge promising four layers of safety checks in terms of testing
and outside audit, a board review and more. And it's voluntary. But the tech says only that turning it into law eventually may make sense. So that's a one page, you know, executive order companion to the executive order that Donald Trump signed, which requires that all federal agencies now abandon use of the term artificial intelligence AI and instead use super intelligence SI. So we'll have to make a decision later as the daily AI show whether we're going to become the daily SI show. But I for one, and very happy that they've given up AI because I'm going to take it over. And I want to share with you now what that really means animal animal intelligence. Now it anti intelligence was right there. And I've got for only $50 you can buy a hat with horse and heart, which is a nonprofit here that does animal rescue and animal intelligence training.
And now you too can be an AI expert, but you don't have to do do any coding to do that. Listen, this is more exciting than thumper. I really like I really like this idea. I want to, you know, let's get those up. Let's get let's get donations going out to horse horse and heart.org. You got that for real. This is not available today, but maybe it is. I don't know. The ladies are putting it up. So animal intelligence is a real thing. It's not it's not a hype thing. It's a real thing. And you'd be amazed at what some animals are capable of doing. So join the join the world of AI now. I do what I believe this I do want to pause really quick to make sure that anybody listening does does understand I think they do, but in case we have new we have new visitors all the time. Andy really quick. Please sell for some don't or because I mean it is a real. It is what you and your wife and your amazing kids do and stuff. So give give a quick shout out for that because I would love people to support it if they can't. Well, you can
actually see where I live and where I work. If you go to horse and heart.org there are pictures and videos of of our operations here, which is what we call consensual horsemanship. So this requires no tack you you interact with the horses and you do not use coercive bit in the mouth or spurs or any of those cowboy kind of things is what is required is that you develop a relationship through understanding and practicing the actual that sort of interpersonal cues that horses in a herd use until you can actually work in their context as a leader as a leader in that herd or with the horses that you develop a partnership with. Only then can you at the horses invitation can you get on board and ride that horse and so it is a very safe consensual environment for horses
and we also do we have nine horses here. Many of them are old horses that work previously sport horses, but unfortunately in the world of sport horsemanship once they are no longer competitive they get not put out the pasture but they end up going into stalls somewhere and have you know a horrible life. Well, we've rescued a few of them here to use in this program and we also do equine therapeutic kind of services for foster kids in the Monterey Bay area and also youth at risk plus many many adults who are going through transitions of their own and who are looking for the equine facilitated learning that helps them understand and progress through their own issues by understanding how horses behave and so that's all happening right up here in Soquel California. Go to horseenheart.org and check it out and apparently soon you'll be able to buy the animal intelligence hat and be a member
of AI. And even if they don't want to like let's say they don't want to wait for that but I'm sure there's people out there Andy who just love to support what you guys are doing. Oh great yeah there's big donate buttons on there because it is there was but I wonder if I got you. It's a labor of love it does not nobody here takes any compensation so it's I mean it's not available the costs of maintaining those horses is more than it more than all the work that is done collecting money from students and from programs that put people into the equine facilitated learning programs all of that it just just barely covers what it costs for veterinary bills and farriers and feed and now hey which is transported by diesel trucks is you know rapidly rising and it's going to be cost prohibitive at some point for many people currently own horses to keep them fed. I think that is very very cool I'll be right back but
very cool I'm glad you tell people about that Andy but I'll be right back. Okay cool. Now Beth do you have some news or shall I go on? So one of the things that I have seen is some reporting that Opus 5.5 has been nerfed I believe so there's there's some theories that most of you have used open A.I. I'm sorry if you use anthropics models you'll see sometimes you get the little message that says hey how is anthropic how is a claw doing for you today and could we talk to about your use case. So it seems very much that this is another one of those places where really great when it was immediately dumped and then slowly pairing it back trying to calibrate what's the lowest level we can do that still gives people what they you know like they're still happy but I want to
say that if you're using a clawed model and you feel like wow it was really good yesterday but hmm a big problem today go ahead and say no clawed is not doing well and pray that everyone else answers the same so that when they re re what retweak things retune it they release it to a slightly better thing and it's like the codex people are having that same kind of conversation so that is my that's this news and it just it's a pattern that people are it's a pattern that people are seeing and it's more visible because things are just released right so it isn't like hey it was doing really well six weeks ago no it was doing well two days ago and now it's less so I don't know if you guys because of course we don't know and I'm not to
circle back on news it would ever because I know I guess I'd check with the figure two but you know we were talking and you were talking a lot about the Gemini for argon at a scroll up to find out what the name was and I will tell you although there's a lot of stuff on there I mean they talk about a working work quantum computing and stuff like that I only really saw one mention of multimodality multi multi-modality geyserine and to a sustained long multi-step task enabled to excel across arrange a random prize workflows so it talks about that but then it really gets into cyber security and it's a leading there and blah blah blah all of which other stuff so I guess on the on the benchmarks LV benches a multimodal understanding thing the previous top mark was with Astra at 87 and LV bench for argon is 92 okay well that I'm glad you brought that up because I missed that part and I was like I was was you know I've been saying for a while I've been looking for this next Gemini model specifically for video and multimodality and so I'll be interested
to see obviously now they're saying it's only going to roll out through API which is fine and then ultra subscribers first whenever that happens but that's where I think you know we were just talking about this on the show that's where I'm the most excited for this particular model is to see what its you know improvements are over 3.1 pro you know in terms of its ability to whether it's more frames per second that it's able to ingest or I don't even really know but you know that's that's really what it would be at this point because I feel like that's the next leaping the next leap that the models need to take in order to do you know really really amazing things with like shows like this now we don't move around a ton we use our hands and move our heads around the stuff but we don't like we're fairly stable here but you can imagine for a lot of other types of video content where there's a lot of dynamic things going on at once audio visual the spoken word whatever the case may be plus plus citations and non-spoken I think that's where I guess really excited so we'll see I guess but I just want to swing back on that really quick because
that was the one thing I was looking for and I didn't necessarily see so thanks Andy for bringing that up so have we seen anyone talk about whether it's just getting faster at processing one one frame at a time right like it used to be okay it's only taking frames every I don't know point three seconds point three milliseconds whatever that is um and and then we're getting to a place where oh no no it can watch video just like a human but is that still it's just faster at processing is it still doing frame at a time or for like is it still selecting like every fifth frame every 15th frame that's anything I've used has been of some variation of that even the the actual thumbnail um builder that I've been talking about that I used for the show which you know now it's been like what about three weeks two weeks of building the thumbnails out of um it's not just cloud code it's it's it's a multi multi-step
process uses other models as well but yeah that's it's it's like it's only looking only looking every so many seconds for us to check if I you know Garrett just had a really good look up to the upper right or left depending if he's there or not yeah looks very like he's he's thinking hard you know so it's like going through and picking those out but obviously that's not the same thing is watching it you know 24 30 frames per second even 12 frames per second for that matter so I would imagine to me I mean video is nothing more than frames per second so yes it probably is a processing problem at the end of the day which is why I know it'll be it will be solved for because there's nothing really there's no like void where black space between where we are today and what we would require for full video processing I had 30 frames per second like I mentioned like there's not there's no leak in technology there's there's a void with what they can do but
it's not like none of us here can conceptually think about what it would take to process video you just it's kind of what you said Beth like oh you need a higher processing power and ability to rapidly I mean you know rapidly rapidly um uh analyze and and produce outputs on 30 frames per second over over two hour movie or something like that and you go yeah that's a that's a capacity problem that's a speed problem that's not a this technology doesn't exist problem you know but then that is for me like okay so are we at a point where the jav model whatever we're using if it's still jav or if it's lunar uh I don't know what jeman is going to call it but is there something that can't analyze the visual enough as part of the model and then just output nah nah keep right yeah so uh to be garath was looking beautifully off up and to the left um
but uh something was fine we're just using you sorry I'm reading Andy's website all right yeah uh as you should as you should but like there's some it's too far it's not the three quarter turn enough or whatever that and then your actual assessment can be like oh not only is that uh like here's face okay but it is also uh yeah that's a key point yeah listen you're talking about exactly what I deal with it I try not to just nerd out on the show every day talking about this thumbnail builder but here's exactly what happens now the majority of us Beth you're kind of looking more straight at my camera is definitely about a foot below I'm looking about a foot below where my this is me looking at the camera and this is where I'm normally looking and there's definitely a difference well this this I'm just going to stay in there like this right here is not great for it's on now looks like I'm uninterested I'm currently looking at all three of your faces but I look uninterested so okay now follow me on this this is exactly what you're talking about Beth and it like
it's kind of nerdy but it like whatever as so Carl we usually like poke fun at Carl but Carl's like I have two screens and I'm looking over at the I'm looking at you guys but I look like I'm looking away right and so whatever so Carl the fee turns his face and now I'm looking at this I'm looking at a yellow duck on my as as a matter so I clearly should not have my eyes over here that's weird right that's like that's like wise Brian's eyeballs right but but to your point Beth like where between here and here am I actually looking ahead because a lot of times our heads are like this because we're turned sideways in our chair your head is turned slightly to one side right now Beth but you're looking at the camera you're still engaged with what's going on here it's the silliest thing but like that's this is what these are the conversations that I have sometimes with okay what lesson can we learn because Garris head was turned a little too much and so in the thumbnail Garris should be looking off meaning you know longingly at Andy's website up here somewhere you know he's he oh he's got his hand on his face at his hand on his face you know and like yeah exactly like that's
okay that's that that's the image we should grab as opposed to me looking down and then what I have to say is like no no if I'm looking below the camera you have to like AI my eyes up right to look more and also you know if I have a really resting face and it goes through all 900 whatever it looks at and all it can find is this it should give me this just just band it just a little bit it does it does because like who wants to look at for okay we'll take Andy out three other faces that look boring you know or like they're bored on the show that doesn't make for a great thumbnail so it's like it's it's a whole it's been a really weird fun process to work through but I think what you're saying Beth is like they will come a time where not only will it read 24 12 frames per second and understand it but what really needs to happen right Beth is that it has to have an understanding of what happened before in context and what happened after the framing context because just just
getting a frame of of your car will just join this Carl throwing his hands up in the air all that says in that frame is yeah why can you play it with a balloon man from from the Simpsons all that says in that frame is Carl had his hands up in the air but it doesn't say why Carl threw his hands up in the air and so I think that's the big divide here is you not only need that bound volume of data for AI it has to stitch it all together and understand how does this one frame out of this 24 or 30 in a second compare or matter to the rest of the frames around it or even seven minutes later in this show or previously I guess where I made an inside joke and seven minutes later Carl has a visual reaction to it what you really need is an AI system who fully understands those two things are linked together and can do something with it and and correlating that to the audio is really
important as well and so the 11 labs sophistication that just came out and its latest release that actually can understand the different inflections for a single sentence and and construe the correct meaning out of that is really getting toward that objective which is correlating what is being said with this with the energy and emphasis that's contained in the audio stream with the actual physical presentation in the video yeah I use the 11 labs model yesterday actually for 11 for what yeah 11 for yeah yeah yeah and it's good it's better than I have a custom voice in there or trained voice in there and it sounds a lot better as robotic it sounds a little bit more excited than me but it does like it's good I'm just not as excited as it makes me sound but it's fine it
you wouldn't know unless you really knew me yeah but yeah hi everybody I'm Karen and I have a great set of tips to give you today I would love it let's see if that comes yeah there we go that's fine you should know metham fedamine for for for Jared to get him go yeah yeah okay I want to I want to jump on a couple of weaves here we were talking about the new accord that was signed by all the tech executives at the White House simultaneously you know you have vans and others saying you know the existing infrastructure of our government can can address the risks to the human population and one of the ways that they're expressing that is that the federal trade commission has already underway and just a released after that accord was signed yesterday that they have an ongoing investigation of open AI and throp and others over AI's risks to consumers and there will be
formal compulsory demands like subpoenas and so on in the coming weeks and that's already underway started in the summer and it's it's being practiced so that the government can say look the executives are saying trust us we're going to do this right we've signed a little piece of paper that says so and this FTC probe is going to be the monitoring and demonstration of their compliance with those things and then the question is does that really even match up with what the labs are really doing behind the scenes you know so is it papering over what's happening and is are there really internal safety documents and practices that will align with what the FTC and other government regulators might do so I wanted to just throw in one other thing which is the other weave which is
we've talked about the distillation practices the so large-scale data extraction and distillation that's done by the Chinese developers the main labs in China and so open AI said yesterday that there was a summer campaign that they've been able to trace back to China's moonshot AI which is the Kimi line and so that is very interesting that you know the most capable cyber offense models are the open ones and distillation can copy the closed ones does even asking the frontier labs to slow down make any sense when the level of reasoning that's already been captured by those open source models which are openly available and can do the kinds of cyber attacks that sort of a
mythos capable model can and now there's a whole suite of them including Gemini 4 which by the way is not available publicly yet and we talked about that at the beginning of the show but is in the hands now of cybersecurity researchers because of its extensive cyber capabilities that is it can mount attacks and it can also identify vulnerabilities and develop defenses against those attacks so all of those things work together but the open source field is so close now to the frontier in terms of the sort of possible negative uses or you know rogue and nefarious uses of these advanced intelligent models that it may not it may not be possible to contain them. Hmm. Well, it's one of those situations where you don't actually need to be
as consistent in the negative area right you can try 50 times and it only works once and it's a success if your ability to defend fails 49 times and only works once that's a deep failure so it's at par but it doesn't even need to be at par to be successful right. I wanted to bring up some new bids you guys might have talked about it yesterday so we don't have to if not but I found it interesting one of the things from DevDay was decisions API did you guys dig into that a bit yesterday? A little bit we've got to show it but not too deep. Well I don't have anything deep to go into but I'm just curious what anybody's take is on it being a Jeff like API do you refuel like decisions was a sort of a direct response I may could have existed obviously I don't want to assume that they built it into each sort of but
this this idea of having a systems system one thinking that sort of plays interactively and I think they said they were it was I don't know if it's out yet but it was built off of their Luna now which would be what the smallest right of the three which you would which you would probably only need for something like a decision API but I find this really interesting I'm curious where the the Jeff differentiator here is or exists perhaps as I think somebody on opening I said on joked on X saying that you know these are the clone wars the beginning of the clone wars was D.A. I'm sorry a former open AI engineer Diego no Diego all media is a CEO he was talking about how it would be the clone wars so I don't know I was just curious what you thought if it was I it wasn't a huge part of the Deb Day I mean I remember seeing it on screen when when Sam was talking about it but I guess I didn't put two and two together in my head that like oh this is potentially a baked in Jeff solution directly through open AI yes 100% it is
and what we haven't seen yet and I'm not sure frontier lab can can match is we don't charge you for the output because a frontier lab is trying to capture as much value that that they can right they're trying to squeeze everything out but yeah we haven't seen anything about the pricing as far as I know I think it's hard though because like to implement Jeff by itself it's a little bit more complex than implementing the decisions API to whatever you're building so I think just the ease of use of that versus no mind you can use your whatever your large language model of choice to implement it but like right I think we're talking about how fast it was and you like I can't tell the difference between this yes this fast versus this like it's already less than the second so it wouldn't really matter um I guess it's interesting it's interesting because there was a lot of
complaints about codex slowing down yesterday and jokes like well you can't in your five hour limit if uh if you don't if you only go at this number of tokens per second and uh we're we're not getting near the limits of the five hours it's taken so long I think bring that up the bet codex well sorry I don't want to switch over but codex um Tibo said that they were switching things over and moves it sounds like they moved more resources to 6.1 sole yesterday so that's why they could have been slowing down for a bit and Tibo's making a call for uh like you put an expo out what would you like to see improvements for dots and space spaces yeah maybe um and spaces are different the page correct I put some of my suggestions in there like being able to customize your own voice with dots um because I think that's what makes it more personal um and then being like I think what being
able to when you're on the phone with dot on your mobile being able to do other things because if you switch out of it it hangs up because it's not like a real phone call um I will share my dot experience yesterday that's was kind of mind blowing kind of like holy crap like this is this is what I needed um and it was like one of these moments I don't know how to describe it I was in awe um I was in the middle of a conversation with my dot throughout the working day working on projects trying to work think we were just talking things through and all of a sudden it interrupted what we were doing she said hey gareth just want to let you know you just got an email from your school or your son's school saying that they've gone into secure protocol because there was someone seeing with a gun on the campus and I was like holy crap like it's crazy that I mean I get emails
all day long and it shows that email yeah you brought it to my attention right away and then when there was a follow-up email it followed up with me and then was just like hey just want to let you know everything is fine they've apprehended the person um and then we just continued to go back work like it was wild like wild but like I wouldn't have seen that email previously I don't read emails I hate them to be honest and so that was like the game changer for me like I it kept me more informed it was very helpful it found and decided that that was something that I urgently needed to see and interrupted our conversation just to tell me that yeah what which agent was that gareth dot yeah that's right oh yeah it was the probably one of the coolest AI experiences that I've ever had yeah just because it decided that was a high in new it was a high priority item compared to all
the other emails that I get hundreds of emails I get a day and it needed to tell me that and I was just like holy crap um I don't have anything that cool but uh muse um actually um found the uh car that I just bought yesterday so that's cool right I mean it did the it did the work I mean there's other tools that do it as well I just happened to be using muse and um a couple nights ago I was waiting for my daughter uh to pick her up and so I was just sitting down in and I played with you see you know and then I was like oh that's right I built the the nero or it did I didn't do anything it put together a nero tracker I was looking at um uh key in the eros I've mentioned this before it doesn't matter um it is a new evee did so that my daughter can have our old evee when she turned 16 and anyway it was I'm not saying this was like magical I mean it was just it was doing the tracking and stuff but man it brought one to my attention that I hadn't
previously and I thought oh that's different that look you know price and and uh all the things and it brought it right to me and anyway long story short um after you know going on site and um you know validating everything uh didn't affect in the being the the uh nero that we the used here was a new you used nero that uh we bought we'll pick up on sunday so kind of cool you know like the agent was involved in it I there there are frankly are other ways that I probably could have handled it exactly the same so I'm not giving all the credit to views but I did happen to be on use when I found this particular uh vinn you know so that's you know kind of uh I've never heard of that I've never heard of that evee maker nero oh it's uh neson yeah isn't kia isn't neson yeah oh yeah no kia it's a kia my uh yeah my friend has been like she you should get it before I got the ionic she's like she's impressing me that because I was in the market for like three four years I just I dig forever to to do these things hey can I um I wanted to ask
this question it's been weeks since I wanted to ask this question to everybody um and it's more about how we are training I don't know if I've talked about this but it's always been a kid thought of mine is how we are training um let's say businesses employees people because there is before all the agents and personal agents was really there I remember the training was you know how do you prompt how do you navigate all the features personalizations customizations those kind of things like the base knowledge and then now um everything's moving into you've got personal agents you can um how you would work with agents so I guess the question is is it to the point where you just bypass all that foundational stuff and go right into um hey you know what let's super like go right into
agents agents training and ditch all the you know the foundations or basic stuff that um we're doing because that's the the thing I'm grappling with in our business is like we offer that training but like I was like eh do I cut it do I change it do I hybrid it like what does that look like now because things are accelerating even faster that I don't know what are the skill sets people should have and if we give them the base skill sets you're going to have to train them on the other thing are we just delaying the inevitable and are we just actually putting them backwards um and there's more and more people using AI but like even the base capabilities still not but I'm like do we even need to give them that or do we say hey you know what forget those the agents can do all of that for you know I yeah no it's I think first of all I think uh Ethan Mollick touches on this with his latest one useful thing um a blog that came out yesterday
I think or no the today actually it's called the dot in the swarm but he calls it he talks about the bitter lesson actually there's a really cool video that Claude made with um Suno in there uh that he apparently one prompted and it gave out this video that explains what the bitter lesson is it's actually really good so go check that out but anyway Cara what I what he is talking about too is like he he feels like he's pretty he's not pretty good at guessing you know what's coming I feel like I'm locked in and you would agree I'm sure but he felt like he got that sort of what you're talking about like that problem wrong which was oh we're still gonna need to in the video it's like sort of lay out the train tracks with all the switches and all the all the pieces in parts and now we're finding out oh forget that just give it to the AI it'll find its own intuitive path to the answer and it's like shoot should I teach people to lay down the track and the switches because which goes back to should I teach people how to prompt I mean six months nine
months a year ago I was still like look there's value in learning how to prompt because you're understanding the uh that would be like learning grammar you're understanding how to write a hamburger paragraph essay right there's an intro there's three parts to the body and there's a close and it's a hamburger right still helpful I can I can conceptualize how the thing is is put together and it makes me better writing an essay let's say but I agree with you now it's like well where do you introduce people if they if they really are just dipping their toe into the water of AI now and just starting to use these models more often where do you bring them in because do they is there any competitive advantage to them knowing some background information if it's not really going to change how they interact because if anything the AI is getting easier for them just to talk more plainly more more of stream of thought and the AI goes I got I got it you don't even have to wait out for me just give me what you got yeah one thing I know Beth you're going coming in I just
the argument there's one argument for saying hey all of us have gotten to where we are about using because of all the calluses we've developed testing doing all that stuff so if you bring someone brand you into this you they won't have developed all those calluses to know hey here's when we switch models why do we switch models understanding the token implications understanding hey I can orchestrate this stuff and also understanding there's certain things that I have to mention to ensure the whatever I'm doing gets to the output I want because of everything that all that callus and experience we have all developed but then people who aren't using it on a regular basis I'll give you an example there's one company I was talking to their consultant and he was like yeah we're only allowed Gemini and we can't use spark or anything so we can only use gems which
gems is going away too so he was like I don't I I can't even remotely get people to use anything other than what we're locked down to and that lockdowns are the capabilities of lockdown so I can see where everything's going but I can't do anything about it nor can I even practice my muscle on it because I'm locked into this unless I do something on my own so yeah I think that I really like Ally Miller's approach she starts with like walk around and complain right it's just like throughout your day you got a complaint go ahead and like to talk about your complaint and then for me I think we come into what is the what is the success criteria when it's successful because you can start to move through that kind of experience like solve the complaint this is what a solution looks like the other piece is predict what do you think I'm going to score it as right
like do I will I think this is successful because it may be that the initial success criteria in likelihood the initial success criteria is not sufficient to be the test about whether it's successful and those two skills learned by doing it I think gives you a solid platform so that then when you're like hey this is really expensive yes now let's talk about going down to a lower model but I don't think front loading is is helpful anymore plus the recipe is going to keep changing right yeah it's changing all the time because people generally outside of the AI bubble they only have so many so much hours in the day because they don't play with this all regularly so they don't have so much hours in the day and so how even if their employers give them time what do they spend their time on because if they're still in the oh I've just learned how to use deep research I'm going to use it cool although it's like you know you're like your agent can probably do that
for you and they do all these things for you so maybe we should just learn how to use personal agents and agents and maybe not that far but it's like but do you need to understand the basis of like how to call them or even the basis of like hey you need to develop some skills first because then your agents can use the skills to do this and and and so on like because skills used to be the advanced but I'm thinking skills needs to be like base now yeah you should learn how to do skills and that's a big piece not very many people actually have really graphed what like building skills would be building plugs and would be even though everything's being pushed in that direction or what an FCP is I think that AI technology advanced so quickly that for the general population it's going to be that old you know saw about any advanced technology any sufficiently
advanced technology will appear to people as magic and I don't think it's there's no foundation of knowledge in the progression that we've followed that you can easily impart to somebody who's just coming in a new and working with an AI agent it's just going to be magical that they and in many cases as I've demonstrated something as simple as chat GPT voice they're like yeah isn't that a person you mean that's a robot they just have a very very simplistic view of all of that and so it's going to be magical yeah so I I'd like to just hit a couple of meta things that I that I found interesting before we go um first is we didn't talk about it but Metta Zuckerberg who was at the event in Washington yesterday uh he poached the CEO of MongoDB fellow named CJ Desi just less than a year after he had become the CEO of that significant company in the world
of databases MongoDB is a pretty important company out there it's a non-relational database system that is used you know around the world and so he's putting that person in charge of uh met an enterprise platform business so this is signaling Zuckerberg's intention to go after that enterprise market that's so valuable to anthropic and open AI and the question is Ken Metta which is profoundly consumer in his general orientation actually succeed in that market well he's hired a person away from a company and they were bitter by the way the the employees at MongoDB are very bitter it's vocal about wait a second you know this is just completely mercenary you know obviously Metta put a very large compensation package making sure that CJ Desi will be a billionaire which he probably would would not have become if he stayed at MongoDB so he's buying
into the billionaire process now speaking of billions here's here's another look inside Metta that I think everybody should be aware of the New York Times reports that Metta classified its data centers which are just the compute centers there've been data centers around forever as experimental facilities to claim federal research tax credits worth four billion dollars of course thanks 2025 taxes so you're paying in our US budget as a taxpayer uh your taxes but Metta is claiming that their data centers are experimental and getting a four billion dollar pay from the tax system and that just you know there's going to be congressional scrutiny on that for sure but I don't know how they can claw that back this is a akin to you know just the shenanigans that
can happen at the billionaire class and in the corporate class that really really tip the scales in favor of their tax treatment compared to the tax treatment of regular working people yeah that's it the the story that shouldn't that surprise shouldn't surprise anybody it's like you said then I was like of course they are of course of course they are well same per same per Zuckerberg same for the offer that Zuckerberg put in front of the Mongo CJ Desi yeah right because I mean that's what Zuckerberg has to offer yeah but you know it's okay so it's worth the end of the show so it's maybe it would just I'll just pose this for tomorrow's discussion or something but I would love to just talk about you know like there's people going I mean I'm sure there's a lot of people that maybe he offers things to turn them down you know I'd say thanks but no thanks suck you know um but I would imagine CEO of Mongo I haven't I haven't
thought about MongoDB since probably 2013 I know they're out there but I just haven't thought about that term you know in quite some time it's like you know throw snowflake into the conversation and like all of a sudden in my my data analytics brain from years ago um so no no no slight towards MongoDB maybe they're doing great um but regardless I'd be curious maybe tomorrow we can dive into this or whatever and people are going they're leaving you know me here to CEO of a fairly large database company and you're still leaving to go do what exactly maybe it's all money maybe that's the that's the very short short answer to this but maybe there's something more maybe it's a it's a feather in the cap maybe it's a pride thing maybe it's uh I got I got picked by one of the big boys type thing I mean you know you never know where the ego is involved in this so yes of course the money but um you know at a certain point I think the money becomes less the thing if you're already stable and secure with the money you're making as the CEO of the no I think I think going to the
billions is well it's very compelling for a new one that's that's a generational wealth at that point right you know right now I think that's where people are at right now I think especially in that world is like I need to I want to retire or this is my maybe they get the sense of like a things are getting a little crazy in the world I'm going to grab up as much money as I can so then I don't have to then I can just sit back and relax and not worry about things later yeah yeah which will also be not accurate yeah correct yeah it is potentially true that grabbing as much money as you can does not ensure you have to remember that even once you're a billionaire that does not mean that that you're a happy person uh you know money you know assuages a lot of anxiety about you know personal security and so on and it will buy you big toys but I've worked for some
very miserable billionaires yeah well as Bruno Mars once said I want to be a billionaire so freaking bad I don't know if I want to be a billionaire to be honest with you but that's a that's a story for another day I don't I don't other than maybe the good I could do in the world I'm not quite sure I have any desire to that's what I'm billion I want that was one of Dolly Parton's biggest legacies right for sure yeah she could have been a billionaire and she chose actively not to so what did she create instead yeah right where you got a point to end on all right listen we'll wrap it up here for today so we don't go an hour and a half today um and uh thank you for the free train bed um but thanks to everybody uh thanks Garrett thanks Andy Vicks bad they Carl had a jump but thanks to Carl as well um we'll be back tomorrow for our Friday show um and yeah see you then thanks everybody for hanging out with us in the comments and online bye everybody
More episodes
More from The Daily AI Show

OpenAI Has Dots and Space To Share at Dev Day
The Daily AI Show

What Does AMD Want With Dr. Fei Fei Li and World Labs?
The Daily AI Show

AI Agents Continue to Escape Their Sandboxes
The Daily AI Show

The Personal Publicist Conundrum
The Daily AI Show