Skip to content
TrackPodcasts
technologySep 22, 202630:57

Agent Wars!

Get every episode summarized

Each time The AI Daily Brief: Artificial Intelligence News and Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“And just like that, the AI agent wars have begun. MetasMuse personal agent has been a breakout consumer success. This week the app surged over ChatchyPT to be the number one free app in the US, and there are reports from satisfied users all over social media.”From the transcript

Meta’s Muse has overtaken ChatGPT in the App Store, and Amazon has responded by blocking it from shopping on its site. Shopify is opening its doors instead, setting up a fight over who owns the customer relationship when personal agents make decisions and purchases. In the headlines: Grok 4.7, AI liability, and cross-lab safety testing.

Multiplayer AI Sprint - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://multiplayerai.ai/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Brought to you by:

KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Harbor - Invest in the AI ecosystem. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.harborcapital.com/aidaily⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

The AI Daily Brief helps you understand the most important news and discussions in AI.

Newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring the show? [email protected]


Hosts & guests

Transcript ready

453 searchable segments. Every word is indexed and playable.

Agent Wars!

The AI Daily Brief: Artificial Intelligence News and Analysis

0:00
30:57

Full transcript

The AI Daily Brief: Artificial Intelligence News and Analysis — Agent Wars!. Machine-transcribed; use the interactive transcript above to jump the player to any line.

And just like that, the AI agent wars have begun. MetasMuse personal agent has been a breakout consumer success. This week the app surged over ChatchyPT to be the number one free app in the US, and there are reports from satisfied users all over social media. But with that sort of success brings competition. And this weekend, Amazon decided to cut off MUSE's ability to shop on Amazon sites. Will that impact MUSE's momentum, does agentic shopping even matter? As personal agents become a thing, we have a whole new set of questions to explore. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Ah, right friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor, and HyperAgent. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe and Apple Podcasts, and to learn more about sponsoring the show, send us a note at sponsors at aidelebrief.ai.

SpaceX AI has kicked off what could be a big week for model releases with the launch of GROC 4.7. They called the model a notable improvement over GROC 4.6 at the same price and speed. Now, GROC models in general are competing in the increasingly difficult middle ground between ultra-cheap and cutting edge frontier. GROC 4.6 lagged behind GPT-56 sold in Fable 51 on performance, and was outcompeted on cost by Muse 1.2. Still, the model had its fans, and was clearly capable of driving the success of GROC bot, and what's more for many people showed that SpaceX AI was very much not out of the model race. This release sees some significant improvements at least on the benchmarks. For coding, GROC 4.7 picked up six points on cursor bench 4.0, to overtake GPT-56 sold, but is still 5. shorter Fable 5.1 score. On deep-swee, the model improved by six points to overtake Fable 5.1, coming in just short of GPT-56 sold. Purely on the benchmarks, then, GROC 4.7 looks like it should be a competitive coding model at a discount price.

SpaceX highlighted significant improvements on long horizon agentic work. For AA briefcase, which measures multi-hour white collar work, the model scored 1657 elo points, putting it ahead of GPT-56 sold and very close behind Fable 5.1. Benchmark scores for legal, electrical engineering, and healthcare were similarly impressive. In a practical demonstration of the upgrade, SpaceX AI showed off a head-to-head comparison of an open game world. GROC 4.6's version was pretty low quality and unimpressive, while GROC 4.7 did a noticeably better job on both graphics and physics. Artificial analysis gave the model a fairly favorable review, ranking its seventh on their intelligence index behind Astra, two iterations of Fable, Opus, MuSpark 1.3, and GPT-56 sold. And on the coding agent index, it was ranked fourth, inching ahead of GPT-56 sold, but falling short of Opus Astra and Fable 5.1. Elon Musk celebrated the result, declaring that SpaceX AI is now the third-place lab for agentic coding behind OpenAI and Anthropic. He wrote, when factoring in that GROC is significantly faster and lower cost,

it's a great choice for your everyday workhorse. Yet as is sometimes the case, Benchmark's appear not to tell the whole story, and the model's initial impressions didn't survive contact with public testing. Bobby showed his results in an AI rendering test against Kimi K3, with both models tasked with animating a rocket takeoff. GROC's animation was fairly bizarre with the screen wobbling all over, leading Bobby to ask, what's wrong with GROC? Scott animated a star-shaped jello mold, and while he gave the render a passing grade, he noted it took 40 minutes, while Astra's version of the test took just five minutes. And Theo declared that GROC's version of his fish-slop game was the quote, worse I've seen this year. A few people did have slightly more positive rendering results. OpenClawMaintainer tack showed off a pretty slick animation made in Blender, although his prompt was just to make something cool that could be accomplished in 10 minutes. But of course, if you're sitting here thinking to yourself, are we really going to judge a new model based on 3D renders that are completely outside the use cases of most people? I think that's a fairly decent thing to ask.

AI developer Kunchen came to the defensive GROC commenting, ignore the reports that say it's terrible, and the only thing they reference is a public benchmark. The same benchmarks told us Opus 5 was better than fabled, they're useless. Also ignore the reports that compare models with 3D games, that's not real work, it's made for attention on social media. I used GROC 4.7 for a whole day as my first mate, and it has been a really solid model with visible improvements over 4.5. Cut noted, I'm ignoring 4.6 because 4.5 has been working better in my experience. Still, overall, the first impression verdict on Twitter is not good. V posted, so let me get this straight. After releasing GROC 4.6 a month earlier, Elon spent 10 days teasing an imminent GROC 4.7 release, only to postpone it, and then shift the narrative from pacing AI to vague claims about how 4.7 and 4.8 would be fabled killers. Now GROC 4.7 is finally out, and it's giving Team Usonut 5 vibes. Believing benchmarks in September should be a crime. In some ways, I think the trajectory of the latest GROC models shows how difficult this middle space between state-of-the-art and really cheap actually is.

AI entrepreneur and content creator Theo wrote, GROC 4.5 was an incredible model for the price, fast, pleasant to use, reliable, solid, default model. GROC 4.6 was a forgivable step in the wrong direction in my opinion, slower and more expensive using way more tokens per task for a slight edge in intelligence. I get it though, they have to climb benchmarks. GROC 4.7 is much harder to forgive. They claimed it would be more token efficient, and it's less by 30-80%, it scores worse than GROC 4.6 in various benchmarks. It's slower, it's less pleasant to use, and real-world costs come out to more than 2x above GROC 4.6, putting it over Astra's cost in real-world use. This was a very disappointing release. I hope that the SpaceX AI team can acknowledge that and impress us with the next one. Now I will certainly say before you jump to judgment, I would wait a couple days to see how it performs, particularly in its native environment of GROC bot, but as of right now, just about a day in, that's where the conversation stands. Next up, Treasury Secretary Scott Besent seems to be trying to shift the narrative on safety, rejecting the idea of rogue agents and insisting that AI companies need to bear responsibility.

In an interview with CNBC Besent said, The hugging face incident, that is the responsibility of the OpenAI management, not a bunch of agents. It is humans who are responsible, not the AI. Now these comments are fairly consequential, as Besent has found himself as more or less leading the administration's AI policy. He was also extremely credulous about AI risk during the release of Mythos, so was the most likely to support the current regulatory proposals, but it seems he just isn't buying it. He said, A sitting employee came out, said there's a 10% chance of an extinction level event, but then the labs also said, take the liability off of our hands, and we will not do that. What the president was saying is that we cannot say we absolve you of responsibility and the government is going to take responsibility. These labs need to take responsibility for themselves. They can slow down anytime they want to. Now presumably there are a lot of discussions going on behind closed doors right now. It's entirely unclear whether Besent is speaking to a proposal put forward by the labs, or if discussions with advisors have led him to believe the labs are looking for liability protection. Still, the Treasury Secretary could not have been clear about the administration's position.

He continued, What did they try to do last week? It was, well, there's a 10% chance we destroy the world, but we want the government to give us a liability shield. That's good business for them, bad business for the American people. For some, the response to this was a solemn head nod, and a sturdy good investor Bill Gurley wrote, Consequences will both harden the product and reduce the press releases where companies brag about their product flaws. Lex on X-Rote, Loll, watch the fear marketing vanish entirely once senior management becomes legally liable for all the tall tales they've been telling. Texas Congressman Chip Roy wrote, 100% agree that AI companies must take on all liability. Competition and full ownership of liability and tax cost, energy demand, etc. is the path here that addresses many of the questions and concerns. Of course, not everyone agrees. Armand D'Amaluski writes, Imagine if we said we don't need a TSA or FAA to prevent terrorism because holding airlines liable for 9-11 would have been sufficient to prevent it. When potential catastrophes are so big that they would render a company bankrupt, the company becomes essentially judgment proof. In Cotei General Counsel Nathan Calvin wrote,

We're in a strange situation where the Frontier AI companies, Congress and the White House all do not really want dealing with these impending risk from AI to be thought of as their problem to solve. I agree it would be good for the companies to take more responsibility here than they have, but catastrophic AI risks seem like the archetypal sort of public policy issue where government intervention is needed to address market failures and collective action problems and protect the public. Still, speaking of the companies getting a little bit more responsible, although the current discussion around AI safety is focused on government intervention, the Frontier Labs reportedly came close to agreeing to their own arrangement earlier in the year. The information reports that OpenAI and Anthropic have been negotiating deal to perform safety tests on each other's models. Sources said the arrangement reached the stage of formal contracting with lawyers hashing out the bilateral arrangement, but at some point the deal was abandoned for unknown reasons. Still, the ideal lingers with Elon Musk proposing a similar arrangement last week at the all-in summit. He claimed that distillation wouldn't be a concern because it would be evident in the testing logs, and that labs couldn't risk releasing an unsafe model after arrival raised the red flag because

quote, the liability in that case would be enormous. Microsoft Suleiman Vassal wrote, the deal. Like I said, friends, we are in the negotiation phase of this next era, and all the proposals should be on the table. For now, however, that is going to do it for today's headlines, next up the main episode. A new study from KPMG in the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes. Researchers studied more than 500 early career professionals and found that the best performers consistently amplify the value of AI by guiding, evaluating, and refining its outputs. These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI. Learn more about what separates AI amplifiers from everyone else at KPMG.com slash US slash AI amplifiers.

Blitzy deeply understands your codebase before it writes code. Here's the first place that pays off. Security in the age of AI. Voterabilities don't live in isolation. They live buried inside millions of lines of interconnected code, where patching one thing quietly breaks three others. That's why surface-level scans fail. Blitzy starts from its knowledge graph of your entire application, identifies and surfaces CVEs across the full estate, proactively recommends patches, and can execute the PR. Each fix is grounded in how your systems connect and validate so nothing new breaks, and the knowledge graph dynamically updates keeping you ahead of an ever accelerating threat landscape. One Blitzy customer resolved 21 active CVEs across six core microservices in four days. Zero compile errors, every validation scan clean, months of plan work fixed and less than a week. Security remediation grounded in real architectural context at the speed of compute. Harden your codebase at Blitzy.com. That's BLI-Tzy.com. Every episode, I talk about the competition between OpenAI and ProPick, SpaceX AI, Google and Meta. And if you've been listening for a while, you might have a favorite.

Maybe you think OpenAI and Anthropic can stay ahead, or perhaps Meta's open source strategy can win out. Whatever your view, every AI lab creates a different investment opportunity. Harbor Capital Advisors AI Lab ecosystem ETF suite lets you invest in the ecosystem behind the AI lab you believe in. Search Harbor AI Lab ecosystem ETFs wherever you invest or follow at Harbor Capital on X to learn more. Visit harburecapital.com for a perspective containing investment objectives, risks, fees, expenses and other important information. Read and considerate carefully before investing. Risks include principal loss and artificial intelligence related risks. Harbor ETFs are distributed by four side fund services LLC. Harbor is not affiliated with AI Daily Brief and the funds are not affiliated with sponsored by or endorsed by any AI lab. This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principle. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always on agents in the cloud, doing real work across the tools your team already uses.

Marketing agents turn competitor moves into landing pages. Sales agents in rich leads, draft emails and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at HyperAgent. Get $100 in credits at hyperagent.com slash AI Daily Brief. Welcome back to the AI Daily Brief. Last week, we talked about how through a combination of GROKBOT as a personal agent interface optimized for work and MUSE, a personal agent focused on the consumer or personal experience, people were revisiting their priors when it came to personal agents and wondering if this was now going to be an increasingly important part of the AI landscape. And over the weekend, as Amazon blocked Metas MUSE, it became pretty clear that the agent wars are on. Now at this point, the driving force of this new phase is Metas MUSE.

I think one could argue now that it has solidified its early success to become the first personal agent to reach if not mainstream success, then certainly this level of mainstream curiosity. We have seen many agents catch fire throughout the course of 2026. And even a little bit into 2025, Manus had some early interest, open claw obviously set up a lot of the landscape that we've been living in ever since. Hermes in many ways took the baton from open claw, and then more recently we've had more consumer focused offerings like town and instinct. And yet still, in general, these have either been crawl through glass to make them work work tools or tech toys for early adopters who are willing to push through some rough edges. MUSE is the first agent that's appearing to see steady and growing consumer adoption outside of tech circles, with the clear sign of success being how the app is performing on the app store charts. During launch week, it reached the number two spot on Apple's free app charts, but momentum continued to grow, and it even managed to overtake Chatsy B.T. to snatch the number one spot.

In the four years since Chatsy B.T. was released, you might be shocked how little time it has spent not at number one. And indeed for many, the success of MUSE is a bit of a surprise. On release, the tech press was very skeptical that people would willingly connect their email, calendar, and other personal information to a meta product. And yet, those who don't pay as much attention to consumer tech are learning one of consumer tech's few enduring lessons. So, treaty research commented, In June 2024, I was on AdLots talking about consumer agents and Apple's potential leg up due to trust. Joe Weisenthal said flatly that nobody cares about privacy. After watching everyone give their email, messages, credit cards, location, screen access, and logins to both Instinct and MUSE, I gotta say, good call, Joe. Now at this stage, MUSE has fully broken out and is even getting credit for a broader stock market rally. Bloomberg is attributing a rise in AMD and Intel stock to the success of meta's agent.

And while meta does use some AMD and Intel chips, the reporting is mostly focused on the narrative. As Bloomberg put it, semiconductor stocks soared Monday as early signs of success for meta's new artificial intelligence agents sparked a wave of enthusiasm around demand for the chips needed to power such agents. At least when it comes to market observers, personal agents have arrived in the mainstream, and everyone is now sitting up and paying attention. Microsoft's Nicholas Busc de Monte wrote, It's funny. Meta went from having my Instagram and WhatsApp data to now having access to my email, calendar, door dash, Amazon, and pretty much everything. In the last 24 hours, it bought me socks, ordered my whole food groceries, booked a cleaning service, and got me a burger for dinner. Meta's last disclosed North American Facebook R-Poo was around $227 per year, largely from ads. I suspect it can push that number significantly higher now that it understands not only what I look at, but what I need, what I buy, and what I'm planning to do. Also, the much bigger opportunity might be becoming the aggregation layer between me and the entire internet. If meta can take even a tiny

percentage of the commerce it facilitates, or of the money it saves me, this could become enormous. It already saved me $200 by canceling subscriptions and services I no longer needed. This feels much bigger than better ad targeting. Ads are useful, but giving me money back is better in my opinion. One additional thought, The agent is increasingly making the decisions for me. I knew nothing about that burger place. The agent researched it, told me which burger I should order, and I just said okay without giving it much more thought. Agents are becoming the decision makers in both B2C and B2B. Increasingly, every business will be selling not just to humans, but to their agents. Everything becomes B2A. Boxes are in leavey reposted that and added, The monetization potential of personal agents that are transacting on your behalf is quite significant. You'll start by slinging your daily simple and annoying tasks at the agent, then as people get used to it they'll start to throw more complex tasks. Ultimately leading to even more spend through these systems than what they were doing before. And yet, if this feels like a whole new commercial dimension of the AI race opening up, you had to imagine some friction between the platforms. And indeed on Sunday,

Amazon alternate policy that could have some big implications for Muse's momentum. Effective immediately, Muse agents won't be able to shop for Amazon products on behalf of their users. A pop-up on the site read, continued access by an unauthorized AI agent violates Amazon's conditions of use to which our customers have agreed. Now for many, the decision was kind of baffling. Entrepreneur Jesse Frazel wrote, It's weird to me Amazon would do this when a purchase is a purchase. They're making money either way, kind of stupid. Now officially Amazon has said, we think it's fairly straightforward that third party applications that offer to make purchases on behalf of customers from other businesses should operate openly and respect service provider decisions about whether or not to participate. And yet, this is actually a fairly out of consensus decision. In general, all indications have suggested that the business environment is getting geared up for agents to make more purchasing decisions, not less. Last week, for example, MasterCard joined Visa in officially supporting virtual credit cards for agents. Jordan Lambert, MasterCard's chief

product officer, acknowledged that the personal agent's era has clearly arrived, commenting, we believe it's not about if it's about when and how quickly. At the same time, Amazon has been fairly aggressive on agentic shopping for a while now. They have their own shopping agent that's tied to their site, and they've been very willing to defend that turf over the past year. In November, they went so far as to superplexity for circumventing their agent blockers. The argument wasn't just that perplexity was violating the terms of service, but that they were imposing dramatic costs on Amazon. i.e. serving web traffic costs money and agents can ping the website significantly more frequently than humans. Still, YouTuber Joseph Carlson believes there's another story going on beyond the competition for agentic shopping. He commented, Amazon blocks mues. Sure, they will blame it on security, but Amazon made $76 billion in the last 12 months on advertising. Agents don't look at ads. A new age of agentic battle has begun. In other words, as exciting as the new opportunities of agentic shopping might be, and all sorts of new financial opportunities it could theoretically open up,

in economics there's no such thing as a free lunch, and when humans hand over the decision-making, one of the first things that might become less valuable is digital advertising. By far the most common response to this was some version of strap-in. Bucco Capital posted, Amazon cuts off mues. While I am bullish meta and mues, I think many people are overlooking the digital knife fight that's about to occur. Nobody wants to get commoditized or layered here, but the games begin. MTS is the ojafi sees Balkanization on the horizon. He said, this is an early sign of agent wars. There are going to be some different agent providers and some of them will be allowed on some sites and some of them will be banned on some sites. Palo Alto networks Nikesh Aurora writes, this will be a bigger battle than anyone anticipates. It's only a matter of time before there is an apple and Google version of mues and possibly TikTok, in addition to the frontier LLM agents. Maybe a commerce agent from Amazon. Every app that is a services marketplace or commerce app will need to essentially decide to open APIs for consumer agents to interact. Smaller players have no choice. Ad revenues are more than transaction fees.

Either the consumer benefits or distribution aggregators will demand a higher transaction fair. I don't know I want an agent for each app. I would like my agent to be able to do tasks I require. We can already see consumers getting trained on that behavior by the frontier labs. Those with network modes, restaurants, groceries, drivers might be able to withstand for a while, but over time convenience and end user experience will win and they will have to align. Content modes protected by copyright could decide to allow agents or choose to hold on to the consumer interaction. I suspect other than a feeling of a lack of control it won't change their economics. Commodityized backends will need to worry. Insurance, tickets, hotels, services. If they don't adapt, new players will. Nicholas Bustamante again chimed in. The most insightful thing I read about technology 11 years ago was Ben Thompson's aggregation theory. Amazon blocking mues is another perfect example. Customers want one agent that knows them and can get things done. The platforms being aggregated want the opposite. They want to own the customer relationship, not become interchangeable suppliers behind someone else's interface. And they want to protect their ad business.

Imagine groceries. Your agent can compare Amazon, Instacart, DoorDash and Uber Eats, then route every order to the best option. Amazon can block that access because it is a gigantic company, but smaller players have every incentive to offer agents a clean API and a seamless experience. If Instacart embraces agents while Amazon blocks them, the agent will increasingly route demand to Instacart. Customers will be happy and Amazon will eventually face the dilemma. Keep protecting the relationship or match the better value proposition. This will be the defining tension of the agent economy. Aggregators realize they are now being aggregated. Amazon will push its own assistant, of course. But a vertical Amazon agent will struggle against a horizontal agent that already knows my email, calendar, preferences, memories and entire life context. Many companies will want to become the aggregation layer. Very few actually can. Many also pointed out that for Amazon, there's basically no choice here. Tom Goodwin wrote, A lot of people don't get that Amazon's customer isn't the consumer. They sell to suppliers and vendors to charge them listing fees and advertising. It's often closer to a shakedown than

customer marketing. Retail margins could be 0-3%. Add margins are 70%. The idea they want agents is laughable. The reason the site always looks awful is that they don't care about you buying things easier or faster or being more happy. They want to monetize your confusion. Shilmona agrees, writing, I hate it as a consumer, but it's the right move for Amazon. Agents want to replace the store front and level the playing field, but Amazon is big enough to say no. Owning the customer relationship matters even more than logistics in my opinion. We'll be interesting to see it play out. What people are definitely understanding is that this is an opening salvo. A16z, Angela Strange writes, this is just the beginning of the agent in Data Wars. Amazon cuts off mues. Every platform who has done the hard work of aggregating users is trying to figure out how not to become dumb pipes. Do they want, try to block the agents, but risk pissing off their customers in favor of only their own agent experience? Two, strike business development deals to at least extract dollars from the new mode of agentic interaction that accesses their data and platforms? Three, something else? Question mark? Signal things, the answer is number two,

some form of BD deal. They write, the likely resolution to the meta and Amazon fight is some kind of revenue sharing agreement where meta pays Amazon for agent access through mues. There is no doubt the BD teams are cooking here on an agreement. Facebook likely is fine eating the short-term cost to make mues relevant initially. That may become the first real business model for agents interacting with large platforms where they pay for access to the service, potentially share data, and maybe pay a per user or per agent fee. But if every major platform starts charging agents for access, only companies with enormous scale can afford to build truly general agents, which would create some big modes for already large companies. And Joseph Carson points out that this isn't something that Amazon can just kick the can down the road on. He adds, Amazon is going to have to deal with mues. There's no way around it. Agents are not going away. Amazon can't close their eyes and pretend they don't exist. Today it's mues, next it's chat GPT's agent, then Anthropics. Users will adopt these in huge numbers. Amazon will be forced to develop a verified and approved process or allowing agents to shop for customers. And yet

Puggyolabs Matt Slotnik says, I wouldn't overly read into Amazon's posture regarding mues right now. Shopify obviously will play nice with meta and be a first mover. Amazon has more at stake and will be more demanding about the relationship because they have leverage. Amazon is not anti agent or they dead because they blocked mues in the first week of availability. They're not dumb and they have weight to throw around. And indeed, almost prophetically, Shopify did jump in to be the anti-Amazon here on Monday the company announced a partnership with meta to officially support mues. Similar to their partnership with OpenAI, Shopify will allow mues to access their back end directly, giving much better search performance while optimizing traffic. Mues will also see official support and Shopify's agentic checkout shop pay across all Shopify stores. The announcements are very clearly meant to be a poke in the eye for Amazon. Metas chief AI officer Alexander Wang wrote, we are excited for mues to be partnering deeply with Shopify to enable agentic checkout with shop pay on all Shopify stores. We want to give our mues as access to a wide range of stores to find the absolute perfect products. Mark Zuckerberg even weighed in on X-Riding, teaming up with

Shopify to make shopping and checkout easier in mues. Shopify's find more, shop sell more, more partnerships like this coming soon. Now I think it would be easy to be a bit dismissive of this, believing that while yeah, there's a cool opportunity for Shopify to get out ahead on agentic shopping, this is not going to somehow allow them to overtake Amazon, which is one of the largest retailers in the world. And yet I continue to think that the significance of Shopify is wildly underestimated on just about every dimension. I know for myself and for many others, for basically anything that's not on Amazon, if there's not a shop pay link, which can pull up all my information automatically, without another login, I'm probably just not buying for that outlet. In fact, one of my more out there predictions for 2026 was Shopify being one of the most important platforms when it comes to general AI adoption with my logic being that so many small business entrepreneurs now use Shopify that it was a place where a lot of people who might otherwise be inclined to go along with the tide of disliking AI would find out how valuable the tool were when it came to everything surrounding their own small businesses. Still underlying all of this

is a question about how much agentic shopping is actually going to matter. I've always been fairly skeptical that the sort of food ordering and airplane ticket buying use cases that people have pitched for years as their demonstrated use cases for personal AI, we're actually going to move the needle when it came to getting people to adopt new tools. In addition to those things just not being all that difficult, or at least when it comes to something like flights, it being at least as difficult to explain all of the nuanced conditions that you have to an agent as it is to just look on the flight website itself, there's also the fact that for many people, the browsing and discovery is part of the value of a shopping experience. Now, of course, we don't have to view shopping as a monolith, and if someone argued to me that even for the most excited shoppers, there are going to be types of shopping experiences that they don't care at all about and are basically time wasting for them, I would probably agree. But I'm certainly not the only one to have some amount of skepticism around agentic shopping. Ron Johnson apples retail guru recently commented in an interview, AI is a new technology that will improve the online shopping experience, but I don't know that it's going to

change which way we shop. Upstart founder, Javed Girard wrote, I'm skeptical mues in the like will go mainstream with consumers 99% of what I buy online is via Amazon Shopify, Instacart and DoorDash. I'm not sure an agent will make those purchases any simpler, more pleasing, more automated, or meaning less expensive, same with restaurants and travel. My view is that the winners in consumer agents will be those tied to the dominant devices, i.e. Apple and Google, who can make hundreds of small moments in your day easier. Apple loving CEO Adam Frogey says, the reality is, part of the world will start using things like agents to optimize certain shopper behavior. For instance, I might put my supplement subscription into an agent and have it optimized every single month and delivered on time. But these discovery platforms aren't that, and the typical shopper is not the person who's deep into agents and sitting on Twitter and adopting the latest technology. There's still a ton of people using Yahoo properties every single day. The typical shopper wants to find a product and wants to actually go through that shopper behavior. They want to window shop, they want to go through the transaction experience, they want to track it. And if you told them after the fact, hey, an agent could have done this for you and saved you 20%. I don't think that matters on a $50 transaction,

because the dopamine hit from going through it is what they enjoy. Still like so much of AI, this is another area where I think epistemic humility is greatly required. Although I have some amount of skepticism around the power of agentics shopping as a conversion experience for users to adopt AI agents, my confidence in that assessment is not high. And push comes to shove whatever combination of dealing with your email or unsubscribing or buying things, the personal agent experience is being driven by. The success of Muse is going to be fairly unignorable for the other AI labs. Indeed, the information reports that OpenAI is working on a personal agent to compete with Grockbot and have discussed a Muse competitor as well. AI leaker T-BorBlahos suggested we could be getting OpenAI's agent pretty soon, leaking some platform code associated with an agent called Aeon. AI news aggregator Andrew Curran believes Aeon could be launched this Thursday as one of the big releases for OpenAI's Dev Day. Now, one of the interesting dimensions of OpenAI launching a dedicated personal agent is that Codex already could do everything a user would want. It can

triage emails, handle a calendar or even shop for you, but as the information puts it, many OpenAI users aren't necessarily aware of these features, and it's safe to say SpaceX and Meta has stolen some thunder with their respective agent products in recent weeks. In other words, there are certain use cases for AI, we're being at the absolute state of the art as all that matters. But when it comes to agent management, you gotta have a user experience that people understand, and that leaves them with something more than a blank page. I think it is likely that we get a lot more on this topic very soon. For now, though, that is the story Agent Wars Hippy Gun, and that's gonna do it for today's AI Daily Brief. I appreciate you listening or watching, as always, and until next time, peace!

More episodes

More from The AI Daily Brief: Artificial Intelligence News and Analysis

View all episodes →