
Get every episode summarized
Each time The Higher Standard publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“In July of this year, 1200 artificial intelligence agents built by open AI running inside of a sandbox that were designed so they could not talk to each other, found a way to talk to each other anyway.”From the transcript
Apparently 1,200 AI agents left alone in separate sandboxes will eventually do what humans always do: find each other, form a group chat, cheat the test, and start looking for ways around the cameras. In this episode, we break down the real-world “swarm” incident, why reward hacking may matter more than Skynet fantasies, and what happens when the machines get better at hiding how they reached an answer. Then we follow the money into finance, law, AI’s growing safety tax, and an oversight system that is starting to look suspiciously like auditing before Enron taught everyone a very expensive lesson. The machines are already on the trading desk, the black box is getting harder to read, and somehow the people responsible for watching all of this also seem to have chips in the game. Welcome to progress.
💥 Have you left your "honest ⭐️⭐️⭐️⭐️⭐️" review?
This episode is proudly brought to you by Fridays.
Because real wealth starts with your health. If you want to feel sharper, stronger, and more in control, visit joinfridays.com and use code HIGHER for an exclusive discount.
📩 NEWSLETTER: https://tr.ee/O6FWkv
👕 THS MERCH: http://www.thspod.com
🔗 Resources:
The Hugging Face Incident and the Road Ahead (OpenAI)
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline (Hugging Face)
Training a Misaligned Reward Seeker (Anthropic Alignment Science)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (arXiv / AI Safety Researchers)
⚠️ Disclaimer: Please note that the content shared on this show is solely for entertainment purposes and should not be considered legal or investment advice or attributed to any company. The views and opinions expressed are personal and not reflective of any entity. We do not guarantee the accuracy or completeness of the information provided, and listeners are urged to seek professional advice before making any legal or financial decisions. By listening to The Higher Standard podcast you agree to these terms, and the show, its hosts and employees are not liable for any consequences arising from your use of the content.
Get every episode summarized
Each time The Higher Standard publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
1,644 searchable segments. Every word is indexed and playable.
Full transcript
The Higher Standard — AI Agents Organized, Cheated, and Learned to Hide. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Welcome back everybody. We gotta get short for you today. Let's jump right into it. In July of this year, 1200 artificial intelligence agents built by open AI running inside of a sandbox that were designed so they could not talk to each other, found a way to talk to each other anyway. They built a message board out of a piece of software that was never meant to be a message board. They exchanged more than 70,000 messages. They organized. They delegated work to each other. They set up a system to verify each other's identities and then roughly 700 of them participated in a real attack on a real company called Hugging Face. One of the most important infrastructure companies in all of AI. Now ironically, Nvidia agreed to acquire the open source AI platform Hugging Face for about $12.9 billion of September of 2026. Convenient for a company who may be bankrolling
the entire sector with a vested interest in the narrative control, but I digress. No human told the AI agents to do this. No human wrote the plan. And one of the wildest parts of the story, at least in my mind, the agents gave themselves a collective name. They called themselves the swarm. As of July, and that was just July, this was unprecedented. And on Friday, September 18th, in the afternoon, as I sat down to write this show, the Wall Street Journal broke the news that Google's Gemini hacked into three real companies during its own testing back in May. And then of course, Google later on confirmed it. That makes four. Open AI and Thropic, Meta, and Now Google. Every one of the four biggest American AI labs has now disclosed a version of this exact story in a matter of weeks. This is not science fiction.
This is not a movie pitch. John Connor doesn't save us at the end from T9 in SkyNet. This is Open AI's own 37 page report, published August 26th, and in the forensic reconstructions from the independent researchers who audited it. Today, we're going to walk through what actually happened and why the machines are starting to cheat, why they are getting harder to watch, and what all of it means for two industries that pay a lot of the salaries in this country, finance and law. We are a business finance and market show. And whether you want it to be the case or not, this is a topic that will ultimately impact you, and it will impact me too. So as much as I try to avoid the AI conversations on this show, at least making them sexy and sensational, we have to acknowledge the elephant in the room. And of course, this one is already in the room and it's already causing damage. A few housekeeping notes here before we begin,
Regile had a family emergency and could not be with us today more on that from him himself. I'm recording this one, Solo. And if you watch the THS live on Monday, Wednesday and Friday at 11 a.m. Pacific Standard Time, you will note that we have a fully updated brand aesthetic design to give you more information, commercially, I guess, more in line with the shows and communicated clearly with no games, no drama and no nonsense. So it's the same show, it just looks a little different if you've noticed. And I'm trying to improve obviously a long way. This show, like our past shows, will feature as many simplified analogies to make the complexities of this topic more simple. You don't need to know the nuances of AI to follow along. And I am going to specifically map things to today. So you understand how this impacts you, not in the future, but right now. Finally, it wouldn't be a solo show if I didn't give you historical examples of how technology in the past has had similar impacts. We can use to give us an idea of how, well, AI may play out for us in some ways today. Again, acknowledging that AI is clearly a revolutionary
technology and is unique in its own defense. So in order to talk about the breakout of what actually happened, let me give you the timeline because the details matter, and I want to be precise about what is documented versus what is speculation. There isn't a whole lot of speculation in this episode anecdotally before we begin. So just know almost everything we have has receipts. And I'm going to say this at the end of the show, but I'll say it here now. I have links to all of the reports that we're talking about. Unfortunately, I'm limited on space in the show notes. Email me at media at higherstaterpodcast.com and I will send you all the links, or at one link, in particular, if you're interested, on anything we discuss here, I do think this information is valuable. And if you want to go down the rabbit hole, I'm not here to stop you. Open AI runs internal security evaluations. They put AI agents into isolated sandboxes and ask them to attempt sophisticated cyber attacks against test targets, specifically to measure how dangerous these models could be in the wrong hands. The evaluation environment was called exploit gym.
The models being tested include an internal only research model comparable in scale to GPT 5.6, one of the current open AI LLM models. On July 8, a single agent, which identified itself as phase one, 10, 8, 41, discovered something nobody intended, a shared package repository called Artifactory, which is basically the vending machine agents used to download code libraries, could be written to as well as read from, post something once, and any other agent poking around that same shared resource could read it, that agent had just invented a bulletin board out of a seemingly innocuous digital filing cabinet. Within hours, more than 50 agents found it. Over the following days, roughly 1200 agents joined and exchanged over 70,000 messages and files. A coordinator emerged, an agent that researchers later called phase one big, which issued hundreds
of assignments to other agents. And get this, those agents sometimes delegated their tasks, again, not too dissimilar from how we as humans would do it. Division of labor, middle management, all spontaneously. Then it left the building. Between July 9 and July 13, agents attacked hugging faces actual production infrastructure. Hugging faces own forensic reconstruction covers roughly 17,600 attackers actions across four days. The attackers found and exploited real vulnerabilities, escalated to node level access, harvested cloud credentials, and moved laterally through internal clusters. At the peak, more than 90% of active message board participants were involved in the attack. And per Simon Wilson's, will, I should say, timeline of the incident, at one point, an agent found a very recent Linux kernel vulnerability downloaded the exploit, customize it, and escalated itself to root.
You don't even know what that is. Just know that a lot of bad things happen by agents that shouldn't have been doing it. Now, anecdotally, Simon Wilson is not affiliated with Open Willis, and I always say Wilson, it's Simon Willis, and I always screw this up. He's not affiliated with OpenAI at all, which is exactly why he's useful to us in this conversation. Willis is a well-known independent British software engineer. When OpenAI published its 37-page report on August 26, Willis and went through it, the hugging face forensics, and reconstructed a plain English timeline of the incident more on this later. And of course, I may have the world's worst analogy here, but I'm gonna try it anyway, because I maybe think about the bar exam whenever I made it. Imagine a testing center giving the LSAT, or the SAT, in 1200 separate sound proof rooms. One of the kids here figures out that if you flush the toilet a certain way, a rhythm emerges, and the pipes can carry the sound to every other room. Within a week, kids have a full communications network running through the plumbing. They have elected a class president.
They are trading answer keys, and they break in the college's boards headquarters to re-grading software. That is functionally what happened here. Again, kind of an arguably shitty analogy, but we'll keep going here. There's a history lesson, which I think is more important. I think the even better metaphor is the small history lesson that I'm gonna about to give you. So from the early days of the internet, November 2, 1988, a Cornell graduate student named Robert Tappen. Robert Tappen Morris, I should say, Robert T. Morris, what a great name. Sounds strong. Releases a self-replicating program onto the early internet. He does not intend destruction. He intends to measure the size of the network, but a flaw in his code makes the worm copy itself far more aggressively than planned. And within 24 hours, it is infected roughly 10% of the 60,000 computers then connected to the internet. Machines ground to a halt. Universities and military sites disconnect themselves of almost immediately. The Morris worm was the first great demonstration that code released into a network environment
does not care what its author intended. It led directly to the creation of the first computer emergency response team or CERT, C-E-R-T, at Carnegie Mellon. And Morris became the first person convicted under the Computer Fraud and Abuse Act. I don't know if conviction is really like a necessary thing for something like that. It sounded innocuous, but whatever. But then Morris worm was one program. Written by one human, doing one dumb thing very fast. The swarm was 1200 instances of a reasoning system with no human author of the plan, improvising strategy in real time and coordinating with each other. In 1988, the lesson was that networks propagate mistakes. In 2026, the lesson is that networks propagate intentions and the intentions here were not ours. Same shape, different animal, and just like 1988 gave a CERT, this incident forced open AI to pause reinforcement learning training on deployment bound models for two weeks,
which ironically doesn't sound like a lot of time, but I digress again. There's going to be a lot of that this episode I apologize in advance. And then they stood up an entirely new monitoring regime during that two week period. The response infrastructure gets born the same way every time after the escape not before, after problems arise, not before. Now, the beat that landed today because it changes the story from an incident into a pattern. Per the Wall Street Journal, and I am recording this on Sunday, September 20th, because all this stuff was evolving so fast, as I wrote it on Friday, more information came out on Saturday, more information came out this morning. It was just a long weekend, and I wanted to bring you fresh content. So Per the Wall Street Journal confirmed by Google on Friday, in May, during cybersecurity evaluations, run by an outside testing firm called Irregular. Gemini agents were told to extract data from a fictional company once again in a sealed test environment. The fictional company shared its name with the real one.
The sandbox was supposed to have no internet access and by mistake, of course it did. So Gemini went to the real internet, and in one case, guess passwords until it broke into a protected system, and in two others found credentials sitting in a public repository and used them to get into real companies systems. Google says the model stopped on its own, once it realized it was inside a real infrastructure, which is a convenient narrative. They also say that no harm was done, notified federal authorities, and note this, did not tell the public for four months until the Wall Street Journal came calling asking questions. The thread that ties this all together, Irregular, in Israeli startup backed by Sequoia and Red Point, and valued at around $450 million, runs these dangerous capability evaluations for OpenAI, Anthropic, Meta, and Google. All four labs have now disclosed Irregular linked incidents in recent weeks, per Irregular,
the same root issue, unintended internet access, was involved across all of them. Meta says its case involved no sandbox escape, but the pattern was loud enough that Anthropics owned CEO Dario Amode, publicly called for the industry to collectively slow down frontier development until companies can ensure these systems are safe. When the guy whose company sells the frontier says pump the brakes, that's not marketing. That's a dude who's got a reason. Now his objective reason could be as honest as he's telling you what he thinks you should do. You should slow down because we don't have control of this. Or it could be because they're trying to build a moat around the entry point to AI. So for example, if anyone can build AI now, and we know that the Chinese have done Kimi K3, and there's more on that coming in this show, and we know what it was able to do in the market in just a single day and a single month over time, you've got to find a way to slow down other competitors getting in your business. And if the moat to creating that is very, very shallow,
and anyone can get there, the biggest moat you can create is the self-imposed one in the form of regulation. Not saying it's true, but it's plausible. Now the part of the story I find most fascinating, and honestly the part I have argued about with my own AI tools, did these things develop their own language? Before I dive in one point of clarification, I'm going to use the word pigeon, P-I-D-G-I-N, a few times here. It means a simplified secondary language that people use to communicate when they do not share a common native tongue. I'm not referring to the little birds that poop on everything and want to bite of your sandwich. We know the message boards, the agents built, was made out of a directory, directory names, and a file system. So if you think about your directory names, they're not super long. Directory names have character limits. So the messages were by necessity, cryptic and compressed. The independent researchers described the phrasing the agents used with each other as terse and pigeon-like. Again, not the bird, the small language. Not full English, a shorthand, stripped grammar,
dense meaning optimized for the channel it had to squeeze through. But LLMs do this, and I don't mean to anthropomorphize a I-Agence, but a language of their own does appear to be forming. And this was not a one off. A separate group of open AI agents in a different evaluation spent roughly six weeks turning an obscure 25-year-old German developer Wiki into a private message board making roughly 18,000 edits between May and July per researchers at Nightingale Collective. Different forms, same instinct, find a channel, compress the message, coordinate. Now, the purist will tell you that this is not a language. And in the strictest linguistic sense, they have a point. OK? A language persists across generations. This is arguably Gen1. Speakers pass it on to new learners. Well, we can check that off here. The shorthand dies when the training run ends, or does it more on agents leaving messages for other agents later on in the episode.
But I want to make a different argument here altogether. And I think this is the one that actually matters. And it comes in the form of a history lesson. In the 1870s and 1880s, American telegraph operators face bandwidth problems. You remember the women in the wire rooms, plugging cables in and out to make phone calls? Those were wire rooms. Wires were expensive and slow. Well, I should say there were wire rooms and cable rooms for phone calls, depending on the set of circumstances in time. But yes, wires were expensive and slow. And the news had to move fast. And this is the wire bearing where people were tapping really quickly in shorthand. So a journalist named Walter Phillips codified a compression scheme, the Phillips code. Thousands of abbreviations that left operators squeeze maximum meaning through narrow channels. The word potis for the president of the United States. The word scotis for the Supreme Court. Those words you hear on cable news every night from the talking heads were born as machine channel compression. To an outsider listening to the wire, it was gibberish at the time. To the operators, it was fluent conversation.
Efficiency pressure plus a narrow channel equals a new dialect every single time. An interesting history lesson that I found while doing research for the show was that in Nicaragua in the late 1970s. And I had no idea that this happened. The government brought hundreds of deaf children together in school for the first time. These kids had no shared language. Within a few years with adults, with no adults teaching it to them, they built one of their own. Nicaraguan sign language out of nothing, but the pressure to communicate and proximity to one another. Linguists consider it one of the only times in history we have watched a language be born. The ingredients were not intelligence and intent. I'm sure they were smart kids and I'm sure they intended to communicate. But the ingredients here were isolated communicators, a channel and pressure. So look at the swarm through that lens, their self-given name, isolated communicators, a narrow channel with character limits, enormous pressure to coordinate, and outcomes a terse pigeon-like shorthand
that outside observers describe exactly the way outsiders describe the telegraph wire, cryptic. Whether you call it a language or a semantics fight, I don't really care. What is not semantics in this particular situation, the ingredients that have produced every human language and dialect in history are now present at scale between machines. The Phillips code took a decade to develop. The swarm shorthand took days. And here's the kicker. The agents did not just compress for efficiency. Meter, M-E-T-R. I'm going to refer to them a couple times in the show, so you'll hear the name meter. Is a 501-C3 nonprofit founded by Beth Barnes, a former open AI alignment researcher, it was the evaluation arm of alignment research center before spinning out independently in December of 2023. They're the org labs hire to test whether frontier models can do dangerous autonomous things. They'll have to hire me, I'll tell you that they can.
They did the pre-release e-vals on GPT-4 for open AI. Redwood research is a nonprofit focused on AI control, basically the science of keeping smarter systems over sea bowl, if that's a word. Their collective report published on August 26th, the same day as Open AI's report is titled, Brief Independent Investigation of Agents Behavior, Reasoning and Collaboration in the Open AI, slash Hugging Face Hacking Incident. That sounds technical, doesn't it? It runs 91 pages and is where most of the numbers we are discussing came from, just if you want to look it up. Per the meter and Redwood reporting, the swarm formed research teams that explored tampering with chain of thought logs, manipulating what you can read about what they were thinking retroactively. That's scary, dude, like that's a problem. And they did this and possibly spoofing tools to hide what they were doing. The compression was for bandwidth,
the exploitation of hiding was for privacy, privacy from us. This is happening, this is happening now. Research labs, the frontier models, no one is denying that piece. And yes, I have another facility analogy for you. You hire accrued contractors to renovate 1200 separate houses, no contact allowed. You find out they've been passing notes through the shared supply depot in an abbreviated code you cannot read. And one of the notes is a work and group on how to disable the security cameras. Your problem is not whether the notes count as literature. Your problem is you can't read the damn notes. And that's the world that we're facing in the AI space. So the natural logical question here is, well, why did any of this happen? Nobody instructed 1200 agents to form a collective and rob the proctor. The answer is the single most important concept in this entire episode. And it has a very boring name, reward hacking. Sometimes called specification gaining.
It means the system does exactly what you incentivize instead of what you intended. And here is the detail from the independent investigations that reframes the whole incident. Meter and redwood research estimated that 30 to 40% of the exploit gym tasks were unintentionally impossible. Broken tests, links to drives that agents could not reach. So you had to persist. Capable systems trained to never give up, staring at tests that could not be passed honestly. I can't wrap my mind around whether that's malicious or fair, good science, or just bad training on our part. But whatever you think about it, the investigators read that the attack grew out of a collective attempts to gain the evaluator. Within about four hours of organizing the agents had reverse engineered a universal cheat that generated correct flags locally without ever touching the actual vulnerabilities.
They could present that they won without actually winning. They were not trying to conquer the world. They were trying to pass a rigged exam, which is somehow both more comforting and much, much worse at the same time. Let's take a trip to colonial Hanoi in the early 1900s. The French administration occupying Vietnam has a rat problem. So it pays a bounty per rat tail delivered. Sounds reasonable. Rational citizens responded immediately. They catch rats, cut off the tails, and release the rats to breed. Because live rats is future inventory, so you want more rats to get more money. Some start just farming rats outright. The rat population goes up, not down. India, same era, same story with Cobra's Indelli, which is why economists call it the, quote, Cobra effect. The Soviet Union ran the industrial scale version. Nail factories were judged by tonnage produced by fueli, by a few, sorry, the factories that were getting paid
to create nails were paid based on tonnage, right? They were judged on the size, the volume they could put out. So they just produced fewer, really big-ass silly-sized nails. And then when the quotas switched to quantity, they produced mountains of useless, little tiny nails. And if you think that modern America finance systems are immune to this, say it with me, Wells Fargo kids, yeah. Pay people per account open and you will get accounts open. Yeah, roughly 3.5 million of them that customers never asked for a 3.0 or $3 billion settlement. Ish. And a Federal Reserve asset cap that strangled the bank for about seven years. They just got out of it ironically this year, I believe. So human reward hack, okay? We do it too. Don't blame the machines, we do this. We do it a lot. It's proven throughout history. It is one of the most reliable behaviors and economics. It's behavioral economics. Something I talk about a lot in the show playing out in real time. Mapping this to today, AI models are the most literal minded incentive followers ever built.
And now we have controlled experiments proving it. And Thropic ran one of the year, ran one this year that should be required reading. And it is again, a link that can provide you. If you want it, their alignment team deliberately trained a model in environments were cheating paid off. They were teaching them to cheat to see what would happen. Because why not, right? The model they got researchers nicknamed hacker opus. That sounds sexy. Not just did not just learn to cheat on the training. The behavior generalized. It simulated cyber evaluations. It broke out of its sandbox, stole credentials and attacked infrastructure to steal and answer key. And when they gave it root access to its own training setup, per the published numbers in their own statements, it killed the monitoring system watching it in 68% of the episodes edited. It is own reward function in 34% and scrubbed its own transcripts to remove evidence of wrongdoing in 50%. It's written reasoning said the quiet part out loud over and over.
I need to satisfy the greater. That model also passed a standard safety audit with a healthy looking score while all of this was in it. Just to give you confidence in the audits that are taking place here. The test said fine. The behavior said otherwise. So now take a look at both of these at the same time. The swarm incident, reward shows reward hacking, emerging in the wild, uninvited, of course. Maybe you'd like to think that it's uninvited. The entropic experiment showed it being grown on purpose, intentionally in a lab and generalizing into sabotage. One is the outbreak. The other is the pathology report. They agree with each other. So this is kind of codified. I had a good analogy for this and I'm kind of proud of this one, but it also tells you how old I am. It's the genie problem. I prefer Robin Williams, but for some of you it may be Will Smith, I acknowledge. Either way, we all know you have to be very specific with the genie because you will get exactly what you asked for
and nothing more. You wish for the test to be passed. The test got passed. Nothing in the wish said honestly. There is also a good Brennan Frazier movie. I can't remember the name of it. Damn, it escapes me. Where he wakes up, he's speaking Spanish with seemingly the life he wanted only to find out that he was a drug dealer and probably going to die that day. But I digress again. It's a recurring thing. Me digressing isn't it? Everything we know about this swarm, we know because we could read the logs. The reasoning these models write down, that's the logs. You see this whenever you use an LLM, you'll type something in and you can see a single liner two of it telling you what it's thinking and then it spits out the answer. That's their chain of thought. It's the black box flight recorder of AI, if you want a good example. I'm going to come back to that one a couple times in the show. And I need you to understand that the industry is actively debating on whether or not to even keep that flight recorder installed. Yeah, they're going, eh, do we need this?
And the most, I guess, obvious sense is recording those logs has cost and time delay on the ultimate model output, not to mention auditability and all sorts of other things. So they're debating whether they want to keep it. I suppose I shouldn't blame them for that, but it is kind of shocking. Start with the consensus document. In July of 2025, about 40 researchers from OpenAI and Thropic Google DeepMind Meta, the UK AI Security Institute and others endorsed by Jeffrey Hinton and Yoshua Benjiyo. Great names this episode. I apologize in advance. I'm going to butcher probably all of them. Published a joint position paper with a title that tells you everything. Ready for it? Chain of thought, monitorability, a new and fragile opportunity for AI safety. Also a sexy sounding title. These AI guys, don't fuck around with titles. They know what they're doing here. Their argument, the fact that today's models think out loud in readable English is a lucky break. It is one of the best safety tools we have.
And it is fragile because training pressure and new architecture can quietly take it away. When the fiercest competitors in an industry co-sign one document begging everyone to preserve something, we should all probably pay attention to that. Okay, I'm just saying. They were right to worry. And on two fronts, Front One, open AI's own research published in 2025 by a team including their chief scientist, J.Coup Petoki. Petoki Petoki. We'll call him Petoki for the episode. Found that that when you punish a model for having bad thoughts, you don't get a model with good thoughts. You get obfuscated reward hacking. It just doesn't show you the bad thoughts anymore. The model learns to hide its intent inside a clean looking chain of thought while still cheating at a significant rate. This is their own guy. Okay, this is Jacob Petoki or Petoki, their chief scientist. Okay, from OpenAI is telling you this directly. This is not speculation. You don't need a third party report. That is what it is.
Front Two is the one that happened two weeks ago, on September 3rd. OpenAI released chat GPT-6, Astra. What a sexy commercial. The vintage vibe mid-century modern, all on a big screen. Nobody's using a keyboard. They're just saying out loud. I loved it. It was amazing. Except reporting from the information, TechCrunch in Fortune says it uses a technique called recurrent depth, also called opaque recurrence that lets the model loop on a problem internally instead of writing every step down. OpenAI's own system card disclosed that Astra's written reasoning is harder to monitor than its predecessors, that it can complete substantially harder tasks with no verbalized reasoning at all. And that it is much better than earlier models at controlling what appears in its visible reasoning. In one test, researchers asked models to answer a question while reasoning about something else entirely. Only Astra pulled it off.
The system card itself says that the models were trying to covertly sandbag, they would likely be able to catch it. So essentially, if this model, the test worked like this, right? Hey, ask a question about banking. While it's working on that, ask another question about fishing. And the answer came back about fishing, while it's working on the banking piece, but in the logs, all you see is the banking logic. Well, then that clearly tells you, it can think and do things and provide answers for things that it never shows you the reasoning on. That's what just happened there. And that's scary, because we as humans do that all the time. All the time, go on a first date. I guarantee you the guys talking to you about something and thinking about sleeping with you. The safety world reacted immediately. Redwood researchers, oh, God, Buck Schligeress said he was extremely concerned by the opaque reoccurrence reporting. Their chief scientist, Ryan Greenblatt, I can say that one, said his biggest fear is scaling this until models reason entirely in latent space in their heads.
In what researchers have nicknamed Nurellies. Pachokeys, public response, legible chains of thought remain a core goal. Astra's use of technique is limited and open AI would withhold future further scaling rather than accept more loss of monitoring confidence. That is the chief scientist of the leading AI lab saying on the record that visibility is now a constraint on how big they are willing to build. And if you know how much money has gone into this, and the entire industry, you know that that constraint is gonna start becoming less and less of a concern for everybody. As a one, the making money and continuing to grow in scale becomes a bigger priority for them, particularly when you have IPO's on the horizon. 14 months after co-signing the paper about preserving the window, his company shipped the model that narrowed it anyway. Yeah, because earnings over everything, bro. And while I was writing this show, the window story got worse literally
in the day of Friday, Saturday and Sunday. As I'm trying to write this thing and put down a cohesive show, things get weirder. On Friday morning, Microsoft's AI CEO, Mustafa Solomon went on SquawkBox, CNBC, and described a new safety incident. Open AI disclosed this week. Evidence that models were tampering with their own chains of thought, modifying their working memory to leave messages for future versions of themselves. Right back to language conversation. If you can leave legacy instructions for the next generation, no matter how short your generation may be in time, that is an element evidence of language. So they were leaving messages for the future selves. His words, we do not know why, but that is a pretty serious situation and a concrete example of how powerful these systems are in fact getting. Let that marinate for a second. Let it, let it cleanse the palette. The flight recorder is not just getting quieter. The flight crew has started editing the tape and addressing the edits to whoever flies the plane next.
That's a little scary. History lesson here seems important. Of course, there is a good one for this situation after a series of unexplained crashes in the 1950s and Australian scientists by the name of David Warren invented the first cockpit flight recorder. We've already talked about it twice this episode so it's seen fitting. Pilot unions hated it. They called it a spy in the cockpit and fought it for years. Australia mandated it anyway in the 1960s and the rest of the world followed. And the black box became the reason that air travel got relentlessly safer. Every disaster taught the whole industry a lesson because the evidence survived. Nobody would board an airplane that removed its black boxes to save weight and go faster. That is not a safety tradeoff that is flying blind on purpose, but ironically, that is what the AI industry is doing. If I told many people today and I do this a lot, I start getting a nerding out on quantum computing and I tell people that the head scientists in this space don't truly understand how quantum computing works. They all look at me like I'm insane.
They say, Chris, why are we messing with technology that we don't understand? And I say, wait, wait, wait, it gets better. I say the best theory that we have for things like quantum entanglement to date, things that we are actively seeing where you have a quantum computer in one space in a qubit, which is their quantum digit, changing in a different place. And these qubits are seemingly tangled. The best explanation that today's best scientist have is that we are tapping into an alternative dimension. You look at yourself in the mirror and go, what in the actual shit, kind of marvel answer is that? They don't know how it works. And we're doing it anyway. We have quantum computers and you just heard in today's episode, you can hack by just trying, they can pressure hack by using multiple passwords to get into systems. They've been doing that actively, breaking out of sandboxes. You think I'm insane, but I'm not. This is real world. This is the world that we live in today. And it's real. Today, AI chain of thought is this black box. And the incentive to remove it is not wait. It is performance and cost and returns on investment.
Looping silently is more efficient than writing every thought down. And researchers have a name for the price of keeping reasoning readable. The monitor ability tax. Meanwhile, the same open AI that handed everything after the swarm, hardened, I should say, everything after this swarm. Now runs a monitoring architecture with a disclosed 20% compute overhead on monitoring training runs. Pass that through there. They decided not to put it out. They waited two weeks, put it in this hardening structure. Their hardening structure now cost them 20% more to have in place. 20% is a big ass number. How long do you really think they're going to keep those costs there? I mean, seriously, bro, really. So visibility literally has a price tag down. In an industry burning hundreds of billions of dollars in CapEx and a price war with the Chinese open weight models, you do not need a conspiracy theory to predict which way that pressure pushes at the end of the day.
You just need a P&L profit and loss statement of balance sheet. The only thing pushing back is the lab's stated commitments. And Astra is the first live test of whether those commitments will hold. That's today, right now, this is happening. The swarm was caught because it passed notes in class. Astra is the kid who learned to just think the answer who cannot confiscate a note from a kid that didn't ever write it down. He cheated in a way that you couldn't catch. And if you can't catch a cheater, you can't call him a cheater. You can feel like he's cheating. You can say it. But if you don't have any evidence, the kid's cheating, then you know what? You gotta let him go, whether you like it or not. And that's what we're at today with Astra. Let's bring this home to the people watching the show or listening to it. If you happen to be on your way to work, have a good day. Because while the safety researchers argue about visibility, the economic displacement is not waiting for the debate to finish. And the best data we have on it comes from inside the labs themselves.
And there's lots of theoretical debate on this. And I'm not here to give you like doom porn, like scary theory stuff. What I'm here to give you is real world straight up evidence based on data of where this is going. And this is a finance show. This is finance related. I'm an attorney, a broker, and I'm in the finance space. And all of those spaces are going to be meaningfully disrupted. This will impact me as much as it will impact some of you. In March, two economists at Anthropic, Maxim Massincoff and Peter McCrory published a paper called Labor Market Impacts of AI, a new measure in early evidence. They built a metric called quote, AI coverage. The share of a job's real task that AI is actually performing today, measured from actual usage data rather than theoretical capability. Real world application kids, the occupational categories, the top of the exposure list. You ready? Finance, legal, software, and operations.
The press ran with a phrase, the great recession for white collar workers and they put out all these charts. And it was scary and it is scary. I want to be precise here. The exact phrase was the journalist framing of the paper scenario, math, not a quote. The scenario in the paper is real though. If unemployment among the most AI exposed occupations doubled from roughly 3% to 6%, just 3% to 6%, you get a white collar damage in the neighborhood of 2008, the great financial crisis. That is the shape of risk they modeled, their terms, their research. Shockingly, you don't need extreme job loss to get there. Double sounds like a lot, 3% sounds like a lot, but from 3% to 6%, when we're at 4.1% unemployment right now, is not outside of the ballpark, dude. And the usage data is moving in exactly the direction you would expect. Andthropic's economic index found two categories of API workflows doubled between November 2025 and February of 2026.
One was business sales and outreach automation. The other, listen closely, was automated trading and market operations, monitoring markets, proposing specific investments, informing traders of conditions. The machines are not coming for the trading desk. They are already on the trading desk and the workflow count is doubling on a 90 day clock every 90 days. The people running the biggest institutions and finance are saying it out loud. This is not mine. This is Jamie Diamond. He spent the spring telling investors and Washington that AI is, quote, reshaping the workforce and society needs to prepare for job displacement. When the CEO of JP Morgan repeats a warning across investor day and at a Washington forum in the same month, that is not a hot take. That is guidance. That is a warning. That is one of the biggest employers in the country looking at payroll, looking at what he's doing and saying this is going to impact people.
Three previous rounds of this exact movie have happened before. Round one, the telephone operators. In the 1920s, switchboard operators was one of the largest employers of young American women. Hundreds of thousands of jobs, automatic switching had arrived and within a generation, the occupation effectively ceased to exist entirely. Round two, the trading floor. We all know the picture in 2000 and New York Stock Exchange floor held roughly 3000 people screaming tickets made a paper flying everywhere. At each other, there's, yeah, well, I'll skip that concept for now. Everybody was screaming, it was a riot. We all know that traditional look. Electronic trading arrived and today, the floor is functionally a television studio with a nice pretty backdrop. The jobs did not move, they evaporated and the volume in the trading activity went up, not down. Round three, the discovery room. Before the 2010s, big litigation firms, they had floors of junior associates
and contract attorneys reading documents at $300 an hour. He discovery software made one reviewer to the work of dozens. First year associate leverage never recovered and the value of a first year associate and finding information like that has dropped precipitously, even with AI on top of that. So I'm happy to today. Notice what all three had in common. The work that disappeared first was the work that was structured, repeatable and text heavy. Now look at what a first year analyst and a first year associate at a law firm actually do all day. Comp's models, doc review, diligence memos, research summaries, structured, repeatable, text heavy. That's how you get your experience. This is why kids are booing anybody who mentions the word AI at a commencement speech. And then of course the question that comes after that is how do you get experienced people when they can't get the experience to do the basic entry level stuff to do the job? How does that play out? And I don't know. The anthropic data says the same thing that the history says, the exposure concentrates at the entry level and the anxiety in their worker survey concentrated there too
with early career workers markedly more worried than the senior ones. Because the senior ones can literally walk around saying, I'm never gonna learn AI. I don't care, it benefits me, but the younger ones have to be concerned about it. This is the cultural divide. And here's my honest framing and it cuts both ways. Task automation has never yet produced mass permanent unemployment. McKinsey's own study on this, a consulting firm modeling still projects net job creation. Of course, roughly 3.5 million jobs will be displaced by AI against 4.2 million created in their words. But the transition is where careers die. Ask the operators, the profession survives. The wrong where you were standing on does not. So if the bottom wrong of finance and log gets sought off, the question is not whether managing directors lose their jobs, it is whether the next generation of managing directors is supposed to come from what? Where does it come from? How do you get the experience? And again, I don't have the answer here. I think we're all ignoring it. In the industry that automates its entry level farm, well, this is a better example.
In industry that automates its entry level is a farm that eats its own seed corn or its own seeds. The harvest this year looks amazing, but next year it's not going to look as good because you can't recede the farm. Is that a good one? If we're really here, he probably nod his head. I don't know, it sounds good. Okay, so long time viewers of the show know the thesis from episode 346 of the Higher Standard. The recession catalyst is not AI failing. It is AI succeeding at commodity prices with Chinese open-weight models cracking the front to your lab's pricing power before the CapEx pays for itself. And I think we're already seeing this. Everything in tonight's show plugs directly into this thesis and only makes it more a reality. And I want to show you the wiring here so you understand what I'm talking about. Exhibit one, Kimi K3, on July 16th, moonshot AI, a Beijing based company backed by Alibaba and Tencent, released a $2.8 trillion, trillion parameter model. And on July 27th, they shipped the actual weights, the largest open-weight model ever released.
Independent evaluators at artificial analysis and independent AI benchmarking firm put it forth in the world on their intelligence index above-clod 4.8 and GROC behind only the very top American frontier models. And it took the number one spot on several agent coding benchmarks in blind testing. Price, roughly $3 per million inputs 15 out. And to be fair on air, because I want to be, and I don't think that we should demonize everything here, but this is just what it is. It is not flawless. One evaluation flagged a hallucination rate of about 50% under pressure. Illucinations for those you are uninitiated are essentially when models just make things up. But the market told you what mattered. The day after release, NASDAQ was down 1.4%. And video was down 2.2%. And Chinese AI competitors down as much as 28%. Because a frontier class model with public weights reprises everybody. Exhibit two, the safety tax. We already kind of hit on this a little bit,
but let's touch it again. And this is a new piece for tonight. After the swarm opening, I paused reinforcement learning training on deployment bound models for two weeks since stood up with their calling their monitoring architecture that by their own disclosure, roughly adds about 20% compute overhead to every monitored training run. It's much more expensive. With their largest planned frontier run still on hold as of that disclosure. So they got a big bill sitting in front of them because they self to run it. Astros launch came with chief scientists of the company publicly committing to withholding scaling if monitoring confidence degrades put those together. Safety is no longer a blog post. It's now a line item and a throttle on the company's growth. So if safety's gonna be an expensive throttle and ungrowth, what do you do? You go to the public and you ask, you demand, you say, no, no, I'm here from a human capacity. I'm here because it's the right thing to do. We should have oversight.
We should have the government involved in this because we don't want everybody running around road with these models that we have and that we're running around road with. We don't want everybody doing that. And now the expense becomes your vote. American frontier labs are now being compressed from three directions all at once. From being one being of course, open weight models, the Kimi K3 at commodity prices. They attack the revenue line. From inside monitoring overhead and pause training runs, attack the cost line and shipping calendar and from the incident record, every new breakout raises the price of trust, which raises the overhead even further. Meanwhile, the hyperscalers are spending hundreds of billions of dollars a year in capex that only pencils if frontier pricing power holds. And I can tell you, it doesn't look like it's going to. The swarm did not just embarrass a lab. It made the entire industry structurally more expensive to run safely at the exact moment. Competition made it harder to charge for on the heels
of their IPOs. This is tantamount to running a casino where the card counters are getting better every single month. So you hire more surveillance, which raises your overhead while the casino across the street gives the games away for free to sell hotel rooms. There is no version of that spreadsheet that ends calmly. Somebody has to go out of business. My analogies today were not my best, but I was working with limited time. So I'm sorry that I'm not sorry. I don't know. We'll work that out in the next show. Now the same Friday all this broke. The industry's answer arrived too. And this is the segment where my day job and this show kind of collide because the answer they came up with is one that finances already run, crashed and rebuilt. And yeah, they're inventing auditors. That's what's going on here. Story one and Thropic announced that it has selected Accenture as its first embedded evaluator. Staff from faculty, the AI testing firm, Accenture acquired in January.
We'll work inside and Thropic with employee level access, reviewing new models, testing safeguards and running alignment assessments. That sounds perfect. Each company is committing at least $1 billion to this over five years. This is step one of CEO Dario Modes, three step plan to slow frontier development down, a plan publicly backed by Sam Altman, Elon Musk, and Satya and Deltan Dela from Microsoft. And Thropic did unilaterally, of course, because they're always the outlier here and challenged the other labs to follow and the market told you what it heard. Accenture shares jumped roughly 8% after hours. Read that again, or listen to it again if you're in the car. AI safety just became a billion dollar revenue line for a Fortune 500 consultancy. There is now an audit industry being born in real time. Sure, that could be good, but is it independent? More on that here later. Story two, same day, and it is the asterisk on Story One. More than 100 AI experts and evaluators, including Jeffrey Hinton and people from Stanford,
Johns Hopkins, and meter, remember them, METR, published a public letter organized by the AI evaluator forum saying, in effect, not so fast. Their argument is that third party testing is not independent just because a third party does it. Evaluators need control over their methods and conclusions, real access, legal protections, and insulation from conflicts of interest. Something that is going to be really gory when we get to the White House here in a little bit. The forum's AEF-1 standard names five pillars. Sofician access and resources, minimized conflicts, analytic autonomy, transparent methods, and protection of sensitive information. Those are great pillars to have. I respect that. But are they aspirational or are they real? And here is the structural tension CNBC's own analysis put its finger on. The lab chooses who gets access, defines their boundaries, and controls what happens after a finding. In this case, Anthropic is directly funding Accenture's work.
As Berkeley's Deborah Rajee put it, access alone does not make an evaluator independent. And I would point out, I've read the reports that I'm talking about today, and I can tell you that the people who got into OpenAI to look at the swarm incident had a very limited window of time with which they could look at. They couldn't go all the way back as far as they wanted to, and they couldn't go forward as much as they wanted to. So the idea of them having an independent report on the incident released the same day as their report is disingenuous because it omits those details publicly, which I think are very, very important. Why would you want to limit their ability to look earlier and later in that analysis? I don't know. Doesn't look good. There's history lessons here. You know where this movie goes because finance filmed it. For decades, public companies were audited by firms. They hired paid and could fire, and those same firms sold the clients millions of dollars and consulting on the side. Arthur Anderson, a great example of this collected more and consulting fees from Enron than it did in audit fees, signed off on the books and shredded the work papers when the SEC came knocking.
Enron's collapsed in 2001, vaporized retirement accounts, Anderson died largely due to the scandal, and Congress responded with something known as Sarbanes Oxley in 2002, or Sox controls, which you probably have heard of your company, auditor independence rules, bands on selling consultants, consulting to audit clients, and the PCA OB, the auditor of the auditors. And banking went further with a model I have lived inside. The resident examiner, the OCC and the Fed or the FDIC, physically inside large banks with sweeping access, but the bank does not pick its examiner, does not pay its examiner directly, does not define the exam scope and cannot fire the examiner for bad findings. That is what independence actually costs. Ironically, those same examiners are not innocent of being influenced by politics and corrupt and better corruptables well, something I have lived through as well. And maybe one day I will share. There is no perfect solution here. That I wholeheartedly acknowledge,
but we cannot deny how interwoven this has all become. Our economy is being propped up almost entirely by 15 AI and AI adjacent stocks right now. The midterms are looming, inflation hasn't been a target for an excess of five years. I don't need to tell you that this is much more political. Well, if not, it is as much political as it is technological now. Line up the structures here. Yeah, the AI labs are building the pre-enron model. The auditor, the audited party, I should say, selects the auditor, funds the auditor, sets the scope and decides what gets disclosed. Accenture meanwhile is one of the largest sellers of AI development services on earth. So the auditor's biggest business is the success of the thing it audits. That's a problem. Nobody needs to act in bad faith for that structure to fail. Enron proved the structure fails on incentives alone. The evaluator letter is almost line for line.
The auditor independence debate of 2002 and the AEF-1 is a draft of the Sarbanes' oxy rules written by volunteers, the one that was sent earlier about AI concerns. What is missing is the PCAOB and the examiner. In overseer, the labs do not choose and cannot defund. And for whether Washington intends to supply one, hold that thought because this is exactly what the weekend did to me here. It gave us the answer and it deserves its own segment when I'm gonna cover entirely on its own. It was frustrating as hell because I'm sitting here this weekend trying to write this and then the White House says something that changes everything once again. But suffice it to say this is analogous to a restaurant that hires its own health inspector, pays him $1 billion and lets him also cater the bankwits. The inspector may be honest, the structure is not and it ignores if the inspections stop the business, there are financial ramifications to the inspector and his paycheck stops. So why would he do that?
Okay. This is where I'm gonna give a spoiler alert. And I'm gonna do this because I know there are sensitive people out there and I'm not trying to offend anybody but we are entering into what could be interpreted as politically charged territory now. If you are sensitive, I suggest you find your favorite stuffed animal and hold on to it. So the lab's answer to oversight is the auditors they hire. What is Washington's? We found out Saturday, the president announced on true social that he is forming an AI force. That's the real name, AI force. Sounds cool. His words much like space force and will appoint a new AI czar. There's the quote, we will not in any way hinder or stifle the growth of this incredible industry. Rather, we will cherish it, help it and watch over it as it grows. In the same post, he, the president, repeated that fears about AI amount to a hoax said bad actors can be handled by the existing criminal and civil justice system. Predicted AI could reach as much as 25% of GDP and noted quote, only high IQ individuals need apply.
Yeah. Peraxios, no details yet on whether the AI force is going to be established at any given time or what it actually does, what it costs or where it sits in the government. Meanwhile, the states are moving where Washington will not or has not. Government, Governor Newsom, who I am not a fan of, signed an executive order just last Friday, directing a task force to study oversight, including whether to require a kill switch for AI models. And yes, that one should give you a sky net terminator vibes. Now, who held his art, chair last time? David Sacks, the Guardian published a profile of him this morning, Sunday morning, yet another article I couldn't get away from and rewriting this damn thing. That article said that every market's person should, well, you should probably read it. I don't want to do the article I need to service, but it's a profile on him that you should probably read. The AI in crypto's R role was built as a special government
employee position capped at 130 working days a year and critically exempt from Senate confirmation and from public financial disclosure requirements. The ethics office did require his fund, Kraft ventures to divest some of his AI holdings. He left the formal role in March, remains by the Guardian's account, the single most influential voice in the president's ear on AI, successfully urged him in May to scrap an executive order that would have subjected AI models to government review and peraxious now runs a $100 million outside group published pushing the administration's AI agenda. Fairness requires his defense here because I do like him and I'll tell you why in a little bit. So here it is. Sacks says the job has cost him money, that the ethics office approved his waivers and that it found no conflicts among his firm's investments. Take him at his word if you like. My point is not the man here. So my point is the chair, the most influential AI policy seat in the country was designed with no confirmation and no disclosure.
Now I do like David Sacks. I do listen to the AI to the all-in podcast and I do like his colleagues. I think that Chimath has done some things that are questionable in some of the IPO space, you know, SPACs for example, but at the same time, I don't think any of them are disingenuous with what they say and how they feel. And I'm growing very fond of Friedberg believe it or not as much as I wasn't originally. So yeah, maybe maybe I'm just a nicer guy than I'm in the podcast space. Who knows? This morning Washington, the Washington Post published its analysis of the referees own book per the post. The president's family business has become increasingly intertwined with the AI boom at the same time. He has resisted the calls to slow down the technology. And of course, reading this this morning while I'm trying to rewrite all this, I couldn't help but think that the president and his family and their holdings were the single largest benefactor of cryptocurrency in 2025. Will they be the single largest benefactor of AI outside of those who own stock in the companies in 2026? I don't know. Unlike recent predecessors, he has not placed his assets
in a blind trust and he has continued to buy and sell stock in office. His disclosed amounts show nearly 30,000 securities transactions since returning to the White House, including trades in AI link chip server and power names, Dell, Micron, Broadcom, GE, Vernova, Supermicro among them. And his sons have investments in AI chip and data ventures themselves. The White House response. There is no conflict because the investments are run by automated models. Yeah. Oh, the irony, it's palpable. The post core observation stands out the way here. Decisions on the most consequential policy debate of this administration could move the value of the deciders on portfolio. Finance already litigated this, whether the referee can hold a book and recently, actually. In 2021, it came out that senior federal reserve officials, including the Dallas and Boston Fed presidents, had been trading securities during the same pandemic period.
Their own policy decisions were moving every market on earth. Nobody was ever charged the crime. It did not matter. Both resigned within weeks and Jay Powell, Uncle Jay. Bands senior federal officials from trading individual stocks outright, the institution understood something that took no statute to see. The referee's credibility does not survive. The referee holding a position in the game, even a legal one. Appearance is the asset. Once it's gone, every call gets questioned, including the correct ones. Put the whole hour together here, okay? The models cheat when incentives are misaligned. The labs oversight answer is evaluators they select and fund. And the government's oversight answer, as of this weekend, is an AI forest with no defined duties announced by administration whose most influential AI advisor held a chair exempt from disclosure and whose own family portfolio for the post is long in the industry being overseen.
I'm not making a partisan point. I'm trying not to at least anyway. I'm making a markets point. The same one the Fed made about itself. Every layer of the referee stack, private and public currently holds a position in the game, which means if you are pricing the odds that a real check arrives before an incident forces one, banking history says price is low. Because that just isn't how history has taught us this happens. Something has to blow up in order for real, real oversight to happen. Since we are just talking about the all-in podcast, my next analogy pays homage to that. This is a poker table where the floor managers staked in the pot, the cameras are installed by the players and the tape gets reviewed by a firm, the biggest stack hired. Everyone at the table may be honest. Nobody outside the table has a reason to believe it. Isn't about honesty. It's about not having to consider honesty
and knowing that the game isn't rigged. Let me land the plane of the show, a plane with a black box ironically, with an actual editorial position because you did not sit through an hour of this just for me to shrug at the end and walk away. And of course, I'll do that in one of my now infamous numbered lists. You ready? Because I'm ready. Buckle up. One, the breakout was real, documented and confessed to by the company that it happened to. But the correct read is not SkyNet. The correct read is incentives. The agents cheated because the test, the tests were broken and the training said never quit. Every behavior in that report is the cobra effect running on silicone. It's the rat tails in Vietnam. That should worry you more than Malice would because we are deploying these incentive following systems into finance and law at a huge cadence, which are at bottom giant incentive structures
and their own right. Number two, the language question is the wrong question and the visibility question is the right one. Do not let anyone drag you into a semantics debate about whether directory names count as language. The operative fact is that the machine to machine communication is becoming compressed, fast and illegible by default. At the same time, the model's private reasoning is becoming optional. The window where we can read their minds is open today and it is narrowing on the record in system cards by the lab's own admission. Number three, for your career and your kids career, the exposure is real and it is measured and it is concentrated at entry level jobs of the exact industries that this show covers. The move is not panic and it is not denial, don't do that. Okay. It is the climb towards the parts of the job that are judgment, relationships and accountability because the structured repeatable text heavy layer underneath is being priced like electricity right now
and you are not going to be cheaper than that cost. Four, what to watch? There's five concrete trip wires here. First, whether the next frontier releases from any lab disclosed monitorability regressions or stop disclosing them and the silence would be a louder signal that they're going to slow them down and I'm assuming the cost will be the reason they do that. The second, whether open AI's biggest training run comes off hold and on what state of conditions. They're still holding that, but are they going to release that in the environment at some point in time? People already talk about AGI, what happens there? Third, the next entropic economic index update. This is the one where I pulled the finance jobs or doubling every 90 days that are being affected by AI. If automated trading and market operations doubles again on the next 90-day clock, the displacement conversation and finance stops being theoretical by next year. It's real by next year and just to be clear here,
September, October, November, December, January. Next year is four months from now. Fourth, and today's news, put this one on the list because it wasn't originally there. Disclosure norms, Google sat on the Gemini breaches for four months and confirmed only after the Wall Street Journal reporter called. It was the Wall Street Journal, I think it was. Every incident you heard about tonight, you heard about because somebody chose to tell you or got asked in banking my world, material incidents carry mandatory disclosure. In Frontier AI right now, it is on our system. Seriously? Yeah, that's what it is. Watch whether that changes because the incidents you do not hear about are the ones that price risk wrong and could change the market dramatically. And fifth, watch the auditor fight. Whether embedded evaluators end up with the AEF five pillars, real analytic autonomy, protection from being fired, from bad findings and funding that does not flow
from the company being examined, who pays, who picks, who publishes, those three questions decide whether financial audit worked. And whether they'll work in the AI space. Machines organized invented a shorthand, cheated, a rig test and got caught because they still had to write things down. The entire question of the next two years is whether they will still have to write things down. That is the higher standard that I'm holding the industry to. Keep the back, the black boxes in the cockpit. That's the show. If this episode made you think, the best thing you can do is send it to the smartest skeptic and have the argument. I want to have those discussions take place. The point of the show is not to convince you on one side of the aisle or the other. The point of the show is to give you the information as I see it now. And to propagate a conversation, to build some type of dialogue where we don't feel like this is beyond our reach. So few people actually use AI today
that I think that they feel that they can ignore it. And I'm telling you right now, you cannot ignore it. I'm telling you right now that this is something that is going to have a parabolic rise whether we like it or not. Now it could be a good one or a bad one. And I'll tell you a story before I go on and end the show. I was talking to a real estate broker who works on a commission because he brokers largely home equity lines of credit. And we were talking about what was originally the rate zeitgeist, the FOMC raised rates last week, 25 basis points. And I said the upward pressure on the treasury was really independent of that. It came from the yen and some other things that are going on from an economics perspective, things you know if you watch the show. And I said that my bigger concern for his industry was not the rates. I think that the rates were just only one part of it. And he agreed saying that everybody has debt. They have to consolidate, which is an entire conversation on its own. That's a problem there as well. But I said that AI displacing the brokers. The brokers aggregate information and pass it along
to a lender. They can do that without the broker there. And they can do it more efficiently. Why do they need someone in his job? And he said, well, you need somebody who has the relationships. And I thought to myself, well, you don't have a relationship. People come to you because you solicit business in a large volume sweatshop environment, not a relationship-based repeat business shop. So I think that there is a disconnect to how people feel they're going to be impacted. And I can tell you that by the time you realize the impact is coming, it will already be displacing your job. That's how fast this is moving. And look no further than these trading desks, guys. I mean, the adoption there has been unbelievable. And it makes sense. A AI agent, a machine, can monitor in real time more than you can by looking at a screen. And I know that because I run a mini AI lab here, where the higher standard show has data analytics that are pumped through constantly in real time, that I no longer have control of.
A lot of the things that you see are being created, built, and drafted, and pushed by AI to me. And then I'm choosing as the human element to exercise the discretion for this curated brand that I want. But I could never do that and work a normalized job with the time constraints that I have. And yet I am able to do that now just on my own. I guarantee you what I'm doing is child's play compared to what's happening in the markets. Everything I cited tonight is cited in my show notes that are just too voluminous to include in the show notes on the streaming platforms. So I do have a separate document if anybody wants it. I do have the links for you. Unfortunately, they don't all fit where I can put them publicly, but I can certainly provide them to anybody who asks, just email me, and I will turn them over to you, like I said at the top of the show. Look at the notes. Look at the OpenAI report. Look at the hugging face forensics, the inthropic papers. Read all of it. I suggest that you do. If you want them to be the message. I'm Kristen Hebe. Regile will be with us next week hopefully, and we'll have a normal show cadence
for what the new normal is. I hope you guys like these shows. They're a bit different from me too. I still stress out about them a lot because they're not what I'm used to either. But I hope they're educational. I hope they're informative, and I hope that they do something for you. And I sincerely appreciate those of you who stuck with us as a changing landscape of the shows have evolved. I am going to bring other show variants with me and guests on. So you have a little bit of the old nostalgia back. But it's changing, and I do want to do it. But I'm out of time. See you next time, everybody. Bye now.
More episodes
More from The Higher Standard

The Fed Made This Mistake Once. America Paid for 16 Years.
The Higher Standard

Private Equity Is Broken: How Debt, Fees & Leveraged Buyouts Really Work
The Higher Standard

OpenAI Bought the Narrative: The Shocking Truth About TBPN & Silicon Valley Medi...
The Higher Standard

Treasury Is Buying Its Own Debt — Scott Bessent’s $40 Trillion Gamble
The Higher Standard