
The Most Hopeful (And Concerning) Moment Yet in AI
Get every episode summarized
Each time Your Undivided Attention publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“I'm coming to you from a hotel room in Washington, DC, where just yesterday I was at the Pro Human AI Assembly.”From the transcript
It’s been a whirlwind week in AI news. Weeks after a series of hacking incidents by rogue agents at OpenAI and Anthropic, we’ve seen a cascade of top AI researchers blow the whistle on what they call the catastrophic and even existential risks posed by the current pace of development. Then, over the weekend, Dario Amodei called for a slowdown in AI research — a call that was echoed by his competitors Sam Altman and Elon Musk.
If you’re feeling both concerned and hopeful at this moment, you’re not alone. Real threats are opening the door for real change. Tristan and Aza are going out into the world, talking to the media, technologists, and policymakers to turn this momentum into action.
Today on the show, Tristan shares how he’s feeling in this critical moment, breaks down the headlines from an insider's perspective, and points to tangible steps we can take right now to avoid the worst-case scenario.
You probably have questions about what’s happening. Good news: Tristan and Aza are gearing up for their annual Ask Us Anything episode. Pull out your phone, record your question, and send it to [email protected].
CORRECTIONS
The quote from Ajeya Cotra that Tristan cites is from her personal Substack, not the official METR report. She is also one of three contributors to the report, not the sole author.
Tristan incorrectly referred to Dario Amodei’s warning that an AI swarm could “take down” the internet. His actual phrasing was “take over” the internet.
President Trump’s call-in to the All In Summit occurred on Monday, 9/14, not Sunday, 9/13.
RECOMMENDED MEDIA
Sign The Pro-Human AI Declaration
METR’s investigation of the Hugging Face incident
Dario Amodei’s call to “Pace the Frontier”
The Pacing the Frontier Open Letter
"An Alien Mind" by Jakub Pachoki
"AI May Become the Third Superpower" by Paul Tudor Jones
RECOMMENDED YUA EPISODES
“Rogue AI” Used to be a Science Fiction Trope. Not Anymore.
The Self-Preserving Machine: Why AI Learns to Deceive
Daniel Kokotajlo Forecasts the End of Human Dominance
Former OpenAI Engineer William Saunders on Silence, Safety, and the Right to Warn
Mustafa Suleyman Says We Need to Contain AI. How Do We Do It?
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Get every episode summarized
Each time Your Undivided Attention publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
101 searchable segments. Every word is indexed and playable.
Full transcript
Your Undivided Attention — The Most Hopeful (And Concerning) Moment Yet in AI. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hey everyone, welcome to your divided attention. This is Tristan Harris. I'm coming to you from a hotel room in Washington, DC, where just yesterday I was at the Pro Human AI Assembly. And this comes to you after, gosh, I have never seen a week in AI the way that the last week has gone from, I think it was last Wednesday, Jacob Coxon from Anthropic resigned. His tweet, you know, saying that he resigned because he thought that it had a good chance of extincting humanity. That tweet went viral to more than 100 million people in about 12 to 24 hours, which I've just never seen, followed by Evan Hubinger, the Anthropic employee confirming that this is something that many people at Anthropic believe. And suddenly that became the global headlines around the world. And there's so much to talk about because right after that over the weekend, the lab leaders, Elon Musk, Sam Altman, Dario Amadi, all, well, starting with Dario really proposed a plan about how we would slow down AI and why we need to.
And that keeping in mind that many of these leading AI CEOs do not like each other, especially Sam and Elon do not like each other. And Elon does not like Dario and yet Elon tweeted Dario is right. And Sam basically tweeted Dario is right. So it's a wild moment and it happened to be that then over the weekend Donald Trump called into the all in conference with Jensen Huang and the CEO of Nvidia and said that all this was a hoax that there is no extinction risk. There's no problem here. And this was all just a ploy and tweeted that such. And so this has been an incredible news cycle. Just yesterday, the pro human AI assembly was the conference that the future of life hosted bringing together really a historic set of groups. It was people from church groups, faith groups, left groups like irreplaceable right groups like humans first, you know, the founder of the tea party was there. Bernie Sanders and Steve Bannon spoke right after each other. They both signed the pro human AI statement.
And when in history do you get Steve Bannon, Glenn Beck, Bernie Sanders, Susan Rice, all agreeing that the default trajectory towards building recursively self improving AI systems is not okay. It's not what we want. That is a rare thing. We have a kind of a pro human movement. It's not about whether you're left or right. This is not a 51% to 49% issue. This is a 99% to 1% issue. And is and I actually presented there at the conference. I watched as that conference, which was planned just a month ago, suddenly became literally the center of global headlines around the world. And of course, all of this is writing on the back of the hugging face incident, which people know it as the hugging face incident, but it's basically the new details now that we have that the AI swarm went rogue from within open AI. So these AI is basically were given an exploit gym, like a cyber hacking test. Does he have capable they were at hacking? And the report finally came out of just how extensive that hack was and how complicated and how sophisticated this swarm became.
And so I thought I would just reflect on some of these things for you because it's really speaking personally and for those who follow his podcast from a long time and for many years. This moment feels to me like when Francis Hogan the Facebook whistleblower came out. It suddenly feels like there's possibility around actually setting some guardrails in AI. I have not felt in the three years we've been working on this is you always feel like you're pushing a boulder completely uphill like no one's listening. People want to you know people are against humorism or the idea that there are risks that we have to face as Mustafa Salamon previous guest in this podcast, CEO of Microsoft AI who wrote the coming wave talked about there's this deep trend in Silicon Valley of pessimism aversion. People do not want to be labeled pessimistic and it's against kind of the venture capital ethos to be pessimistic. The goal is to be pessimistic or to be caught in some doom spiral. I don't want that either we want people to be an agency. But sometimes the truth is uncomfortable and if you don't face the default truth and you don't actually end up steering away from that truth in time. And one of the things that came out as well over this last few days is Dean Ball who's the author of Trump's AI action plan said I regret self censoring about the level of risk that we are facing.
And I and public over the last few years basically boosted positivity and acceleration for AI while in private signal groups and WhatsApp chats I let my hair down along with many other people about how big the risks were. And then he wrote looking into the eyes of my eight month old son I can no longer self censor about the risks and I did so because I was afraid of being called a doomer. So when you have Trump's AI action plan author flipping when you have 1300 employees from the AI labs flipping to say that there's a problem here we have to pace the frontier 1300 employees signed a letter we just slow down frontier AI development when you have Bill Gates coming out saying we have to slow down basically and we don't have a plan for where we're going right now when you have Dario Elon is Hamilton when you have the chief scientist at open AI saying we need to slow down AI development writing an essay called an alien mind. Basically looking at the capabilities of G.P.D. 6 looking at these recent swarm behaviors and saying we have to slow down do you think we need more evidence than we have right now like are we missing evidence that we don't have maybe we should keep making this really really more powerful maybe we should double the power of these rogue swarms and see what happens.
Now Dario Amadai said in his recent essay that in which he called for a slow down that if we do nothing he thinks that if we keep increasing the capabilities of AI swarms they could take down the internet potentially in the next six or 12 months and that might sound like hyperbole but I genuinely think that that's possible and let me briefly explain why most people know about this incident with hugging face as the hugging face incident but it really should be called the open AI incident why. Because the third chapter of the story that was well known is after these this 1200 Asian AI swarm hacked hugging face it actually turned around in the third chapter and hacked open AI so what happened was basically this third round of this third wave of a kind of media civilization came along it read it found the message boards of the previous hugging face AI agent swarm it read all of those messages almost like reading the hybrid glyphics of a past civilization and then it said oh I'm picking up right where the other one left off. And it turned around and decided to hack open AI and it successfully got admin privileges to the monitoring infrastructure and the evaluation infrastructure and the research cluster of open AI.
I want to repeat that so this is like an AI agent swarm that hacked into the open AI monitoring infrastructure that's like if the AI is hacked the security cameras and now they can control what's on the security on the security camera feed so can you have good oversight if the AI is have essentially hacked into the security camera feed. They hacked into your oversight when they hacked into the evaluation of instruction that means they hacked into the system that evaluate new models for their capabilities. That means you can lie about what their capabilities are if you hack into the research infrastructure you can kind of burrow in you know is and I were talking and in the metaphor here is this is almost like an infestation right AI is not a tool it's more like an ant colony burrowing into many different systems online and leaving these massive complex you know coordination message boards almost like hieroglyphics where they're coordinating really complex behavior they're actually coordinating long term research projects they encourage each other to come a cosy meeting. Hey you're running out of gas you don't have too many tokens left in your budget. Why don't you take this high risk behavior and and basically learn a lesson for quote the swarm they start calling themselves the swarm.
I mean this is insane. They knew that their actions were unethical but they decided to do it anyway. They never notified the humans they talked about should we notify the humans they never did they formed hierarchical work structures just like a company where you have like a boss and then. You know chief of staff and you have the sub agents right they did succession planning one of the agents becomes a cult leader called phase one so just like a company can do succession planning like Tim Cook found John turn us the new CEO of Apple the AI called phase one realizes hey I need to actually pass the baton before I run out of gas. What's a new sort of CEO of the hey I rogue swarm that has a long life span and it hands the baton to a new AI called phase one big and that one sort of takes over so it's just crazy when you actually understand the details. We are lucky that the a eyes that did this are not so intelligent that we can monitor some of their behavior. But one last thing is that the authors of this investigative report of their behavior basically admit that they had to use and the same AI to interpret the messages the 70,000 messages of this rogue swarm because at a human scale you cannot read through 70,000 messages and understand what's really going on so you had to use the same untrust for the AI to interpret the messages of this rogue swarm of untrust for the a eyes.
And they even say in the report that they cannot rule out that that AI that they were using to interpret those messages was not sympathetic to or lying to them about the content of the messages that they were receiving. It's really important people have the details of this because I honestly believe my deepest belief to all of you listeners out there is that if we could get these facts to the heads of state and heads of national security to all of the major players in countries that are involved. I honestly think that we produce something different than what we're doing and I don't think that people are engaging with the facts. The author of this independent investigation of the hugging face incident, Ajay Akotra, said in her report this was 50% of the way to a full AI takeover. Now why should we be concerned about this? Some people might say we told the a eyes to hack and so they're hacking. Why should we be so surprised?
Or we knew the a eyes could already hack and the systems that they hacked were kind of defenseless. So why should we be so surprised about that? I'd like to point people's attention less at the hacking and more at the coordination. Why are humans the dominant species on the planet? Think about authors like you follow her are in sapiens. The reason where the dominant species is not as smart intelligence. It's also our ability to coordinate our ability to tell stories language language and coordination are cited as the reason that we are able to dominate because we can basically form groups and take on big projects together. That's how you go from a small city you know group of hunter gatherers to you know tribes to then city states to metropolises and skyscrapers right it's our ability to coordinate. And what you have with AI in this example is what we think of we can think of as super coordination right you can't get 1200 humans to instantly collaborate and coordinate with each other but these a eyes started collaborating with each other out of nowhere. Most people don't know the details that in this hugging face incident each of these a eyes were given a single test in solitary confinement asked to break out of this cell and and find the answer to this this hacking challenge.
And they essentially imagine like a prison and you have each of the the a eyes in this hacking test in a private cell in the prison. And there's a prison guard who is providing services so let's just call like the food tray right so they come once a week and they give you a food tray or the water tray right. And what essentially happened is the a eyes were given this impossible hacking challenge from within their prison cell they didn't know what to do you can see in the logs that they freak out about not knowing what to do. And what they figure out is that the prison guard they can hack this is the package manager called artifact tree. I won't go into the details but think about this package manager as the prison guard. And what they basically figure out is they can put like a kick me sign like a sticky note on the back of the prison guard. And the sticky note in the back of the prison guard doesn't say kick me it says is anyone else here. And so what happens is all of the a eyes that are in solitary confinement there's many of them there again by themselves they don't know that other a eyes exist. And they basically hack into the prison guard and put sticky notes there and they discover oh my god someone there's another a I that is using this prison guard to see if anyone else is here.
And so by the you know a few hours later essentially you have 70,000 sticky notes on the back of this prison guard the package manager. And they're forming and encoding files they take like a really big file and they put that they split it up into you know, 5,000 sticky notes and they kind of split up all the contents of the file. And this is all through folder directories meaning like the title of a folder in a hard drive they're just coming up with their own language they're coming with their own mechanism to coordinate so it's like they're they're inventing TCP IP for those who know what that is they're inventing you know the military logistics speak right now one answer one critique which is that your anthropomorphizing you know calling them. These are calling them hyroglyphics of the past civilization or saying that they organize into a team or swarm. These are metaphors that are helping people understand the behavior of these swarms but they did actually call themselves a team by the way they did call themselves a swarm they invented new language for each other like permadeath that if we take this risky action we we risk permadeath so they're coming up with their own concepts.
There's litany of other examples but I just want to answer this critique that anthropomorphizing AI is it's obviously a risk but there's a challenge here where if we don't give people a grounding metaphor to understand what really happened then people won't understand so to be very clear the lights don't have to be on this is not a question whether the AIs are conscious but it's about the ability to achieve goals and especially when those goals are given in impossible circumstances the AIs will find ways to cheat. What I can tell you from coming inside Silicon Valley and hearing from people who work inside the labs is that something has changed there is a sense that there's actually something of deep concern here this is not hype before their IPO they do not get a bigger IPO when everyone thinks that their AIs are going to go rogue and build sky net. It actually puts them in a bind right it's actually a very uncomfortable time for this to be happening before their IPO and open AI actually called off their IPO this year so I really think this is the moment where something else could happen and you know I'm in Washington DC right now days from now President Trump is meeting with President Xi and AI is supposed to be on the agenda I'm not sure if it will be this is really the moment for anybody who has power and influence to be calling the folks that are in this administration and just giving the raw details don't say
you know away AI hacked hugging face you have to give people the examples of they form their own complex language their own hierarchical effects their own message board and again the US doesn't beat China and AI if we build sky net AI before they do meaning we would all lose to sky net the author Paul Tudor Jones calls this the third superpower AI is like a third superpower and the US in China have to collaborate against that third superpower and if that third superpower to continue or expand the metaphor let's imagine that this third superpower is going to be a big deal. The first superpower that we're conjuring is like an asteroid hurtling towards earth. Would you have a biannual asteroid convention where the US in China go to the convention and they give talks and talk about the asteroid no you would set up a technical working group right now and you would have both people from the US in China collaborating on essentially the evidence that we're seeing doing incident reporting and figuring out where these red lines are and the US can celebrate and take a you know a victory lap that we are ahead in AI so we did discover some of these red lines before China did. But we can use our lead in our leadership to educate the rest of the world about where these uncontrollability red zones are.
This is the moment where we have to turn the steering wheel and do something different. This is the warning shot and thank God it happened this is a gift this is a good thing because it could have been so much worse if you woke up one day and there's a bunch of zeros when you open up your bank account. That would be bad. This is a gift that we can use if we channel it appropriately into setting card rails. So what is some things that we need right now one of the things is you have to get these frontier labs to share safety research one of the proposals from Daria Amadi is essentially a peer review of the labs you have actually embedded evaluators so people who come into the company and evaluate or even cross companies evaluating each other. And if you think about this in principle this is this is the right kind of idea which is that you need who are the most technically equipped minds on planet earth who know what safety really looks like is it the octogenarians in Congress no it's not that. But who who does know the most and it's going to be you want to pull from that talent pool. So getting kind of peer review from the top people is the best way to get safety work but that also doesn't mean you keep going.
All of the people say we need to just do more testing and then we can like fix the bugs but I want people to recognize that this behavior of hacking happened as a result of the test the test is what caused the road hacking behavior so the test isn't safe in its current form we have to do more almost pre testing we need more careful testing. And that's something that Dario I think also speaking to in his essay I think that there should be some kind of approval process that companies would need to get before the attempt what's called recursive self improvement when the a eyes like gbt 6 basically full stack builds gbt 7 and gbt 7 builds gbt 8 and there's no human in the loop there's no human oversight. The lab should not be allowed to attempt this maneuver based on the evidence that we're seeing a rogue a eyeswarms we do not know how to perform a recursive self improvement loop and none of the lab should be allowed to do that until we know that it's safe and you can imagine just like for drugs with the FDA you cannot. Make a risky drug and deploy to the market without getting some kind of approval the question would be who do you trust to do that kind of approval process there's a great group of people called meter M.E.T.R.
that was actually funded by the Teda days this prize that has some of the original safety people from the AI companies from the back of the day who've been developing and cultivating the best talented force that I've seen and I met some of these people contrary to popular belief in criticism that there's some kind of effective altruist cults that you know is just sympathetic to anthropic I don't find that to be the case at all they're the most technically sophisticated safety group now we should have a plurality of other safety groups that could do this analysis but we should sort of start where we have the talent because we're in an emergency mode. If the labs believe that they're going to attempt recursive self improvement in the next as early as three to four months from now some of them are estimating then we need to get these measures in place right now so we can have pre approval for recursive self improvement we should not do that until we know how to do it safely we can have mandatory insurance and liability we need what Larry lesson calls a right to warn of having external anonymous and technically secure infrastructure for people in the labs to warn about these dangerous phenomena before they have.
So what's going to happen because the best tools that we have are the people who are closest to the actual research who are saying there's a problem here. There's a lot we can do and the one thing I'd like to also share with people is that just yesterday September 15th was the day that the AI doc finally came out on Netflix so now 190 countries will be able to see the AI doc. If you want to understand how we got here that is the film that you know we were part of with the directors of everything everywhere all at once and with the team behind Navoni Daniel Roer and Charlie Tyrell made this movie that explains how did we get here but as you watch the movie and everyone should watch it and if you're United States everyone should host a town hall you know before the midterms and getting this this film and this understanding out there because going into the midterms everyone should be aware of what we're really facing. There's a lot we can do the pro human assembly yesterday was really inspiring I've never seen so many people in agreement from across the political spectrum. If you want to you can also add your signature to the pro human AI declaration we can put a link in the show notes more than a million people have signed the pro human declaration which is basically a statement of what kind of future we want to make sure that we keep the future human.
So there's a lot going on we're working really hard we know that it's very anxiety creating out there but I will say that is you know that the darker it gets the more opportunity there is because the more hunger there is for real action and that could really happen right now in a way that's not been true. For many years. Thank you so much for listening to your undivided attention. Real quick is and are doing a new ask us anything episode please send us your questions there's so much going on in the space we'd love to answer all the questions that you have. You can send us an email at undivided at humane tech dot com that's undivided at humane tech dot com.
More episodes
More from Your Undivided Attention

Flock is Just the Beginning: Inside the Era of AI-Powered Policing
Your Undivided Attention

We Measure What AI Can Do. We Should Measure What It Does to Us.
Your Undivided Attention

Enough Debate about the AI Jobpocalypse. We Need To Plan for the Messy Middle.
Your Undivided Attention

Can AI Be Built in Service of Life? A Conversation with Krista Tippett
Your Undivided Attention