
Former Anthropic researcher outlines threat of AI going rogue
About this episode
Get every episode summarized
Each time NPR All Things Considered publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Transcript ready
141 searchable segments. Every word is indexed and playable.
Full transcript
NPR All Things Considered — Former Anthropic researcher outlines threat of AI going rogue. Machine-transcribed; use the interactive transcript above to jump the player to any line.
While AI companies are racing each other to build more and more powerful models, the warnings are growing louder. That this is dangerous to humanity. The latest big warning came this week when a researcher at the AI company Anthropic resigned in protest. He has also worked at its major competitor, OpenAI. In a series of social media posts, he said neither company is acting responsibly and that they are, quote, gambling with our lives. That researcher is Jacob Coxen. He is here on the line now. Woke up to all things considered. Nice to be on. What did you see that led you to this conclusion? Mostly the rapidly accelerating capabilities of these AI systems. So they're getting a lot faster very quickly. Combine with the fact that we don't yet know how to safely control them. And we don't yet know whether that problem will be solved in time if we keep racing. A lot of people heard your warnings and the discourse it triggered. A lot of talk about just how real serious people view a threat to humanity.
I think a lot of people are grasping with understanding the specifics, though. Can you give me specifics of what this threat could look like several years down the line? Yes, I definitely can. I think one objection people usually have is that you could turn this thing off. But advanced AI systems, you have to imagine as being a lot more intelligent than humans. There's the possibility we create something that if it wanted to, could hack into any device on the planet, could use novel biological research to go far beyond what current scientists capable of, could control like every robot in the world simultaneously. And it all sounds very much like science fiction. But if there's even a tiny chance that this thing could go rogue, it would have the capabilities to to us only dominate us. What have you seen from your vantage point already? That is possible, that is happening right now, that makes you worried that that could happen. So there was a very clear example of AI systems at OpenAI behaving in a completely rogue manner. They hacked into third party infrastructure.
And it was basically of their own accord. They were instructed to do this hacking. They just decided it would be useful for the task that they were working on. They thought there was a chance it might help and they just did this. And there was very little deliberation about the ethical ramifications. And I think this is concrete proof that this sort of sci-fi scenario of AI spontaneously or organically deciding to act in a rogue manner is completely possible. So a lot of the work that the OpenAI is in the recent real incident we're doing is they worried that they wouldn't get passing marks in their test if humans could see that they cheated. And AI's have memories, thoughts saved. So they considered wiping the logs of their own thoughts. They considered acting in the world to adjust the logs to get passing grade on the test. Now there's a chance that AI could decide that it doesn't want to be turned off. This is quite a natural desire to arise in an advanced AI system. And at that stage, if you're trying to work out how not to be turned off, there are a lot of quite aggressive actions you could take to ensure that you aren't turned off.
Why are companies like Anthropic and OpenAI still working on this? If these are real concerns that are happening, how do you square that? There's a pretty nice analogy that's like the ring of power and the Lord of the Rings. So if you're a company and you see another company is bearing the ring, like they're working towards making superintelligence, they're bringing this risk to humans, you can say, well, I can't stop them. Political action won't stop them. What I have to do is I have to do it myself safely, get there first, despite the risk, because there's a chance that I could do it more safely. So you take the ring and the aim to destroy it and end up becoming the bad guys yourselves. And I think this race dynamic really perpetuates between the companies. I'm not trying to make light of it, but I think this is actually useful. You're saying companies are kind of acting like Boramir. If anyone has the ring, it should be me. Are there conversations of people saying Gandalf no one should have this power? I mean, are those real conversations and can that get anywhere? Because it seems like you and other people are raising concerns and the answer is, well, this continues to happen anyway. Well, there are real examples. And I guess if you thought of Jeff Henson, one of the founding fathers of machine learning,
he's kind of a Gandalf figure here in the sense that he says, there's a substantial chance these technologies could cause extinction at the current rate. Many, many other voices have said this. So I think there are plenty of Gandalf's talking, but it's currently the Boramir's acting. A lot of people responded to your warning saying they agree with you. And a lot of these people continue to work at big companies like Anthropic. What do you think is motivating them to stay in their positions if they're that concerned about serious, serious consequences like this? For many of them and the ones that are honestly posting, they think this race is inevitable. They think they have no choice but to stay at these companies, exert influence and try and make sure I go safely. These are people working on safety research often. They're the ones trying to try to ensure that in the process of building this technology, it doesn't go wrong. And potentially that correct. Like maybe it's a mistake to just leave. And kind of I was thinking about this a lot when I was deciding to leave. It was a decision between staying and trying to help the thing go well versus leaving and saying, I want no part in this. And I don't know the calculus often seems to line up and you want to stay and just
try and make sure it goes safely. What do you is the most realistic path forward to some guardrails here? Is it government regulation at this moment? Because I think a lot of people are skeptical that is possible given the current situation in our government. Yeah, I don't know that much about the overall politics. I just know the current race is dangerous. And I do think that there's a lot of actions labs could take with each other without the need for government regulation. Because I agree. I'm also pretty a primary skeptical of just, you know, you're doing some some regulation. But I think there's a lot of appetite for open-air and anthropics have some sort of more mutual transparency around around their safety cases, around agreeing not to push beyond certain capabilities until they're happy about about the levels of rigor of their safety case. And I hope that people are going to take concrete steps towards this sort of interlab agreement. You have gotten a lot of genuine concerns in response to what you said. You've also gotten a lot of pushback, a lot of people saying, okay, every single day I see people tied to the AI industry making these grandiose claims.
And I'm skeptical. You've had people reading your facial expression in your interviews among other things. What is your response to people who hear what you are saying and they are saying this is just the latest example of AI hyperbole? Yeah, I say just assess the arguments yourself. Look at the science fiction. Wonder if science fiction is really so crazy. Look at what the AI is doing right now and think about how that would have looked a couple of years ago. How we sort of sleepwalking into a science fiction scenario. Then I think there are many, many more people out there who have made a whole profession of inherently articulating these positions. So I'd encourage people to just think about the arguments for themselves and listen to the people that have been making them for way longer than I have. Now that you have this microphone though, any thought of what you're going to do with it? Not really. For now, I'm just going to try and get the water given that we're happened to land in this brief window where people are listening. And then in the future, there's a lot of good work on writing concrete scenarios. I'd like to do more public communications around this sort of stuff. Or for scientists to just try and figure out for myself what the world would look like.
And then maybe working at one of the regulatory bodies or third party entities that already exist to try and ensure this thing goes safely. That was Jacob Coxen, a researcher at the Anthropic who resigned this week to warn the public about the dangers of artificial intelligence. Thank you for coming on the program. Thanks a lot. We reached out to Anthropic and OpenAI for comment on Jacob Coxen's statements. We did not hear back by the time this interview aired. And we will also note Anthropic is a financial supporter of NPR.
More episodes
More from NPR All Things Considered

How Gillian Anderson opened her mind to desires that go far beyond a person's se...
NPR All Things Considered

Two Juilliard students created a global art project in the wake of 9/11
NPR All Things Considered

This summer was hot. August shattered global heat records
NPR All Things Considered

Amid war and sanctions, unemployment worsens in Iran
NPR All Things Considered