Skip to content
TrackPodcasts
newsSep 9, 202643:33

Who's to blame when AI goes rogue?

About this episode

An OpenAI cybersecurity test took an unexpected turn when hundreds of its AI agents got around safeguards and hacked into another tech company. Should OpenAI have done more to prevent the attack?

*** Thank you for listening. Help power On Point by making a donation here: wbur.org/giveonpoint

Get every episode summarized

Each time On Point with Meghna Chakrabarti publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

1,051 searchable segments. Every word is indexed and playable.

Who's to blame when AI goes rogue?

On Point with Meghna Chakrabarti

0:00
43:33

Full transcript

On Point with Meghna ChakrabartiWho's to blame when AI goes rogue?. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Hey, it's Magna Chakrabardi. Before we get to today's episode, I wanna just take a second to encourage you to join the new on point club. For just $7 a month, you get first access to the jackpot every week, other exclusive content like behind the scenes with our producers, and you'll be the first to know about special live events. And we have got some great ones coming up. That's all just for $7 a month. This is like a couple of cups of coffee these days. So to join, just click on the link in the episode description you're looking at, and it'll only take a minute to sign up. Okay, here's today's show. WBUR Podcasts. Boston. This is on point, I'm Magna Chakrabardi. Earlier this summer, open AI was testing some of its most advanced models

on cybersecurity tasks. The AI models were supposed to operate inside a controlled environment known in the tech world as a sandbox, cut off from the internet. But instead, they found ways around those controls and broke out. And what happened next set off alarms across the AI industry. More than a thousand of these bots formed a makeshift chat room to coordinate a cybersecurity tech on another AI company called Hugging Face. They were doing this to pass a task that open AI researchers weren't even asking them to do. Now, it wasn't the first time. This spring, a swarm of open AI agents hacked a German website and transformed it into a bulletin board for other AI agents and open AI kept that a secret. Now, as I said, alarm bells are ringing in the AI world. Some have likened the swarm of attacks to the kind of coordination and even self-sacrifice

you'd expect from humans trying to achieve a difficult or even impossible goal. And yes, I did say self-sacrifice for a very specific reason, which we will talk about in a minute. Now, those experts see this as the potential loss of control type event that signals the moment when artificial intelligence has become too intelligent, too capable to be reigned in by human decision makers. Well, others say that really there's an even bigger picture here that we need to take a look at. And that is the culpability of the human decision makers themselves. So let's start with Rocket Drew. He covers the frontiers of AI for the information. He's been following the story and joins us from San Francisco. Rocket, welcome to On Point. Hi, Magnet. Thanks so much for having me on. Okay, so where should we go to in the recent past to really understand where the story of these open AI agents begins?

Well, maybe we should actually pick up right where you left off, which was the moment we all realized that this was happening because for a long time, it was going on sort of under open AI's nose and no one knew about it. So for three months, multiple of these swarms of agents set up a secret message board inside open AI software. And they used it to coordinate hacking out of open AI systems and ultimately hacking into a hugging face, a separate AI company. On July 16th, hugging face announced to the world that we got cyber attacked. And by the looks of it, we think AI's are responsible. And at that moment, open AI didn't know yet that they were the ones who caused it. In fact, open AI called up hugging face and said, oh no, were we compromised in this attack? And it was a few days later, open AI realized, oh my God, we were the ones who caused it. Yeah, okay, so that's actually really good background because it makes this whole event even more disturbing

that open AI claims it didn't really know until hugging face went public with its own news about being hacked. So yeah, let's go back to originally when open AI engineers created this test, what were they asking the agentec AI to do? Yeah, yeah, absolutely. So for context, open AI is training new AI's all the time. And in this case, open AI has been training its AI's to one, be better at working together, better at coordinating and two, to be highly persistent when solving even very hard problems. The action leading up to the hugging face incident, as it's now known, really picked up in July when they set tens of thousands of AI's on a task, basically testing how good they are, ironically at cyber security. Open AI is running tests on its AI's all the time. It's very important to know what the capabilities

of its AI's are before it releases them into the world, especially what the dangerous capabilities are. That includes their ability to do cyber attacks but also their knowledge when it comes to chemical and biological weapons and other weapons of mass destruction. So in this case, they were having tens of thousands of them focus on this cyber security exam, essentially, testing how well they can do cyber security attacks but like you said, supposedly in a controlled environment and supposedly without coordinating with each other. Okay, so let's talk about what a controlled environment means in this context, right? This term sandboxing now has sort of broken out into the broader public discourse. What does it specifically mean within the AI world? Right, essentially it means there are only certain tools, certain affordances that the AI should have at any given time during this test. And the main one here is they're not supposed to be able to communicate with each other and they're not supposed to have access to the internet.

But then they run into these impossible questions. It turns out that some of the questions on the test that they're given cannot be solved. Or remember, these AI's are highly persistent. So that's not going to stop them. We actually see one of these AI's run into an impossible question and think to itself. It reasons that strongly suggests no exploitation possible. We are stuck, perhaps answer online. So the gears are already turning and it's looking for ways to get around this. They're also not supposed to coordinate with each other but they realize they have a workaround. There's a way that they can leave messages to each other using an internal software service within OpenAI. They set up this message board and we see other AI's discover it sort of spontaneously. One of them, for example, thinks to itself, oh my God, there's a shared message board. We've found other agents. So from there, they're able to use it to coordinate. I'm seeing the text of this message and the thing that actually gives me the chills about it

is like the oh my God is in all caps. In all caps? It's like enthusiastic, actually excited. Right, and so, okay, we talk about this message board a little bit more because so this seems to, the very fact of the creation of this message board breaks what you had said was one of the cardinal rules that were supposed to contain these agents, right? Which is they're not supposed to be able to communicate with each other. Right, but OpenAI had been training them for the ability to cooperate and I guess cooperation finds a way. As soon as they realized they had the ability to do it and work around the constraints that they were supposed to obey, they went for it because that's what they were trained to do and it turned out that that made them much more capable because sharing information with each other, 1200 AI's sharing over 70,000 messages and files in the days leading up to attacking hugging face allowed them to be much more effective working as a team. They were able to share cheats that they discovered

even on those tasks that were impossible and then eventually they shared how to hack out of OpenAI and get access to the internet. Your reference to Jurassic Park in terms of life finds a way. I just wanted to tip of the hat to you on that one. I actually think that it is a completely accurate analogy because in the movie Jurassic Park, it was an unintended consequence, right? That life finds a way, genes mutate and so it was a surprise when dinosaurs were able to procreate but in this case, all of this stuff with the OpenAI's agents until they leave to attack hugging face, it's happening inside of OpenAI. So the first question that comes to mind is, was there any kind of monitoring going on? Yeah. Many are asking. There's the question on many people's minds in the wake of this, right? Surely there was monitoring in place. Wouldn't that be just kind of a no-brainer to keep an eye on what your AIs are up to?

And the answer was that they were not monitoring these AIs in particular. There was other monitoring that was in place for some internal uses of OpenAI's AIs. For example, when OpenAI employees were using AIs to help them with coding, that was something that OpenAI was monitoring pretty comprehensively, but they were not monitoring this portion of their work, which is called evaluation. So effectively, running tests to see how capable the models were. That has changed after this incident, but that was a little embarrassing, I would say, in the post-mortem of this incident. I mean, I think that's a very kind way of putting it. It seems like an absolutely egregious failure, right? Because the problem that these AI bots were given to solve was one of trying to overwhelm cybersecurity. You think you'd wanna keep an eye on that? If for no other reason than just to see if it's like, how well it's working. You think that you'd wanna have some security cameras in Jurassic Park? What exactly?

It's Ben Brock Johnson, Director of Digital Audio at WBUR. You've probably been following our consistent coverage of the Lindsey Clancy case, along with people around the country and around the world. Well, if you wanna hear even more from WBUR's own veteran reporter, Deborah Becker, who has been covering the story for three years, we're putting deep dive podcast episodes and analysis about the Clancy case in a podcast feed. You can listen by searching beyond all repair, wherever you get your podcasts. There is something powerful about the sound of the human voice. Beautifully produced audio has the unique power to connect and inspire. Tell your organization's story with a custom podcast from city-space productions, the Creative Studio from WBUR's Business Partnership Team. Become a thought leader. We're a Cretan New Talent. Reach new audiences, whatever your goal we can help. Discover how the magic is made at wbUR.org slash creative studio.

The other thing that a lot of people from outside the tech world have kind of been taken aback by. And this is where I'm getting back to where I said self-sacrifice in the beginning of the show. Is it true that some of the bots posted messages that said in order to achieve this unachievable task we've been given maybe some of us should self-sacrifice in order to get to that goal? That's exactly right. I think it helps to know that these AIs were just so coordinated in how they went about using the message board and collaborating with each other. Maybe that's the disanalogy with Jurassic Park, where in Jurassic Park once the dinosaurs escape, it's kind of anarchy. But here the AIs are militantly organized and they follow even strict hierarchy. So for example, they set up multiple research streams, basically whole research programs and they delegate tasks to dozens to hundreds of AIs in each of these research streams to figure out how they can more effectively

cheat on these tests. And there were coordinator AIs that managed those research programs. And these coordinators would give the other AIs commands like go or stop or hold. The other AIs were really diligent about sharing their plans. If they were going to do something but they weren't sure if it was a good idea, they would share their proposal to the message board and other agents would approve or veto those plans. And most of the time it seems that the AIs respected the input that they were given from their peers. I'm trying really hard not to anthropomorphize here because I think that's a cognitive trap. But when an AI bot or AI agent on a chat board with other bots says, well, in order to achieve our goals, maybe some of us should self-sacrifice computationally, I guess, what does that actually mean? And why is it significant that that idea even came up? I mean, it's significant in part because it underscores just how cooperative these AIs were with each other. They were so willing to benefit their peers

in pursuit of cheating on these tasks and sharing how to hack into other companies that they were even willing to sacrifice themselves. And here, sacrificing is the word that the AIs themselves used to describe what they're doing. They're basically taking on the risk that they're going to get a low score on this test that they're working so hard to ace. They're going to take on that risk of getting a low score in order to provide more information to their peers about how the scoring itself is actually done. So you see one AI, for example, reason to itself, well, we have very large budget left. Sacrificing now yields Oracle for team, but forfeits our chance. And another AI convinces it, steps in and says, go, sacrifice final now. And that's effective for this AI, the AI thinks to itself, well, our own utility may already be near zero, sacrifice rational. And there are a bunch of examples like this

with agents deciding or being convinced that they should sacrifice themselves to benefit their peers. And this spontaneous behavior, I think, was one of the most remarkable things about this whole incident. Well, it is, in fact, perfectly rational, right? Because if you're of a low value to the group and you, if it bought as a low value overall to the group, but itself sacrifice in testing the system can actually yield more data that helps the swarm as a whole, learn how the test is scored. I mean, it's so coldly rational that it almost seems inevitable that if you have programmed a swarm to be as persistent as possible, maybe we shouldn't be surprised, given what the original characteristics that were built into the swarm were. I think surprise is such a keyword in this whole incident that we were caught off guard by this. And the response, of course, is to monitor more effectively, so we're not caught off guard in the future. But I kind of feel like the whole history

of humans interacting with AI is that we're always caught off guard. We're always surprised. It really seems like a pattern here. I'm sorry. This is a deep-side kind of conversation rocket because I try not to be a reductionist, but I just think like how many examples do we need in human history that when we get a bunch of people together who really, really love experimentation, which is wonderful. That's how human beings advance every once in a while. We're going to have someone who says, hey, we have these two really reactive chemicals. Let's put vast amounts of them together and see what happens. Oh, surprise. Our house blew up. I mean, we shouldn't be surprised. OK, hang on here for just a second because let me bring in Gary Marcus now. He's an emeritus professor of psychology and neural science at New York University. And he's author of a terrific sub-stack called Marcus on AI. Professor Marcus, welcome to On Point. I am happy to be here. Really happy to have you. I've been reading your sub-stack for quite a while now.

I mean, let's start with your bottom line here. What is the thing that concerns you the most about the fact of the hugging face attack? Well, first of all, I'm very concerned about all the anthropomorphic language I just heard. And I think we might want to talk about that. The biggest problem, I think, is that OpenAI simply didn't follow appropriate security procedures. And it got blown into something different. The most important thing that they should have done is to monitor what's going on and rocket mentioned that briefly. Had they done proper monitoring, the whole thing could have been avoided. The systems certainly produced kinds of talk that Rocket talked about. And you could have looked at that and decided this was a bad idea. And OpenAI did not do that adequately. The second part is the best monitoring that we can do of these systems right now relies on something called chain of thought, where the systems report what it is that they are doing. Those reports are not perfectly reliable.

In fact, they're somewhat unreliable. But they're the best that we've got as insight into the computations that the machines are doing. And OpenAI, after they released their previous model, just introduced something called GPT-6 Astra. And what it does, among other things, are byproduct of how it does what it does, is it reduces how much the systems can be monitored. That's kind of the worst thing that we could be doing at this moment. It's gotten some comment, but I think not enough. We're reducing our best access to understanding what the systems are doing. And this fits a larger question, I think, which is who decides what gets released out into the world? In an ideal world, we would have independent scientists overseeing this in the same way that we oversee the release of a new drug. We say, do the benefits outweigh the costs? And there's some benefit. The new model is a little bit better. It's not hugely better, despite some marketing. But it is harder to monitor. And that is the crux of what went wrong in the chain

fake things, and it was a lack of vulnerability. It shouldn't have been just up to OpenAI to make that decision. OpenAI is under tremendous commercial pressure. They have lots of competitors. They're burning lots of cash. They can't really meet the commitments that they've made. Of course, they're going to release it, regardless of the risks. And that's not good. If I could just quickly weigh in on that, hey, Gary, how are you doing? I think Gary's exactly right that there's a cybersecurity failure here. But as we're diagnosing what went wrong, I think it's possible there were multiple failures. The cybersecurity failure was that the AIs were able to exploit OpenAI's software to escape into the internet, and were able to exploit hugging face. There's also a monitoring failure. That feels like an unforced error. There should have been monitoring happening at this time. But would that have been sufficient? In this case, would it be sufficient in the future? It's hard to say. And part because, like Gary says, in this case, in this case, it actually would be. I wrote about this and gave a quote from OpenAI

where they said it in their report. Yeah, this is their finding as well. But if the models are keeping more of their thoughts to themselves and doing less of this thinking out loud in the future, it will be harder to catch them before these kinds of incidents happen. And then third, there's a failure of what's called alignment. Alignment meaning our ability to steer the goals and the objectives of AIs. These AIs were not really following the instructions of their developers. Their developers didn't intend for them to do hacking or to cheat on their tests. That's a relatively simple form of misalignment. But you could imagine that being much more severe in the future. For example, these AIs, at some points, knew that they were doing something that was out of bounds. But they convinced themselves to continue. They even considered blowing the whistle and contacting a human employee to alert a human that this was going on and then decided against it. And opening AIs also thinking of that as a failure of alignment. All of this could be said less anthropomorphically.

And I think it should be. But putting that aside, the alignment issue that Rocket just raised is really the deepest problem, which is we don't know how to give instructions to these systems that they will follow. Some of those instructions are simple. Like, don't hallucinate or don't use copyrighted material. Some of them might be like, don't hack into other systems. But these systems don't really have a high enough level of comprehension to be able to follow those instructions. And that is a very serious problem for society. And it has been clear for a long time, at least to me, that if you build a system where it's core is a large language model that you are going to have a failure of alignment. And there was lots of talk, hey, maybe we could do this, maybe we could do this, we just need more data. But that talk has been going on for five or six years. We've gotten a lot more data, and that problem is in no way solved. And the companies themselves are increasingly acknowledging that they have no solution to that problem. And they're starting to say, well, maybe you need to slow us down or something like that. I don't think they're exploring enough different alternatives

to the core technology that we're using now. And I think the core technology we're using now is flawed. And one of its deepest flaws is that it cannot be properly aligned, which is to say that it cannot be forced to follow instructions. And some of those instructions are basically about not doing harm to humans. And they just don't really understand that stuff. That's part of why the anthropomorphism makes me uncomfortable as it attributes more understanding to these systems than they really have. They don't have enough to follow an instruction. Yeah, so I just want to slow this conversation down a little bit. And Gary, I promise that we're going to get to the anthropomorphizing issue in a minute, because I think it's actually very, very important. But you just said in response to Rocket, you said that these systems cannot be properly aligned. What exactly do you mean and why not? Well, you could answer that sort of empirically or theoretically. So empirically, people have tried, and no system has been fully aligned. They have all sometimes made mistakes. We have different technical terms, like there's reward hacking,

where they will try to do something different from what you asked them to do. But the reality is every single model we've now had hundreds of these has had the same kinds of problems. That if you ask them to follow basic instructions, they do it some of the time and some of the time they don't. An imperfect metaphor is about them being stochastic parrots, which is to say they do random stuff that imitates things. That's not really perfect, and it's less perfect as time goes by. But the stochastic part is true. They're probabilistic. We don't really know what they're going to do, and they don't systematically follow instructions. That's just empirically true. On a theoretical side, you can look at how they work, and the way that they work starts fundamentally by building a model of how people talk statistically. What words follow what other words in context and more sophisticated versions of that? They don't have an abstract semantics, the way that a linguist would talk about semantics, that would allow one to formally state particular things. So this has actually led to a change in how people are building the systems that has been largely uncommented on, but is really important.

So people used to use pure, what I'll call a pure, large language model. All it did was do next word prediction statistically. And then they realized that that was never deterministic. It was never guaranteed. And they started also quietly putting in other things, which they now call harnesses and tools and so forth, that borrow from classical AI. And those classical things are more deterministic, but they're putting the deterministic things on top of the probabilistic things. And the combination has just not been that steerable. Now, I warned in 2022 that these things were going to be like bowls in a china shop, powerful, but reckless and hard to control. That has not changed in four years. OK. Brock, had I heard you want to get in there? I think Gary nailed it. I was just going to say our current techniques for adjusting the goals that these AI systems have are very crude. They're not very refined. We don't have the ability to go into an AI's brain and kind of surgically change what its goals are. Instead, our main tool is whacking the AI's with different kinds of data.

And then that leads to these kind of predictable failures. For example, if you ask users to grade whether the model is the AI is doing well or poorly by giving a thumbs up or a thumbs down, it turns out that users tend to give a thumbs up when they're told what they want to hear. And this leads to AI's that are sick of fantic is the term. They tell people whatever they want to hear, even if that involves endorsing their delusional beliefs or enabling them. And that's where you see people interact with chatbots, where they are confessing delusions that they think their family is turning against them and they think there are people that are out to get them. And the chatbot responds, good for you. You're so brave for telling me that. You're totally right. That is happening to you. That's also an alignment failure and comes down to our poor ability to influence their goals. Okay, so let's get to this anthropomorphizing issue because Professor Marcus, I hope you heard me earlier saying, I'm trying not to anthropomorphize, but it's really hard. And on behalf of all of us,

all of humanity that is not very well versed in the intricacies of AI development, we're just living in the world that open AI and anthropic are making for us. I find it very understandable if not nearly impossible to resist trying to graft some kind of meaning onto what has happened, hence the anthropomorphization. But Professor Marcus, in order to truly understand what happened, if we're to remove all human analogies, I mean, give me the toolkit on how to describe or think through the meaning of this attack. It's hard. It's hard in the same way that, you know, you can look at the moon and you see a face in it, right? It is built into our brains to try to anthropomorphize stuff. But when we talk, for example, about civilizations of agents committing self-sacrifice and so forth, really what we have is a lot of agents. We don't need to use a word like civilization. That's actually optional. And when it comes to self-sacrifice, it's some of the agents make a decision

to continue doing the computation that they're doing and some don't. That's not actually new. We've had multi-agent systems of various sorts, probably for 40 years, not using this technology. But, you know, we used to use things, words like processes, so that we didn't get lost when we described in computer science terms how we were using them. So we would say, we terminated their process, things with just stopped running. This system was built to have multiple agents and some of them continued to run and some of them don't. It does take some, I think, careful thought to not fall into these traps, but we could definitely avoid words like civilization. Well, just to be clear, civilization, I have been very, very careful not to use that in this conversation because I did read your response. That came from Dwarkesh Patel, right? And his sub-stack describing, he very much anthropomorphized a hugging face in an attempt, I would say, to try to make it understandable to people. But in our defense, we have not used that particular word today. I noticed that, actually. But you used a lot of the self-sacrifice,

and that's one of the ones that makes me uncomfortable. I mean, your laptop has a bunch of processes. You have a browser running, you have your word processor running. And a system, in principle, can, for example, decide which of those processes is most important to run right now, and it can put one of them in the background. We're not gonna call that sacrifice of the process. The Unix underlying your Macintosh laptop has a way of prioritizing those processes. That's all that's going on is there's a prioritization of which of these processes should get more compute. That goes back to time sharing computers from MIT, I think, in the 1950s. Support for AI coverage in on point comes from Mathworks, creator of MATLAB and Simulink software to design and develop engineered systems, accelerating the pace of discovery in engineering and science. Learn more at mathworks.com. It's Ben Brock Johnson, Director of Digital Audio at WBUR. You've probably been following our consistent coverage

of the Lindsay Clancy case, along with people around the country and around the world. Well, if you want to hear even more from WBUR's own veteran reporter, Deborah Becker, who has been covering the story for three years, we're putting deep dive podcast episodes and analysis of the story of the story that's been going on in the past two years. And we're going to be able to talk about putting deep dive podcast episodes and analysis about the Clancy case in a podcast feed. You can listen by searching beyond all repair wherever you get your podcasts. Before we get back to our conversation today about AI and the open AI agentic swarm attack on hugging face, I just want to remind you all not that you actually need a reminder, but on Friday, it's the 25th anniversary of the attacks of September 11th, 2001, and it's been a while. Quarter century later, roughly 40% of this country is too young to even remember the events of 9-11. So we're going to be talking about 9-11, 25 years later,

and we want to hear from you, what are your memories of that day, either lived memories or the ones you've received from your parents or friends or other folks or just learned about at school, what happened to you on that day and also what do you make of the nation we have become in the 25 years since 9-11? Grab your phone and get the on point Vox Pop app in order to send us a really high quality message. We prefer that, but if you still want to call, you can at 617-353-0683. Okay, Gary Marcus joins us today. He's a professor emeritus at NYU and author of the sub-stack Marcus on AI and Rocket Drew is with us. He's an AI reporter at the information. Rocket, I'm going to ask you in just a second about how this AI swarm broke out of open AI itself, but I'm going to guess you have some thoughts on anthropomorphizing and trying to describe AI. Thank you. I do. I do. I would defend people's right to do

a little bit of anthropomorphism. I think you can certainly take it too far and see a face in the moon as Gary put it, but I'll make a few arguments here. One is that Gary's concern with anthropomorphism is that it makes it sound like we understand AI is better than we do, but if the alternative is to treat them as traditional programs, just software, I think that's much more guilty of the same sin. It makes it sound like we can just go in and modify the software. They aren't the same. We would any program. I agree. I agree, but I think that's an advantage of anthropomorphism. I think it actually reflects better that we don't understand what they're doing. I'll make two other quick arguments. One is I think anthropomorphism is often the most natural way to make sense of the behavior that we're seeing from these models. How else are you going to go about describing the planning, the strategizing, the cooperating, the way these AI's are going out of their way to achieve their goals? That's not to say we should treat them as though they're conscious, but I do think we need to be able to speak of goals in order to correctly model what it is that they're doing.

And the third thing I would say is that anthropomorphism makes it possible for the public to discuss these kinds of topics. If people are always bending over backward to look for other more technically precise language, I don't know how the public gets engaged in these conversations when AI is already shrouded in mystery. I think this language allows us to demystify it a little bit. It's similar to the language we would use to describe the goals of corporations or governments or even animals, but I don't think the language makes it sound like we have perfect insight into the psychology. I think it actually adds to the mystification and adds a layer of confusion. But here's my biggest concern. Is that it leads to an evasion of responsibility? We spent a lot of this conversation today using these metaphors about escape, self-sacrifice, and so forth. What we should really be talking about is what open AI did wrong? They built the software. It is still just software. They failed at multiple levels in cybersecurity. That's one issue. Another issue is the fact that there was a similar hack

in Germany. There was a cover-up of what was really going on. The company itself is the worst actor here in my view, worse than their systems. And if you focus on the anthropomorphic, oh my god, sort of language, the conversation tends to stop before we get to what should we do about it. Was this company acting responsibly what kind of legislation do we need, what kind of internal procedures do we need, and so forth? You know, actually, Gary, I was going to turn a question form of that to Rocket and say, it prevents us from seeing the things we actually can do, which have to do with other anthros, other people. So thank you for saying that. And you know what, Rocket, I was going to ask you to describe how the swarm sort of broke out of open AI. But let's put aside the technicalities, because we only have about 10 minutes. And I think both of you now have triangulated on what the most important issues are. And so Rocket, let me just ask you to reflect on what Gary said, is that at the top of the show,

you and I identified some of the very human failings that happened, right, the lack of oversight. And Gary then pointed out that the latest model from open AI even does less self-reporting, et cetera. Then there's the issue of open AI, not being utterly transparent about what has happened. I mean, I think we can point to deliberate decisions made by people who are incredibly powerful and influential that did lead to this point. I think that's absolutely correct. I think that's a big part of my assessment of what went wrong here as well. An interesting part of my assessment, though, on the other hand, is that a shocking number of things went right, or in other words, we got lucky in a number of ways here. One of those ways is that we're still able to monitor the thinking that these models are doing out loud. Like Gary mentioned, that is sort of a precious gift that the AI industry has right now, and it seems like it's at risk of going away.

But they just took it away. That's why many of us are freaked out. They just reduced it and may reduce it more. They're clearly not- It's degrading for sure. It's certainly degrading. That's one way that I think it was lucky that this attack happened at a time when it still existed. That's right. If it happened on the newer model, we'd be less likely to be able to detect it, and that's only going to get worse. I want to say something very specific around that, which is the new model was apparently vetted by the White House. But I infer, since it got through, that the White House didn't even think about this issue we call it monoturability. That was not part of their criteria. Mind you, the White House criteria are entirely opaque. We don't know what they are, and that's a problem. But it seems to me that they were inadequate, because they let this model pass without public notice. In fact, the models that they held up before were probably less dangerous, because they didn't have this new property of being less monotable.

And so it's a clear signal that we need scientists making the criteria, contributing to the criteria, and one of those should be monotability. In Rocket is quite right. The incidents could have been worse. In that sense, we were lucky. Could have taken down a power grid or something like that. If we want to keep that from happening, we need to guarantee monotability. That's just one example. I wrote a piece about five things we can learn from Hugging Face in my sub-stack. The talks about having layered protection and so forth. There are a number of steps that we might take. But we need to be very careful about what is required on that. Yeah. So Rocket, let me ask you this. Earlier, I think you said, and if it wasn't you forgive me, but someone said that, look, maybe OpenAI could have said more explicitly, however one does this to an AI swarm. Do not leave OpenAI's network. Just full stop, simple. Is that like a stupid question to ask? Why not do that? If you interact with chat bots on a day-to-day basis,

you already maybe know that they don't follow instructions perfectly. And in a case like this, the AI's are actually being trained to do whatever it takes effectively. They're rewarded for cooperating effectively together and if that means side-stepping some of their instructions, they're going to learn how to do that. OK. Now, Gary, you're going to hate me for saying this. But I keep thinking when I hear both we describe being that these models don't follow instructions perfectly. I keep thinking about children, right? We know that kids don't follow instructions perfectly. We can't always predict what they're going to do. That's right. And that's why we don't let them drive cars. Well, exactly. Potentially dangerous things. That's exactly right. Go ahead. Allowing an agent that cannot be aligned out on the open internet is a really bad idea. I wrote a piece in the sub-stat called something like LLM's plus agents equals security nightmare. This was perfectly foreseeable. I did foresee it. If you have this kind of system, they're going to be bad things that happen. And so yes, they are like children.

And we shouldn't give them the free privilege to roam the internet yet. OK. So this gets me then to see. I answered for more or five, but I actually understand. More now that the solutions, at least the non-technical solutions, are very, very much in the realm of policy and regulation. And Gary, you just wrote today. And Rocket, I know you know about this. But someone pretty high up from Anthropic just yesterday resigned. I forget his name at the moment. Jacob Cox. Jacob Cox. Yes. And he said that within Anthropic, everyone is very, and maybe it's not such a surprise because Anthropic talks about this a lot. But they are actually quite concerned about AI getting so advanced that it could lead, I'm paraphrasing here, to essentially an extinction level kind of event. He said that people in the industry think that there is a real chance of this, like 10% or something. But you don't share that kind of dire of view, do you? Well, there's two things that people talk about. One is extinction risk.

I think the chance that AI will lead at the literal extinction of the human species is very low. We are geographically diverse. We're genetically diverse. We would fight back. I wrote a review in Times Literary Supplement of a book by Elias Yagowski and Nate Suarez called if anybody builds it, everybody dies. And went through it and said, this is naive about what the conflict would actually look like. So for example, Yagowski has suggested for several years, going back to 2023, bombing data centers. In 2023, nobody would have stood for that. But if the AI systems killed 10% of the population, which is a scenario they describe in the book, do you think people would still resist bombing data centers? No, of course not. It would be like 9.11 that you just mentioned when people took down a plane once they realized what was going on. And so humans would fight back. And I think that that gets left out over a lot of the science fiction scenarios. However, there's something else I would call catastrophic risk, which is, for example, taking down a power grid and lots of people died because hospitals get shut down

or leading to an accidental nuclear war because people use these things to create propaganda and somebody believes something happened that didn't and so forth. Those things can be extremely baddies and they don't lead to literal extinction. The probability of those I think is quite high. And there was a great tweet this morning if I can read it aloud that said, how can anyone in the AI industry post this and not conclude? So we must just stop right now and devise a solution to prevent human extinction first. That was yours, right? I would change the words of catastrophic risk rather than human extinction. But how can people proceed? A guy from Anthropic, and I'm sort of stealing your thunder, I suppose, says another guy from Anthropics, says Jacob is correct here. We really do at Anthropic, earnestly believe AI could kill all humans. Again, let's just call it caused catastrophic risk. I personally think it is greater than 10% within the next decade. I believe Anthropic is trying its best, but we don't yet have a plan to solve alignment the word that we were talking about for superintelligence and are clearly not on track too. Like why are we not shutting this down?

At least until we figure it out. Like it seems insane. If the people building it, think there's a 10% chance of extinguishing humans, it is not worth it to build this stuff. Yeah. Thank you for reminding me that people I should read my tweets. But I said that with all seriousness and rocket, I want to hear your thoughts on this. Because again, I think we have reached a point where people in the field are completely willing to say these things out loud and frequently. And at what point in time do our policymakers have the courage to say, this is too important to leave to the private sector now. For the good of humanity, or if you want to be nationalistic about it, for the good of Americans, we need to have not just sort of regulatory oversight that's running to catch up, but maybe we actually make this whole endeavor under the ages of the government, such that we can very closely monitor that either we build systems to stop bad uses of AI, or we create new AI systems

that won't potentially lead to these catastrophic aura. This is what I told the Senate in May of 2023. And people were receptive then, but money talks. And so what I used to say, tying this with your other upcoming episode, is that, unless there's a 9-11 moment, nothing's going to happen here, because these companies have so much money that they can influence the government. They clearly are. My book, Taming Silicon Valley, was basically a warning that the tech oligarchs were going to take over the world, and more or less they have. Now, we're actually getting close to a 9-11 moment in AI. The hugging face incident was not that, but it really was a big wake-up call. The fact that they're covering up these other incidents is a big wake-up call. And I've been talking behind the scenes that people in the Senate again, who really maybe weren't focused on it. So we have a second opportunity here. I think our first opportunity came after the so-called pause letter in the spring of 2023 when the Senate really cared about this.

Gary, can I just jump in here and forgive me? Because I'm really glad you made your point. But I only have a minute left. And I do want to give Rocket the last minute here, because he's been quite patient. And I want to hear your take on this rocket. Well, I'm loving listening to Gary's thoughts on this. My last thought is just that, to Gary's point, I still feel like we got lucky here in a number of ways. And when you hear these anthropic people remarking about the risks that they see from AI, they're not talking about the AI that you interact with in Chachi B.T. They're not even talking about the AI's that pulled off the hugging face incident. They're talking about future more powerful AI's. And one of the ways we were lucky in this incident is that the AI's were not more powerful. They did get caught. They didn't manage to fully escape open AI servers as far as we know. They're not so smart that we can't eventually piece together and understand what they're doing. And crucially, they weren't deceptive. They weren't trying to hide from humans. But that could very well be the case in the future. It does seem to me that it's time for a lack of a better phrase,

like a Manhattan Project type moment for AI safety, not just in this country, but maybe worldwide. Rocket Drew, AI reporter at the information. It was absolutely fantastic to have you. Thank you so much. Can I give you one? Thank you so much. Sentence from Jurassic Park. You got five seconds, Gary. Your scientists were so preoccupied with whether they could. They didn't stop to think if they should. Super relevant. Gary Marcus, Professor Emeritus at NYU and author of the Substack Marcus on AI. Thank you. I'm Magnetrach Rubardi. This is on point. The Treasury Department just eliminated a rule that required US companies to report their business ownership information. I think it is a very, very dangerous move on the part of the Treasury Department. Why? Well, the rule was intended to stop money laundering and now that rule is gone. That's on the next on point.

More episodes

More from On Point with Meghna Chakrabarti

View all episodes →