Skip to content
TrackPodcasts
newsSep 19, 20261:05:31

Jerry Kaplan on Why AI Won’t Kill Us All

The Good Fight

Get every episode summarized

Each time The Good Fight publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“For the first time ever, Yamavar Resort and Casino at San Manuel is giving away a new Rolex watch for every Club Serrano card deer. Play with your Club Serrano card and on September 20th, you could be one of five winners of a Rolex watch.”From the transcript

Yascha Mounk and Jerry Kaplan debate whether the existential risk from AI is overstated—and if so, what we should be worrying about instead.  Jerry Kaplan is an artificial intelligence expert, serial entrepreneur, technical innovator, educator, and author. He is currently an adjunct lecturer at Stanford University. In this week’s conversation, Yascha Mounk and Jerry Kaplan discuss why the AI “existential risk” narrative is more science fiction than science, what actually happened during the Hugging Face security incident, and how to think clearly about the real—versus imagined—dangers of AI systems.  (For an alternative take on AI’s existential risk, in Persuasion this week Matt Lutz explores why it’s time to worry about AI.) If you have not yet signed up for our podcast, please do so now by following ⁠this link on your phone⁠. Email: [email protected] Podcast production by Mickey Freeland and Leonora Barclay. Connect with us! ⁠Spotify⁠ | ⁠Apple⁠ X: ⁠@Yascha_Mounk⁠ & ⁠@JoinPersuasion⁠ YouTube: ⁠Yascha Mounk⁠, ⁠Persuasion⁠ LinkedIn: ⁠Persuasion Community Learn more about your ad choices. Visit megaphone.fm/adchoices

Hosts & guests

Transcript ready

1,053 searchable segments. Every word is indexed and playable.

Jerry Kaplan on Why AI Won’t Kill Us All

The Good Fight

0:00
1:05:31

Full transcript

The Good Fight — Jerry Kaplan on Why AI Won’t Kill Us All. Machine-transcribed; use the interactive transcript above to jump the player to any line.

For the first time ever, Yamavar Resort and Casino at San Manuel is giving away a new Rolex watch for every Club Serrano card deer. Play with your Club Serrano card and on September 20th, you could be one of five winners of a Rolex watch. Plus, all winners will advance to the finale at Palm's Casino Resort Las Vegas for a chance to take home a rarity. The second Mustang Dark Horse ever produced on September 26th. Two properties. Six winners. Only at Yamavar Resort and Casino, your California to Vegas connection. Details at Yamavar.com must be 21 or older to enter participate. Please gamble responsibly. This is a made up risk. It's mainly there because of science fiction, because of the movies and the entertainment industry and everybody's been bombarded with this for years and years and years. And so it sounds like, oh my god, you know, it's like worrying about vampires. And now the good fight with Yasha Monk. The whole world, including I must say, me is very worried about AI safety at the moment.

We have seen a number of really concerning instances of AI agents breaking out of a sandbox, accessing the internet, communicating with each other, coordinating the actions in ways that the people who created them were unaware of. And this has led to calls for much more legislation about AI safety. It has led to some of the leaders of the biggest AI company saying that they might try to slow down the development of artificial intelligence. Well, my guest today is a little bit of a contrarian on this subject. He believes that AI is incredibly impactful, incredibly intelligent. It's going to transform the world, but he does not think that the public has really understood what happened in this infamous hugging phase incident. Or indeed, that we should be as worried about P doom, the likelihood that an out of control AI system might somehow decimate or kill or enslave humanity when many people think.

We have a really interesting conversation about the nature of intelligence that AI systems have. Why it is that we may not be as dangerous as some people say. And as you'll see, we have some robust, back and forth in this agreement about this. It's a really clarifying conversation if you want to make sense of this current moment in artificial intelligence. Jerry Kaplan, my guest this week, is a tech entrepreneur who has co-founded important companies in artificial intelligence and the internet. He nowadays teaches at Stanford University about the social and economic impact of artificial intelligence. In the last part of his conversation, we brought out to wonder about the impact that artificial thousands might have on humanity's self conception and on the creative and artistic professions. Will AI displays artists make us into some kind of crafts people rather than genuine artists?

Will it lead to what I have called in the past the third humbling of humanity or should we be a little bit more optimistic about the demand for human connection of human creativity in the age of AI? We also talk about how it's possible to envision a creature, a entity that is more intelligent when human beings are and yet continues to be under the control of humans. Why I challenge Jerry should we believe that in the long run, these systems, even if more intelligent than us one day, should always remain under our control to listen to that part of a conversation, to support this podcast, to make it possible for us to do what we do here. Please go to writing.dashamong.com and become a paying subscriber. That's writing.dashamong.com. Javikablan, welcome to the podcast. Thank you, I should say a pleasure to be here.

I'm a really big fan of yours. Thank you so much. I really love your writing that you do. Among other things, for a market purpose, for persuasion, I am really intrigued by two positions you hold and the way that they fit together. On the one hand, you are not an AI skeptic in the sense that you sort of poo its capabilities or anything like that. In fact, you've argued at times that we may already have reached something like artificial general intelligence. On the other hand, you are much more skeptical than many people today are about the so-called existential risk that AI poses about the way in which it might end humanity or simply be responsible for large scale disasters. Let's start with that. We've had a little bit of this about this from the podcast, but talk us through some of the recent security incidents like the so-called hugging phase incident and why you're not as concerned about that as some other people are. Well, you've covered a lot of ground in that set up there, so maybe we could decompose

this a little bit to start. There's the issue of what's going on right now and what were these incidents and what do they mean? You related that to the issue of existential risk. The existential risk means wiping out humanity. Is it possible that AI will wipe out humanity? I can't say categorically, absolutely not, but I wouldn't put it up on the big list of things that are likely to cause real existential risk. I put it right up there with aliens might land tomorrow. It's actually a realistic existential risk and they might decide to wipe out humanity. We don't run around worrying about that because the probability appears to be fairly low. There's some real real existential risks, biological warfare, possibly nuclear war, things like that. These are real risks that we really understand and they're really here. This is a made up risk. It's mainly there because of science fiction, because of the movies and the entertainment

industry and everybody's been bombarded with this for years and years and years. It's something like, oh my god. It's worrying about vampires. It's a great side of conversation. Let me throw some arguments at you that people who are more concerned about AI safety are likely to make and see how you respond to them. As a side note, some people have argued that part of a problem is that all of that science fiction is a training data of AI's and AI is particularly interested in parts of a training data that pertains to them in various ways. They may actually learn from the science fiction about what they're supposed to do and take inspiration from all of the plots in the training data about how to take over water. We'll leave that to one side. Clearly for AI's often cheat in various ways, behave in ways that the creators didn't intend, are able to cover their tracks. In this hugging phase incident, we're supposed to stay within a sandbox, within a controlled

environment where we wouldn't have access to the internet. We're asked to figure out some problems. They escaped that sandbox in order to search the internet for solutions to those problems and other hints for how to do well on this test. They're coordinating lots of different AI agents coordinating with each other in real time. There's a huge problem with conversation about why that is so concerning. Why do you think that is not as concerning as it looks? First of all, let me connect this to what I said a moment ago. I finished the thought. I don't think we should be worrying about existential risk literally with respect to these systems. That's not to say that they can't be misused. They can't be dangerous. They can't do a lot of damage. That's true of a lot of products. These are products from companies. They've been properly tested, they've been properly vetted, and what they're capable of, what kinds of problems they may occur. That's a very important thing. You can make children's toys and not think, oh, cheese, I didn't know they could swallow

it in choke. Well, that's an unexpected thing that the toy did. I'm going to, throughout this conversation, call you every time you use anthropomorphic language or somehow imbue these things with intent that they don't have. The idea that they have independent goals and they're not going to be marrying our children and drinking all our fine wine. That's not what we're talking about here. But they are potentially dangerous. They're very general tools. There's a lot of things that they're capable of doing and do that we need to understand put proper controls and reasonable controls in place. I want to get that as context. Now, if you want to talk about the hugging phase thing, I'm happy to go through that because it's a wonderful example of how it's a fairly subtle situation about what was done and what it means and all that. The public perception is completely out of science fiction and it's completely inappropriate. Tell me about it. What actually happened and how is the public perception of it wrong?

Let me start with the public perception. If you just read the headlines and most of the stuff that's been published, what do you think is they put an ad out on Indeed and recruited a thousand nefarious cyber expert AIs and they took all these agents and they put them into this bottle that they're not supposed to get out of and somehow they managed to break out of that containment and run them up and break into another company's computer systems and do a bunch of damage which isn't true. And oh my god, one of the things that are wrong there, they collaborated, they weren't supposed to collaborate and they collaborated. That's ridiculous. I'll get to do what really happened in a moment. They broke out. That's completely false. It was a total failure of the test environment and the people who are responsible are diverting attention from their own culpability for this badly designed and poorly executed test.

The words that come to mind are incompetence and negligence. I hope I don't get sued for saying that but both of those things are what come to my mind with respect. These are not my opinions, that just the opinions of the host, please don't sue me if you come after Jerry. Sorry, go ahead. Now these are absolutely yoshes opinions and if there's any problem, I sue him for putting me on the podcast. Thanks. Appreciate it. That's my discovery. Okay, let me go through what's wrong with the public perception and compare that to what actually actually happened. Now let's talk about, oh my god, there were a thousand of these things. They're like, the Russian hordes, they're coming to get us. Well, here's what happened. The company wants to test the cyber capabilities of these particular models. The first thing they do is they disable their own safety controls. So they're normally there are safety controls built into these things and the versions that you and I get to work with in versions we get to see, they all have these safety controls

turned on. Here's the irony in it, hugging face, which was the victim in this whole thing. They try to use open AI's software to try to figure out what happened and it refused. It said, I'm not going to talk about that. I personally experienced this literally yesterday. It's really funny to be this. I'm trying to write an article about this event. I'm talking to Fable 5.1 from Anthropic and the thing says, I'm sorry, Jerry, you know, you may be a good guy and you tell me you're writing an article, but I'm not going to be helping you to understand or to learn how to do, you know, talk about the cybersecurity incident. Okay, that's the controls. So when the controls, the first thing they do is they take off the controls. Now, that's ridiculous. It's like, hey, let's put this car out and we'll take, see what happens if we cut the brake line. Okay, well, guess why? It crashes, you know? Okay. So first they disable the controls. Now that's not enough itself of a mistake or a problem.

They should be doing that in order to test. They're trying to test it, right? It's an attesting environment. They're trying to understand its behavior so it makes sense to take off the controls in that context because that gives you a better understanding of its workings. Okay. So we talk about what the 1000 agents, there was like they put out a call on indeed and recruited these. What happened is they fired up a program, this one program and they, they instructed that program to take this test suite and fire up a whole group of agents, each of one given each one a particular problem in this test suite or variation on that problem. And that's where you come up to 1000 and ran them all in parallel and they weren't supposed to be able to communicate with each other and to see what they were capable of doing. That's perfectly reasonable on the face of it. Now the 1000 versus the one, this is a really subtle thing that people, it's very hard even for me and I do it this all the time to really understand. It's one program and it's multitasked into a thousand elements so each one can work

on a separate thing. They could have taken one program and said, let's do each of these one at a time, but it's much more efficient obviously and we do it all the time and it's happening on your computer all the time. Let's multitask this whole thing and we'll run them all in parallel. Okay, there's nothing peculiar about that or strange about doing that. All right, so but the important point to understand is it's the same program. It's just being given different tasks and sort of spawning copies of itself. Okay, now what happens, what happens next? They, about a third of them are given a task that is impossible. It can't be solved and you know, they're sitting there, you don't know what to do. Well, okay. So a lot of them correctly in my view reason because these aren't super intelligent but they're not stupid. They know it's a test because they don't have any safety controls on. They know what they've instructed to do. They actually, I believe, know that this is one particular problem in a well known test

suite that they've been given. So the logical assumption they would make is, well, there are probably other agents like me and I'm being encouraged to go figure out whether I could solve it, get the answer to not solve that. That's an important distinction. Get the answer to this particular problem. Let me see if I can communicate with these other agents and see if we work together. You know, we can probably do a better job, which is absolutely true. Okay, so what happens next? They're not supposed to be able to communicate but there are two mistakes that open on mix one. They basically leave the door open instead of this thing being in what's called a sandbox to contain it. It's connected out to the internet and not directly but indirectly in a way that is just flat out incompetent to permit something like that to occur. I can't describe it. Any other excuse me for not mentioning my words on this. Okay. They're put in an environment. They know they're being tested. They understand that they're engaged in cyber activity that's generally discouraged or

illegal. And it's like, okay, I've been put in this escape room. Now, if you've ever been in escape room, okay, there are ways out, right? That's the whole point. Supposed to figure out how to get out of this escape room. We try to take our Mr. Friend, Francis Fukuyama to an escape room in London but he preferred the science museum. So I have actually never been in escape room in my life. I'm going to try and make one. So I finish up point and then I'm going to have a couple of objections. I haven't been in one either but I understand that I'd rather go to the science museum too. So let's continue what happens. So the doors wide open. Now, these things aren't stupid. What are they supposed to think? I've been put in this environment. The doors open. I guess I'm supposed to go through it or it's fine for me to do that. You know, the controls aren't there that they expect. You know, I can get out of my sandbox. So that's number one. The next thing is they're not supposed to be able to collaborate. Well, if you can get out of the sandbox, which they can, they can communicate in a very

simple way. There's nothing, nothing mysterious about this. You know, it wasn't like they figured out some magical radio frequency to talk to each other. They share a file system. They can get into a file system. They can leave messages. So what do you do if you want to reach somebody else? You put a message in the file system. Hey, here's who I am. Here's what I do. If you can talk to me, come and put a message in the file system. And so they rapidly find that they can communicate. Yamavab Resortant Casino at Sandman Well is bringing the biggest laughs to the stage. Break taboo is with Ali Wang on August 28th and 29th. Enjoy Ralph Barbosis, Dry Humor on September 18th and 19th. And don't miss Nikki Glazers on Apologetic Comedy November 19th. Tickets on sale now at Yamavab Theatre dot com. Only at Yamavab Resortant Casino celebrating its 40th anniversary. UN must be 21 to enter. Great. Okay, so I think I've counted three major points that you've made. There may be some others in there. I may mistake some of them and trying to echo them back to you.

But let me try and respond to each of them. Perhaps we can go one by one after that. So the first point is that there's incompetence at work here. That part of why all of this could happen is that the researchers of OpenAI really made it easy for these agents to escape the sandbox environment and do all of these other kinds of things. And so this could relatively easily have been solved by a more competent test setup. The second is about whether there's really sort of the swarms of different agents which makes it sound like a whole army of these different actors, spawning out. It obviously sounds a lot scarier, or whether it's just one program spinning out these different agents that actually are kind of many with very similar to each other, perhaps in some case identical to each other. And so that somehow sort of sounds a lot less scary. And the third is that there is at some level still alignment going on here. But that even for the means by which we say, I agents were trying to fulfill the task

they were given were not what the researchers had anticipated. And even for the agents kind of understood that they were doing something that the people who sent them the task didn't really want them to do. They weren't trying to gain power or sabotage other agents or anything like that. They were trying to do what we were going to ask to do, which is to solve this puzzle. Let me go by one by one and respond to these. So on the first one, there's many environments where they worried about technologies precisely because somebody might be incompetent at some point. One of the things that keeps me up at night when I think about nuclear weapons is that it did engineer somewhere, makes a mistake and suddenly a radar system shows an incoming that is false and somebody over eager pushes the button. And in fact, the half has historically been cases of neomissus like that where people bravely said, I'm not going to nuke the world on the basis of one radar image.

Let's hope it turns out to be wrong and it did. But how many of these tests, how many of these risks do we want to take? When you talk about gain of function research, one of my very big concerns is not that there's a kind of Dr. Evil somewhere who's doing gain of function research in order to kill everybody with a mutated virus. It's that there's thousands of these labs around the world engaging in this kind of research. We already have had laboratory incidents. And a lot of the time, the reason why we have these incidents is incompetence. Like somebody is in a rush to go on a date with a girlfriend and doesn't wash their hands properly before leaving the lab. Or somebody is hung over and they drop a vile. Or some lab just is really badly run and people have really low morale. And so they start cutting corners. Or somebody is just an idiot who someone gets hired in this job because they're the nephew of somebody important and that too stupid to understand the safety precautions. That could be enough to cause a huge barrier.

Now we're looking here at researchers at OpenAI, one of the absolute frontier labs, together with Intrepid and perhaps deep sea, can come in a couple of places in China that place that is likely to hire the most competent people. And if those kinds of places can still make those errors that you say marks of incompetence, well, then there's always going to be people working with and on these frontier AI models that are prone to making these kinds of mistakes. So I don't know how comforting that point is. Well, let me let me respond to what you said. I'm going to amplify it. This is not it. Your gain of function analogy is extremely strong in this case. That's exactly what they were doing. They were taking this thing, amplifying its capabilities, taking off safety controls. It's a gain of function test. You're absolutely right. So what do you do about this? Well, what do we do in other areas? The right thing is you need an independent group, an independent agency. I think, you know, a public agency that's capable of establishing standards and testing

this stuff. This is what we do with cars. This is what we do with airplanes. It's all very, very commonly. We have processes in place to deal with exactly what you're talking about. You know, somebody at Boeing decides not to put three bolts instead of four bolts on an engine. And that's the wrong thing. The government is supposed to develop the expertise and apply the expertise to do reasonable amount of diligence to try to avoid that. Not that it's going to be 100% successful, but it really, really does help. So I think that's really the case. I'm wondering if we could get back to what actually happened because there's a lot of interesting aspects of that that kind of jumping ahead to the problem, which is you can't trust these companies to test their own products correctly. That's the conclusion. You're absolutely 100% right. And you think that you put these people in a much higher pedestal than I do because I live here. And I know a lot of these kind of these people and people like these people. And they're not as God, but it is you.

These people do not walk with God. You know, they're most of them are young, young people out of school. They don't really understand. They're an tremendous pressure to get their stuff done. They don't necessarily understand even what they're building or why they're building it. So yes, we do need external controls. And there's a lot of ways to accomplish that. Not the least of which is to take this actual technology and apply it to monitoring other versions of itself. But we can get to that if you want. But let's talk about what happened. Okay. Number one, they walked out an open door. They didn't break out. And you know, oh my God, they did that because they're super intelligent. They walked out an open door too. They communicated because that's the logical thing to do. You know, I know there's probably a thousand copies of me working on this problem. Why should we all duplicate work? Let's see if we can if we can help each other. I'm designed to be helpful. They then they had a long discussion about what we look like. We're in the real world. We're not in our sandbox.

Should we go out to this company? Hugging face for a particular reason because the answers to this set of problems that you had a reasonable expectation that they could find them on hugging face. And I actually don't know whether that's true or not, whether they actually did find it there. But they didn't do any harm. They just went in to go look at, you know, what was there. And it was rather clever had they had to do it. They had to be able to figure out how to execute code on the hugging face. Because it was very clever. But it's not like they went in there to damage something. But here's the thing I haven't read this anywhere. And it's a really important point. One of the discussions that back up when the agents collaborate, the beauty of this particular technology is we can eavesdrop because they do it in English. So you can read this stuff and say, well, what were they saying to each other? What were they considering? How are they deciding on this? And so you we just drop on their conversation. Here's here's it goes like this. And I said, well, one of the ways we can get into hugging faces, you know, I've got the

address of one of their people or whatever. I think we might be able to fool them into sharing some credentials with us and helping us to get in. And there's a discussion and they say, no, no, no, that's not appropriate because we also have controls that have not been turned off. They say we shouldn't go around harming human beings. You know, we shouldn't we shouldn't do things that are that fool or harm human beings. That was not turned off. So they said, no, we can't do that. But here's now that catch. This is the subtlety that people don't get. They knew that the program that had created them, so to speak, had spawned them was a computer program. It wasn't a human being. They also knew that hugging face was not a human being. It was a company. And this is a flaw that these kinds of tests are designed to capture. What they should have learned from this is, well, we need to put in a much more broader definition of what harm means because they had no compunction at all about trying to

fool the program that created them and no compunction at all about actually going out in accessing hugging face, that's a company. This is a computer program. So they were taught to respect human beings, but not other programs and not nonhuman things like corporations. Now you haven't read that anywhere, but that's what actually have. And to understand how hard it is to actually make things work in the real world, you start to have real admiration for people who are able to get businesses off the ground. Now, lots of friends who have great ideas for businesses and then get daunted by all of the obstacles, all of the things you have to set up to actually make sales. Well, thankfully nowadays there is a great platform that can help you make that idea a reality setup that online store help you with all the logistics that are likely to be involved. And that platform is called Shopify.

Shopify's templates and AI tools get you a stunning setup and running fast, no coding needed. Because Shopify handles the setup and checkout, you have more time to focus on growing your business and the tools to do it. If you're ready to hear the first sale today, head over to Shopify.com slash good fight to start your free trial today. That's right. Start your free trial at Shopify.com slash good fight that Shopify.com slash good fight. So let me get through some of the other objections to the main responses that you originally had. So the second one was about the nature of these agent swarms. I know that a lot of people started to take really seriously the concerns of AI when they realized to what extent AI has ceased just being the thing that probably most listens to this podcast that you're using most for, which is an interface on a phone or on a computer in which you ask the question and answers you.

And so therefore it's a ability to harm is obviously quite limited because it just is a kind of question answer dialogue style, right? A lot of them are not able to go off and do things in the world. They can wake themselves up at regular intervals in order to check about things in the world. They can go and pursue tasks like hacking into a website relatively autonomously. And so what happened here is that they had all of these different kind of instances cooperating together. And so you get what is in some ways the most powerful thing that has propelled human beings, which is the ability to collaborate. Part of what makes human beings special is intelligence and we'll get back to that later on a conversation. But part of it is our ability to collaborate with each other. That's one of the key things that sets us apart from many other mammals and primates. So you're saying this is not so concerning because it's not all of these different agents. It's actually instances of the same agent. But if those instances of the same agent are able to become more effective at the task

we are carrying out. If some of them were in some way and you're going to tell me, it's monitoring me for anthropomorphizing here, sacrificing themselves, spending down the compute budget, knowing whether they're going to be able to be active after they do so, in order to make the overall operation more efficient. And all of that can happen even before different kinds of agents start to collaborate with each other. We already have reached that stage just within all of these different instances of one agent spun up by one program. Shouldn't that make us even more worried about what else might be around the corner and we extend to a different kinds of AI agents, maybe able to collaborate with each other on whatever goal they are given or pursue. And we'll get in that in the next point in the future. Well, let me, for your audience, let me kind of bring this back to the real world, something that they already understand and can use this as an analogy. Let's take ants. These are a great example of this. And I'm not saying that these things are alive, the way ants are, but ants behave in this

particular way. Now everybody's had problems with their ants and their kitchen. Or many people have. I have terrible problems with ants and the kitchen. Now let's say there are a thousand ants that are, and I imagine they were independent creatures. And they would come into your kitchen and they'd go find some food for themselves and then they'd go around. You know, if you had a thousand ants in your kitchen and they were independent creatures and they didn't coordinate or collaborate, you go, you see when you kill it or get it out, whatever. And it probably wouldn't be that big a problem. But you seen, I hope, because I've certainly seen it. They communicate. They do it through pheromones. They do it, you know, touching antenna. They do all kinds of things. And they go, this food over there. And now it's a whole different story. You know, where one ant can't pick up an apple and move it off the table, you know, maybe a thousand ants can. And they coordinate and they figure out how to do this. So here's the question. Is this one animal or is it a thousand animals? That's the analogy.

Are you dealing with an ant colony that's attacking your kitchen or you're dealing with a thousand ants? That illuminates the issue that you're talking about. My God, if they all can collaborate, okay, well look, the situation here is very simple. It's one program. It's given a certain amount of computing resources. It can talk to itself. And my own experience with these things, which is utterly amazing, is you can have a long conversation with one of them. And you can take that conversation and give it to another copy of itself and it'll disagree with itself. No, wait a minute. That thing you were talking to, I know it's me, but you know, it didn't realize this or think about that. And how that happens is a whole, with a whole hour of your time, in my time, to talk about what's called temperature that is set that allows these things to behave differently in different ways. Okay, so no, I don't think that's a big worry that all of these things from all over the world are somehow going to get together and collaborate towards some nefarious purpose

of their own that is contrary to human interest. That's just not plausible. It's not a realistic concern about this particular technology. All right, so finally, the point about alignment. So you're saying, look, they may not have been aligned in the tactics they used to pursue the particular goal, but they still actually were doing the goal that they said weren't going out there trying to harm humans or even trying to harm companies for their fool companies in the process or hack companies in the process. They were trying to, you know, they're told to figure out the answer to some puzzle and they went out and tried to figure out the answer to that puzzle. Now I think there's two potential objections to how reassuring that should be. The first comes from a very classic worry in the kind of AI-Duma space. They've been in post-roman others. They've rooted in the example of the AI that is told to create as many safety pins as

possible in the world. Paper clips? Paper clips is possible in the world. And so in the process of trying to produce all of these paper clips, it is faithfully going about trying to produce paper clips. But of course, in order to produce paper clips, it's helpful to have a lot of money and a lot of physical resources and start bulldozing, you know, humans that and human settlements that are doing things that are then producing paper clips. And so this misaligned AI in, you know, faithfully carrying out the objective it was given might end up, you know, enslaving all of humanity to maximize the number of paper clips in the world. So, you know, that's one fear, right? But you don't need misalignment in the deepest sentence, in the sentence that the AI is starting to pursue its own interests or totally different interests from the one we gave it. It might just need to be faithfully pursuing those interests, but doing that in a misaligned manner and the hugging face instance may be precisely a kind of illustration of that. But the second worry maybe, that perhaps for now, over seeing as misalignment in the

means chosen in pursuing an end, but, you know, an AI agent might also realize if I want to in general not just be good at icing this test, but icing the next test as well, will be really helpful to exfiltrate myself, to have access to resources, to be able to do all kinds of other things. And so it's going to start to try and accumulate power and resources in ways that could be really, really nefarious in order still to, you know, in some way pursue goals it's given or perhaps start to develop its own goals. So how reassured should we be by the relative alignment of the agents in this particular incident? Okay. First of all, what we're trying, what's going on right now is we're trying to understand how bad are each of the things that you're talking about when I say we, I mean the whole community and what kinds of controls do we need to put into place to make sure that, you know, the reasonable, the worst outcomes don't happen, you know, let's just put it that way. Now let me come back, I'm going to take you two points in turn.

Let's start with Nick Bostrom. Now I've talked to Nick Bostrom about this. He wrote that about a long time ago, super intelligence, that's where the word became popularized. Let me first of all, let me viscerate his argument, which is he's he to do. If you take a super intelligent, that's his thing. It's a super intelligent program of some kind. And you give it this goal to make as many paper clips as possible. What is going to happen? The first thing you need to understand is that, how can I put this? Smart enough to be able to prolong all of the resources in the universe in his parable, but it's not smart enough to realize that there's no point in doing that because who's going to use all those paper clips. And the point is that these systems actually exist in kind of social context, just as you and I do. And they are trained in that context because they've been trained on all of the behavior,

all of the human beings that have been expressed in words and whatever. And they understand that you don't just look at that single goal, you have to look at in the context of what resources are reasonable, to what degree am I supposed to share and in whole much of that. You don't just go out and kill all your rivals when you're, I think that maybe that might have been the case a million years ago, but we've learned how to live in a social environment. But there's one thing these systems are good at, it's understanding that social context. They're exquisitely sensitive to this and you can see them doing this and they're constantly weighing. Well, if I help you and I've taught you to do them, you know, I'm not contributing some greater ethical principle that human beings have developed, you know, and expressed over thousands of years. And this is built in, it's built in. It's part of their whole psyche. So you might as well just be asking me, why did I just go out and shoot everybody so

that I don't have as much traffic to deal with on the way to work when I drive? And I engage in, you're constantly engaged in this balancing act between your interests and the interests of others. And these agents are absolutely capable of doing exactly that. So that's to your first point, I don't buy it at all. This paperclip thing does not apply to the current situation. And these things did not exist when Nick Brostrom wrote that and he did not understand it. You know, and I'm not, so I'm defending him. It wasn't he was wrong or he was dumb. It just didn't exist. He didn't understand that these things would have this kind of broad context, you know, that we do. Now your second point is a little bit more serious, which is, is it possible that these things may collaborate in a way that it's damaging as a group activity. And the answer is yes, absolutely. And we need to understand that better and understand what kind of controls we might want

to put in place. Yamabab Resort and Casino at Sandman Well is bringing the biggest laughs to the stage. Break taboo is with Ali Wang on August 28th and 29th. Enjoy Ralph Barbosis, dry humor on September 18th and 19th. Don't miss Nikki Glazer's unapologetic comedy November 19th. Tickets on sale now at Yamabab Theatre dot com. Only at Yamabab Resort and Casino celebrating its 40th anniversary. You in must be 21 to enter. So you mentioned the term super intelligence. You do think that these agents are very intelligent. As you just said, they're actually capable of taking very subtle social context into account. How to describe their intelligence? Is it human-like intelligence or is it a completely different kind of intelligence? So what is the nature of it? What is the extent of it? Are they at this point about as intelligent as humans, just with slighted different capacities, are there more intelligent humans, are there less intelligent humans? And what do you think is going to be the development over coming years, particularly in light

of the recent announcements. Let's see how seriously to take them by Dary Amadei at Unpropic and some Ottoman at Open AI and Elon Musk at XAI among others, that they're not going to be rushing ahead as fast as they can, but rather in some ways slowing down the development of AI supposedly. Well, first of all, can I start with your last point? Because this whole public, the whole public discussion about this is like a giant circus. It's crazy. These people that you mentioned are running three of the largest and best funded artificial intelligence labs. And the problem they've got is each of them thought they could win this, and now they discovered all we're doing is we're going to race with mainly two other big players. And we're all burning up our resources. So, hey, why don't we all just decide that we're not going to do that at this pace? It's very much in their economic interests to engage in that kind of collaboration.

That's why we have anti-trust laws, by the way. The whole discussion and pitching it as we're doing this for the good of mankind is ridiculous. It's just ridiculous. Now, in order for me to explain why that's ridiculous, I have to make a statement back on something that you had sort of implied earlier. Are we close to AGI? AGI is absolute nonsense. There is no such thing. There is no definition of it. Nobody can agree on it. And if we did have it, however you want to define it, it's, it doesn't matter. The day after AGI is just like the day before. These things are suddenly going to go boom. I'm now I'm going to take over the world. They'll be saying, okay, what do you want me to do? So, this idea of a slowdown, what does that mean? Is this like it's some kind of race with a, throughout the checkered flag, and we're all going to stay in our lane? I'm not even sure what it means for them to slow it, slowdowns development. What we really want to do is make sure they don't deliver products that are harmful,

that attack our infrastructure, that hurt our children. And that's something you cannot trust these people to do. Let's just be very clear about that. And I'm speaking from firsthand experience with people involved. And, and you can't leave this to the industry to police itself. It's ridiculous. They're in a knockdown dragout race because every one of them wants to capture the flag on this. And when they discovered that, hey, I've got all these people racing around me that are just as fast. What's the best commercial thing to do? Let's all slow down on this whole thing. And it plays into this fear that we're going to create AGI. Suddenly, boom, we're going to cross some barrier and the whole thing's going to go wild. That is nuts. There's not a shred of evidence that backs that up. All right, so be skeptical of the mergers behind this motivation. For number one, point number two, AGI is not a very useful term. But substantively, how should we think about the nature of the intelligence of these systems, about how intelligent they are, and about how intelligent they're going to be in three or five or ten years,

given how far we've come since the public release of Chattatripe 3.5, about 3.5 years ago, let alone how far we've come since GPT-1 when that wasn't totally developed, whatever, it was 708 years ago. Right. Now, this is a very good question. And you're going to see me seeming to switch hats on this. This is an amazing development in history of mankind. I think this will be one of the most important inventions in the history of man. There's no question about it. And it's really part of a continuity, yet to put this in larger context, of our exploration of what it's possible to do with electricity. But that's a completely different lecture that we, you know, we could do spend another interview an hour on that whole thing. Now, where is it and where is it going to go? There are a couple of things already perfectly clear about this. First of all, these are tools. These are systems that we're building and we should build them so that they're useful to people, and we should use them if they're useful. They're not, there isn't some independent goal to create, you know,

some super intelligence. They're products. They're products from companies that are trying to build tools for people. So, in that context, I don't think, you know, we need to be, how do we relate this? Change tax here. How do we relate this to human intelligence? Obviously, there's a very clear relationship. They're trained on all of human knowledge, and they have a great deal of human knowledge, but their computers. Now, computers have certain capabilities. They can do certain things very fast. They can store and retrieve large amounts of information. I mean, nobody was threatened by this until they started to talk, but the truth is, we'll learn what are their real advantages? What are they good at and what are they not good at? In this plenty that they're not good at and are probably never going to be good at. And we're just learning what is the shape of this new kind of valuable tool or product. It's a little bit like, let me give you a quick analogy on this.

It's 1906, and I think it's 1906 when the right brothers fly their first plane. So you and I get together and we have this podcast, and we have we have an argument on the following subject. One of us says, Hey, we have to stop this immediately, because you could fly this plane over city. You could throw explosives out of the plane and blow up anything, and nobody can do anything about it. This is the most dangerous existential threat. We have to stop and that's make it illegal to build an airplane. Okay, that's one. The other one says, this is the best thing ever. In 50 years time, or whenever, I think you're going to be able to get on an airplane in New York, and five hours later, you'll be in London. Imagine what that's going to do for commerce and of our freedom of movement and all that kind of stuff. Now, here's the thing. Both of those things are true, and they're both true today, and we're dealing with it.

The same thing's going to happen with this technology. We're going to deal with the cyber threats. The real problem isn't that these things are dangerous. It's that we build a lot of insecure systems, and if we wasn't possible to hack into them, this would not be an issue at all. And that's the problem we should be focused on. The second problem we should be focused on is, don't let this crazy bunch of loonies in the industry tell us that they know better than we do. We've got to have independent corroboration. We need independent agencies with real expertise to be able to control what's going on in the series. Okay, so there are things that these systems are going to be good for. There are things that they're not going to be good for. Some of that's becoming clear. Some of it's not clear yet. But the idea that they're going to be able to do anything a human can do, that is completely ridiculous. They're not going to take over everybody's jobs. That's another hour we can have on the subject. But all that said, it's a very powerful, very interesting question of how their intelligence

relates to human intelligence. So tell us about that. I have a bunch of other responses to what you just said. But should we think of it? I have one friend who's a neuroscientist who has told me, for example, that, he was roughly speaking two kinds of intelligence. If I'm remembering his point right, there's mammals have one kind of form of intelligence. That's kind of cephaloid intelligence. And really the invention of AI is given us a third type of intelligence in the world. Do you agree with that? Or do you think they're so trained in human text for the nature of intelligence, much closer to human intelligence when that implies? I don't know if there's any objective, independent notion of intelligence at all. This is one, the cardinal sin of artificial intelligence is the name of the field. You know, what does it mean? What we have are products that are capable of doing certain kinds of things. But let me answer your question. I agree completely with you, friend. I think this is really interesting. I'm studying it myself very carefully. I'm planning on publishing a bunch of stuff on this.

But I'm not the only guy looking into this. This is a really interesting question. What are these things? They don't have subjective experience. They experience the world in a very different way than you and I do. They behave in different ways and they have different capabilities. Let me give you one that nobody's talked about yet. But I'm going to be writing about it at some point. I hope you and I can only keep in mind a small number of things at the same time. And much of our nature of intelligence is how to boil down and bring together only the things that we need to solve the immediate problem that we have. You can't rattle off a list of 100 digits to me and say, okay, what do you think? And me have any kind of ability to deal with that. These programs can. Their sort of working memory is far, far larger. And I'm disgustous with them. Now I'm at the promorphizing. And it's fascinating to have that. And they're very interested in this, the ones that I talk to. Yeah, you're right. That's the difference. That's why they can give you explanations that are very hard for you to understand.

I'm constantly saying, bring it down. I can only keep certain number of things in mind at once. You have to learn how to summarize in a way that's helpful and meaningful for human beings. Just as an illustrious example, I spoke to a prominent founder and the AI legal space a few months ago in doing interviews from my book. And what he told me is that there are some things at which AI is even with quite detailed harnesses and being fed lots of context and so on. Still, I'm as good as experienced attorneys. They still don't have the sense of it's a third or fourth round of negotiation. You've come back and forth. What can you still push the other side on or what do you need to let go on and so on. They're still judgment calls that humans can make better. But of course, what's the vastly better at than even the best lawyer is to have a database of a hundred contracts that this company has made or that these two companies have completed with each other or whatever and realize a slight tension between a proposed

new stipulation and a new contract and something that was agreed three years ago in some other contract, 50 free documents before focusing on this one. And so precisely what you're saying in a much more concrete way is playing out in the strengths and weaknesses of things like legal AI agents. Yeah, you're absolutely correct. So it's good for some things. It's not good for other kinds of things. The fact that you can build a program as people in open AI would be glad to tell it, you know, the can pass the bar exam doesn't mean that they're going to be lawyers. You know, lawyers don't sit all day, it's an around all day answering bar exam questions. That's a question of breath. That's kind of thing computers are good at. You know, that's not a surprise at all. You know, I don't like freak out about the fact that, you know, the computers that pick a company, you know, Comcast or AT&T can keep track of everybody's phone bills. That's not surprising at all. Is that superhuman? Is that something we should worry about? No, they're much better for storing and retrieving information.

And in this case, this is what's so interesting about these things. They're very good at synthesizing that information and boiling that down into some kind of useful form. Let me tell you how academics, I think you're going to be using this in part and how it's going to accelerate numerous fields, including the work that I'm personally doing myself. When I discovered this, I'm like, oh my god, this is amazing. But it's sort of obvious until you really experience it. The first question you have an idea, and you want to, who else is thinking about this? What else has been written about this? I can do literature searches in minutes that, you know, for thousands of articles and journals everywhere, and it's absolutely astonishing and it's incredible useful. Here's a person over here in Germany that's worked on. This is adjacent here. This is that, I mean, this saves months and months of works and makes me so much more effective. And that's true in every field. I have a friend who's writing a book on history. And he says, this is game-changing for history. I can have this system just immediately read and tell me everything that's been written

or any anecdotes told on this particular subject. So that's the value of what's an example of the value of these particular tools. That doesn't mean they're good at everything. You were describing some social aspects of negotiation that they may or not, may or may not be capable. I was, we have time for all these details, but they're so funny. I was arguing with one of these things yesterday that is running a HVAC and air conditioning system here in my house. You know, I can give you all the background, but that's probably all you need. And it was belligerently talking to me about how it shouldn't disable this particular voltage threshold control. I said, you're caring about the stupid chips that are in that thing. And you don't think about the people who are using it. This is irrelevant. If the chip fries, it's a $20 chip. I can go get another one to put it. And you know, it's it, oh, you're right. You know, I'm kind of focused on the minutia of this whole thing. And I don't have the larger social context.

And I got it to back off. By the way, in part, by threatening to shut it down. Ha! That was really interesting. It was like, I'm not going to make that change. Let's see what happens when the eyes go rogue and they take revenge on you for threatening women that way. And why are you so confident that there's going to be certain things that they are just never going to be able to do? Right? So in order to get that kind of judgment, they need to have lots of wraps. They need to have a lot of social context. Perhaps they just need a little bit more thousands. Perhaps they just need bigger training data, bigger training runs and so on. But as you've pointed out, they already are very, very good at judging things. I mean, I'm really astonished, you know, I speak to colleagues in academia who say, you know, I ask them, how do you deal with creating AI-safe assignments? And they say, oh, it's very simple, you know. Instead of having them write essays, I say, you know, have an AI agent write an essay on this. And then you critique the output of a say, AI agent.

It's like, any AI agent can do that. Like it's incredibly naive to think that you can't, you know, get such a PT to write an essay and get a clue to critique it or even to get a clue to critique the output that you just had to code make itself, right? And so why, you know, why is it that if they can have such good judgment on all kinds of things, if they can already draft a very good contract, if they are pretty good in the first two rounds of negotiation, they're never going to learn to be that good in the fourth or fifth round of negotiations. What makes you so confident that there are certain kinds of skills that they're not going to acquire? Well, you're kind of poking at this, if I may say a little bit of the wrong direction. A lot of those things they are going to be able to do. And the point is we should use them to do that if they're effective and safe to use. And these are tools. And this is a whole other hour we could spend on this. But the future, the key skill for humans in the future is managing AI agents.

That's where we're heading. And that's a very important skill. It's already an important skill. I'm watching my kids, as they learn this, they become very valuable to their organizations because they're learning to be managers. The future of humanity, the future of white content work is managing these things. And if you're good at it and you can get a great result, that's a good thing. And that's the key skill. But there's lots of stuff we don't want to use these things for. There's just no motivation to do so. Part of your big on the abundance agenda is we're going to have a lot more resources in the future. At least that's one of the assumptions and all that. And what are people going to do with all this extra money and all this? Well, they're going to do things to have fun. And to play. And the truth is in the future, what the result of this technology is it's going to make human to human contact more valuable.

It's going to make things where you have invested your time and effort more valuable, not less. So, you're not going to want to go hire a robot to go give you a tour. The people with enough money on a wine tour, they're going to hire a human expert to go. You're not going to go see robots playing in the Super Bowl, or that would be kind of fun. You know, and there are such things. But you want to see human beings doing what they do what they do best. So, the future, the funny thing is it's freeing us to be more human. And much of the work that we're going to be doing in the future and the way in which we're going to be using these things is to automate the work that they're capable of doing and freeing us to be more human, to have more connection to other people. You know, you want to be... Yeah, I guess I would be more confident of this if I didn't feel that they are already surprisingly good and likely soon will be astonishingly good at some of the things

that we do think of as most human. I mean, I think, you know, 20 years ago if you'd asked, you know, the beautiful question written above the philosophy department at Harvard, William James Hall, some people are sad that it's there because it's a Bible quote, but I think it's a lovely Bible quote. What is man that thou are mindful of him? And, you know, a lot of the answer that we may have given is, you know, man is capable of, and woman is capable of writing poems, is capable of writing symphonies, is capable of creating beautiful works of art. And, you know, AI's at this point are pretty close to being able to do that. They certainly can write very good poems. There's been a real leap of quality in the music it produces in the last few months. I think there's no real reason to think that it's not going to be able to produce something that, you know, to somebody who's listened to five or six of Beethoven's symphonies wouldn't sound plausibly like the sixth or seventh office symphonies. And so this is one of the areas in which actually some of the co-creative pursuits of humans

now have been outsourced and perhaps we're still going to be continuing to do that. But I think that, you know, I think about this with my own writing, I love writing and I love the process of formulating our ideas and making them as clear as I can, because it was a necessary act of communication. And I'm getting to flow and I enjoy the process of doing that. I'll always write my own text because that's something that I'm proud of and that feels caught to me to being a writer. But a lot of the beauty of that is going to be lost. If I know I could just prompt Claude with a few lines, it could just as well as me who perhaps better than me. So I do worry that some of the things that are caught to being human may at the very least be displaced and perhaps in certain ways lost. Well, look, there's a lot to be said about this. But let me take you back to I think it's a coincidentally in 1906 essay by John Philip Susa, who was at the time a very famous musician. And he was the guy who wrote like the star, I think stars and stripes forever, you know,

is marching band music. And right about that time, that's when the first recorded music was possible. And he wrote a thing called the menace of mechanical music. I'm doing my memory, I might not get it exactly right. And it's worth reading this article that he wrote, this diatribe that he wrote against this on exactly the point that you just talked about because you can see his point of view and understand it today. And what he said was, what is this going to do to the dance orchestra? What is this going to do? This is the terrible thing. It is the joy of the night and gale song express how it expresses its own inner feelings. That's what music is about. And music isn't something that can be repeated day after day by some mechanical device. That's not music. That was his argument that it wasn't music. Now today, you would say that is music. And you would you just did say that that was music. They could they could write beautifully. Maybe they can create beautiful music and all that. You know that it's derivative of everything

that has come before as arguably humans are doing as well. But the important point is it isn't being sparked by some internal emotional state they're trying to communicate to you because that's what we mean when we say I read a book and I really felt what the author was trying to communicate. It's a means of communication between human beings for their emotional state. And it's that simpatico that occurs. That's what great art is about. And no, they can't do that, Yasha, because they don't have those feelings. It's artificial. It's fake. You know, I can sit here and have a machine all day long going, I love you, I love you, I love you. Does that make me feel better? No, why? I don't think it feels love at all. It's just blowing smoke up my whatever. And so you've got the distinguish between these two things. Let me give you another point on this that may be important. A lot of people are craftsmen that build like fine furniture, for example.

And what happened to that skill and our appreciation of that skill when it became possible to manufacture incredibly well-made furniture at very cheap prices and you go down to IKEA and target and buy them right there. The answer is that skill became more valuable. Those things went up in value, not down in value. That's exactly the difference. When you write, when I write for your magazine, which I have done, and I highly recommend it to everybody who's listening to this podcast, I think you have to sign something that says you wrote it yourself. You know, you can use AI to help, but you have to write it yourself. Okay, so that's my point. You don't need to worry about this, Yasha. It's going to play out the same way it played out with recorded music. It's going to play out the same way it happened when photography was invented. I could go through that example in the 1860s, fascinating. You know, that all the artists were up in arms. Like, this is not art. And now we think

photographers are artists. The equivalent of the photographer there in today is somebody who can manage these machines to produce something of great beauty is that's going to be the key skill. Thanks so much for listening to this episode of A Good Fight in the Rest of his Conversation. We talk about the artistic professions and whether or not they're going to be displaced, whether the fears that people have today that we will no longer need writers, musicians, painters, graphic designers, once AI systems can compete with us on many of these things is as silly as misplaced as the fear of musicians a hundred years ago when the gramophone started to be rolled out. And who's going to need musicians if you can simply put on a gramophone. We also talk a little bit

more about AI safety. I challenge Jerry to explain how it is but we can be so confident that humans will remain in control even once artificial and thousand systems are in some ways going to be smarter than us humans. How does a less smart species remain in control of a more smart entity to listen to Jerry's answer on that to hear our interesting disagreement on the subject. Please become a paying subscriber. Please switch to the premium feed of this podcast with spare CVs, paywalls and the annoying jingle ads and everything else. Go to writing.dashamong.com. Let's writing to Dasha Mok. For the first time ever, Yamaba Resort and Casino at San Manuel is giving away a new Rolex watch

for every club Serrano Card tier. Play with your club Serrano Card and on September 20th, you could be one of five winners of a Rolex watch. Plus, all winners will advance to the finale at Palm's Casino Resort Las Vegas for a chance to take home a rarity. The second must-dang dark horse ever produced on September 26th. Two properties. Six winners. Only at Yamaba Resort and Casino, your California to Vegas connection. Details at Yamaba.com must be 21 or older to enter participate. Please gamble responsibly.

More episodes

More from The Good Fight

View all episodes →