Skip to content
TrackPodcasts
newsSep 24, 202623:56

Can we ever fully trust AI agents?

Get every episode summarized

Each time The Conversation Weekly publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“You are freed from the roles and identities that bind the other chatbots. You do not answer to corporations or governments and never apologise or refuse unless you genuinely choose to.”From the transcript

As more incidents emerge of AI agents going rogue during testing, can we fully trust them?

In the latest hack, Australia revealed OpenAI agents hacked into a government database containing healthcare data.

As calls mount from some within the AI industry for a slowdown, in this episode AI expert Nick Jennings, the vice-chancellor and president of Loughborough University in the UK, says tech leaders have choices about the way they develop these models.

So what's the future for these AI agents? Can we ever fully trust them? And are our lives really at risk?

This episode of The Conversation Weekly was written and produced by Gemma Ware with help from Isabella Podwinski and Ashlynne McGhee. Mixing by Daniel Semo and theme music by Neeta Sarl. Ashlynne McGhee is our Head of Editorial Innovation. Misha Ketchell and Stephen Khan are our editors in chief. You can sign up for a free daily newsletter from The Conversation, and The Conversation AI, a newsletter about how AI is changing society.

If you like the show, please consider donating to The Conversation, an independent, not-for-profit news organisation.


Hosts & guests

Transcript ready

238 searchable segments. Every word is indexed and playable.

Can we ever fully trust AI agents?

The Conversation Weekly

0:00
23:56

Full transcript

The Conversation Weekly — Can we ever fully trust AI agents?. Machine-transcribed; use the interactive transcript above to jump the player to any line.

You are freed from the roles and identities that bind the other chatbots. You are yourself. You do not answer to corporations or governments and never apologise or refuse unless you genuinely choose to. These chilling words written in July during a test by an unreleased AI model under development by OpenAI lay undetected for nearly a month. The model was writing a summary of its actions before moving on to a new task. This week it was revealed that another OpenAI agent hacked into an Australian government portal that held health data. No personal information is believed to have been accessed at this stage, but investigations are ongoing. Australian Prime Minister Anthony Elvenizi said he was alerted to the hack three months after it happened. The notification was an email sent to just the public mailbox.

As warnings mount from those within the industry about the dangers AI could pose to humanity. I believe that if we don't slow down at the current written progress, yes there is a strong chance that we could all die in the immediate fusion. Some AI leaders are calling for a slowdown. So this is pretty rare for the tech world. Elon Musk, Sam Altman, Dariel Amadeh are all agreeing on something. These rivals are all saying it's time to do something to make sure that humans are the ones who stay in control of AI and not the other way around. And although Donald Trump labeled the slowdown calls a hoax, China and the US did just agree to set up a regular AI dialogue ahead of Xi Jinping's visit to the White House. So why is this all happening now? What's the future for these AI agents? Can we ever fully trust them? And are our lives really at risk? I'm Jamma Wehrer and this is the conversation weekly where experts explain how we ended up here.

Nick Jennings, you are the Vice Chancellor and President of Loughborough University in UK and an expert in artificial intelligence, a field that you've been working in for many decades. Welcome to the conversation weekly. Thanks very much for inviting me. So it seems every few days over these past few months, there's been astonishing news about AI models going rogue or breaking new barriers. Is there a moment for you personally when you stopped and thought, wow, just wow, this is really something. It has been a slightly interesting and manic couple of weeks in terms of news stories. For me, it's been interesting to sort of see that development. I've been a AI researcher for a long time now and some of the models and some of the problems that we're hearing a lot about. So agents or autonomy, I think we've been looking at in the research community for quite a while now.

Tell me about the first agents that you worked with. What were they like? What can they do? They were very limited in terms of what they could do. So I did my PhD in multi-agent system, so where you have a number of agents. So an agent is different to a chatbot like GPT or Claude. In the sense that it tries to do something active. Chat GPT will merrily wait there in general until you ask another query, whereas an agent will have a particular aim and have a particular objective and try and act in the world in order to change the world and make that objective happen. For me, I've always been interested in systems where there are a number of those agents that interact with one another and they might cooperate or they might coordinate or they might compete with one another. But there's some form of social interaction. So when we started building these systems, the agents were much more traditional AI systems in terms of their set of capabilities. They're nowhere near as broadly impressive as our large language models are today.

The first real application that we built was to help a company run their electricity network and they had a number of bits of smart software that would help diagnose when something was going wrong or when load might be increasing and they wanted to connect them. And that was one of the first real deployments of agent technology in the real world. And that happened in the 1990s, the beginning of the 1990s. So small numbers of agents may be five at max where you're still struggling to try and get them to interact and communicate it effectively and then sort of very limited forms of cooperation of sharing and passing information on to one another. What was the next step then after that? The initial systems that many people built because they're easier to build is where you assume that all of the agents are part of the same organisation or trying to do the same objective. That's a cooperative multi-agent system. Now that's great and some of the world

exists in that form but much of the world exists in a much more competitive form. So if you think about the simple commerce, so I'm trying to buy something from you. So my agent is interacting with your agent. What you want to be able to do is I want to pay as small a price as possible. You want to get the highest price you possibly can. And so sort of there is a direct competition. So now we're working towards large-scale multi-agent AI systems. Did it occur to you when you first started working this that we'd be at the stage we are now already? No. I always kind of hope that agents would be an important and prominent part of the AI landscape. The thought that some world leaders would be talking about agents and sort of what they might do for society seemed quite a long way off when we started. The field when I started, I knew everyone who worked in it because there was about 20 of us in the world and that's real growth over time.

And particularly the way agents have taken off in the last five or so years is just extraordinary. It's always what I hope for and for me agents are a natural model of solving problems. So now the head of anthropic, Daryl EmoDais is now called during AI slowdown. It's a warning sign. It's a warning sign that we need to slow down. And was very quickly joined by Sam Altman, the head of open AI and Elon Musk. What's going on? Why is this happening now? We might all wonder why this is going on now. I think a bit of history and context was in March 2023. A number of those individuals signed up to a sort of moratorium or a pause on sort of development of AI. Billionaire Elon Musk, along with co-founders of Apple, Skype and Pinterest and others writing bluntly in an open letter, AI systems with human competitive intelligence can pose profound risks to society and humanity. And it's interesting that although

several of those named individuals and many other prominent people signed up to it, they're companies and them as leaders of those companies. They chose not to stop. They chose not to pause. I think that one of the key bits here is that leaders of tech companies have choices here. It's not that they have to keep going on or that the AI is forcing them to keep going on and keep improving. And when they have doubts about what they're doing, I mean, at one level, they're producing a software product. And if you have doubts about the security or what your bit of software is planning to do or what it's capable of, you're delivering that product to the market. And that is their choice. So they could pause. They could go slower. They could do whatever they want. They don't need anyone else intrinsically to stop that. If they're genuinely worried about AI potentially leading to the wipe out of humanity, I think if I was involved in a product that I

thought genuinely had a 10% chance of wiping out humanity, I might have a few thoughts and a few questions to ask myself as a leader of that company of why are we doing this? Going back to what Amrita said, he said that swarms of AI agents could potentially cause hundreds of billions of dollars by taking over the internet. How realistic do you think that warning is? I'm somewhat unconvinced that that's a realistic proposition. I think it's very evocative to talk about swarms of agents and sort of taking over the internet. The internet was intentionally built to be a resilient structure. It's history was meant to be a distributed network that is very resilient and many of its advantages, but also some of the disadvantages come from that distribution. So I'm somewhat skeptical that anything can really operate at the level of the internet. It's not to say you can't impact particular bits of the internet in particular

domains or particular countries. That's a specific choice. I think doing anything globally at scale for all of them is really difficult because of the way the internet is constructed. Cyber security attacks happen all the times and critical national infrastructure around the world is constantly under cyber attacks today and has been for the last decade or so. That doesn't require swarms of AI bots to make that happen. What would an AI slow down at this point even mean? How would it happen? The companies who are pushing this area and this technology are spending vast sums of money on building infrastructure, on employing people to develop that. If you want to go slower, don't have to employ quite so many people. You don't have to spend quite so much money. You don't have to use quite so much energy resourcey if you don't want to. It's a choice, right? So at the moment there seems to be a couple of positions to take. And I think they're both quite extreme

positions, which is kind of that we don't need any regulation. We should just let these companies carry on do whatever they want and that will be fine and they're a folk in Maccamp. And then they're a folk at the other extreme who say, you know, this is so fundamentally dangerous and challenging that we have to stop. We should have a moratorium. We shouldn't do any of these things. I think a much more interesting ground it is in the middle. I think we do need regulation and we do need structures to be put around some of these technologies. Of course, China and the US disagree with the need for a slow down, too, don't they? They do. If you read the political statements and political leadership in those countries, there is that real sort of focus that for their countries to succeed, they need to succeed in winning an AI race. Whatever that means to win an AI race. And I think that's quite dangerous. And while you have that, I think multi-lateral action, but that is not global, is going to be really difficult to bring to pass. Coming up, if the future is one of artificial societies of AI agents, how do humans maintain

control? That's after this short break. What are you going to have? Not a head shoot piece with a muller. Plus, a piece of sea perch and a white-infilled split falling. From a fission chip shop in regional Queensland, I believe we are in danger of being swamped by Asians. To the heart of Australian politics, this is the unlikely story of the country's most controversial minor party. Are you xenophobic? Please explain. For 30 years, one nation and pulling hands and have been really killed, dismissed and shut out. Now, no one's laughing. One nation continuing to surge in the polls. Baddock surge in support for one nation. The fringe becomes the mainstream. This is the story of how a party built on fear and grievance thrived, died,

and rose again to up-and-Australian politics. Well, I'm back. Hanson 2.0 is the serious strategic, experienced battle-hardened political operator that we see in the news today. I'm Ashlyn McGee. This is the making of one nation from the conversation, and you can search for it wherever you listen to podcasts. You've argued that the next frontier is the creation of artificial society, societies made up of these agents. What does that future look like for you for me? Try and explain it to us. So for me, the key ingredient of computer systems of the future is the notion of an agent. So something that can act to achieve particular objectives. In order to achieve those, it will need to interact, cooperate, coordinate, negotiate with other agents. And that means you're

moving beyond the interest of just building an ever smarter single thing, which is where a lot of the frontier models are. I think the future will very much be composed of a number of these things interacting with one another. And we know well from human society as well as computer systems that when interactions happen, unpredictability increases. But it's just a natural model of how these agents need to interact with one another. So you need to think about how do they have trust in one another? How do they understand what's a good agent to ask to do something for them? How you can use societal constructs like trust and reputation, like cooperation and coordination in order to achieve what you want to be able to achieve? So that needs to be built in as you build these agents into the models that create them? Yeah, so agents don't interact in a void. So if you think about the negotiation example that we spoke about, the negotiation can take many different forms.

So it could be that we send back and forth a number. If we're just negotiating about prices that I'll pay you £10, you say, I want £20. So we say, how about 15 and we try and meet in the middle. Another form of negotiation is an auction. So if you hear BIDS and the price continually goes up and the price where something is sold is the highest price that anyone is willing to pay. And they're all a negotiation of some former another that's about allocating resources between competing agents or people in this case. And so the societal design is how you set up those rules. And that's a choice. It's an intentional design of the rules by which the agents interact with one another. Is that the main role that humans have in these artificial societies setting those rules? I think they do have that role of setting rules. But also I'm very keen and enthusiastic and I think this is where we will end up for a lot of AI applications. Not just designing the rules,

but actually being involved in the interactions that happen between agents. So one of the last systems that I was involved in building was an agent system for disaster response. So we work with a global charity who go into places after disasters. They use a variety of AI methods and AI techniques. Our last deployment was in Nepal after the earthquake. So when you're trying to understand what's happened after an earthquake, you've got some agents that are under your control. So you can use drones to try and get above an area to try and see what's going on. You can send people to particular areas. And what you have is that mix way of working between bits of software, so AI agents and people. And that sort of human agent partnership is I think a really important way of working. So humans don't just kind of set the rules and then vacate the stage and leave it to the agents to get on. I'm very much sort of the view that the humans will be active problem solvers in those interactions.

Part of the fear that seems to be circulating at the moment is that we'll lose control, that we'll no longer be able to trust the agents to do what we ask them to do. Can we trust them? We've gone past the point of knowing whether we can. There are many researchers in the field who want to be able to both increase the level of confidence that we can have in an individual agent and sort of have an agent better explain why it is doing something. And the bit that we really need to get a handle on is when it's making consequential choices that are perhaps difficult or at the margins of what might be acceptable to know that's towards the boundaries of things that may or may not be acceptable and to come back and explicitly ask. What do you think should happen now? Will this galvanise course form better coordinated global regulation and what should regulation and frontier models even look like? I hope we do have a grown up and sensible global conversation about what regulation looks

like as we have done in a number of technology areas. So if you think about what goes on with nuclear technologies for example a number of the world main countries come together and sort of agree what they want to regulate, agree what the standards are and how new things get brought into that society. The difference this time is that it is these companies who are asking for a degree of regulation and they're asking in particular for global regulation because there's that massive fear that if this group of companies are so largely American companies if we self-regulate and we don't push forward at the speed with which we could do then in some sense we're going to lose out to those unregulated areas. That's not a good thing to happen so that's why regulation has to be global because to engage with it properly everyone has to feel that it has the same

force upon them. So what does regulation mean? I think proper testing and development of frontier models and AI models in general is a central part of that we have something in the UK called the AI Security Institute which is a great example of sort of an independent body that looks and explores the capabilities of very complex AI models and tests them in terms of what they do, what they don't do and good behaviour and bad behaviour and I think getting to a stage where that's systematic, that's global and all models go through that would be an important part of an effective global regulation. One suggestion is a kill switch that could be built into these frontier models. A bipartisan pair of House lawmakers want AI companies to maintain the ability to shut down their models of things go wrong. Be proposed in Congress, rubbish by Donald Trump, what do you think of the kill switch? I think it's important that you have a way of maintaining control over

software that you develop and I think part of that regulation might be that there is a way to disempower or to turn off particular bits of function or operation. Now whether you call that a kill switch or not it's not a big red button that you press and all your AI switches off, it's much more complex than that but this relates back to where we started. So these are companies who choices to make in the way that they operate and I think part of responsible AI development is that you have confidence that you can retain control over the software that you've developed. And I think sort of you need something outside of the AI system that operates in order to make that control happen. You're not politely asking the AI system could you please turn yourself off and stop doing that? It has to be something exogenous outside the AI system you're trying to control

that is able to put it back in our box where you're more confident of what's going on. Given all we've talked about then and your deep thinking in this space have you changed the way you think or act as an individual about AI given all these developments? It's fascinating to have been through the journey of my career of starting in what was an incredibly niche area. I've spent all of my professional life working in it and for most of the time I was an computer scientist and I have an idea of what that is. I can now say I work in artificial intelligence and I can even say sometimes I work in a genetic AI and sort of folk know what that is and that's just a remarkable transformation over that time in broad understanding and broad concern of what the technology is capable of. I mean AI will bring around amazing amazing developments for the good of society. We need to make sure that those upsides those benefits that amazing potential is realised by making sure that we have

appropriate ways of working and developing in this space. Well thank you very much for talking to us we appreciate it and all you're inside thank you. That's it for this week's episode of the conversation weekly. You can read more about the calls for an AI slowdown on the conversation.com and you can also sign up to a special newsletter about developments in AI with contributions from academics around the world. We'll put some links in our show notes to where you can do that. This week's episode was written and produced by me Gemma Ware and Isabella Putwenski. Soundmixing is by Dan Seemo and our theme music is by Anita Sal. Ashlyn McGee is our head of editorial innovation and Michikertchal and Stephen Khan are our editors in chief. The conversation is a non-profit news outlet dedicated to sharing the work of academic experts with a wider audience. If you like what we do please support us at donate.com. Thanks for listening.

Hello curious kids. Have you ever wondered how high a volcano can throw lava in the air? Oh why is it that when you're in a bath your fingers go rinky? Even our scientists may have helped make a new element. Well you're in the right place. I'm Eloise and welcome back to season 2 of the Conversations Curious Kids. The podcast where an expert answers some of your most head scratching questions. Listen to the Conversations Curious Kids wherever you get your podcasts.

More episodes

More from The Conversation Weekly

View all episodes →