
What happens if your AI agents get 1% better every day?
About this episode
An agent that gets better every day compounds into a very different agent within months. That's the idea behind Wren, PolyAI's new release: an agent whose job is to make your other agents better. Some of what it finds is a small refinement, some of it is a big fix. Either way, Wren watches every real customer conversation, learns from what happens, and keeps improving your customer-facing agents.
Nikola Mrkšić sits down with Arkadiusz Kwapiszewski, Head of Agent OS, to get into what that changes, how Wren is putting agent-building in the hands of people who could never touch it before, and why 95% of PolyAI's own production changes now ship through it.
Listen to the full episode, and learn more about PolyAI at https://poly.ai?utm_source=youtube&utm_medium=podcast&utm_campaign=podcast&utm_content=podcast
Get every episode summarized
Each time Deep Learning with PolyAI publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
850 searchable segments. Every word is indexed and playable.
Full transcript
Deep Learning with PolyAI — What happens if your AI agents get 1% better every day?. Machine-transcribed; use the interactive transcript above to jump the player to any line.
You tell, okay, build me an agent based on those specs, based on those conversations, and it will configure everything on the platform for you. You don't have to do it manually. You have that option. Everything is readable. You can go and edit everything. You want, but RAN really is the main entry point. And it handles everything for you. And that everything gets better over time. It watches all the interactions that are happening and makes the system better. Adapt system, right? It's an adaptive loop where your agent gets like 1% better every day. And that obviously compounds like that LinkedIn meme, right? If you get 1% better every day, you end up being quick math. Hi everyone and welcome to another episode of Deep Learning with Pali AI. Today with me, I'm Gartar Kadjusz Kapuszowski, who runs our agent OSD, and we're here to talk about RAN. Our coding agent and so much more.
But RAN, I thought we could maybe start with just a bit of your background and how you ended up doing this stuff. Yeah, so I started Pali four years ago, which feels like a lifetime in this industry. I was originally part of the deployment team. So I was working on actually configuring and deploying dozens of agents to production, designing them, integrating them, making sure that everything went smoothly for the users. But at peak time, how many reports did you have? I think it was about 40. So it was the whole agent design team, which we still obviously have, like really brilliant people, lots of linguists, lots of people with computer, so human computer interaction backgrounds. So yeah, they're still doing fabulously. But now I moved to the project team to actually work on sort of automating some of those tasks that we were doing, making them work at scale. So they're not just confined to human experts, but all that knowledge, all that expertise
is now available at scale to all of our customers at the click of a button. Totally. I think that there's a lot of talk of FDA. I think a lot of our competitors who previously didn't believe in forward deployed work are now rushing to talk about it to hire people. I think we, as always, are two years ahead. I think we're a volunteer gutter I would be building foundry after many, kind of like forward deployed deployment, is like people forget the second part is for the platform to become something that doesn't require hugely motivated, intelligent, people problem solving ad hoc for every client individually. Like, it shouldn't be that way forever. If it's like that forever, then you're just a consulting business. Absolutely. I was always like a believer in this deployment, product feedback loop, right? Because all that expertise from the front line, you don't have to bring it back, you have to platformize it, so it then is available at scale, right? Like reinventing the wheel, doing things individually on each deployment that just doesn't scale.
We've definitely been guilty of that too, right? I think that, well, you're not about very motivated for less of that to happen. So, okay, tell us about run. So, run is our coding agent, but it's so much more than just a coding agent, right? So, it has basically two ways of interacting with it. You can think of it as your platform assistant helper, or really the main interface to the platform, right? It comes to a platform, agent studio. You say what you want, like, you know, tell me about the trends for the last week, build me a new flow for making bookings. Ran will do it for you, but Ran isn't just sort of passive waiting for you to ask a question. It's also proactively review all of your calls, review all the interactions, and bring the recommendations so that every morning, you open the platform, and you see like a list of things that can be better, right? Here's a list of suggestions that you can accept to make your agent like a couple of percentage points.
I mean, one of our customers just posted yesterday saying that booking conversions went up 30%. But not in my wildest dreams. Did I think, as you said, I thought it was gonna be like, I love what percentage I'm getting. Yeah, and he just went through the whole list, engaged with each recommendation, thought about it, gave a bit of feedback, pushed the production, and just saw the stats go up, and then was so excited he posted about it on social media. I think like, you know, as I think about like the role of AI and who gets to implement it in an organization, I was with a very large customer yesterday talking about this, and, you know, big dog, dog nine in the room, started asking questions, then had a few really sharp observations about this, like, application that was designed by our pre-sales team. And his insight was like, viscerally accurate. I was like, well, no offense, but there are five people in this room from your team who did not communicate that to our team. And that's because they don't have your context, right? So, I'm similarly to how you had all that context and were able to start building this.
I feel like making good quality use of time by those people who are empowered to make those decisions and just agree to them, is where the organization gets like. I'm not saying it becomes a one person organization, but the improvements don't go through like five committees and debates and risking governance and just someone who's fully equipped to understand what they mean can just say yes. Now, it is true that sort of like organizing human to human collaboration is kind of the bottleneck, right? Because passing all of that context, making sure that the person who's executing understands the vision, all the context, it doesn't really scale. Whereas with AI, the person who has all the vision, who knows what to do, like can just say, say that and watch the results happen. Totally. Well, you know, you promised some religious references to me and in the run up to this one, but I think one of my main Courtney jokes in the company now is that I refer to my co-founder Sean as the Holy Spirit in that I think at this point, he will only be really meaningful to engage with you. If your agent can fully outline your worldview
around something and then he will like mix it with his and we'll take the next step in that it's easier than talking about it because good talkers don't dominate conversations. Absolutely. And to go back, I think we should talk a bit more about like RAM and what it is, right? Because it is really sort of the next evolution of water coding agent can be especially for conversational AI platform. It really is the first time when we can generally call a platform like fully agentic, which is very exciting. So, you know, until a year ago, the platform allowed you to build agents and those agents were interacting with your customers, right? So you had agents for different use cases, you had hotel booking agents, a concierge, right? Or maybe an agent, customer service agent that allowed you to check on your order, lots of different use cases, but it was like a single agent talking to single customers at the time, handling millions of conversations, right?
Then we have sort of the first iteration of RAM, really, which was a coding agent, right? So again, here are the religious references that I promised. It was like a god of daism, right? It builds a system for you and backs off, like leaves it running. So you can come to it and this is all like testable, like a platform is open, so you can go and play with this. You tell, okay, build me an agent based on those specs, based on those conversations and it will configure everything on the platform for you. You don't have to do it manually. You have that option. Everything is readable, you can go and edit everything. You want, but RAM really is the main entry point and it handles everything for you. So it sets up the system, the system then interacts with individual users, but it doesn't follow up on those interactions. Whereas now what RAM is doing, it's like an intervention is gut, right? It's, it watches, it still cares about the world it's created, about the system it's created. It watches everything that's happening and it basically performs miracles
in order to make sure that history sort of tends towards improvements to logically and that everything gets better over time. It watches all the interactions that are happening and makes the system better, adapt the system, right? It's an adaptive loop where your agent gets like 1% better every day and that obviously compounds, like that's that LinkedIn meme, right? If you get 1% better every day, you end up being quick math several times better over the year and that's basically, that's basically what's happening, right? Performing these little miracles, if like finding, finding interactions that can be improved, suggesting better design solutions to certain problems, finding knowledge gaps for you, implementing the changes for you. So all you need to do is click like a proof and then your agent and agent gets better, right? And we will also do also deployments in the future. So you want, you want, you want even need to be there to reap the rewards.
I mean, really you are just as team working with Ren, your job is to provide the right context and maybe direct the plumbing of that data in a slightly better way. I think one really interesting thing is the project context feature in there where it's kind of like, you look at the recommendations one day and you go like over obsessing over like, I don't know, handoff routes or something like that. Instead, like, I really like to see more ideas around how to keep people in cancel subscription flow by whatever, or you know, recently, I've just seen some really cool stuff with like, agent memory then being used and targeted by Ren to figure out what we could do and kind of like loading up your information, figuring out how to talk to you and then running tests about it. But just all this stuff that was so hard to coordinate, implement, execute and then evaluate to just run in the background and compounds is essentially miraculous. Yeah, I mean, I love the fact that there's no latency cost, right? Because the user, whenever you have like your standard coding
assistant, you ask it to do something big, you have to wait for like half an hour for it to complete with Ren. Now, everything's kind of ready for you, right? So it really is as simple as kind of swipe left swipe right. Yeah, I mean, it's like, when you have an idea, you should absolutely go for it. But the truth is, and I think this is where like the magic of what we do is maybe hardest to explain, but most important is there are people who live in cloud code like you do, I do, although more and more of a living around as well. But do you know how many messages you send back in four of a cloud code in a month? Well, I was on Patelief last month, say, three send more than usual. I said more than you should. I think it'll be easily in the thousands, right? Like I think about it a lot, like how quickly we like, adapt to this because like my job was completely different like a year ago, whereas now it's all several cloud code terminals running in how many, how many do you actually use? I think about like, I limited myself. I found myself like, you know, the human attention,
I was the bottleneck, right? So I have a rule where it's like, okay, three or four things at once. So I can actually actually follow what's going on. Yeah, I have five, I renamed them over a few days, just to kind of keep it going here and thread up the things I'm working on. Now everybody will have these personal strategies, but the reality of work has changed a lot. And I think this is also like an interesting thing, right? Because we're trying to kind of take that experience that way, like frankly, like addict it to now, right? It is like, I love this new way of working of having an idea and being able to like delegate it at speed and just sort of be in charge of a team of agents. And but it's still not like super accessible. Like there's only, you know, these, these most devoted people and tech companies who are really deep into this. Yeah, yeah. The world is like a few, maybe point one, maybe one percent, probably not even that. Or like heavy addicts. And then there are the rest of the world is just like, oh, cool, it can read an email. Yeah, exactly. And I think what I see as part of our mission with Ren
is to make that experience like more accessible, right? To all these people managing custom, like call centers. So they also get, they also feel the magic, but we make it sort of tailored to them. And we, you know, the things that we find, we experience as friction, like the overwhelm, the slop, right? The making sure that validating all the changes manually, like we want to build it into the harness, into the workflow. So it feels even more magical, right? It doesn't require much skill. And they can get hooked in the same way on an experience that is even better and just tailored exactly to the needs and to the workflow. Yeah, I mean, one thing I think about a lot is people talk about slop. And I think a lot of people who are not bought into the new world kind of use it as a disparaging term. It's true that it's tiring to kind of work through these things. It gets, like you're holding a lot of, you know, our context windows are limited to, right? But I think what is wonderful about the Ren and the way that it works in live deployments now
is that the customer calls and the data actually ground and work as a regularization function for the slop. So it's actually not like the slop is guided very heavily by enough like forces and input that you could never provide on your own trying to contain it is just too hard. You can't see enough code. You can't think about enough things. You can't remind about enough things. Whereas there's this full of things coming in constantly. And like it's reflected in the tickets it creates. Now, it takes us back to basically this sell idea that part of our role is to make, to give brand all the data that it needs, all the constraints. And then surface some of that to the user, right? So so we can build trust incrementally, right? They see a recommendation, rent suggests a change. And you know, we link relevant calls so they can see at a glance they can like just understand intuitive. They're okay. I can see the issue. I see it like a one sentence description. I see how it affects real colors, right? These are not made up. These are real calls. You can listen to them.
I see that RAN actually already made the change. It used simulation, converted simulator testing, right? To actually replicate the issue. And then, and of course, they're for posterity. So the same thing doesn't happen again. Exactly, right? So it shows you like how it's simulated, the same conversation after the change and how it's now working way better, right? So you have the confidence that the change is real. And then, you know, once you, once you approve it, it keeps monitoring the calls. So then, you know, shortly after you get like, okay, like I can see that this type of calls is now improved. You can see the stats going up. So that's exactly the constraints you're talking about. It's not slop. It's just intelligence, right? Within that workflow of improving your call center assistant. Totally. I mean, it's super human in a way that's just unbelievable. So I was in a bit of a customer tour over the past week and just demoing things around. And, you know, I think you and I are deep in this rabbit hole with probably your team and a few other people.
And I just like, would show people something new? Like, hey, what's coming next? And the whole meaning we had hijacked and overrun by an hour. Because people start seeing these like recommendations and then they're like, wait, what is that? Click on that. And then like these is the terrible things that someone might have mentioned to them once or twice and then they were forgotten. Because they could never really quantify the impact of that change. They'll like, wait, wait, wait, hold on. Click on that. Click on that. Enable the post call tracking. Because they're like, I want human data in here as well. Because like, I just want you to surface these complex like phenomena happening in a colon. The agent, AI agent does this and then human does this. Why did that happen? Try this instead. T up this piece of information for the human. People just get so creative. And I think to me like the, you're right, a moment. And I mean, you remember kind of like how we went from having like a studio assistant. Somewhere in the corner behind the feature flag for internals to like this and out the home page. And that's it. Like I think that was a big moment because it's really a gateway
drug to having a lot of fun. And that as I look at the data by our customers now, there are about 10 people who have 2,000 messages back and forth with Ren every month. Now that's addiction of this. Yeah. And we have had that feedback, right? They love it. They trust it. They see it as co-worker. And they have a lot of fun because it shouldn't feel like work, right? Like with AI at its best, you just like 10 ideas into reality. It doesn't feel like work. There is an affriction. You just, you know, you just see your vision happening in real time. Yeah. So run. Why run? So we went on that name like back and forth for a while. But Ren definitely had the most internal support. This is a, this was a name we could imagine ourselves saying like every day we could use it as a noun. We could use it as a verb, right? Can you run that change? Can you ask Ren to do it? It's a proper name. So the original inspiration is Christopher Ren, the polymath,
physicist, astronomer, architect, who rebuilt London after the fire of London. Also has some incredible buildings in Cambridge. So does that local connection? So I mean, simple cathedral where he's buried is like one of his many other collectible London ones. I spent at least like three years of like heavy exam terms in the round library in Cambridge. And he wasn't a Cambridge man. He went to Oxford. But he designed the round library eternity. And like it just had this like, well, I guess, anglophilia, I feel that poly is very much about. And as you said, he was really a polymath. Yeah, polymath. So that's the sort of historical connection. And obviously when it comes to branding, we sort of also, it's cute, right? I mean, it's the name of a bad, like, and there's no hiding that fact. And we have that bird icon in the platform, which again, also helps, I think, to have like a little mascot, right? We see that with a lot of coding.
I think we've had many, we have had many a bird, you know, I think the previous version of our LLM was raven, the speed tracking, as a whole, TTS, Macau, right? I think like code names for our initial kind of like vertical agents, the kind of like non-interventionist ones were things like Finge and wow, I'm forgetting now. We might have had a starling and a few others, but yeah. What I kind of like about the bird metaphor as well, in this case is that, you know, you basically do get notifications, right? Like you get that notification chime when you come to the platform, it tells you, I have a hair or some things you can do. And that also is kind of like being working up by bird song, right? It's supposed to be a positive thing. So it's not alerts. It's just, you know, you get like that positive association of like a nice sound to come to. And that's also like this first experience of coming to the platform is also something I wanted to talk about, right? Because in the past, I think most people just looked at the conversations
review page, right? So most must have are external users call center managers. They look at analytics, they look at the conversation review page, that's sort of the entry point. They review some calls and then sort of, you know, how many calls can you really review? No, more than 50 a day with a busy people, right? Like, it's not even just that. I think there's just a limit. Anyone who's worked on, I mean, you know this, but anyone who's worked on dialogue, like 50 calls in, whoever can do more is like superhuman. That you have to be boring. It gets boring. And then I think the problem with that is that then you look at the performance of your agent through this like super narrow lens. Okay, there's a sample of 50 calls out of 3,000 that happened that day. You see one example of friction and then you sort of like, you know, because humans are foldable like that, right? It's anecdotal evidence. But you're like, well, I want to fix that one example of friction. And then it turns out that it literally happened only once and it will make no difference to your stats, right?
Whereas now the sort of homepage, as you said, is Ren, right? You come to the platform and Ren is your gateway into the platform. You see the recommendations and the recommendations are based on like thousands and thousands of calls like AI can see patterns in data that no human ever was, right? Like we had all that data before. We've had it for years, but it's only now that we can make it properly, like, actionable and we can make it understandable to humans. And I think that's, that's the beauty of it, right? It is that AI can sometimes lead to overwhelm, but also it is like the cure to overwhelm. Well, suddenly, suddenly you can make sense of the one. Like through the slop until you see the next film. Yeah. And you suddenly allows you to make sense of all the data in the modern world, right? And this is also what we want to, like the experience we want to provide. So you don't have to start with conversation review anymore. You start with the recommendations. You click approve, you know, swipe left swipe right and your agent gets better.
You can still listen to calls as well, but it becomes way more targeted. Way more scientific, right? Because we can assess the impact of every change. So that I think that to me is what feels super magical as well. Yeah. I mean, like I think it's just really interesting. The whole kind of like, I mean, maybe for the general audience, right? People used to kind of like listen to a sample of 2% of the calls. Sometimes they would listen to calls by the best agents and the worst agents. And they would have like these look very brittle. Algorithms that would be like, oh, when people speak to a caduce, they say, nice words like happy and thank you. And they speak with Nikola, they're frustrated because they say things like, you know, unbelievable or whatever, right? And that was just like not great because the best thing they could do is like performance manage me out because I'm a worse agent than you. If they had enough agents to begin with, right? Now I think like with this, it's just possible to ask anyone of these questions and get an answer. And I also really like the kind of like drift of any individual deployment
worth through the context. They're able to guide like, okay, like for the next few weeks, we're going to work on optimizing this and this and that. So I would really like you to dig into the potential things that can be done here and there. And then I think, you know, there are people who were just deployed. There are people who will test anything that that's how do you see the future of that? Well, so I definitely want to make it more autonomous, right? So there is, there are some really, really interesting decisions, design decisions to be made in terms of when to engage a human in this AI loop and when not to when you think about what ran, ran it's like our customers in the past. Sometimes they watched, watched looked at a conversation. They didn't like anticipate when the agent did something clever. They hadn't realized and they would ask us like, is the agent like self learning? Is it learning from the interactions? And, and you know, again, until ran, the answer was was no. And now now it's, it's yes, it is learning.
It is changing. Every interaction run has changes a little bit and it optimizes ran for the actual production calls. So we want to make that up as optimization process better and faster, right? So your agent is really the perfect assistant for your color base. If the color base changes for whatever reason, I mean, like, you know, we have that. Right. Right. The agent will change with it, right? People start asking things about, like asking about things you haven't anticipated. The relationship with your brand changes, like we can catch it, we can optimize for that. In terms of like when to engage a human in that process, right? Because it's that loop can be like fully autonomous, right? You detect friction, you repair friction that could happen without human supervision. But obviously there will always be moments where you need to engage a human to talk about, like business rules, right? If there is a, let's say, does the user is asking a question? We don't have any answer to. Well, ran might be able to find the answer.
It might be able to find the answer on your website. It might be able to find an answer in some of the documents that you upload it to the platform. If it can, it absolutely should ask you and will ask you. It won't, it won't make it up. Right? Sometimes you just, we just have to defer to you to help us. Improve your agent and fill in the content. There is like a really interesting distinction as well between self healing and self improvement, which is guiding like a lot of the design decisions that we're making. Self healing in a way is easier because it's all about fixing friction and friction is easy to detect, right? There's some frustration on the call, some gap, a bug or something and and and run, goes in and resolves that. Everything goes smoothly in the future. Self improvement is very different and it's very much KPI based. You have some kind of north-star metrics, so you want to optimize your booking rate because obviously this is how you get your return on investment by by agents making bookings
autonomously. You could have an agent that has no defects. It's performing exactly as designed, but it was just designed not in the most optimal way, right? So suddenly it's not about fixing defects. It's about finding better ways to engage your color based, to represent your brand in order to get those metrics to go out, but maybe finding ways to integrate your agent with your systems better to make it more useful. More useful objectives as well, right? I think that like what I found is often the difference between our deals. They're like mid six figures and deals to get to look mid seven figures is the level of sponsorship you have. So really what that means is you move up from just the context center and into like the wider image of a chief commercial officer, CMO, someone typically more in charge of revenue as well as just the bottom line. And what's interesting then is like the metrics they look at are different. And then if you can connect your agents, in this case, your fully autonomous context
center, to those objectives, then for instance, the value of cross-selling when people can book a restaurant here, like no, I don't have that time. I've got half an hour later. I've got something that's four blocks down, same time. That's like really valuable. And then you experiment with it and you book and then you break through the ceiling where previously you were trying to, you can't really do much if it's not available, it's just not available, right? And in theory that some of those things were always possible, but that takes connecting it to the wider or aggressive. That's where that human context center would have been a cross-selling one. Now again, the superhumanity of our agents is that they've got time, right? So they're not comped on how many calls did you take? Why did you have a channeling time too long? Well, if it's longer because I convince you to show up at a different location, I'm a hero, but the metric typically would just show that Nicholas low on his shots. And well, then I'm not really in center, I used to do it, so I'm not really going to
do it, right? And then I think when you go up through the org, then the design of should we incentivize people and AI agents to do that or not? And was the propensity of that thing to succeed? It's just impossible to quantify what humans do. I think that iterating at the speed of dialogue where you just have this idea, you don't even have to code it up in the way that you will call code or run where you put it in context and run comes back to you the next day with five suggestions. It's got a whole team of analysts that went in, came up with these proposals, then a whole IT team that implemented it, then a whole team of, you know, like, continue to optimisation people running it and evaluating it, which is really just calling them the ha, conversions up, 0.8%. Tick, except done. And like, we're not there just yet, but at the rate that this is all going, I don't think we're more than three months away from someone just going over auto mode, like, accept all and go. Yeah. I think all the decisions that can be made autonomously should be made autonomously.
And then we need to delegate to clients when we just need to integrate with the systems better or where we need to do something unusual in terms of something out of the books in order to really move the metrics in a more dramatic way, like what you said about the next question. Yeah. So this is like, you know, you asked like, what's the future, I think, I think improving the self-improvement part of it to make sure that we're not just picking the low hanging for it, but we really think strategically, we really think out the box about like, how can you really like change the design, the flow, the conversational flow, the setup of the agent to engage, engage people and convert more users and find these solutions. And we have a system now if we can, if we, like, we're getting more and more users to our platform and we're getting all these decision makers as well, right? So we now don't have to necessarily go through these loops of human approval.
People who have the full business context can interact with friend directly to make those decisions and they will know what's best for their business and for their brand and then they can make it happen and we can be aligned with these top level metrics, right? Yeah. And so, people always get like completely petrified when the CEO has a full request, right? And I've done it a few times, it gets people to just reimplement it in the right way because the risk is too high. But where I think it's really interesting. With a friend, I mean, that's kind of what we want, right? The CEO has the context and they have the decision power. Yeah, and I've done like that with the number of people on my cook. I think that's a good idea. Share screen. Dad? Yeah. Cool. And I also, one can expect to see that thing. I was like, five minutes. And I think that like at that point, people are just like, their minds are blown because it's like, oh, wow. This is now like a transformation project where there's an invoice. And then they're like, I think the dopamine hit goes, I did this, right?
And then it's like, oh my god, let me do another thing and another thing. And as you said, you know, one percent better every day. And like, also like, they have the context, right? So I think that putting the future as of like all these roles and like, because you know, we make it seem like there's just one person at the top, like, you know, pulling all the strings, but I don't think that's really true. No, I mean, it definitely is easier to be like one person at the top again, with what we talked about with AI, like allowing you to actually make sense of all the data, have better overview, right? But I think you still need people to like go deep and have all the context of specific areas. I do think that this flattening of roles is very real. I mean, I've experienced it myself where I mean, my role is like part engineer, part product manager. And that just means you can carry the context and the expertise you have through through the entire workflow and just make things happen. Again, it's all about having that context in your head.
I think this is some of the most valuable thing about workers now is just having all the context and knowing what the direction is, knowing what needs to happen, knowing what needs to be built. And then you have all the tools that you're disposals to make that vision into a reality, right? And I think really organizationally, the only thing you have to do is find the people willing to, you know, have their skulls collapse under the pressure of the ever expanding context window and empower them because they'll just, if you empower them to make those decisions, like, they'll iterate and they'll get their way faster than they can explain. Because I feel like the iteration has become a lot cheaper than the proceeding the bait. And the bait used to have to happen because it was like, you know, or Mike, I can be better working on this for a month or are they working on this? And, you know, we've got five sessions each. We've got a lot of sessions each. And, you know, I think, and that's, that's super true, where you can just, they make things happen. And again, with Ren, it's so cheap as well.
And a big part of it is experiments as well. So, you know, we talked about self-improvement. You don't necessarily always know what will be like the best for your users, right? Like, sometimes the way you phrase a question, like, how many questions you ask, it makes a, like, which voice your users will respond to, right? Like, you can't know a priority, like, how that will impact your metrics. So, you basically do have to do experiments. And Ren being able to conduct these experiments for you, saying, okay, like, I have a hypothesis. I think this user journey can be like, cut in half. And just, you know, just make the questions like, let's verbose. Like, get more snappy. I think this will really help users engage. How about we test it out? Like, we can launch it for like one percent of course, like, percent of course, whatever you want, right? We'll watch the metrics. We'll see how we'll review the calls. See exactly what's happening. It's like, so the friction is so low, right? For you, like, the cost of that experiment is so low. But then you don't have to debate it anymore, right?
You don't, that's, it cuts through the debates. I remember one of our largest customers, which will not be named. Had a few voice options. And a very senior stakeholder one, one of them. I think the other two had like 98% of the votes between our team and theirs. And like, for years, I had to kind of like prod and like, and with like experiments, it's just, well, everyone contributes their idea. And then you let the data decide. And honestly, but attempt to throw in their idea, I feel a lot of the egos are satisfied. And then like, you know, the data will vote. And you kind of forget about it. And you know, like, you feel good with your idea, works, but like, really, if it didn't, you just compete harder to add more ideas to it. And then, this is part of the optimization loop. Right? Like, I mean, this is, this is what we're doing as a company, right? Like, we're iterating on a platform faster. We're pushing to production. And we're letting, like, we're putting all the new, like, UI, all the features in front of users to validate them, right? Because we want to, we want real feedback. We don't want to like, see you rise this. And we want our clients to be able to do the same, right? Like, make changes, see how real people react,
validate that against data. And then make, I mean, that's it. And you know, scientifically, right? Yeah, I mean, I think that a lot of people, kind of like, try to de-risk this for companies going, we'll learn from how your humans talk to this platform. And like, the only thing I've been able to come up with as an analogy is, you know, we see Waymos around London now, they've been around the Bay Area for a long time. That took time because people behave differently around Waymos than they do around other drivers. So Waymos have to like, collect this data and gradually adapt. And humans over time change their behavior on the tool, right? I think initially they're a bit aggressive, a bit afraid in a later. It's just kind of comfortable because they know it won't do anything unexpected. In fact, in many ways, it's maybe an always, it's a better driver. So if you like the levels of autonomy there are just like increasing. And I think the fact that we have so many advanced workflows just allow you to go look, all right? Can we now do the highway? Can we maybe try the highway at 120 miles an hour? Because we all know it can be done, right? It's just a matter of like, you know, we need to trust it to be right.
And I feel like in a fully agenteic world, you can do that a lot more easily. And I think we've crossed the chasm of where we now have enough data and the abilities to put it back in the hands of our customers to do it. Because they're really always the ones with the context, right? Absolutely. And again, cut to the middle man. Like, they have the context. I have a friend. The middle man. They have the context. They have a run and they can just make things happen. No, I am super excited about running and everything that we're doing with your team, I think that our customers should just expect this thing to change twice a week, going forward. And I think that's one of the greatest permissions that we gave ourselves here. So, yeah, I just, I'm thankful to our customers for being so excited about it as well. Yeah, I mean, it's like the privilege of my life really working on this. And it's also just so much fun. And again, as I said before, like, my mission is to just share that fun with our clients. So they can also like experience, experience this, these feedback loops, experience the growth,
experience like immediate results with like zero friction. And just like love using the platform. Like, we want the platform to be to be part of people's routine, right? They come to it every morning. They click on these recommendations. They watch the stats go up. And they just, you know, it's such a good feeling, such a, every like hit of dopamine in the morning with your company. I look at it every day and like, I can't believe how fast it happened, right? I thought it would be just the gradual up there. And then I'm like, oh, wow, like person from Telco, in the Middle East and then someone in a software company in America and then in European bank have all done more with rent that I've done with cloud code. I'm starting to feel inadequate. I mean, I mean, I mean, I mean, easily like, I think 95% of pushes the production and I've done through brand and it will just go up. There'll be more and we'll see better and better results for customers and for everybody calling customer service, right? We have, we have a mission to fix customer service and it actually is looking really realistic. I think we finally have lightsabers, right?
I think that we work with for the fight and so are customers. So I guess we'll report back and again tackling this problem in two ways, right? We have the harness like we allow you to improve the harness and we have the model layer as well with Dialogreason1. So really, I mean, those two improvements are happening at the same time, but we see them both just pushing all the stats up. Yeah, I mean, yeah, Dialogreason1, I think we'll have a whole different push these weeks, but you know, it's the world's fastest reasoning model. It's the only one that's smart enough to do these tasks well while being, you know, both quick, reliable, not hallucinating and, you know, the more people build complex use cases would run and get them into production. The more data we have to make Dialogreason1, two, three, I don't know what. It's feedback clip. Yeah, feedback clip. All the way down. Looks, loops here again. Yeah, awesome. Larkadius, thank you so much for today and we'll check back in like three months and see how much further we got in.
Thank you. Thank you. Thank you. It's been a pleasure. Thank you. Thank you all. As always, like, share, subscribe and we'll see you in the next one.
More episodes
More from Deep Learning with PolyAI

Who's coordinating your army of AI agents?
Deep Learning with PolyAI

Why should CX leaders care about MCP?
Deep Learning with PolyAI

Can AI really hear a call the way a person does?
Deep Learning with PolyAI

Is word error rate just a vanity metric?
Deep Learning with PolyAI