Skip to content
TrackPodcasts
technologySep 9, 202656:30

1037: WebMCP is here (and you should care)

About this episode

Scott and Wes talk about WebMCP with Sarah Drasner and Dominic Farolino from the Chrome team, which is the new W3C standard that lets your site hand real tools to an agent. They cover the security model, headless and multi-tab agents, and how to wire it into an app you’ve already shipped. Show Notes 00:00 Intro 00:46 Welcome to Syntax! 01:18 What Is WebMCP and How Is It Different from MCP? 02:32 WebMCP for UI: Enabling Agents to Use Websites Directly 06:05 Brought to you by Sentry! 06:43 Is This Just a Chrome Thing? W3C Standards and Open Development 11:49 WebMCP Security: Annotations, Hints, and Multi-Layered Defense 15:26 Best Use Cases: E-Commerce, Video Editing, and Complex UIs 16:27 The WebMCP Challenge: 5,000 Entries and Creative Applications 19:20 Headless WebMCP: Running Agents in the Background 22:43 Multi-Tab Scenarios: Using WebMCP Across Multiple Sites 24:23 WebMCP and Security: Safe Origin Policy and Agent Containment 28:09 Performance and Core Web Vitals for the Agentic Web 38:11 The Future of Web Monetization: Ads, Commerce, and Agent Payments Universal Commerce Protocol x402 40:53 Building an Agent Platform: Extensions, Harnesses, and Browser Vendor Roles 45:19 Support from Anthropic, Safari, Firefox, and Other Players 49:11 Getting Involved: Origin Trial, Community Feedback, and the W3C Spec Ora 52:45 Sick Picks + Shameless Plugs Sick Picks Scott: Wes: Dominic: Steven Pinker: The Sense of Style Sarah: Vintage Story Shameless Plugs Scott: Wes: Dominic: X account Sarah: 25% discount with code Community25 at AGNTCon Hit us up on Socials! Syntax: X Instagram Tiktok LinkedIn Threads Wes: X Instagram Tiktok LinkedIn Threads Scott: X Instagram Tiktok LinkedIn Threads Randy: X Instagram YouTube Threads

Get every episode summarized

Each time Syntax - Tasty Web Development Treats publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

667 searchable segments. Every word is indexed and playable.

1037: WebMCP is here (and you should care)

Syntax - Tasty Web Development Treats

0:00
56:30

Full transcript

Syntax - Tasty Web Development Treats1037: WebMCP is here (and you should care). Machine-transcribed; use the interactive transcript above to jump the player to any line.

And I think Wes, you make a really good point about the multi-modality of it. Not everybody's going to be using the web the same way. Like we've known this forever. I think some people are like, oh, the web will go away. It'll just be people using it ahead this way. And then people go, no, nothing's going to change. And I think the truth is like these experiences are going to be done all sorts of ways. Like even a shopping, I'm shopping as like a leisurely activity. Then I'm going to be a human in the loop in that activity because that's part of the point. But if I'm just trying to get some shirts for my son that he knows the size of and he knows what he wants, then that, you know, I'm okay with an agent doing that for me. Welcome to Syntax. Today we're talking about WebMCP. We got two folks from Google here to talk about WebMCP. And I personally am really excited about this. It's starting to make its way into clients. And I don't think a lot of people really understand what it is or as a developer,

how you should be integrating this into your site if this becomes a thing that everybody implements. So I'm stoked to talk about it. With me, we have Dominic Ferreino and Sarah Drazner. They are both from Google working on Chrome. Most people listening to this, I would say everybody listening to this episode, knows what an MCP and MCP server is. Everyone's slapping them into their agents. Can you just give us a rundown of what is WebMCP? And how is that different? Yeah, I mean, I think people get confused a little bit because of the name. WebMCP is on the client, not the server. And a lot of the kind of similarities are when we are thinking about the principles, like exposing tools to agents, things like that. But they're actually pretty dissimilar. The governing bodies are dissimilar. One is the AgenteGai Foundation in Linux. That's the MCP for WebMCP. It's to W3C, like other standards bodies. For WebMCP, you are really only working on it on the client. But that doesn't mean that it can't negotiate or create server actions,

like kickoff arrest or GraphQL API interaction. And then also the way that people are using it agentically. You do have headless scenarios that you can use WebMCP. But we're also thinking about a new surface area of cobras. And cobras is where we're spending a lot of time, which is when you have like a human and agent on the loop collaborating on a web experience together. Domicide. Yeah, I think that was perfect. I think the way I try to describe it to people is when you give MCP tools, traditional MCP tools to your agent, it allows you to just enable the agent to do better capabilities inside the chat. But we don't have a really like an MCP for UI for using the site. Like I'm on Google Sheets, I want to do this complicated thing in the UI. I can't remember where all the 10 clicks that functionality is hidden behind. Sorry, my cat's on the background here. And so I want to be able to talk to my agent and say, OK, I really want to add conditional formatting to this column. And can you do that for me?

If we've ever seen agents interact with websites like Dom and UI right now, it's really slow. They're like screenshoting. It sucks. It's awful. It's such a terrible experience. And it's expensive, actually, too. It reenters the inference loop every time you got to click a button and do anything. But the agent is thinking in terms of atomic tasks. Like I want to add a video track to this YouTube editor. I want to do something of this column at Google Sheets. And all that code actually exists in the client. And WebAppCP is sort of like finally a way to let the agent get access to it and actually get to that code. And translate its atomic tasks that it's thinking of into actual tasks that the platform or site that the client app is capable of handling. Yeah. If you look at where we're at right now, like you said, if you have an MCP server for something and you want to interact with it, you're typically just getting like chat back and forth. And a lot of cases that that's awful. So now we have like MCP UI where like you can like have these little widgets that make their way into the chat.

And that's okay in some cases. But in some cases, like you know what the best UI is? The existing app is the best UI for that. So in that case, or if there's not, then you're like, okay, well, like, hey agent, use this browser. Or like the site owner will strap some crappy little chat box onto the side. I don't want that either. And like I've been like beating my drum saying like this is like clicks and clankers. Meaning that like I just want to use the UI myself, like I'm normally doing. But then I also want to just be able to do agentic stuff with that application, with that website, with my own. With that same functionality. Yeah. With the same function. And I want to see it. I want to say, like another example is like if I have like six expenses in my bookkeeping software, I don't want to click them all and edit them all and select the taxes for every single one. And in some cases, I do want to click things and drag them into wherever they want. But if it's like a batch job, I just want to say like find all the expenses from my

cell phone carrier and mark them as having Canadian taxes. And then we'll just go off and do that for me. And then I can visually see what is happening in the UI. Yeah. Well, and I think like one thing that people don't often compare it to, which they should is DOM actuation. Like if you don't have something like WebMCP tools exposed, the agent is inferring what it should be doing from the DOM, from the accessibility tree, from screenshots and things. This actually gives tools to developers to be able to say, Hey, this is what I want an agent to be doing with my site instead of it just happening to them. It definitely lets the developer kind of express the capabilities that it knows an agent is going to want to get at from its app directly to it so that it doesn't have to constantly do this like inference thrashing and the DOM actuation and the accessibility tree reading for every repetitive action and check if the button moved and all it did. And then what happens? It can kind of let it let it think in atomic actions and also act in the tool. And if you want to see all of the errors in your application, you'll want to check out

Century at century dot IO forward slash syntax. You can sign up today and get two months for free. Century is just a really incredible tool for not only tracking your performance, making sure your application has no bugs, but even just seeing what goes wrong when something goes wrong because things go wrong all the time when we're coding. And you don't want a production application out there that while you have no visibility into in case something is blowing up and you might not even know it. So head on to Century dot IO for slash syntax. Again, we've been using this tool for a long time and it totally rules. All right. Yeah, that's such an important thing. You one thing I when I've talked to some folks about Web MCP and they're arriving in Chrome, people say, is this a Chrome thing? Is this another Google thing? You had mentioned that it's involved with the standards board. So can you touch on the status of all that? And is this just a Google thing? No, it's definitely not a Google thing. We're working with with other folks in the community, other model vendors.

We're working with Microsoft on this. We're developing the API as an open web standard in the W3C. It's managed in the what is currently the web machine learning community group, where some other agentec kind of APIs are for the web platform. And we've been in close contact with model providers and also developers who are going to have to use the API and other browsers that we'd love to have implement the API. So there's been lots of discussions in the normal standards bodies, gauntlets of consensus and controversy with other vendors and developers to get that pushed forward. And so I think that's, we're still early stages and all of that sort of thing. But so far, it's going really well. And we've gotten a lot of good developer feedback. And clearly, I think the community has proven the need for something like this. Every time we present this to a model vendor or an extension provider that augments Chrome with some agentec capabilities, they're always like, this is what we want. Like we want an actual clear capability layer to talk to the site instead of have to like, burn the users wallet through inference loops and screenshots and all this stuff that

takes way too long. And I know you guys work at Google, but the idea was, is that like whatever client you're using, right? Like if you're using, like it's in chat GPT desktop right now. If you use the built-in chat GPT desktop browser or I think the Chrome extension or whatever, it will surface those tools via like, we'll talk about the API in just a second of how you as developer can implement it. But it's any client, right? This hopefully eventually will work in like the cloud desktop or Gemini or like any client that wants to be able to do this. It should be able to just open up a browser and say, ah, here are the tools that I have available to me. Yeah, if people haven't seen already, chat GP, chat GP and open AI came out with a big announcement that they're supporting WebMCP as a first-class citizen. There's a big hackathon. We're involved in that even though, you know, it's across many different companies like Fursal and CloudFlair and a few others. So yeah, it's more an ecosystem thing.

Edge is definitely proponents of the spec. They're worked on the first versions of it and things like that. And I think that the chat GPT sort of hackathon that they've announced with WebMCP is going to be a really good way along with the origin trial that we're running in Chrome. It's a really good way for us to pressure test the API and understand like how developers are actually using it. What granularity of tools are they writing? And then it also puts a lot of importance on the ability for the ecosystem to adapt and come up with like real evaluation metrics and determine like, I made this tool. Did it help the agentic path on my site or did it hurt it? And like, I made this change to my tool. Did that help or did that hurt? And so we're thinking in terms of like ecosystem activation to understand what it takes to write a good tool and what the API needs to do to expose tools, you know, in the most effective way to the agent brain so that it can have a really effective experience on the site. And so we've seen a lot of feedback from our origin trial and this hackathon coming up, that's going to help inform the API shape and all of that stuff. So the API is, and the last time I implemented it was both an HTML API and a,

like, a declarative and an imperative API and like a JavaScript API. So how somebody has a website, they are a JavaScript developer. How do they implement this into their site so that agents can use their website? Yeah, I mean, I try to think about it in a couple of different ways, but this is also why system design is really important. You want kind of task level milestones. I think a failure mode can here can be that people kind of attach to every little micro interaction on a DOM element. That's not exactly what's intended here. We're thinking about more like stateful or potentially effectful kind of tasks that an agent can use. We're actually collaborating with some large frameworks to investigate how we might infer tools just from your existing application authoring. So you could just write JSX, TSS, and then there's a separate potentially like reconciler pass that infer some of those tools that you can look at. But if you're working with frameworks like most of us are, you want to expose tools potentially via hook. I made a Web MCP tool hook that's on NPM that us in Chrome will keep up to date with the

API. So don't leverage the existing DOM structure, think in agent actions. I would also want to leverage existing APIs in your application like rest or graph QL. If you have those, you can, those can be invoked with a tool with a really good description. We have a lot of annotation hints and those are also really important. So one really important part of the spec for me is that you, the ability to keep your site secure. So you kind of also want to go through and look at, are you exposing user generated comments somewhere on your site? If so, you want to read only hint, you'd want like user generated content hints because the agent can then say, okay, I shouldn't look at any of those and take action. And that can guard against things like prompt injection attacks or, you know, agent traps or things like that. So you might want to think through if you're a site developer, like, where are you taking content that you're not sure somebody might be doing some of those harmful actions?

Yeah, the security side of it is, is pretty, pretty vast. I mean, like, you know, historically when we work on Web platform APIs, we kind of have like two consumers. So it's the browser and then there's the developer, which is, you know, writing experiences for the user. And Web Apps APIs unique in that, like whenever we adjust the API or add new things, we have to make sure it works well for developers. Like, does it make sense for you to add this attribute on your tool or like, is it okay if it tools unregister at this point? But then we also kind of have to ask the inverse questions to models, whereas, you know, it's like, can you, does your security harness benefit from consuming this attribute? Does, does having a tool market self as read-only or untrusted is that useful for security harness and an agent? And, you know, if a tool unregister here is not too much stashing, or do you guys have a caching mechanism? And so there's kind of like two different halves of this API and getting feedback from both sides has been, has been super important. And like, like Wes mentioned, there's the imperative API, which is just ordinary JavaScript allows you to basically pipe in a JavaScript callback as the implementation of a tool so

that the model can sort of call directly into JavaScript that the developer chooses to expose. And then there's the declarative API, which is like annotations, like, like, attributes on existing forms. It's probably no surprise we find developers most interested in the imperative API so far, where most of the site code lives. But I think declarative also may have an important place because it's an easy lift for your, like, you know, crappy local, like, government car registration website that is like from 1995 and has a bunch of forms on it to just like sprinkle a couple things on and then kind of become a little more agentified and allow, like, those kinds of use cases. So yeah, we're trying to collect feedback from developers on both API shapes and see what sticks. Yeah, and so many of us think in interfaces anyways, right? These are the interfaces that the users are interacting with. And if we view an agent as another user, you know, just being able to sprinkle that declaratively on top does feel like a big win. Yeah, I see like, they're being just like a 10 stack query plugin for this type of

thing, you know, because like I already write all my, all my mutations, all my queries of like get items, update item, delete item, you know, all that, all that crud that, that we're, we're used to. And then I, I, in most cases, I probably want like 90% of those also to be available to the agent so they can just programmatically call it. So like, I see a place where you maybe aren't doing a whole lot other than just throwing in like a plugin or a hook or something like that. Yeah, and I mean, it also allows us to, you know, if you do exactly what Wes is talking about, you can also return structured errors back to the LMS, you know, you plug in exactly the what success happens and the reason and so that the agent knows what failed. So you already have existing, you know, mutations, actions, dispatch, kind of logic, you can leverage that existing logic for some of your tools. Um, what, like what kinds of apps are, are like, you see like the most obvious benefit from this type of thing? I know it probably can be used for, for anything, but is there any like specific use cases

where yeah, that makes a lot of sense. Well, we're collaborating with people like Shopify has integrated. So we're seeing a lot of use cases for Shopify as, you know, obviously e-commerce is a big one. Yeah. Right now we're talking with YouTube about their in-page agent being able to surface, like if you ask YouTube, um, all of the things that people might not normally get exposed to that are on the YouTube page, like playback settings. And I mean, that's one that people usually find, but there's all sorts of settings in YouTube that people don't necessarily find easily easily. So anything where discovery could be limited and an agent can kind of surface those discovering mechanisms, um, Instacard is another one where people probably want to like agentify or make their shopping experiences a little bit easier. Things like that. We're seeing a lot of benefit from. So yeah. Yeah. And I'm sure you're going to uncover a lot too with this Web MCP challenge that's going on. So we have this new Web MCP challenge. The deadline is September 3rd. So by the time you're hearing this, the deadline will have passed, but it seems like there's a massive

amount of submissions or at least people signing up to be potentially submitting this already. I would imagine you're going to see some really cool stuff. And Sarah, you're a judge for that, right? That's right. Yeah. And we're looking at like creative uses for a Web MCP. We're looking at for the good of humanity. We're looking at do you structure your tools well and make sure that everything is, you know, kind of copacetic and easily surfaced by an agent. So there's a number of factors we're taking into consideration with the judging. I am really excited. So far, I think last I checked it was like 5,000 entries that already gone through. Oh, yes. Yes, one of those will be mine. So hopefully my tools are set up for an end of the exciting to use the deadline. Well, one of the demos that you gave was like a 3D modeling software, which 3D modeling software and video editing software, I think are two huge use cases for this because those are UIs where like I'm not using a chat for that no chance. But like we, all of us on syntax, we use DaVinci Resolve and we all of the DaVinci Resolve

MCP running because in some cases it's way easier to just type in the box what you want it to do. Then to figure out what the crazy clicks and whatever that you need to do. And I love I've been calling it you're calling it co browsing. I'm here to tell you it should be called clicks and clankers. Meaning that you know, because we had we had bricks and clicks when the like the web was coming, you know, you have the store, but you also have the online thing. Now we have the human and the agent clicks and clankers. Nobody seems to be it's not catching on. So I'm reaching out to you to for that help. What make it happen? But no, I agree with that that general use case. Like I do a lot of like makerspace stuff and like, you know, CAD modeling and like, there's only so many times I'm going to like relearn how to do the same thing and like fusion 360. And like I'd really just want the chat to be able to like, okay, I can describe it exactly what I want. I can kind of probably get there on my own, but it'll be quicker if you could just, you know, take my intent and map it to actions on the actual site. And so I think that kind of use case where these complicated configuration UIs,

but I also still need to be involved because I need to like double check the end result or see the measurements or whatever that that's I think the most, the most like immediately useful. But like we said, it's also pretty useful for headless scenarios. Like a lot of times if you're just messaging your, your, you know, codex bot on telegram and telling it to do something, it's going to open up a headless Chrome in the back and talk to it through CDP. If every site it uses has WebMCP tools or can talk to the service worker tools inside WebMCP and do some background executions, it's going to make even all of the headless scenarios, like just as easily if they're, if it's not talking directly to the server. So I really, I think it's a unlock for a lot of different use cases. Click sync, Lankers and headless. That headless one because I was talking to some people earlier on in WebMCP and there wasn't like any headless thing now, but now you're saying there is. So like you're saying that I could technically just have like a, like a box running somewhere with like, or or my, just leave my laptop open. And if I have my websites open or they can open them, then I'll be able to access those sites that have WebMCP.

I think so. Yeah. Like if you're messaging your, you know, your bot on some VM somewhere and it's really doesn't present a lot of UI to you unless it really needs your intervention or something like that. And I go tell it to register my car or update my license registration or something like that. And it navigates to the site, you know, the state of Massachusetts does not have like an MCP server. I'm sorry. And so, but maybe they, maybe we convince them to drop a couple attributes on their forms on their site. And, and now, you know, my bot, wherever it's, you know, living headless or not, you know, behind a hologram VM or not. If it can, if it can actually the site through WebMCP tools exposed through CDP, it feels headless to me. And it doesn't really make a difference from the agent's perspective because it's just interacting with the site in any, in any way. So I think that's kind of the idea. I mean, if you never have to go to a site or whatever to register your license or like, I would love that. Yeah, exactly. Sarah promised me a promotion if we can get all the chance to register. So that's all.

This is a perfect use case for something I'm working on because right now, every single time it does popped open dev tools MCP specifically, I have a Mac mini right here behind me that you might be able to see it flashing occasionally. That's because it's running dev tools MCP to export a video from this application. And it's like such a perfect that if I would never need to see the interface for this, the process should just be able to do it headless. So that's really exciting to hear. Yeah, I mean, the dev tools for agents stuff is really exciting. If you're working with WebMCP, like sometimes people, we get, you know, a lot of people asking us, like, how should I be debugging this? I think dev tools, Chrome dev tools for agents or dev tools MCP is a really great way to like send an agent off to do a bunch of things. It also invokes Lighthouse for agents. It also can perform audits for you. And so if you're not using that already, it can integrate with a lot of different models, including the frontier models. So that's a really good debugging journey. I think some people, maybe not everybody knows that there's also a Chrome extension that you can use.

So I also tend to use the Chrome extension with if you look pop open dev, dev tools and the application tab, you can see all of your tools listed. And then you can invoke the tools on the page and automatically see right in the page feedback for how those tools are getting executed in some observability. So like there's a number of different ways you can do it. Some are like headless, some are like directly in page and those are cool. And what's that Chrome extension called? Is that just the Chrome dev tools MCP extension that's been rolled into that? Or is that a set is that the separate Web MCP extension? Yeah, there's one that's right directly in Chrome dev tools. That's the one that I was mentioning in the application tab. And then the one that Scott is talking about is a more headless model that you can just run in the background. And you know, it kind of operates sort of like playwright and puppeteer, which are also good debugging tools as well. Well, that's awesome. Another question I had, I don't know, maybe six months ago when you first started talking about this was like, like what about multi tab?

And like I assume that you have two tabs open. And both of them expose Web MCP tools. Your agent would be able to use both of those, right? Like one example I had is I built a shopping list application where you could add stores and you can add items to each of those stores and you can mark them off yet, yet, yet, right? And then I had another recipe website open. And I wanted to get all of the ingredients from that and put them into my shopping. Right? So that's two totally different websites, two totally one web Web MCP one was simply just scraping. But like I should be able to do that, right? I think yeah, like the, you know, the agent will be able to use Web MCP tools where they exist and and make use of that site functionality when it can. But ultimately, yeah, it's the agents sort of prerogative to understand what, what origins to reach out to what tabs it makes sense to interact with to fulfill. Kind of the user journey. And this is actually one thing we've been, we've seen a lot of confusion about with Web MCP. When folks are reviewing it from like a traditional web platform perspective, it looks kind of like a wonky API because it's, it's sort of like the first of its kind, like really facilitating

agentic use of traditional web content. And so I think a lot of people mistake it for like, oh, this is the agents on the web API. And I think from our perspective, it's really like, no, this is the like, let's give developers a chance at presenting something that's sensible for agents because agents are already on the web API. Like it's agents are using the web, you know, regardless of whether whether Web MCP exists or not. And so, so we've been like really trying to understand, you know, like there's been a lot of confusion, for example, for security. It, you know, when people think through what it means for an agent to use a Web MCP tool, kind of like, oh, like an agent can use this tool. But what if it has like some, some stuff from another origin, like in its brain and it wants to like share that information with this tool is that's kind of like violating the same origin policy, right? Like that's that's kind of scary. And that's that's violating chords. Like what does that mean? And I think it's a little hard to think about because we don't mind users violating the same origin policy. I'm the user. I can see all my cookies. I can see all my tabs. I can see all my origin data. But that's because the product is kind of serving me. And so in one sense, the agent is kind of an extension of the user and sort of punches through the traditional like web sandbox security model.

At the same time, like users are not comfortable, just like, probably giving the agent its entire identity and letting it assume it's full persona like on the web and just browsing to whatever it wants. And so there's only so much of this we can control from the platform perspective, which has like the same origin policy and cores and all that. And we're starting to think about what it might make what what it means to produce like an agentic platform like in the product. And instead of thinking about like the same origin policy, what is what is a safe origin policy look like some some agentic model browser so that the agent can know like, yeah, I should be able to assume the user's identity on these four sites related to this task. But I shouldn't be able to do everything. I can't read all their cookies. I can't go to their bank and start making transactions. And so we're starting to think through what it what a what a capable agentic web harness looks like in a browser that integrates somewhat with the web platform and some of the product. To actually make agents on the web secure and also facilitate the use of web mcp through traditional web platform content. So it's a complicated model. But

Yeah, and like to zoom out for a second. So there's, you know, as Dom said, there's only so much you can do on the platform side and for site developers, but we're also talking to agent developers and our own agents about what's potential there. And I think one thing to get people to really understand is that we this won't be solved by one thing. It has to be a multi-layered event defense strategy because agents can be somewhat non deterministic and also you need both the site side and the agent side. So some things that we're thinking we're doing within Google that we're thinking about open sourcing and making more of a standard are things like prompt injection classifiers. So they can identify attacker instructions in content before instructions kind of go out critique LMS. A lot of people know about like a secondary judging LM. You can imagine that being applied to like web surface areas. And then also just like restricting origins like you may want to have like that LM. That's a judge say, OK, you were supposed to go book travel for me and you can go to Expedia and United. But why are you going to my bank?

Why are you going my health site like to make sure that we're not going off to origins that they shouldn't. And finally, there's like this kind of special agent containment layer that we're thinking about exposing to the community so that agents can have like more of an identity that separate from the user. Like so far, the agent is you. But you could imagine that you also you might want to have an agent identity that separate from you that like can only spend $20 a week only has access to some information about you only have some passwords things like that. Yours $20. Yeah, I gave my shop of my 20 bucks. They have a shop. If I have like, if you go to any shop of my website, they have like a 4th slash agents.md with all the information about how how to to communicate with it. And I you have to like explicitly give it a little bit of money and let her rip that that was funny. Um, can we talk about something that makes me sad about the web is the performance where we went in like the process of like a year and a half. We went from it really matters how quickly your key ups happen and you should not block the thread and

that we have all of these web vitals about making everything super fast. We cared so much about all of that stuff. And then these agents came around and it takes like $4 and three minutes to click a link to do something. And I was like, like this experience sucks. If you're if you're if you're out somewhere and you don't really care, but when you're waiting on it to do work. That experience is absolutely awful. Is there is that obviously we have to see people get that better, but will we bring vitals to agents. Yeah, I am actually kicking off a like what would core web web vitals look like for an agentic web thing. So there's things like, you know, if you're using, you know, Claude or Gemini or Chatchy Keepee that time to first token is a new metric that everybody's kind of looking at. And that kind of those kind of streaming delta's because you don't have just like the second it goes, but you also have the second that it like streams all of the possible input. But in co browse or wait, what was it clicks and clankers clankers like second clankers. Okay, I got it. In clicks and clankers approaches.

You the thing that's fascinating about it is that some of the old meals and things still are like are still some things that we're seeing here. So like one second per tool call. Five seconds before you want to see an entire action go through if you're not familiar with that. Those are some of the earliest human computer interaction things like people don't wait longer than five seconds for a web page. That actually has worked in like clicks and clankers approaches as well. But the thing that we didn't have in there is that when you're watching co browse. You're not just judging the agent and the time. You're seeing how fast it is compared to you. Because you don't want to be doing an experience on the web where you're like watching it and you're like I could be clicking that faster. I could be doing this whole flow faster. And so that's the first time we're seeing metrics that might be comparative to a human. And the real trick is exposing that to developers like in the past like we could let developers do things on their site.

And we have like performance observer and that kind of that kind of thing that let's let's the developer get access to how long that navigation took and how long that animation transition took and they can understand when they make changes on their site. If it affects the user experience through real user metrics. The agenda side of things is a little trickier because like a lot of times the the success metrics are kind of locked up in the agent's brain. It knows like how many model turns it took to fulfill a user a user journey and how many user journeys. You know, it you know, we're completed with a web mcp tool and how many tools were involved and how many tools confused the agent. And so we're also trying to think through super early on this, but we're trying to think through some ways to expose some of those metrics to developers as well. Whether it looks like some agent performance observer or some reporter API or something like that, where we can let developers actually measure the effects of the agent to targeted things that they're doing on their site. So they can know if they're actually having a positive or negative impact because right now, like it's it's slow for a lot of reasons. And web mcp tool calls make that a lot faster, but some of them are still slow. But one of the like really big challenges here is like the whole things opaque.

We kind of really have a hard time measuring like how how a changes I make to a tool impact the success rate and how they impact the latency and what's confusing the model or not. We need the developer to transparently understand that kind of stuff on their site or else they're flying blind and it's going to be impossible to make changes and measure against them. Do you also have a way to like measure just for like regular people using it as well of like success rates and what not. Like I think back I was booking a hotel on Expedia a couple of weeks ago. And like I wanted to buy I want to book like a suite that a separate room and that wasn't like a filter on Expedia. And I was trying to do it entirely agentically and I was like this is awful. I need a map. I need photos of it. I need like all of I need the UI. It's not a very good experience, right? And like all these tech bros are just like, oh yeah, I booked a flight for me and I bought red shoes online and like that. That's not how regular people do do their work, right? So like is there some sort of like measurement that you're doing with like regular people as to like whether this is something they use in a sticky enough.

I do think that for some of the like crux things that we're working on for agents. We are invoking like we have a bunch of UX researchers who are looking into this so that we do these types of studies both with like real users and then also by doing analysis of the web and like we have a lot of data because of Chromium being used by so many people that we can kind of leverage here. I would say that we're pretty early and like full stop. This is the way that everything works. We do have some targets and like Dom said, I don't think that anybody's at those targets yet because these experiences are so new. People aren't even used to building out a product experience that might incorporate something like this flow like that. Even just like for PMs of a site to like think through. Okay, what does that look like if somebody is going through a cobras flow. Sometimes we've had these like deeper partnerships, but we do need like site developers and product managers and everything to incorporate that kind of thinking as these.

You know new agentic surfaces evolve and I think West you make a really good point about the multi modality of it. Not everybody's going to be using the web the same way like we've known this forever. I think some people are like, oh, the web will go away. They'll just be people using it ahead this way and then people go, no, nothing's going to change. And I think the truth is like these experiences are going to be done all sorts of ways, like even a shopping experience. If I'm shopping to like, you know, as like a leisurely activity, then I'm going to be a human in the loop in that activity because that's part of the point. But if I'm just trying to get some shirts for my son that he knows the size of and he knows what he wants. Then that I you know, I'm okay with an agent doing that for me. And I think a lot of these things are going to evolve as we, you know, XR, whatever experiences involve with it. Like, oh, computer, show me what I look like in this shirt. Okay. Now this shirt, you know, I mean, like who knows what all that's that it's going to evolve as we go. Yeah, totally. It's definitely one of the challenges of API design and this kind of initial error.

Like all of this agenteic web space is super nascent. And so it's it's hard to tease out some of the patterns, you know, we're seeing and, you know, derive what experiences we can from from actual MCP and see what makes sense over in web MCP and, you know, everything's everything's new and moving so fast with the developer side and from the model side. So it's kind of hard to pin this down and understand exactly what what makes sense. Which is why the origin trial and the hackathon and that kind of stuff is a really good way for us to collect experience. Have you heard from any site owners who are like resistant to this type of thing because like that's like another weird spot is like if I'm an airline. I don't know. I certainly know and this is more like the MCP server way. If I'm an airline, if you ever tried book a flight, they try to hard up sell you on absolutely everything. They, oh, what if you get sick? $20 for for the insurance and all of like these like like black tactics. Yeah, GPT about to be defensive about that right?

Oh, you might need this. So therefore we added this. Yeah. So like I'm I'm wondering like the airlines doesn't just want to be like a like a straight up utility for just vending out the cheapest flight for this type of thing because they want to be able to make make more money as well. So have you heard any like push back from people who own sites are like, we don't want this. Yeah, there are a couple of cases that I probably can't disclose on a podcast. I do think that there's, you know, when I look at the interest, it's far more people wanting to expose Web MCP and tools because they want the agents to be able to discover and not fail. And like I think what they're really worried about is like, oh, okay, if it's just don't don't actuation and things like that, then we can't guarantee that they're going to have a good experience. They might go to some other place or something like that. So mainly it's interest, but there have been a couple of outliers of people wanting to be like, maybe we just say everybody go away and like abuse in that kind of direction.

I think in cases like that we're still like Dom said, we're so early on and trying to figure out what those loops and experiences might be. Do you empower the user? Do you listen to the site owner? Like if the user really wants to be using an agent, are you going to flat out tell them no like I don't think that the industry has a collective answer for that yet. But typically we try to like put the user first. And so that you know is a little bit of attention. Thankfully it's not that common. That's good. And of those ones that are common. I wonder how many of them are just trying to protect their business of like like like either like come along for the ride on the agent world or or like be left behind a lot of people are saying. So I don't think that's all of them. I think there's certainly a lot of people who can like like I guarantee Amazon could say, nope. None of this and that would be a big problem. Same with like Apple Pay at Walmart Walmart just says, you know, you know, and like that's a big deal.

They're big enough to do that. But for a lot of people, they're not big enough to actually push people around like that. I did not know that Walmart does not take Apple Pay. That's that's just to mean they haven't Canada for years, but Canada is a great country. But apparently US is just getting it now. Yeah, just getting it. Yeah, end of 2026. Yeah. It's a hard space though, right? Like should my agent be like watching ads for me? If that's like what the site wants, like, you know, probably probably not or like should it be clicking on ads that I think I'm useful. Like there's a whole monetization model that's like totally naced in here about like what what does agents on the web look like for the traditional funding and attribution and refer and all that kind of stuff. Model like that is that's like a whole new space. It's it's beyond web MCP. I mean, it's web MCP. Like there's a role in it. But I think, you know, we see we see a lot of different corners of the industry rallying to answer similar questions. Like this is kind of, you know, also related to the USCP, the commerce protocol spec.

You know, like there's a lot of upselling there. Like how does that integrate? You know, there's got to be answers with that. I don't know, maybe maybe someday they'll be like an ad viewing spec or you can view your ad through MCP. No, I'm just kidding. But like, you know, it's there's a lot of a lot of new things here. And then obviously that the natural thing is for like large kind of like business conservative enterprises to to maybe resist it and keep their traditional model. But I think ultimately most most vendors of that sort will end up figuring out a way to integrate industry solutions to, you know, enable new ways of monetization and new ways for agents interact with their sites. But exactly how is it's unclear yet, but yeah, it's going to be an interesting future for sure. We are working on some things internally, but I don't think that they're totally ready for prime time. Maybe we, you know, send you all links in the future. I think. Send it a star way. I didn't realize this. So like for people listening, that's UCP.dev universal commerce protocol. And then I've also been keeping my eyes on the X402 project, which is like,

agentic payments. I know Cloudflare is rolling out wallet soon. I know Stripe has their wallets in the US, which your agent can spend money on. So it's that's obviously not part of web MCP, but it is kind of related as the how agents use the web without bankrupting everybody. I mean, we do examine this as part of health of the web, because in order for people to keep the web healthy, they need a way to make money off of it. So if you have like beyond commerce, if you have a content site that makes money off of ads in order to show and display content, then you're kind of going towards this subscription. Why can't I say subscription? Subscription model for those things. Yeah, totally. Yeah, I mean, because the AI agents are just lurping all that stuff up now and there goes your income. So yeah, yeah, it's a different world, I think, for a lot of sites that are trying to work on that model right now.

Yeah, and I think it's definitely, I think it's important that the web evolves to try and meet the moment. Like I think there is a lot of, you know, understandable and natural pushback in general about the web evolving too fast, writing too many AIs that are tailored to a gentle experiences. And you know, these things have a necessary long tail of controversy that we have like the Web MCP and the PromT API and so forth. But like I think it's ultimately a good thing that we're focused on trying to figure out what the web's real place is in this kind of a genetic world because like I'd certainly rather us all argue about a healthy web that is still relevant than a dead one that died because we didn't keep up with any technologies that are coming down the pipeline. I think that's like a really important thing to worry about if the web doesn't really meet the AI moment and figure out what it means to present kind of kind of two platforms now, right? Like traditionally, it's always presented the web platform to developers into the user. This is things like readable streams and module scripts and anchor positioning. But like now we probably need some like agent platform side of sort of things where you know I kind of envision a future where you can plug any agent brain into your browser.

And then that agent brain can sort of get like whatever containerized access of you know to your to your site data or things that are useful to you as the user. And be constrained by the guardrails that the the browser's agent harness kind of is able to provide. And then whether that is partition credentials or read only views of certain you know parts of the users personas that they can actually on their behalf. And so on like all of these kinds of things like you're going to have to know we're currently working on them and we're currently trying to figure out what they're going to look like in the future and understand what it what it means for the agent to kind of have a sort of the browser to play a role in sort of like this agent platform like space. Because that's kind of that's very new like we've only had platform and product. And now like product is is kind of containing like a bunch of agent primitives and we're actually thinking about this from the extensions point of view on Chrome as well. Like it's it's really widely known that like a lot of agents that are living in extensions on on Chrome. Like they need like full access to the page and so they get like accessibility tree and screenshots and all this kind of stuff.

But to do that they kind of end up tripping over the debugger API and then it kind of flashes this like scary banner and enterprise clients don't don't like it because the agent has so much direct control over the page through an API that was never really designed for agent usage. And so we're kind of stepping back and being like what does it mean to like re factor and redesign all of these things to provide like an actual agent platform for agents to plug and play straight into your browser. But but meet the safety standard that the browser is known for through whatever safe origin policy or harness or tool tool set or read only view of the users state or data that makes sense. And I think that that's like one really big thing. I think it's important for browser vendors to focus on it. It's actually it's the reason why we see a lot of these new agentec browsers sort of come and go from other companies. Like they're you know spinning up their own their own binaries and trying to encourage users to use them as their browser not because they're different. She themselves on the web platforms perspective, but it's because they bring with them a bunch of agent tools that traditional browsers might not have thought about you know from from day one.

And so I think it's important for all browsers to understand like what it means to kind of be an agent platform and a harness and I imagine like a large suite of plug and play. Configurable tools and security policies and any kind of browser that manages an agent. I mean it's a doms point if you don't think about it then you can't secure it then you can't make a private like whether or not you want agents to be on the web or any individual browser wants agents to be on the web they are on the web. What security has been talking about is this fourth actor right like you have the platform the user and the site and then all of a sudden you have the fourth actor which is AI agents and they call that you know in security we call it a trust diamond. And that you can't think through those pathways and actually make them safe and secure unless you're actually paying attention like the you're not here me I can't see you. It's not necessarily the approach that's going to allow us to make things safe for users.

So we have an CP supported by open AI Chrome shop five or sell Cloudflare have you heard any peeps up or down from the two boogie monsters and Thropic and Safari or Apple Firefox and and Safari or or and Thropic and well and Thropic because this would have to be in Claude in order for it to be like everybody to use it right there's such a big player at least right now and then like Apple and and Mozilla as well I guess yes. Yeah, we have heard from all of them like we've been talking a lot with with Mozilla about the imperative API they're pretty interested in it and they've expressed some public support to the imperative side of things from web up to be which is awesome. We've been in discussion with yeah other other model providers like Anthropic and browsers like Safari I think probably we can comment on some of the stuff that's already public. I think there's a lot of enterprise interest from from both parties because this is where like a lot of knowledge work happens.

It's making sure that there's widespread adoption among developers that that can help enable WebOp in these in these products makes a lot of sense. And yeah, Sarah do you have anything to add to that? Yeah, so we are talking to all of the people mentioned they are investigating what it means for them and so as part of that investigation are doing due diligence on what they want to be supporting what they don't want to be supporting and formulating thoughts and so I would say we're still in the kind of like meeting with people and talking through things stage of things and not in the like here's what we can formally announce stage of things. But it is covered territory. That's cool. I had used a like a Web MCP B extension which basically turned my Web MCP websites into like a like a proper MCP server and that was really cool because then I could just I could take that MCP server and put it into anything that supported MCP and then I could I could just chat with it right.

That was cool because I I did I slapped it into Claude I typed into Claude but I could see it controlling my browser. So I thought like maybe that will be an experience at one point as well even if they don't end up supporting it. Yeah, I just talked to him last week. He's really like I don't know if you've talked to the creator of Web MCP B. I think he does a really good job of like thinking through what people might need because he used to be a consultant for all of these kind of like Salesforce and other companies. So he's kind of good at those like glue layers. We are incorporating a polyfill that you know he worked on previously and things like that. So he's a really good community member and I like the work that he's been doing. This is Alex Neha. Yeah, he's been great. He's been an awesome partner. He was one of the original folks that came up with kind of the original shape and idea of what Web MCP might look like and he's pretty active in the community group. It's been great working with him and building poly fills alongside of him and so on.

But you can definitely imagine that any model vendor that has a Chrome extension or really wants to be able to perform knowledge work tasks and user for users really wants something like this. It's amazing. I mean every time we bring it up to people and their eyes light up and they see all the benchmarks go green and they're kind of like oh my gosh yes this is saving time. Dollars latency everything and there's already been a bunch of public like benchmarks about this that actually integrate they kind of build their own harness sort of like MCP that sits in between Web MCP tools and the cloud code harness or the open AI extension and so on. So you can actually get a feel for like what it would what it would look like for these extensions and model providers that don't support Web MCP today. Like what it would look like for them to actually do so from a performance and a usability perspective and so far it's been it's been really good feedback. Yeah. Is there anything else that we haven't hit that you all want to make sure we cover. I think I think the one thing I'd love to to encourage is the community to keep tabs on the W3C spec.

You know every every day we get new issues filed bugs or proposals or additions and so on straight to the repository. It's a really good way of getting real world developer feedback and real world model vendor feedback like we have extensions. You know folks that they're build kind of like community. Dom actuating extensions. They love to chime in on the repository and help us understand what parts of the API makes sense for their harness and what doesn't. And so we would love people to just stay stay in touch and keep in the loop with the API and also provide their feedback to us because it's all super useful like we mentioned a few times. This stuff it's a really really early space and we're figuring it all out and we're trying to make sure what what ways we can impact the security and the usability of the API. And so I would encourage people to stay in the loop and start experimenting. There's been some really really great tools being built. There's ORA.AI which is a tool that has been built by one of the MCP apps MCP UI co-creators. And this kind of like helps you measure how agentic journeys are like actually happening on your site and how WebMCP tool calls are being used by real world agents.

It's a good like benchmarking kind of framework that we've been looking at and thinking through as a way to understand how useful tools like really are. And so the more feedback like that we get from the community tools being built to measure how good WebMCP tools and agentic experiences are. All of that really helps the industry create something measurable and effective from the use of not just WebMCP but anything agentic on their sites. So we would encourage developers to get as involved in that sort of space as possible because I think it's super useful feedback for the browser engineers and model vendors as well. That's great. Yeah, like seriously folks listening to this try build something slap it in your site try build like a little to do app or whatever and like give your feedback now because like one of my first pieces of feedback was like I want this to be headless as well. And like I don't I don't think that was me but like the now it is right and like provide your feedback now so that we can like nail this because even if you look at like the the journey of MCP.

It's had so many high highs and so many low lows people have been said it's over like six times since it's been released and it's just because we didn't know what it needed to look like. So like chime in and then that's super helpful. That was so funny I was at the MCP Dev Summit conference in New York City in March and you know the whole this was like when skills were happening and CLI pop it off and everything is MCP is dead and so they actually threw an after party on one of the nights at some like kind of like dark nice cocktail bar in Manhattan. It was called MCP is dead the funeral and they had like a quartet and they were like singing sad songs and it was like they went so far. And so I'm like oh my gosh but yeah it's entertaining to see all the community hype and unhipe about random things as they fluctuate. But yeah otherwise thanks for giving us the opportunity to chat about Web MCP and kind of talk through what we think the future of all of this stuff might look like. We're marching forward as fast as we can but also a lot of this stuff is pretty early and speculative and so we'd love to stay in touch and keep keep an idea on what developers are building with.

With these tools and how they work in public amazing so now it's a part of the show where we give you the opportunity to share something that you're really interested we call them sick picks these are things that are just in general that you're enjoying in life right now could be literally anything from a TV show a podcast or a what did we sick picked I picked some sanding paper so you can pick whatever you want. So Dominic sir do you have sick picks for us today. What grit? Yeah oh he's got he's got all the grits. Yeah a wet sander. Oh it's yes oh I got like nine different grits so it's for a sanding 3D prints I had a little 3D printed device that I sanded down and it's so smooth you can't even tell it's 3D printed it's beautiful. I thought yours would be more about like dancing and break dancing. Oh we have tons of these I've had sick picking for eight years now. Yeah many many dance competitions absolutely yes. I'm happy to go. Yeah go ahead. One thing I was I love really really good writing and so one thing I would love to

recommend to people is this book I've been reading called The Sense of Style by Stephen Pinker. It's a book about writing and it's a kind of a style guide to the classical sense of writing which is like its own some style on its own and it's really really really really well I like to write and I try and you know be as kind of a writer as I can I love reading really persuasive succinct concise impressive prose and I think this book is filled with that so I'd recommend it to people for sure. Well say I kind of feed this into my prompt into my into my clanger. I think mine is if you all like Minecraft or if your kids like Minecraft there's a game called vintage story it's a harder version of Minecraft and it's sandbox but it's very survival style so you're the first human you have to survive the winter it has some like love crafty and horror vibes and elements and you can mod it to your liking but it's really hard and really fun so if you're like you know into Minecraft and you want like to level up

and some difficulty me and my kids are just having so much fun playing that. Is it can you play multiplayer or you just taking turns? Oh okay yeah yeah it's actually really fun multiplayer because then everybody can like help with different tasks and things like that you do like a kind of collaborative effort. My kids are both into Minecraft right now so perfect perfect opportunity yes. Nice 24 bucks too I love I love that you don't have to pay monthly for this thing you just buy it once that's great I'm so sick of all the subscription things yeah me too I'm over it's too much. All right next thing we have is Shameless Plug this gives you a chance to plug as many and whatever things you would like you guys bring your plugs today. Yeah I would say for me mine's pretty boring just feel free to follow me on Twitter or X just at Dom Ferrelino should be pretty easy to find I tweet about random web platform API things that I'm working on and stuff in specs and web mcp and so I'd love to hear from people over there. I'm Sarah Edo on X and other platforms.

One Shameless Plug I do is that we're both speaking at agent con in San Jose that's run by the Agente AI Foundation of Linux Foundation in October so they gave us a discount so you can get 25% discount with community 25. Wow awesome thank you for sharing that we'll make sure that's all linked up to. Cool all right well thank you both for coming on appreciate this let us know down below in the comments what you think of web mcp and we'll catch you later. Peace thank you.

More episodes

More from Syntax - Tasty Web Development Treats

View all episodes →