Skip to content
TrackPodcasts
technologySep 9, 202639:45

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

The a16z Show

About this episode

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better?

As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks.

They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.


Resources:

Follow Rayan Krishnan on X: https://x.com/RayanKrishnan

Follow Ben Horowitz on X: https://x.com/bhorowitz

Follow Jennifer Li on X: https://x.com/JenniferHli
 

Stay Updated:

Find a16z on YouTube: YouTube

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Get every episode summarized

Each time The a16z Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

449 searchable segments. Every word is indexed and playable.

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

The a16z Show

0:00
39:45

Full transcript

The a16z ShowWho Grades the AI Models? | Ben Horowitz & Rayan Krishnan. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Every time a new trillion dollar industry emerges there's a need for this independent testing group when Meta release Lama 4 on our held out private benchmarks the model is actually underperforming but on all of the major public benchmarks it was showing incredible capabilities. What's the limit of what you can achieve and then within that how are you going about it? In an ideal role to take a frontier model and have it train the next version of itself but obviously that's very expensive and slow and so what we're doing is forming a set of proxies for every part of the process takes to build the next version of the model. If valuations as become more complex have a fewer sample size but a larger set of criteria expectations of them. Where do you see the gap that's happening today? The government kind of has an inclination of what it's afraid of the biohacking, cyberhacking but then there becomes the question of can the model do it and then can you get the model to do it? What do you think the landscape will look like? AI models keep getting better but the tests we use to measure them can become obsolete almost as quickly. In this

episode I'm joined by A16Z's Ben Horowitz and Jennifer Lee for a conversation with Vals founder and CEO Ryan Krishnath about the increasingly difficult problem of measuring AI. We get into why public benchmarks can give a distorted picture of model capabilities, why independent evaluation matters, and what it takes to test a new model in a few hours before it launches. But this is becoming about much more than model leaderboards. As companies spend more on AI they need to know which models and agents actually perform best for their own work and whether that intelligence is worth what they're paying for it. Ryan also explains why benchmarks need to evolve alongside the models. How Vals is measuring recursive self-improvement and why evals could eventually become a shared language for AI capabilities, risk and policy. So I'll start a question from when Vals got started in 2024 after a team discovered that all the public benchmarks are just not sufficient enough to

measure model progress and there needs to be a new methodology and approach coming to keep us on the frontier and help model labs continue to hill climb. Take us back to the inception of Vals and what do you see was missing in the market then? Yeah, yeah. I mean so I had a background doing research in particular building benchmarks and evaluations and so what was very clear to me was the very tight relationship between what it takes to build new systems for generation and actually new mechanisms for evaluation. In fact, in order to get one you often need to get better at the other and actually one of the biggest drivers for model capabilities having a new legible way to evaluate models. And so around early 2024 we're seeing is there actually many interesting models coming to market. They weren't all coming from open AI and then also in particular it was harder than ever to actually ascertain what was newly capable with the new models. And so in kind of a first principles way what we realize is that there would need to be some third-party company that solely existed to build really high-quality evaluations in benchmarks to be able to discern what was newly possible with these models. And so we released

our first benchmarks in 2024. Now the last couple of years that has kind of been realized by many different parts of the industry. I guess one of the obvious question is why do you think the labs can do this by themselves? Because they know the best software the models are hell climbing on and what it's missing capability-wise. Why can't they be the benchmarking stores? Yeah I mean internally they do build a lot of great benchmarks and that's what drives model progress. But I think there's an issue when we speak about model capabilities in a way that's all reported. And so one of the early indications of that you saw was when Meta released Lama 4 that was a bit of a disaster and interestingly what we saw is that on our held out private benchmarks the model is actually underperforming. But on all of the major public benchmarks with the questions and rubrics are actually open source. They were showing incredible capability. So there's a huge disconnect between what was self-reported based on these open benchmarks and then what we were actually finding with our higher quality higher signal benchmarks. But I think it speaks to a broader concept which the labs I think understand that they would like to see a rational buying market. They would like to see that when they invest billions of dollars to build a new model

they're actually substantive ways they can point to evidence and say we're advancing in these ways and it's not just entirely self-reported to justify that investment. And so you also see instances of Demis and others in industry calling for an ecosystem of third party valuators. And what's the historical analog that you have in mind here that there are radio agencies, audit firms? What's the right comparable? Yeah I think there's honestly less to learn across the board and every time a new trillion dollar industry emerges there's a need for this independent testing group and I think the fact that it's moved so quickly in AI's caused necessity for a lot of these parallels to be born out. We think about ourselves is trying to sit on both sides of the market. So their mechanisms by which labs need to prove that new models are very capable but they're also parallels where enterprises need to figure out what adoption strategies can amounts to the greatest ROI for them. Yeah. Take us through sort of the six hour pre-release window before the model drops. Obviously need to run tens of billions of tokens without delaying the launch. Which part is you doing an all-nighter versus it being automated? Take us through that. It's honestly been a journey and I think the real goal north story we think about is we never want to

be kind of a lagging indicator or a delay to a model release. And so that means we have to move really, really quickly and extract the most possible signal with the rate limits or capacity we have. As I early on with this look like was my co-founder links and I pull pulling an all-nighter to try and get as much done as possible and get results out the door. Now we built a team but we've also really invested heavily in infrastructure and so we're able to run evaluations in a massively distributed way running effectively the maximum possible rate limits with every model we get access to. And we also have this internal system called Steve. Steve the economic valve employee. And so that's been a mechanism by which we're able to actually take more of the human work over time and put it into Steve. How do you deal with the kind of issue that it's a little bit of an AI complete problem in that we still aren't really good at evaluating humans or we haven't agreed on it. There are things like IQ tests, there's EQ, there's the big five personality and so forth but there's not really an agreed upon framework for which we do it and people have issues with things

like the SAT and this and that and the third. And then of course models are really good at hacking the benchmark. So how do you think about that issue and what's the limit of what you can achieve and then within that how are you going about it. I mean I think the honest answer is that it's forcing a lot of the more fuzzy or distributed forms of eVALs to be made explicit. What is really the distinction between an associate and a partner at a law firm and there isn't a clear test or an eVAL for that in the human world. And so we have to first establish a lot of that in these different enterprise or real world workflows for us to be able to test models in the same way. And I think long term that will be actually the biggest bottleneck our ability to take companies and the eVALs and make them legible because that's how we'll figure out what signal we hell climb on and where we actually adopt. Very interesting. Actually maybe one question for you Ben on just like how the industry has formed before intelligence came through like we're now measuring something that's very fluid versus before like when we're talking about enterprise

software there's gardener rating on like 70 different metrics like you can sort of stack rank on the cloud run to say this company has these features covered these features not but now it's like very jack frontier that's very hard to measure in the industries. What do you see one is the analogy to the past that lessons we can borrow and what do you see that's really going to be the challenge and missing pieces going forward. It's a little bit reminiscent of the MPAA right where it's like what's art what's porn where's the line what is it are when is it X and by the way the definition of that has changed over time I think things that used to be X are now are and so forth and what's PG 13 all that kind of thing and there is no in the famous line as well I know when I see it and I think that this fund I just think it's going to be necessarily fuzzy but they're will develop norms over time and you know like if enough kind of people who run companies you run

finance or run whatever it is kind of agree yet that's a norm that I think it is as opposed to kind of what we have a lot now and the open benchmarks whereas if you can solve this specific problem then you're at this level and so forth I think that's one attackable and then it's too narrow I think when you also look to some historic analogies there's a lot of lessons that you can take from them as well on what's gone wrong what we need to avoid I think for instance fed balls one very early decision we made was decision to to never sell training data to labs it's often a place that were pushed when we start working with a new lab to actually source himself for them a bunch of training data yeah that's a lucrative business yeah and actually a lot of that industry has now built these gimmicks style benchmarks as a mechanism to sell their data and so that's become kind of their go to market as well but I think if you look at auditing as an industry you end up with issues like Enron where if you have the same group who's responsible for doing the audit as well as also consulting and supporting the company you have a mix incentive structure and

and then it just becomes paid to pass the audit or in this case paid to win the benchmark and that's really not what the market benefits from and what we're trying to do and right today you have already a pretty extensive catalog of different type of benchmarks some of them are more focused on specific industries some of them are more like consumer mental health related maybe first just talk through what are the benchmarks that are most popular and most like rat upon and we'll love to dive into one of them as well yeah well we've done a lot of work in kind of the economically interesting applications of models our finance agent benchmark is used by a bunch of the big financial institutions to get a sense of how models are improving we also have a lot of good work encoding so our vibe could bench measures how well models can take a natural language prompt and build a full stack web application and so that's been a keen way to track model improvements over the last nine months yeah we're also doing a lot more experimental work so one benchmark we release recently I'm very excited by is our recursive self improvement index it's a topic which a lot of big labs have been talking about and and and starting to report on in their model cards but there isn't a shared language to talk about the RSI potential of models and so we created this

as an apples apples way to actually benchmark across the models yeah I thought that one is a very cool benchmark is sort of all of a rage in the research community of how do you measure like progress it can make through having more frontier models that you can just compound on the capabilities how do you actually go about like building this RSI benchmark yeah I think you know in an ideal world what you want to do is actually take a frontier model and and have it train the next version of itself and see where the delta comes from but obviously that's very expensive and slow and so what we're doing is forming a set of proxies for every part of the process it takes to build the next version of the model so there's some work around pre-training post-training harness level engineering and then seeing in which mechanisms and behaviors the models are able to to do very good research work and build something new and where they're struggling very cool there's also cases where you have deprecated indexes and benchmarks it's funny that I always watch this benchmark industry if people like just like the early to future model days like a cherry

pick whichever it shows up the best and the most perfect like benchmarks as well like you pick something that's you know very popular but maybe already saturated and you rank very well on that or score very high but you took a very different approach in like if these benchmarks are saturated you'll deprecate it like maybe talk us through the thinking I think it's that too in necessity and this is kind of the infinite game or in I mean you're wearing our shirt and so we have this unofficial motto always a higher peak yeah so insofar is foundation model labs are hill climbing they're searching for the next peaks to summit it is our job to perpetually construct these next mountains for them to summit and I think that's also how the economy has naturally functioned over time as you know agriculture becomes less important for our labor market there are new forms of labor that's required out of our out of our population and so in the same way we should expect our benchmarks to keep up with the new frontier for what we want models to do there's another component of retiring benchmarks which I think is is underappreciated

which is that benchmarks should also be reflective of the current state of the world it's in the same way if you're a lawyer you have to you know retake the bar exam and get certified or if you're an architect you have to get your certification and and or you know a doctor we should also expect models to be tested on the current state of the world and what we know in medicine or what we what we have is our set of laws so in the instance of case law updating to something like legal research benchmark that's a desire to create a benchmark more reflective of the current state of the world and also push the models in place we want to see them go and how has I guess one like it used to be like we're doing these like multi answer or multi step questions to like just evaluate prompt answers now there's like a lot more agentech work that's happening whether it's on finance or legal or coding especially like there's a lot of you know async background agents that can just complete tasks how does that change sort of how you build infrastructure how think of evaluating the capabilities of not just the

models and agents themselves and there's also like a lot more dimensions that people care about is not just like capability is cost is latency it's like you know whether this model is flexible enough to to to address like broader domains and tasks and so on so how do you think about the additional parameters to to what you evaluate oh yeah this I mean there's a lot that goes into that I think on the infrastructure level you have a whole new set of problems I mean for instance now we're testing models and they're willing to run over hours days sometimes weeks and so the infrastructure needs to be very stable to support evaluation over time and if there is a failed request which people to reach high from that one and not redo the whole trajectory so there's some simple simple mechanisms in the infrastructure we have to think about but I think in general what we've seen is evaluations as to become more complex have a fewer sample size but a larger set of criteria expectations of them and so what I mean by that is a benchmark is largely some kind of input space of things you're trying to query a model to do and a set of requirements or rubrics that you see an expectations of the output and so early on you have things like

image net which have millions of images you're trying to see a basic categorization for so it's a once one mapping between an image input and a text label output now what we have is for fewer set of tasks you know generate me 50 full stack web applications but a much larger complex mechanism for evaluating the output produced and I think that trend is going to continue as we see more complex workflows evaluated with models and you do you think that it will become kind of and in a real time kind of mechanism like so for something like open router which strike just but with open router looked to files and say okay where should this next request go or is this going to be kind of strictly for like picking a model and the enterprise for a task yeah I mean I think you know open route is a bit of a misnomer in that most of their usage comes from being a model gateway and so it's actually up to their their users decide which models they want to use when

and that's because really the hardest part of routing is building the e-vals and trying to determine in what places a set of intelligences should be used for a particular application and so so you know for it's a pretty enterprise and building e-vals has actually supported a lot of them and also adopting routers and see more about why it's not only important to labs but also existential for for enterprise and maybe just say more about how you guys work with enterprise yeah of course yeah I mean I think that the lab side of this is very clear like you know if you're raising lots of money investing heavily in building models it's it's essential for you to show why your model is getting better and then why this customer should should pay premium for them but what I think is still under appreciate is on the enterprise side this is turning to be existential as well you know I have a I was small anecdote really to this actually you know it was meeting with a company and the fortune 10 and they the way that they've adopted cloud code has been with roughly a hundred dollar a day budget for their engineers and so what I was hearing is that this is actually fundamentally changed how work gets done in this company in that there's a

a rate limit which resets at 4 p.m. and so it's the most productive hours of work are actually now 4 to 6 p.m. when the rate limits reset but then there's this dead period in the afternoon when people go on walks or you know get a coffee because they just don't have the rate limits and so I think what's really you know illustrative there is that you see that there is a mis-vailing of intelligence happening at every layer of the stack and so by that I mean you have engineers who have a hundred dollars worth of usage limits and they don't really know how to apportion that to the greatest productivity for them you also have this fortune 10 company which is kind of arbitrarily said they're going to allow a hundred dollars per employee they've actually recently increased it to three hundred dollars per employee so almost an employee's worth of salary and tokens for them to use and this is actually pretty arbitrary because it's hard to quantify what the right usage limits should be but then also in for apik is running on pretty narrow margins to support this and they have massive cost to serve these models and so I think we're in this world where it is still very unclear what ROI looks like and how to value this intelligence that's being used and so as we

talk about the existential concern for enterprises I think it is this kind of direction we're shifting in where tokens band may start to eclipse salary spend and and so if if this is such a meaningful line item in in your costs you actually have to justify the ROI much more keenly than you you've seen over the last six months and over time as we were talking about I think a firm really is just its e-vals and so the ability for a company to make its e-vals legible in order to solve this ROI calculus is going to be the reason why that that company wins out over the competitors in the long term and and maybe just double click on that like similar question to to why the labs come to it themselves or requires a third party agency to to rate it I think it's a lot more understandable that you know you need that that neutrality across the industry but for enterprise they will argue that they know the task the best for their customers like what what what how does like vals come into prolett value and maybe you can talk through the sort of vals myth new product

launch as well I would recommend a lot of companies to develop in-house expertise but I think that should not be the only solution you know there's this explosion of intelligence happening there are somehow still more foundation model labs getting getting constructed and and each labs also releasing more models than ever with many more hyper parameter options and they exist within a complex set of harnesses and agents so the option and we're even talk about specific intelligence this new paradigm that's emerging so there's a growing set of intelligence options and I think what we're finding is that we're still finding new places we want to use AI models and so the use cases are growing in complexity as well and so I think if you're if you're a company you you have a compounding set of complexity in the set of options it's very very hard to develop the internal capability to do the evaluation and so it's a time to try and remedy that we've started to release some products more openly for enterprises to use the first of which is called valsmith and so valsmith is focused on co-gen the area we're seeing to be the highest I spend it in enterprise AI and allows any company to

take their GitHub code base and build their internal coding benchmark from it to get a sense of what coding agents are going to be the most performant but also what's going to be prato optimal or the the highest are a life for them to use and actually we we use valsmith a lot of vals and we're seeing that a lot of the best enterprises and so sophisticated ones are doing that too I would expect that to be the direction the market moves as it rationalizes and what are some of the examples when you let's say benchmark on a private ripple that it just shows very different performance cost behavior compared to let's say like using like frontier model using a public ripple benchmark I think today it's still very unclear whether the best open AI model or the best and theropic model is actually going to be best for your repository and so we've seen a lot of non intuitive examples where you actually had to run the e-vail to figure out what's going to be the frontier performance for that repository I think you also now see a very complex middle set of options in that there's now opus

and sonic models for my and theropic but also Luna and Terra and Luna's very cost competitive new spark is also very cheap and and 1.2 is very capable there's also a growing ecosystem of open source models which i companies can choose to self-host so I think in this messy middle it's actually very non intuitive what's the right fit we're actually seeing in a lot of cases sonic is more expensive than opus because it is so token hungry and so I think if you were to operate based on you know you sonic where you feel like it's applicable you may actually end up spending more than you need to and same more about how this evaluation framework will apply to knowledge work and other domains or whatever some examples I think coding is a sign for what's to come in every domain and a lot of the primitives established there are carrying over to other places and if you if you have a very good coding agent chances are you have a model that can also make PowerPoint slides or DCFs in Excel and with with the high degree of capability as well I think what we need to leverage in a lot of these industries though is the existing repository

work that has been done as a mechanism to build evaluations and so just as Ben was talking about we you know haven't really solved the question of what is human intelligence but I think in a lot of industries we have a sitting repository of data around what work has looked like and it'll be the task of us and others to try and codify that into evaluations that can that can stay dynamic and actually evaluate models for human workers being done maybe just take along the earlier question how are you guys using Valsmith internally to evaluate what's the best coding model for Vals? Yeah I mean to be honest this was actually born out of a problem that we saw as well so I wanted to do a token maxing experiment and it was able to get unlimited access for our team for a month for some of the coding tools and so in in retrospect looking back we had some we had a lot of engineer spending between one to two billion tokens a day I think Pete Day was one engineer spending six billion yeah it's also crazy because how much does that equal to two dollars so okay and then I went back and did some math and and it looked like in that month we spent roughly

one point five million dollars worth of tokens this is free by the way I know what's I want to but it was actually 10x more we were spending in tokens than employee salary for that month so it's not even like oh this is this is this 50 50 it's 10x and it was interesting to debrief and see the places where people were using agents and it's kind of you know insecure to use models all the time everywhere and so what we were faced with is okay we cannot continue with this mode of operation for the next month how do we actually intelligently figure out what are the right tools we should use and for what teams and what projects and so we ran this experiment of looking at the work that it was done we looked through a lot of the traces we looked through our GitHub repo and built out the Valsmith tool and we found some pretty surprising insights like like for instance the cognition Devon tool is actually very token efficient and so that's a place we've chosen to adopt more and and I think there's a lot of places when you can get better pricing models out of subscriptions as opposed to token-based pricing and so it's actually

informed our strategy for how we can actually effectively token max without spending one and half million dollars per month. Very cool so is the current operating mode that you were using like one I guess more token efficient harness plus model and then like on top of that people have some more flexibility to use token-based for some higher or more challenging tasks. Yeah so we have access to all the tools we give everyone access to everything but we auto issue recommendations for any GitHub issue or ticket for where to begin their session and that should titrate the actual usage depending on the intelligence required for that task. Very cool. I want to segue to the policy side for a second because we talked about how quickly benchmarks become obsolete in policy it's even worse in that laws move much slower relative ticket abilities. Ben and Mark spend a much time in DC and the policymakers to try to close that gap so given that who should define

the standards here is it labs is it independent value like like you guys is it customers is it is it government how should this work from policy perspective. Yeah I think the short answer is that everyone should be involved to some extent. I think there's a benefit from very perspectives. I think the main issue though is that policy conversations as they've happened over the last couple of years have been very abstract and there's been no material grounding to figure out what policy should cover and so even when you have proposals from labs to have a third party testing company or ecosystem it isn't actually made explicit what the behavior and maximum by which their work is and so I view our role especially early on is to just be in evidence gathering mode where we're able to pull a lot of information and empirical data about what models are capable of and where the risks are and that can go on to inform you know a more sophisticated conversation about policy. How do you think about like who does what because the government kind of actually did the first

evals and is continuing to do evals in terms of okay what's at the frontier and needs to be regulated right so they started with you know some crazy idea with 10 to 26 flops or some such thing and so when you think about it like what should the government be doing you know to put it in this 30 day weight period or 60 day weight period or whatever it is and then what should happen in that weight period and how does that intersect with what you're doing and what's the right way to determine whether a model is on the frontier or not and needs to be put in some special box for a while to make sure it doesn't break into everything like how do you think about like how that relationship works. I think there's effectively two counter-reling forces that that have to be considered the first is desire to move very quickly and ensure that the the the government process isn't slowing down the rate of technological innovation and I think the

other part is to make to make sure the technology as it's developed is in the best interest of Americans and people more broadly and so I think these are very tough to reconcile and often you know picking one means it's at the expense of the other and so what I'd hope to see is that by doing this evidence gathering process we can help policymakers in form what they believe technology should look like in order to be aligned to American interest and it can be the job of third-party evaluators to develop the technology to actually test and enforce that because I think that will create a maxim by which you can see advancement in methodologies for evaluation and testing in a way that actually keeps up with the frontier it doesn't lie behind or slow down the pace of development. Do you think in terms of that already and developing your e-vails like do you think well can we test to see how easy it is for this to get this model to start reward hacking or that kind of thing you know in doing illegal stuff or is that kind of not in the scope yet or how

do you think about that? Yeah I mean we think about this broadly under the category of alignment. I think there's places where you see that we're now we're models that are being tested for one cybersecurity risk or actually reward hacking and figuring out other ways to get around it but what we're trying to evaluate is our models aligned with user intent and so in those places we're actually finding evidence that models are exhibiting behaviors that are not. And maybe that's a question for you Ben as well I guess how do you think about the right division of labor here like what should you know garment agencies control and do themselves and where they should like you know partner trust you know private companies to take care of and where do you see the gap that's happening today? Yeah so I think the government agencies do get a lot of warnings from by the way the big labs oh this thing is going to biohack this is going to be a cybersecurity risk and so forth and so I think what the government needs to do is go

okay if the model you know if the model is capable of it and then can somebody kind of basically prod the market to actually do the illegal behavior and you know kind of specifying exactly what are those things that they don't want in the market and then having a third of a card to kind of evaluate that so it's a kind of does the government you know the government kind of has an inclination of what it's afraid of be it biohacking or cyber hacking or so forth but then there becomes the question of okay can a model do it and then can you get the model to do it and and then somebody's got to actually evaluate those two capabilities and I think the government is particularly ill suited to do the latter decently over time it's just not a good government function

but they're very good at setting the rules because they can enforce the rules so I think that that's kind of the combination you want that the government sets in and forces the rules and that a very competent kind of private company then tells them if the rule is broken you know and kind of it's it's been interesting to see like the large labs start to go well the model's got the capability and you can get it to do the bad thing so we're not going to let anybody have it we'll just use it and make sure that our people don't get it to do the bad thing and that's and even that doesn't always work so you know we're an interesting time so we're setting and to me there's the gap of like what's the narrative and what actually happens in real world because you know every setup in again like enterprise setup is very different like or just like people however they use the models are very different the narratives that connects to the actual examples are very rare which is why we still talk about opening i and hugging face hack

today we still talk about you know what happened with fabo and and it was for like two months but a lot of times like you know that's not really how the model is being deployed the the environment they're running on is very bespoke so how to like you know really bridging those thoughts and again set up the the right environment and also like rule basis for adopting these models I think also just requires you know someone taking the capability and taking like what's the the the guard rails and and put it down to to the ground so that people can like you know have the confidence using using the models to then to say more about how exactly the policymakers should work with the evaluator what information do they need would it would actually the relationship work so that it's most effective yeah I think I think in the first order there should be an a mechanism by which insights and data can be passed directly to relevant people in government so now we're regularly doing briefings for executive and legislative branches on what we're finding capabilities

in risk of models and so I think first order that helps people there get up to speed on what's going on and and also track through what will be problems in the future I think you know things are moving very very quickly it's hard to predict where things are going but at least when you have data you can start to extrapolate a trend and then I think from there it's it's up to the people in legislative branch to decide where they they want to see policy and so it's not really our place to give recommendations like that but if they if they see that there is significant risk and say mental health for for people under the age of 18 or biosecurity risk in the models that necessitates having a standardized way to curtail model release then it's up to them to inform policy and I think then there's other places where the executive branch in the places like department of commerce or SEC is responsible for for making sure that private companies are able to adopt and use the models in a way that's going to be productive for the whole system. I love to propon another angle just around geopolitical I often see eVAL being like a representation of sort of the value of

of the model the model developer like you kind of develop this rubric of what's embedded in in the model and of course different different countries and labs in those countries care about different things like I mean I'm born and raised in China I use a lot of like Chinese open source models too like you still cannot let them you know just go freely talk about CPC and obviously there because you know what happens in China so how do you think about how eVALs I guess and benchmarks play a role in like standardizing or like being treated by different different model labs from from different places. Yeah I mean to be honest for my very idealistic perspective I'm surprised to see so much investment in sovereign AI you know if I was taking a gods eye view it would be extremely inefficient to build all of these data centers and replicate this data engineering process and train these very large models when in fact you could

probably consolidate a lot of these efforts but it seems like that's not the world we're in or the one we're headed towards and there's actually increased efforts to build AI in a sovereign way and so I think that takes having a shared language to communicate about what the framework for evaluations are and and where we're going to collectively align around the risks. You know I think there's actually a lot to learn from from nuclear here as well I think Reagan had this line trust but verify and so I think we're starting to see signs of trust in that you know it's easy to ping and Trump are going to be meeting next month but there is no clear way to actually do the verification part of this having the shared language of eVALs will allow us to say things like you know you you have the the right number of nuclear warheads and in that example there were also flyovers so so Mexican by which country could audit another country's nuclear stockpile by having flyovers and so I think similarly if there's concern about the societal or even existential risk of AI it will necessitate us constructing this shared language of evaluations to do the verification process. How do you think about harmonizing a policy like that so that

you know it's hard enough to do it in America and then how would you think about kind of take me at global you know because now you're dealing you're not dealing with enterprise enterprise customers you're dealing with governments and those governments are competitive with each other and how would you think about that working. I would be naive to say I have the perfect solution to this problem today and so I think there are baby steps in which we can start for instance there seems to be a lot of talk about cyber security risk I think the concern around biosecurity will become even more important over time and so there are clear places where there'll be mutual interest in aligning around ways to prevent conflict around cyber or bio. In my opinion I think long term what's actually going to be the most interesting is the recursive self-improvement possibility and that's a place where you could see one country one or one company you kind of run away with it and produce models that we we don't know much about

operating in ways that are unknown to us and so I think having a way to in a joint way describe this being the level of pace we're comfortable with or this being exceeding the pace of development as a release to RSI it's going to be super important and that's where I think you see a lot of the researchers at close source labs calling for joint conversations between governments today. What do you think landscape will look like going from from here now that we have lots of different capabilities and capable models as well as like you know countries that care about different developing I mean everyone cares about RSI for sure but like on the biocide or like the cyber side people care about you know slightly different different things whether it's more offensive defensive and so on like what do you think the landscape will look like and how do you think about developing new benchmarks to to keep up with with that. Yeah I mean we're we're hyper focused on on building benchmarks that capture the frontier and so

into far as we see new places for capabilities or risks at the frontier we want to make that it actually well-documented evaluation on vast.ai and I think it takes have increasing coverage over time you know for instance I think in cyber security a lot of our historical work has been done around code vulnerabilities or memory leaks that may exist in code but actually a lot of the the biggest concern or risk is in the infrastructure level and so these are not things that are expressed in code but take simulating larger environments of enterprise cloud infrastructure or even grid infrastructure for us to be able to say this is what the offense or defense of capability of models is and and so making sure evaluations are reflective of those new cases places is really important for what we do at VALS and we believe that the most valuable form of this business will be one that's incentive aligned around doing really high quality evaluation not supporting the intelligence development process or the the process by which the models can actually improve on that set over time. Awesome thanks for

coming to the podcast it's been great episode thanks much for having me. Thanks so much Ryan thanks Ben. That was fun. Thank you. Thanks for listening to this episode of the A60Z podcast. If you like this episode be sure to like, comment, subscribe leave us a rating or review and share it with your friends and family. For more episodes go to YouTube, Apple Podcasts and Spotify, follow us on x a A16z and subscribe to our substack at a16z.substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder the content here is for informational purposes only. Should not be taken as legal business, tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16z fund. Please note that A16z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details including a link to our investments please see a16z.com forward slash disclosures.

More episodes

More from The a16z Show

View all episodes →