基本信息

来源摘要/节选

AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:

From being Anthropic’s first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won’t be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?

We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.

We discuss:

Why risk, liability, and trust may become the binding constraint on AI adoption

Rune’s path from reading the Scaling Laws paper to joining Anthropic in its earliest days

What Anthropic understood about scaling, compute, and the future years before it became obvious

Why Waymo illustrates the gap between AI capability and real-world deployment

AIUC’s $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies

AIUC-1: a standard for AI agent security, safety, and reliability

How agents are tested for jailbreaks, hallucinations, and data leakage

Why most AI companies optimize the happy path without seriously stress-testing adversarial cases

Why AI standards may need to update every quarter instead of every decade

The emerging trust gap between frontier AI labs and governments

Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface

Why standards and insurance may need to evolve together

How Lloyd’s of London can insure AI systems and bring trust to enterprise deployment

What happens if a $20 Cursor subscription contributes to a $200M plane crash

The Air Canada chatbot case and how AI failures are beginning to clarify legal liability

Why copyright may be one of the hardest AI risks to insure

Evals, mechanistic interpretability, monitoring, and models becoming aware they’re being tested

The impossible CISO mandate: adopt AI fast, but don’t let anything go wrong

Why robotics will make AI liability dramatically more consequential

Whether AI engineers should have Level 1, 2, and 3 certifications

AIUC’s roadmap across agents, frontier models, robotics, and universal red teaming

Why AGI could become a question of national sovereignty

Why the labs can never fully serve as their own watchdogs

The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?

Rune Kvist

LinkedIn: https://www.linkedin.com/in/runekvist/

X: https://x.com/RuneKvist

AIUC

https://aiuc.com

Timestamps

00:00:00 AIUC’s $40M Round and the Risk Bottleneck for AI

00:01:07 From Scaling Laws to Early Anthropic

00:07:58 Why Trust, Not Capability, Could Limit AI Adoption

00:12:19 Founding AIUC and Building AIUC-1

00:18:52 How AI Agents Are Audited and Stress-Tested

00:25:26 Frontier Models, Government, and the AI Trust Gap

00:33:32 Cyber, Child Safety, and AI-Enabled Biological Risk

00:38:14 Why Standards and Insurance Belong Together

00:41:45 What Does an AI Insurance Policy Actually Cover?

00:50:44 The $20 Cursor Subscription and the $200M Plane Crash

00:53:53 AI Liability, Monitoring, and Earning Enterprise Trust

00:56:21 From AI Agents to Models to Robotics

00:58:29 Copyright, Adverse Selection, and AI Insurance

01:03:28 Evals, Mechanistic Interpretability, and Eval Awareness

01:08:36 The Impossible Enterprise AI Mandate

01:11:52 Prediction Markets vs. AI Audits

01:14:43 Should AI Engineers Be Certified?

01:19:10 AIUC’s Roadmap, AGI, and Who Watches the Watchdogs?

Transcript

Introduction: AIUC, the $40M Series A, and Risk as the Adoption Bottleneck

Swyx [00:00:00]: Okay, we’re in the studio with Rune from AIUC, the Artificial Intelligence Underwriting Company, with our trusty co-host, Vibhu. Welcome.

Rune Kvist [00:00:10]: Thank you. Thanks for having me. Thank you.

Swyx [00:00:11]: What are you announcing today?

Rune Kvist [00:00:12]: We have raised $40 million, led by Ribbit Capital and First Harmonic.

Swyx [00:00:17]: You first came to my attention when Nat and Daniel invested in you guys. Is the story, like, pretty much the same? Like, what are you today versus what you thought you were back then?

Rune Kvist [00:00:26]: When we raised our seed round, we had a hypothesis that at some point risk was going to hold down adoption. At that point in time, that felt kind of hypothetical, and I think that is now over. Clearly, the moment is now with Mythos and Fable. It’s pretty obvious that literally the binding constraint on adoption is risk. And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously it was speculation, now it feels like fact.

Swyx [00:00:54]: And let’s get a list of the customers that you’re highlighting as part of your Series A.

Rune Kvist [00:00:58]: Totally. Yeah. So we are now working with folks like Cursor, Harvey, Lovable, ElevenLabs.

Swyx [00:01:05]: Yeah. Amazing. Congrats.

Rune Kvist [00:01:06]: Thank you.

Swyx [00:01:07]: So you were famously one of the first hires involved in GTM and product. I’m just kind of curious: what was your path into AI? Just recap.

Rune’s Path Into AI: Scaling Laws, Capital, and Anthropic

Rune Kvist [00:01:18]: Yeah.

Rune Kvist [00:01:19]: Late 2021, I sold a company, my first company, an edtech company. I had a bit of time to think about what was next. I came across the Scaling Laws paper, and that just struck me like lightning. I was just like, “This is a big idea.” In short, the Scaling Laws paper just says the bigger the model, the smarter the model.

Swyx [00:01:38]: So this is the Kaplan one, not the Chinchilla one?

Rune Kvist [00:01:40]: Exactly, the Kaplan one.

Swyx [00:01:42]: Yeah.

Rune Kvist [00:01:42]: And the important thing that clicked for me there was, oh, now capital will understand this. If you put in more money, you get more money out, and so that will kick off a hype cycle. And so you get a sense of predictable returns, which is, in fact, what’s played out. And so I just packed my bags. I’d never been to San Francisco. I’d never been there. I just packed my bags, flew out here to find the people who had written it. And at the time, they had just started a small lab called Anthropic. There were around 40 people at the time or so. Drank a bunch of coffee until I eventually got introduced to Dario. And at the time, they were wrestling with some of these questions of, like, should we deploy our models? Should we make revenue? How should we engage with the rest of the world? They’d just broken off from OpenAI, and it’s been publicly reported that they were kind of concerned with how they were dealing with deployment. So they were wrestling with some of those questions. At this point, this is early fog of war, like early 2022. The hottest product at the time was, like, Jasper. Like, there’s nothing out there. So where value was going to accrue, and what the different parts of the stack were going to be, were all open questions.

Swyx [00:02:48]: I want to highlight to people, you ask these questions because you have a PPE background.

Rune Kvist [00:02:52]: Yes.

Swyx [00:02:52]: I actually was in Singapore in one of the sort of feeder programs for prepping people for PPE. So I had a tutor. We learned, you know, philosophy and politics and economics. But, like, I think your kind of background matters. Machine learning people who read the neural, Scaling Laws paper would not necessarily draw the same conclusions that you did. Whereas any capitalist would read that and go, “Holy shit.”

Rune Kvist [00:03:19]: Correct.

Swyx [00:03:20]: Right?

Rune Kvist [00:03:21]: Yes.

Swyx [00:03:21]: Who tipped you onto that paper? Because it’s not a paper that you normally read, right, like, in your circles?

Rune Kvist [00:03:26]: Yeah. I think I’d actually, ever since AlphaGo, had some appreciation that AI was a big deal.

Swyx [00:03:36]: Yeah.

Rune Kvist [00:03:36]: But it kind of felt like it raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go. But it was obvious enough that it was like, this is going to be a big thing if we find the kind of right mechanism to kind of get the techno-capital machine to work on this. But it was just not clear. And so I think there was some way in which, like, that became obvious, and also it wasn’t as obvious at the time than it is now, right? Like, it was just like, wow, this is so interesting. But it still felt, coming from kind of a philosophy and economics background, it felt like if this turns out to be true, you’re going to be wrestling with all of the big questions in society. Everything you’ve learned about politics gets thrown out of the window. Everything you’ve learned about economics at least gets challenged. And so what felt interesting was to be at that frontier that has ramifications across everything. So that’s why I sought it out.

Swyx [00:04:32]: I mean, clearly really good insight. For people who don’t know, the PPE program is, like, where prime ministers are born. So then you end up meeting Dario.

Rune Kvist [00:04:41]: Yep. First Dario, yeah.

Swyx [00:04:43]: Yeah. Well, I mean, like, so did you get extra insights from talking with them that you didn’t get from your original hypothesis?

Anthropic’s Early Conviction and the Scaling Laws Crystal Ball

Rune Kvist [00:04:50]: If you read the Scaling Laws paper, you get this, like, very vague sketch of like, wow, this seems kind of important. There are some lines on a chart. This seems kind of important. And what I think the team at Anthropic had thought more about than anyone was like, what are the implications of this if you really play this out? And back then they had, kind of vision documents for what the world would look like in 2026, and they were kind of in vivid detail playing out how much compute is going to be needed, what the CapEx was going to look like, what some of the societal concerns were going to be, but also what is the amount of economic value coming out here? And so it kind of felt like they held a crystal ball that in hindsight turned out to just be dramatically correct. And they weren’t holding it like they were obviously correct. They were just like, “Take this hypothesis really seriously.”

Swyx [00:05:38]: Think it through, yeah.

Rune Kvist [00:05:38]: And think it through in the same way as the kind of situational awareness that is

Swyx [00:05:43]: Across the street.

Rune Kvist [00:05:44]: Across the street.

Swyx [00:05:44]: Your office, yeah. Oh my God, we’re all living across the street in the same one square mile.

Rune Kvist [00:05:50]: Correct. And that’s now a couple of years old, but also people keep referencing it these particular weeks with Fable and Mythos, and it’s like, wow, if you take this one idea seriously- For the Scaling Laws, a lot of things fall into place.

Vibhu [00:06:03]: And keep in mind, at this point, this is the same team that did GPT-1, GPT-2, and GPT-3.

Rune Kvist [00:06:08]: Correct.

Vibhu [00:06:08]: Which is also, like, it’s not just some experimentation. Like, this is a real model that we just scaled up.

Rune Kvist [00:06:14]: And they had deep conviction in this idea: if you take a big blob of compute and data, it just wants to learn, and out of that will come smarter and smarter models. And all the particulars were not clear.

Vibhu [00:06:26]: Yeah.

Rune Kvist [00:06:27]: And all the implications were not clear. But their deep conviction in this, like, core thesis, and that was kind of dizzying. It was both phenomenally interesting and exciting, and also very quickly you get to, like, the world we know today will no longer be if this hypothesis holds. So it also just felt, like, important in some kind of grand sense.

Vibhu [00:06:48]: What kind of shaped you there? So that was early 2022. Not only had GPT-1, GPT-2, and GPT-3 come out, but, you know, the amazing founders of Anthropic that have never split up, the only ones, they actually had the conviction to leave OpenAI, start their lab. You said there were about 40 people there. What was the time like there?

Inside Early Anthropic: Mission, Deployment, and Risk

Rune Kvist [00:07:06]: It was kind of remarkably like what it looks like on the outside today. Extremely cohesive, extremely mission-oriented, and living in this tension between their two ideas, which is AI could both go really well and really bad, and we want to be part of building it. That creates astounding amounts of tension. And they were wrestling with this incentive challenge where they know they’re in a race that they’re in where you might get forced to cut corners, but it also felt very important to them to be at the forefront of technology. And all of those ideas were just present at that time. It kind of feels like that line has been just very clear, and I think kind of love them or hate them, they have really stuck to their guns. There’s a core set of beliefs that they hold more deeply than most companies hold any beliefs.

Vibhu [00:07:58]: Yeah. Fast-forward to today.

Rune Kvist [00:08:00]: Yeah.

Vibhu [00:08:00]: What does that lead us to AI underwriting company? What are you up to? What motivated you to start this?

From Waymo to AIUC: Confidence Infrastructure for AI

Rune Kvist [00:08:05]: Yeah. AIUC builds confidence infrastructure for frontier AI through standards and insurance. The link from Anthropic to building confidence infrastructure, looking out the windows at Anthropic offices and seeing Waymos driving by. Already back then, early 2022, Waymos were in some ways like AGI for cars. Like, they were superhuman drivers, but you couldn’t take one to the airport. And now, four and a bit years later, you still can’t take your Waymo to the airport, despite now everyone having kind of looked at the evidence and being like, “They’re better drivers than humans.” So in that particular instance, what’s clear is that the binding constraint on AI being useful is not capability, but is that liability or risk or trust. That problem is, general. The reason why right now

Rune Kvist [00:08:52]: Fable is not open for access is not because it’s not a good model, it’s because it’s a very good model. It’s just hard to make promises about what it will or will not do. And this problem gets worse as AI gets better. Basically, more intelligent AI can be more autonomous. That’s more valuable, but also the risk surface grows. And so - what Waymo illustrates is that unless you build the confidence infrastructure to make promises about AI, or at least bring light to the risks, you grind adoption to a halt. Governments, banks, hospitals, militaries need to have some sense of what AI will and will not do to be able to operate for them to incorporate it. And that’s the problem that we’re trying to solve. Now, why standards and insurance? If you trace this problem back through history, every technology wave has had some version of this problem. So if you go back to, like, year 1900, electricity comes

Vibhu [00:09:47]: Ben Franklin.

Rune Kvist [00:09:48]: Cars burn down, sorry, houses burn down, lots of people die. 1930s, cars are a big deal, kill lots of people. 1950s, private nuclear energy is a big deal, poses big risks. In each of those instances, the market runs ahead of regulation to create confidence infrastructure because that’s required to make go/go decisions. That is required for adoption, and the market fundamentally wants adoption. And in all of those instances, common blueprint emerges between standards and insurance. The reason these two components is standards kind of provide the rules of the road, and they also specify, like, what are the tests that need to be run so we can get a sense of how high the risk is. So take in the case of cars, that’s like a car crash. Great, everyone, they inform your insurance pricing today, they inform your purchasing decisions, et cetera. That’s basically the risk framework. The insurers are important because they pick up the bill. So they are the private institution that is most on the side of. That is best incentivized to quantify the risks truthfully and then figure out all the ways to reduce the risk ‘cause that increases their profit. So they’re basically, they help shape the incentives. And these two work really well in unison. Now, how does that show up as a company? Well, one of the things that was obvious even - or starting to become obvious even a couple years ago was that frontier companies, some of our customers today, like Cursor, Sierra, ElevenLabs, Harvey, were going to have a very easy time selling a pilot to a bank. The, like, the demo just sells itself. It’s magic. But bringing that through, if you want to do a wall-to-wall rollout at a bank or a hospital, you have to go through the risk process. These banks have no idea even which questions to ask, let alone which answers are sufficient, let alone, like, how do they go and test whether these agents actually work the way they’re supposed to. And so they had this problem of, like, what can we say to earn the trust? And we think there’s, like, a golden sentence that goes something like, “Hey, I hear you’re really worried about hallucinations or jailbreaks or whatever it may be. We’ve had an independent third party test us against the gold standard. We passed with flying colors. And as a vote of confidence, the world’s most conservative insurers have looked at the data.” And they’re willing to take some of the risk onto their balance sheet.

Swyx [00:12:06]: Yeah.

Rune Kvist [00:12:07]: So if something does go wrong

Swyx [00:12:07]: There’s money behind it, yeah.

Rune Kvist [00:12:09]: Exactly. So that’s kind of like the link between all this. We can get into some of the hard parts related to the technical testing, which is, I think, the crux of the matter, but I’ll pause there.

Swyx [00:12:19]: How did you and Rajiv come together? This– there’s always, like, you come across very confident and, you know, and we’re announcing your Series A and all these things, but I want to see, like, the early initial stages of, like, idea formation.

Cofounding AIUC with Rajiv Dattani

Rune Kvist [00:12:31]: Yeah. Rajiv is actually my soon-to-be brother-in-law.

Swyx [00:12:35]: Oh.

Rune Kvist [00:12:36]: So I’m actually, in a week and a half getting married to Rajiv’s sister.

Swyx [00:12:42]: Okay, now you’re tight.

Rune Kvist [00:12:44]: Exactly.

Swyx [00:12:44]: Now you know.

Rune Kvist [00:12:45]: So - Rajiv and I have known each other for a decade. Funny story, I met both Rajiv and his sister, Hena, at the same time when Hena and I were interns at McKinsey in London, and Rajiv was assigned as my mentor. And so met them at the same time. For the longest time, it was not obvious that we were necessarily going to work together. I was in startups. He was, an insurance partner at McKinsey. Three or four years ago, I think Hena convinced him that AI was going to be a really big thing. And so he quit his job, cushy partner job at McKinsey in London, packed his bags, flew to San Francisco, and ended up joining METR. You guys are probably online enough

Swyx [00:13:24]: CEO.

Rune Kvist [00:13:24]: Exactly.

Swyx [00:13:24]: We’ve, we’ve, we’ve heard of METR.

Rune Kvist [00:13:25]: You see the plot– the chart of the horizons of the tasks that agents can take on is doubling extremely fast. So he was COO at METR, led their partnerships with Anthropic and OpenAI to test their models before release, but also working closely with the US and UK government, to figure out, like, how do you know whether a model can be released? And in some ways, that was, like, the perfect background. He’s spent a lot of time in insurance, knows that world, spent a lot of time with frontier testing of models. And so when I was bumbling around this idea space, starting with some of the ideas we talked about related to Waymo, as soon as we got into the content, we were both like, “Oh, this would be an amazing business to build together.” This is wrestling with the problem that we both think is the most important in the world from a market angle, which is kind of our intuitions is that the market can do a lot, and the faster AI moves, the harder it is for government to solve some of these problems. And then it took a little bit of time to work through what is it like to work with family.

Swyx [00:14:27]: Sure.

Rune Kvist [00:14:27]: And,

Swyx [00:14:30]: Because you were already dating at the time

Rune Kvist [00:14:31]: Yeah. Yeah, exactly.

Swyx [00:14:33]: Yeah.

Rune Kvist [00:14:34]: Already back then, it

Swyx [00:14:35]: Yeah.

Rune Kvist [00:14:35]: We felt like we were a family.

Swyx [00:14:36]: Nice.

Rune Kvist [00:14:36]: And so starting a business together felt like kind of a big step. And, here we are with just immense amounts of trust.

Vibhu [00:14:43]: Yeah. So now you’re a company of how big? How big are you guys now?

AIUC-1 Certification: Agent Security, Safety, and Reliability

Rune Kvist [00:14:46]: There are just 20 of us now.

Vibhu [00:14:47]: 20 of you guys now, have Series A, and you have your first certification out, the AIUC-1. Let’s bring up the certification. So this is the agent certification, right? What goes into the process? I have, like, two questions here. One is, walk us through the certification, and two is, what is the process for a company to get certified, you know?

Rune Kvist [00:15:08]: Great. As it says right on the top, AIUC-1 is a standard for agent security, safety, and reliability. The fundamental design principle is take all of the concerns that slow down adoption, so all the questions, all the fears that keep, security leaders in the Fortune 1000 up at night, and put them into one comprehensive framework. That’s what you’ll see there. You can see the six categories. Two, you want to ground all of this in technical testing. So one of the concerns with security standards that often feel kind of like theater paperwork is that they’re not actually ground out in, does any of this work? Does any of this matter? And so we had a conviction from early on that was going to be the kind of crux, was to pass this, you must get tested every quarter, basically run thousands of simulations to see, well, so can it actually be jailbroken? How hard is it to jailbreak? How often does it hallucinate? How often does it leak data? Et cetera. And then the last, core idea here, if you scroll up to the top here, is to refresh it quarterly.

Rune Kvist [00:16:08]: So the core trait of AI is that it moves extremely fast. Whatever concerns we’re discussing today were not the same ones three months ago, and this will keep changing. Typically, standards update on a, like, a decade cycle is obviously not going to work. But the question is kind of how do you update it? And the core thing here was to basically get the risk leaders of the Fortune 1000 around the table. So if you go over to the left here

Vibhu [00:16:32]: Yeah

Rune Kvist [00:16:32]: You’ll see the AIUC-1 consortium. The consortium is a group of risk leaders who run real banks, real hospitals, real critical infrastructure, who are facing these challenges every day. And we meet with these folks twice a quarter and hear what’s top of mind, what is keeping them up at night. There’s tremendous amount of desire for that conversation. And then we operationalize that into a specific

来源说明

当前保存的是 RSS 或来源节选,不代表原文全文。请以原始来源为准。

本页只呈现已保存的来源证据,不包含基于缺失正文的扩展推断。