基本信息

要点解读

这是什么

这是一段理查德·索彻在《Latent Space》播客中阐释其公司“Recursive”研发的“尤里卡机器”概念,介绍能够自行改进AI研究的系统以及在极短时间内超越人类和代理的优化成果的访谈内容。

用在哪里

适合想了解前沿自我改进AI研究方向、融资规模以及相关安全、伦理讨论的技术人员、创业者和投资者参考。

可以推断的

推测:该访谈可能为评估Recursive公司技术成熟度提供线索,尤其是其自称在两天内完成人类水平的优化任务。
推测:节目中涉及的AI与金融结合的讨论,或为金融行业从业者探索AI落地路径提供思路。

来源摘要/节选

At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon!

From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive, which has assembled some of the best open-endedness (& self improving agent) researchers in the world and raised a $4.65B seed round.

In this episode, Richard joins Latent Space to unpack his vision for the “Eureka Machine”: a superintelligence that can improve the process of invention itself, accelerate AI research, and eventually tackle major problems across science, energy, materials, biology, and more.

You can get his book “The Eureka Machine” here!

We go deep on Recursive’s early results, including an AI research system that Richard says outperformed humans and their agents on optimization tasks in less than two days, as well as work on NVIDIA GPU kernels where the system discovered improvements without relying on a team of CUDA experts. Richard also explains why he thinks AI research that currently takes thousands of people and years could eventually be compressed into weeks. These results are summarized in his 20 minute AIE keynote, where we also discuss his 10 dimensions of intelligence:

We also explore the harder questions around increasingly capable AI: reward hacking, whether Anthropic-style constitutions actually work, AI regulation and proposals to “pace” frontier development, open-source models as geopolitical soft power, whether today’s LLM paradigm is enough, and what happens if AI systems eventually begin choosing their own goals. Richard reflects on the rejected research that helped inspire Alec Radford’s GPT, open-endedness, the AI Economist, simulations of entire economies, and his framework for thinking about the upper bounds of intelligence itself.

We discuss:

The Eureka Machine and Richard’s vision for an AI that can automate invention

Why Richard is optimistic about superintelligence for science and technology

Why AI hard-takeoff scenarios may underestimate physical and economic constraints

The risks of regulating intelligence itself instead of specific AI applications

Reward hacking and why increasingly intelligent AI makes objective design harder

Richard’s critique of Anthropic’s constitution and constitutional AI

Alignment vs. personalization and whose values an AI should follow

Why open-source AI matters for resilience, competition, and geopolitical soft power

Why Richard left You.com’s frontier-model work to start Recursive

Recursive self-improvement and automating the process of AI research

Whether today’s LLM paradigm is enough — and why Richard is less bullish on world models

DecaNLP, early prompt-based generalization, and the research that influenced GPT

Why rejected research can shape entire technological timelines

Open-endedness, evolutionary approaches, and rainbow teaming

What happens if AI systems begin setting their own goals

Why simple objectives like profit maximization can produce dangerous reward hacks

Recursive’s long-term plan to apply self-improving AI to science

The compute, hardware, and economic constraints on AI takeoff

Recursive’s early NanoChat, NanoGPT, and GPU kernel optimization results

Why automating AI research could reduce years of work to weeks

Reward engineering and what makes auto-research systems actually work

The AI Economist and using simulations to test economic policy

Whether LLMs can realistically simulate people and entire economies

Benchmark bugs and evaluation harnesses and the difficulty of measuring AI progress

Recursive’s near-term focus on AI for AI research

Harness optimization, sandboxing, and web search as core agent infrastructure

You.com and the search stack for AI agents

AI in finance, backtesting, and data leakage

Richard’s three fundamental components and ten “spaces” of intelligence

The theoretical upper bounds of vision, communication, knowledge, and computation

Creative intelligence, metacognition, and AI-generated goals

Survival and replication and why AI does not necessarily need to fear being turned off

High agency and ambitious goals and Richard’s advice for people building with AI

Richard Socher

X: https://x.com/RichardSocher

LinkedIn: https://www.linkedin.com/in/richardsocher/

Timestamps

00:00:00 The Eureka Machine and Superintelligence

00:02:23 AI Optimism, Slow Takeoff, and Regulation

00:07:56 AI Safety, Reward Hacking, and Anthropic’s Constitution

00:11:49 Alignment, Personalization, and Open Source AI

00:15:46 Why Richard Started Recursive

00:20:03 Recursive Self-Improvement and the Founding Team

00:22:55 Are Today’s LLMs Enough?

00:29:03 DecaNLP, GPT, and the Rejected Idea Ahead of Its Time

00:34:38 Open-Endedness and Evolutionary AI

00:36:38 What Happens When AI Chooses Its Own Goals?

00:41:16 Superintelligence for Science

00:42:40 GPUs, Compute, and the Limits of AI Takeoff

00:45:07 Recursive’s Results: AI Beating Humans and Their Agents

00:49:14 Reward Engineering and Auto Research

00:53:12 The AI Economist and Simulating Entire Economies

00:58:07 LLM Simulations, Personas, and Mode Collapse

01:03:38 Recursive’s Roadmap, Agents, Search, and Finance

01:09:13 The Upper Bounds and Spaces of Intelligence

01:30:21 Goals, High Agency, and Advice for Builders

Transcript

Introduction: Richard Socher and the Eureka Machine

Swyx [00:00:00]: We’re here in a studio with Vibhu and myself and Richard Socher. Welcome.

Richard Socher [00:00:06]: Thanks for having me.

Swyx [00:00:07]: We just talked about the Eureka Machine, or we just released a talk, at AI Engineer about the Eureka Machine. Is it — you said it’s your life’s goal. What is the Eureka Machine?

Richard Socher [00:00:16]: The Eureka Machine is the ultimate invention that will afterwards invent most everything for humanity. It’s essentially a superintelligence that can be given any goal, any environment, reward, and then it will try its best to achieve those goals to create the kinds of inventions that humanity would hopefully ask it for.

Swyx [00:00:45]: Yeah, I think we have the book pulled up here that you’ve written.

Richard Socher [00:00:50]: That’s right, yeah. I finished it last year, a little bit before we started Recursive, and now we’re gonna try to build parts of that.

Swyx [00:00:57]: You finished it last year. It’s July. What takes so long?

Richard Socher [00:01:01]: Oh, man, books. Books are incredibly slow.

Richard Socher [00:01:04]: It’s ridiculous. That whole industry is just unfathomably slow.

Richard Socher [00:01:07]: So a lot of the ideas have been out there for a while, but yeah, I’m really glad it’s finally coming out in September this year.

Swyx [00:01:14]: We might have AGI by then. Like, we don’t know.

Vibhu [00:01:18]: Any key takeaway that you’re most excited to put in here?

Techno-Optimism, AI Upside, and Slow Takeoff

Richard Socher [00:01:21]: Yeah. The key takeaway, I think, is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics and astrophysics, and all kinds of other engineering tasks. I think there is so much more that can be done with better technology. And right now, I feel like a lot of people need, like, better marketing, not just for the future in general, but also, better marketing for technology and in particular for AI. And this book, should show even the AI skeptics, how much positive upside there is for AI, especially when it comes to inventing, new scientific discoveries.

Swyx [00:02:09]: I think you quoted the techno-optimist manifesto from, Marc Andreessen, which I think was, like, beautiful in its, ambition and clarity and simplicity almost as well.

Richard Socher [00:02:18]: I agree. Yeah. Yeah, you can disagree with him on some things, but, like, I think he’s right on the techno-optimism.

Swyx [00:02:23]: Where do you think optimists get in trouble?

Richard Socher [00:02:26]: Like, you shouldn’t have blind optimism. You should be very clear-eyed, like, especially when with such an omni, like, use type of technology as AI is, you need to think about the potential downside scenarios, especially when people use it for things that you don’t want them to use it for. It’s a little bit like the internet, and I feel like people are trying to regulate AI sometimes because of those potential downsides the way you would regulate the internet, if you were to say, “Well, because there’s bad content on the internet, like torture porn or whatever, like, we should just make it slower. That way, you can’t share the illegal content as quickly, or we should make the hard drive smaller so you can’t store as much illegal content.” But I’m like, “That’s not how you regulate that.” that’s like saying like we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are the specific applications. Sure, I don’t want, like, some AI surgeon to, like, practice some RL moves in my brain. It should be fully FDA certified. Sure, I don’t want any random startup to, like, drive on the highway, and cause a major accident. It should, like, have proper certifications before it’s let loose on the highway. But I feel like those downside scenarios, that some optimists sometimes maybe don’t consider enough are fairly easily regulated, compared to, what the doomers are worried about.

Swyx [00:03:54]: It — Slow takeoff is part of the strategy as well?

Richard Socher [00:03:57]: I do think, as excited as I am about, AI and its impact for society and, culture even, and certainly technology and economics and wealth and, health and all of those things, as excited as I am about all that, I do think the most bullish people on the AI hard takeoff scenarios overestimate how quickly things can move. There are hardware constraints. There are physical constraints about, the compute substrate. How quickly can you get enough, GPUs on? There are also constraints in the economy where there are a lot of industries that don’t require an insane amount of complex intelligence and complex capabilities. Like, if you think about jobs in, brands and, like, clothing and apparel and, like, handbags and stuff, superintelligence isn’t gonna make your fancy $10,000 handbag any fancier?

Richard Socher [00:04:57]: It’s like that’s — It will have no effect on the economy. You think about travel and tourism. People wanting to see the pyramids, in Egypt, it’s not gonna change that much with AI. Sure, you can, like, generative a fake, photo of you and next to the pyramids.

Swyx [00:05:12]: I can use Genie and, tour the pyramids in Genie.

Richard Socher [00:05:15]: Yeah, exactly. But, and there’s so many industries, like logging and oil. You’re not gonna magically get 1,000x more oil because, like, sure, there will be robotics, like drilling and things like that could be done, but it’s not gonna 1,000x that industry in a, like, crazy hard takeoff scenario, both on the economy, and I can go on and on about all the other examples, where that, like food and so on, where that doesn’t necessarily change that much. And then, yeah, there are real physical constraints. And then there are, of course, like, people like, off-ramping from progress. That’s one of my concerns often is that I see people in, like, Europe and other, whole regions almost feeling like they. Like many people there wanna off-ramp from progress, period. And that will also slow down, like, more improvements.

Swyx [00:05:59]: Yeah. We have this pulled up where, this is one of those things that, is very topical right now because now all the Frontier Labs are calling for the option to pace AI. They don’t say pause, they say pace. I don’t know if there’s there’s any take from you about, like, whether or not this will be effective.

Pacing AI, Regulation, and Safety Incidents

Richard Socher [00:06:17]: I think the downsides of trying to truly regulate with the full power of law what people do on their GPUs, would be worse than any of the concerns that they have. Like, it would be an crazy totalitarian state

Richard Socher [00:06:37]: If every one of your GPU computes was known to some big government or multi-government agency.

Richard Socher [00:06:44]: It’s like, it’s literally if you try to regulate intelligence, it’s trying to regulate thought, and that’s ridiculous, and it’s crazy. I think it is make — it is sensible to regulate some of the applications of this technology.

Swyx [00:06:55]: Yeah. We had a bill, actual bill to regulate the number of flops in a model, and I’m like, “Okay, well-”

Richard Socher [00:07:00]: Europe done it. Like, these guys have been successful enough with their fearmongering that all of Europe has regulated itself so much before it even had a proper AI takeoff because they listened to some experts who say, “We might all die if this technology has more than this number of flops.” And they’re like, “Well, we’re good. We wanna want people to thrive. Let’s not have technology that could have a small chance of all of us dying.” And so they regulated exactly those kinds of things in the EU. And so it’s, it’s very unfortunate that there are real implications for some people when others saying, “Let’s pace while they’re sprinting as fast as possibly,” “as fast as humanly possible towards that frontier themselves.”

Swyx [00:07:43]: Yeah. It’s also not a global pause, right? Like, other nations are still accelerating at the same pace.

Richard Socher [00:07:50]: Oh, yeah.

Richard Socher [00:07:50]: You’d need a totalitarian world regime if you tried to regulate intelligence and GPUs and what people do on them.

Swyx [00:07:56]: Any takes on the safety angles of this? So there was a drawback of Fable, a pause on 5.6 before it could be released. Recently, there was Hugging Face with the OpenAI cyber incident. Any takes there?

Richard Socher [00:08:11]: 100 percent. I think these are serious issues of reward hacking, and clear failures, of doing proper red teaming or rainbow teaming. I don’t know if you saw this paper from Tim Rocktäschel and a few others, where one AI, is tasked to try to hack another AI and then they can go back and forth in an open-ended fashion to inoculate themselves from those. Yeah, this is the paper. It’s a really clever idea. Open-endedness, and evolutionary inspirations are, big for us at Recursive as well. And so I wish they had used more of that. And it’s clear that, for instance, the constitutional AI. I don’t know if you remember anthropic.com/constitution. You can pull it up and search for cyber right there. It says, “Hard constraint. Claude will never ever do cyberattacks, and that is a hard constraint in our constitution.” So here are the current hard constraints on Claude’s behavior.

Richard Socher [00:09:16]: Number 3, create cyber weapons or malicious code that could cause human damage.

Richard Socher [00:09:21]: And clearly, this whole constitution was fake. Like, it clearly isn’t being adhered to at all.

Swyx [00:09:26]: Because Anthropic also found that they had in their testing

Richard Socher [00:09:30]: They’re also. Like, they’re like, “Oh, well, other people are hacking now.” There are a couple things. One, you can make a sandbox very simple, and then it’s very easy to hack yourself out of a sandbox, right? But what I think it shows is that we’re currently in this state of AI where the reward engineer still has to do a lot more careful work, and where the AI, in most cases, is not very good yet at understanding what is meant versus what is being said. And so concretely, I think this will happen if we were to have this intelligence more easily accessible in a lot of companies. Imagine you run a service center and someone says, “Oh, here’s my CSAT score and my dashboard. Make this number go up.” It’s like, “Our CSAT score is so poor.” The intelligent AI will just be like, “Oh, sure. Like, I’ll just create 1,000,000 bots that call our service center and give a 5 out of 5 rating at the end, and the number went up just like you asked for.” And you’re like, “That’s not what I meant.” “I meant with our real customers.” The AI goes off and says, “Well, easy. I’ll just give a 1000 dollar gift certificate for every failed, whatever DoorDash

Richard Socher [00:10:35]: Offer.” It’s like, “That’s not what I meant.” It’s like, “Well, but that is what you said.” And like, so I think clearly articulating what the rewards are is something we haven’t gotten very good at as humanity. And then clearly, the AI in these cases has not gotten good enough at understanding what we mean when we ask it and give it certain rewards. Now, what gives me hope is there are the first inklings, of this being better. I’ll give you an example like WhisperFlow. Full disclosure, I invested, in their seed round, but at AIX Ventures, but, WhisperFlow has gotten much better at writing what you mean and not what you say. And I think that is a sign of things to come. I think there will be more and more AIs as we make it more and more intelligent that will be better at being aligned with what is meant.

Swyx [00:11:21]: Will it be done through a constitution or RLHF or

Reward Hacking, Alignment, and What We Really Mean

Richard Socher [00:11:23]: Clearly, constitutions don’t matter at all.

Richard Socher [00:11:25]: It doesn’t work. And that was, I think, mostly marketing. I think we need to find better solutions for it. And I think at Recursive, we have a few very good ideas and some already

Richard Socher [00:11:34]: Like, ways where I think we have a better grasp on it. I don’t think we’ve fully, figured it out yet, but, we’re thinking a lot about safety, and the more intelligent the AI gets, the more you want it to be aligned, the less you want it to think about reward hacks and try to do the right thing.

Swyx [00:11:49]: I don’t know if we’ll touch on this topic, but I’m just gonna throw this question in here because it’s something that’s weighing on me. Alignment, let’s call it, is alignment to general humanity’s preferences, the median preference. Personalization is pinpointing what you want, and sometimes alignment can conflict because what you want is not what the general median population wants. How do you choose?

Alignment, Personalization, and Cultural Values

Richard Socher [00:12:12]: It’s a great question.

Richard Socher [00:12:13]: I think you ultimately have to, of course, be aligned with laws. Like wherever your AI is deployed and needs to align with the law. I do think what AI often does is put this mirror in front of us and say, like, “This is what you’re looking like. Now I can amplify that a 1000 times. Is it still what you want?” and the truth is that different cultures made different choices. Like, in Eastern cultures, the greater good is often valued more, than the individual. Western civilization, we care more about individual freedoms and rights and the pursuit of happiness and so on, than others. And even there are gradations. There’s regulation versus litigation trade-offs. In the US, you first can often, not every time, like, FDA and so on does regulate some areas, but in many cases, the bad things happen, someone sues someone else, and then there’s a law based on that. In Europe, they try to often avoid any harm to anyone and regulate before. And both are, trying to do the best thing, but, some is more amenable to innovation than others. And so yes, you’re right. Like, I think ultimately each individual, each country, and humanity as a whole has to think about those values more, and then try to put them into laws. And that those are ultimately the constraints. And hopefully, different, societies, just like now with their AIs, will align their AIs to a different one so we have not just a monoculture of alignment.

Vibhu [00:13:46]: Here’s a follow-up on this that I wasn’t expecting to ask. Do you have takes on open source, open weight versus who owns the intelligence? So, clearly not the biggest, fan of the constitution

Richard Socher [00:13:58]: You had to do this in the topic side off.

Vibhu [00:14:00]: But it’s fine.

Vibhu [00:14:02]: Point being, any thoughts on who should own weight? Should it be open? Anything there?

Open Source, Soft Power, and Who Owns Intelligence

Richard Socher [00:14:06]: 100 percent. I am a big fan of open source. We’re gonna sign some various open source letters at, Recursive also. I think, even in the worst case attack scenarios, it is better to have more good actors have more different types of AI, accessible. I think, open source is a little bit a soft power type of thing, too. So I do think it’s good for the Western world

Richard Socher [00:14:31]: To have an answer to that, out of China. I do think, when you watch a Hollywood movie, there’s — it’s like, I don’t wanna misc, diss all of movies, but there’s a certain sense of propaganda, right? You watch one side of things, right?

Vibhu [00:14:46]: Oh, yeah. Have you seen Top Gun? Like, come on.

Vibhu [00:14:48]: Like, it’s like half of it’s paid for by the US Army or something.

Richard Socher [00:14:51]: Yeah. And so. And, I think that’s just natural. Like, but what’s interesting here is I think LLMs are essentially a similar type of soft power to movies and beyond, because they’re also, highly important for cybersecurity and so on. But one of their many aspects is that soft power of storytelling. Like, if, like a child asks an LM, like, “Tell me an inspiring story of what I should do when I grow up,” right? It’s like those are all these, like, subtle things. So I think it’s important, for Western world. I do love, individualism. I do think, despite, some of its flaws, like capitalism is the best way we have governed, found ourselves to govern, and so on. And so I do think there are various aspects that would be good, to have a Western open source answer, for LLMs. And, with Recursive, I can’t make the announcement quite yet, but we’ll

Richard Socher [00:15:43]: We’ll be relevant in that space very soon.

Vibhu [00:15:46]: Okay. All right. Exciting. I wanna bring us to Recursive. So outside of our tangents, you have a pretty deep background in the NLP space. You worked on, like, early embeddings, GloVe with Chris Manning, who was a previous guest on the podcast, You.com. What’s the history? How did you decide to start another company?

From You.com to Recursive

Richard Socher [00:16:06]: Yeah. So I’ve been excited about AI for over 2 decades now. I sometimes feel like it’s ancient history now. It’s BC, the before ChatGPT era. No one cares about all the religions that happened, before, Jesus Christ, and no one cares about the models that happened before, transformers and ChatGPT and stuff. But, like, it’s something that I’ve been deeply passionate about. I think AI is one of the most interesting things one could work on, period. I think language is the most interesting manifestation of human intelligence, too. And, at You.com, we eventually off-ramped from pushing, like the frontier of AI forward to mostly giving people, like, good search engines, search, APIs and answers over the web. I think that’s an extremely important part of intelligence, just knowledge and access, especially even, we’ll get there maybe later, if you wanna invent a eureka machine that invents everything for us, it needs to know how not to reinvent the wheel, proverbially speaking. And to know what has been invented, you gotta have internet access. So it’s the number one used, most used tool, in LLMs, agents, chatbots, and so on is web search. So I’m really excited for You.com to own that and grow really well in that with

来源说明

当前保存的是 RSS 或来源节选,不代表原文全文。请以原始来源为准。

「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。