AI researchers debate how close we are to recursive self-improvement

AI researchers debate how close we are to recursive self-improvement

September 11, 2026 • 1 hr 37 min
🎧 Listen Now

🤖 AI Summary

Overview

This episode features a deep dive into the future of AI research and development, focusing on the challenges, breakthroughs, and timelines for achieving recursive self-improvement (RSI) and artificial superintelligence (ASI). The discussion explores technical bottlenecks, the role of reinforcement learning (RL), and the implications of scaling AI systems for real-world applications.

Notable Quotes

- There's this cycle that keeps repeating where a new model comes out, people are blown away, and then it starts to feel dumb after a month or so. – John Schulman, on the iterative hype and limitations of AI models.

- The models are learning all these very discrete tasks, and when you average them together, you get this very smooth loss curve. But on an individual task level, it may look like a phase transition. – Charlie O’Neill, on the emergent capabilities of AI systems.

- Even if you have a few bits of RL, the actual impact on function space can still be super dramatic. – Beren Millidge, on the efficiency of reinforcement learning in fine-tuning AI behavior.

🧠 Why Recursive Self-Improvement (RSI) Might Stall

- Beren Millidge suggests that AI could hit a Marvex Paradox, where models excel at benchmarks but fail to generalize effectively in real-world applications.

- John Schulman highlights the bottleneck of models being unable to self-check or make long-horizon judgments, leading to repeated cycles of overestimation and disappointment.

- Charlie O’Neill notes that current architectures like transformers may be far from the global optimum for learning, requiring paradigm shifts to achieve RSI.

🔄 Reinforcement Learning (RL) and Its Role in AI Progress

- RL is seen as a key driver for improving AI capabilities, but its efficiency is debated. Beren Millidge explains that RL provides high-signal updates by focusing on correct outcomes, avoiding the noise of supervised fine-tuning (SFT).

- Charlie O’Neill emphasizes RL's strength in enabling models to handle longer task horizons, doubling their effective working time every three months.

- However, RL's limitations include reduced diversity in outputs and potential overfitting to specific environments, as noted by John Schulman.

🌐 The Sim-to-Real Gap and Long-Horizon Learning

- The transition from simulated environments to real-world applications remains a challenge. John Schulman argues that sim-to-real frameworks may not dominate forever due to the difficulty of simulating complex human interactions.

- Beren Millidge points out that current models require thousands of interactions to learn effectively, making sample efficiency a critical bottleneck.

- The panel agrees that continual learning and real-time weight updates could eventually close this gap, but technical and economic barriers persist.

📊 Scaling Laws, Data, and Model Size

- The discussion highlights the interplay between data, compute, and architecture in driving AI progress. Charlie O’Neill notes that diminishing returns in data quality and availability could slow pre-training gains.

- Beren Millidge argues that larger models are more sample-efficient, but scaling is constrained by hardware limitations and inference costs.

- John Schulman raises concerns about sparsity in models, which could impact data efficiency and scaling laws.

🚀 Timelines for Key AI Milestones

- Fully General Remote Worker AI: Predictions range from 1 to 3 years, depending on whether the AI operates in structured environments or unstructured, browser-based tasks.

- 10x Productivity Uplift for AI Researchers: Estimated within 2 to 5 years, contingent on advances in automating experimental feedback loops.

- Artificial Superintelligence (ASI): The panel's estimates vary widely, from 3 to 10 years, reflecting uncertainty about solving long-horizon learning and domain-specific challenges.

AI-generated content may not be accurate or complete and should not be relied upon as a sole source of truth.

📋 Episode Description

New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.

Watch on YouTube; read the transcript.

Sponsors

* Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street’s tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to antithesis.com/dwarkesh

* Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at x.ai/bot

* Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at janestreet.com/dwarkesh

Timestamps

(00:00:00) – Steelmanning the case against RSI

(00:18:39) – What’s driving the Chinese labs’ progress

(00:28:06) – How will automated AI researchers be trained

(00:33:51) – Will long-horizon RL elicit AGI?

(00:45:24) – The sim-to-real gap

(01:00:33) – Ho