Why I still haven’t bought into true RSI

Frontier AI labs are scaling inference‑time agents for efficiency, but true recursive self‑improvement remains unlikely in the near term.

Why I still haven’t bought into true RSI

Why I Still Haven’t Bought Into True Recursive Self‑Improvement

The buzz around recursive self‑improvement (RSI) has reached a fever pitch. Every week a new paper, podcast or tweet claims that the next leap in AI will come from an autonomous system that learns to make itself smarter, and that this leap will happen within a few years. I’m not convinced. My view is that the current trajectory of frontier labs—OpenAI, Anthropic, and their close‑knit community—shows a pattern of rapid scaling of inference‑time agents, but that the deeper, systemic bottlenecks that would enable true RSI remain far from resolved.

1. The “agent swarm” reality

Frontier labs are already deploying thousands of concurrent agents to tackle well‑defined, measurable problems: code‑generation, data‑labeling, and even internal research pipelines. These agents are a powerful tool for efficiency—they can iterate on experiments, monitor logs, and automate routine tasks. The upside is clear: productivity in research and product development can jump tenfold. The downside is that these agents operate within the constraints of the existing architecture. They do not yet possess the ability to design new architectures, discover novel training objectives, or fundamentally alter the learning dynamics of the underlying models.

The key point is that scaling inference‑time compute is a different beast from scaling intelligence. The former is governed by well‑understood scaling laws: more compute, more parameters, better performance. The latter requires breakthroughs in representation, learning algorithms, and, crucially, a way to measure and direct that progress. Until we have a clear metric for “intelligence” that is both meaningful and actionable, claims of imminent RSI are speculative.

2. Diminishing returns and resource bottlenecks

Even if we give the labs unlimited compute, the law of diminishing returns still applies. Each additional agent or larger model yields smaller incremental gains in performance. Moreover, the resource side of the equation is not infinite. Training large language models (LLMs) is expensive, and the cost of compute is rising faster than the rate of algorithmic efficiency gains. Labs that plan to go public or compete on price will face intense pressure to keep margins healthy, which limits the amount of compute that can be devoted to internal R&D.

Post‑training work—tuning behavior, safety alignment, and fine‑tuning for specific tasks—remains a highly human‑driven process. Automating these tasks is difficult because they involve nuanced judgments about how a model should behave in edge cases. As John Schulman noted, “It would be very hard to automate the whole thing, just because someone has to think about how the model should behave in this area.” This human‑in‑the‑loop requirement is a hard bottleneck that RSI alone cannot overcome.

3. The cultural temperature and risk perception

The AI safety community has become highly attuned to the possibility of a rapid intelligence explosion. This heightened anxiety is amplified by the competitive, high‑stakes environment of the San Francisco AI scene. When thousands of agents are already working productively in a company, the fear that they could suddenly become superintelligent is magnified. However, the historical record shows that many early warnings about AI risks did not materialize on the predicted timelines. The current anxiety may therefore be more a product of cultural amplification than of imminent technical breakthroughs.

4. What would change the game?

For RSI to become a reality, we would need a foundational breakthrough that changes the way we learn and represent knowledge—something that is not just an incremental improvement in scaling laws. This could be a new learning paradigm, a radically more efficient architecture, or a way to harness online learning at scale. Until such a breakthrough appears, the trajectory of progress is best described as lossy self‑improvement: we can make models cheaper and more efficient, but we are unlikely to see a sudden jump to superintelligence within the next decade.

Practical takeaway

  • Invest in scaling inference‑time agents: They deliver real productivity gains now and will continue to do so as compute becomes cheaper.
  • Monitor but don’t panic about RSI: The current evidence points to incremental, not exponential, progress in intelligence.
  • Support research into foundational breakthroughs: The real leap will come from new ideas, not from more compute.

The frontier labs are doing impressive work, but the path to true recursive self‑improvement remains uncertain and likely longer than the hype suggests.

---

TL;DR: Frontier AI labs are scaling inference‑time agents for efficiency, but true recursive self‑improvement—an autonomous leap to superintelligence—remains unlikely in the near term due to diminishing returns, resource bottlenecks, and the need for foundational breakthroughs.

Source

Read original source

Why we picked this

Funding and policy angle on AI integration in regulated workloads.

← Back to all articles