Research guide / Recursive self-improvement

Recursive self-improvement (RSI): what it is, and how to tell whether it happened

A sourced guide to the definition of recursive self-improvement in AI, its history from I. J. Good to 2026 systems, and the controls an RSI claim needs before it counts as evidence.

External results below are summarized from their authors’ own reports and linked to the source. They were not re-run or independently verified for this guide.

What is recursive self-improvement?

Recursive self-improvement (RSI) is a process in which an AI system improves the process it uses to improve itself. Each accepted change becomes part of the system that makes the next change, so a successful round can make later rounds more effective.

The word “recursive” is what separates RSI from ordinary self-improvement. A model that gets better at a task through practice or feedback is improving itself. RSI requires more: the improver must change, and the changed improver must produce better improvements than the original one would have.

Published definitions differ mainly in how strong a process they require:

  • Strong, classical sense. Wikipedia describes RSI as a process in which early artificial general intelligence systems rewrite their own code, possibly causing an intelligence explosion that leads to superintelligence [1].
  • Reflexive sense. LessWrong defines it as improving one’s own ability to make self-improvements [2].
  • Operational, full-autonomy sense. Anthropic describes the endpoint as an AI system that can fully autonomously design and develop its own successor [3].
  • Bounded, engineering sense. Much recent work studies loops in which an agent edits its own code or scaffolding, and each accepted rewrite becomes the agent that the next round edits [10].

A 2026 survey of 1,250 arXiv papers argues that terms such as “self-refine,” “self-reward,” “self-play,” and “self-evolve” conflate different ambitions. It separates bounded self-refinement, which is convergent, evaluable, and already industrial practice, from open-ended RSI [9]. This guide uses the same distinction.

A short history of the idea

RSI began as a thought experiment about machine intelligence. Since 2023 it has also become an experimental research program with measurable, if bounded, results.

  1. 1965: the intelligence explosion. In “Speculations Concerning the First Ultraintelligent Machine,” I. J. Good argued that a machine able to design better machines could trigger an “intelligence explosion” [4].
  2. 2003: Gödel machines. Jürgen Schmidhuber proposed self-referential problem solvers that rewrite their own code once they can prove the rewrite is beneficial [5].
  3. 2000s: seed AI and takeoff. Eliezer Yudkowsky’s “seed AI” was designed to gain intelligence mainly through RSI, and he argued that such systems would likely produce a hard takeoff [2].
  4. 2023: STOP. Zelikman, Lorch, Mackey, and Kalai showed that a scaffolding program that calls GPT-4 could improve its own code. It discovered strategies such as beam search and genetic algorithms while the underlying model stayed fixed [6].
  5. 2025: the Darwin Gödel Machine. Zhang, Hu, Lu, Lange, and Clune replaced formal proof with empirical validation. Their coding agent modified its own code and reported gains from 20.0% to 50.0% on SWE-bench and from 14.2% to 30.7% on Polyglot [7].
  6. 2025: AlphaEvolve. Google DeepMind reported that its Gemini-powered coding agent sped up a kernel in Gemini’s architecture by 23%, which reduced Gemini’s training time by 1% [8]. In that case, an AI system contributed to the infrastructure used to train its own model family.
  7. 2026: a research field. ICLR 2026 hosted a workshop on AI with recursive self-improvement. A July survey mapped the literature [9]. In September, AIDE² reported seven successive self-discovered improvements to a research agent over an eight-day run, with gains on four held-out benchmarks [10].

Does recursive self-improvement exist yet?

It depends on the definition.

Bounded RSI has been demonstrated. STOP, the Darwin Gödel Machine, and AIDE² each report loops in which a system rewrote the code or scaffolding it uses to improve, then used the rewritten version in later rounds [6][7][10]. In these examples, the underlying model weights stay fixed or are trained outside the loop, and humans choose the benchmark tasks.

Open-ended RSI has not been demonstrated. Anthropic describes full RSI as not yet reached and not inevitable [3]. The 2026 survey concludes that open-ended RSI remains bounded by grounding requirements, collapse dynamics, and compute constraints on every axis it examined [9].

The survey also names governance-grade measurement of self-improvement as the field’s most underpopulated niche [9]. That measurement problem is the focus of this project.

How to test a recursive self-improvement claim

A rising score across rounds is necessary but not sufficient evidence of RSI. Many ordinary effects can produce the same curve. Each confound below has a control that removes it. Several were found in this project’s own experiments.

Common ways an apparent RSI result can mislead, and the control for each.
ConfoundWhat it can fakeControl
Extra computeLater rounds do better because they spent more search, not because the improver improved.Compare k recursive rounds with a single round given k× the compute.
Co-designed evaluatorThe task generator and search procedure fit each other, so almost every change is “adopted.”Freeze the evaluator. Use held-out tasks the loop never selected on.
Wrong comparatorBeating one baseline looks like beating all of them.Report every contrast, including the untrained baseline.
Non-recursive gainContinued search improves the task solution while the improvement process stays the same.Compare with a control denied the self-modification mechanism, using identical random streams.
Aggregate maskingA single composite hides a failing dimension or a missing one.Publish every axis score, the weights, and which axes were supplied.
Selective reportingOnly runs that worked are visible.Publish all runs, retractions, and postmortems.

The first two rows are drawn from this project’s own record. An earlier compounding-improvement claim in gated-self-improvement depended on a 5× compute mismatch and was retracted; the repaired result that replaced it was retracted after an audit found unshared random streams and a data confound. Separately, a 15/15 adoption result in rsi-metaforge-core was traced to a task generator and search engine that had been co-designed, and it was rejected before publication. See the failure log.

Open RSI experiments from this project

The Intelligence Research Project studies bounded, testable forms of recursive self-improvement with open code. It does not claim to have demonstrated artificial general intelligence or open-ended recursive improvement.

Gated self-improvement: pre-registered results

This is a deterministic, LLM-free Python engine with counterfactual controls. Its latest experiment asks whether experience at improving carries over from one problem to the next. Pre-registered and run once on 300 untouched seeds, an improver that kept its improvement memory across problems beat the identical learner with its memory wiped by +1.52 final-holdout tasks per run (p = 0.0011). Learning within a single problem did not beat a frozen policy, and the advantage did not grow over later problems. Read the results and limits →

RSI-Bench: measuring recursive self-improvement

RSI-Bench is a measurement framework with six axes: self-modification depth, improvement trajectory, operator discovery, meta-adaptation, safety and stability, and goal generation. Its composite uses a weighted harmonic mean, so a weak axis cannot hide behind strong ones. Missing axes are skipped rather than scored as zero. See how the score behaves →

rsi-metaforge-core: validation-gated synthesis

rsi-metaforge-core explores meta-meta learning loops and anti-cheat verification inside a program synthesis runtime. Its main lesson so far is negative: the 15/15 adoption result above was traced to a co-designed task generator and search engine, so it was rejected. Read the program-synthesis summary →

Frequently asked questions

What is RSI in AI?

In AI, RSI stands for recursive self-improvement. It means a system improving the process by which it improves, so that its gains can compound across rounds. It is distinct from a model simply getting better at a task.

Is recursive self-improvement the same as an intelligence explosion?

No. An intelligence explosion is one hypothesized outcome of RSI, in which each round arrives faster and is larger than the last [4]. RSI can also level off. The 2026 survey reports that measured loops have so far stayed bounded [9].

Has anyone built a recursively self-improving AI?

Bounded versions exist. Agents such as the Darwin Gödel Machine and AIDE² rewrite their own code and keep the changes that score best when evaluated [7][10]. No published system has shown open-ended, fully autonomous RSI.

How is recursive self-improvement measured?

Measure improvement per round against matched-compute controls on held-out tasks. Check that later gains depend on earlier changes to the improver, and report each dimension separately rather than as a single score. The test checklist and RSI-Bench describe one way to do this.

References

  1. Recursive self-improvement. Wikipedia. Accessed September 26, 2026.
  2. Recursive Self-Improvement. LessWrong wiki.
  3. When AI builds itself. Anthropic.
  4. Good, I. J. (1965). Speculations Concerning the First Ultraintelligent Machine. Advances in Computers, 6.
  5. Schmidhuber, J. (2003). Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements. arXiv:cs/0309048.
  6. Zelikman, E., Lorch, E., Mackey, L., & Kalai, A. T. (2023). Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. arXiv:2310.02304.
  7. Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv:2505.22954.
  8. Google DeepMind (May 14, 2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms.
  9. Chen, M., Wang, L., & Qu, B. (2026). Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv:2607.07663.
  10. Srikanth, D., Zhao, B., Xu, D., Wu, Y., & Jiang, Z. (2026). Recursive self-improvement of AI research agents. arXiv:2609.26457.

Browse all research notes →

Support independent research

Help fund the next experiment.

Each run is bounded by the compute it can afford. In the latest experiment a run gives the improvement process about 50 decisions, where measuring its options would take roughly 400. Research funding would buy longer runs and more of them. Every result will be published, including the failures.