Open-source AI research from the Intelligence Research Project. The central question: can a system improve its own methods, and can that improvement survive an independent check?
Each note connects a research question to pinned code, interpretable results, and a reference you can cite. New to the topic? Start with the guide to recursive self-improvement.
Two pre-registered experiments, run once on untouched seeds: carried improvement memory, and recursion against a compute-matched single round. Includes the nulls and the retracted first headline.
Recursive self-improvement (RSI) research asks whether a system can change the methods it uses to solve problems and improve through those changes. A higher score alone does not establish this: extra compute, an easier task distribution, or a mismatched control can also produce an apparent gain.
Ordinary self-improvement runs the outer loop. Recursive self-improvement also revises the inner one: the process that proposes and keeps changes.
This project studies bounded, testable forms of self-improvement. It does not claim to have demonstrated artificial general intelligence or open-ended recursive improvement.
The featured experiment
gated-self-improvement is an LLM-free, deterministic engine with counterfactual controls. Its latest experiment, pre-registered and run once on 300 untouched seeds, asks whether experience at improving carries over. An improver that keeps its improvement memory across ten independent problems beat the identical learner whose memory was wiped between problems by +1.52 final-holdout tasks per run (p = 0.0011), on task families neither had seen. The effect is small, about 2.4%, and it does not keep growing over later problems. An earlier pre-registered test found that five recursive rounds beat one round with five times the compute by +0.40 of 24 tasks. Read its claim boundaries before interpreting a result.
Why matched controls matter: an earlier compounding-improvement claim in gated-self-improvement relied on a 5× compute mismatch and was retracted. The repaired result that replaced it was retracted too, after an audit found that the arms were scored under different random streams and that the control saw fewer distinct tasks. The original report remains part of the research record, with the correction.
02 / Measurement
Score several axes, not one number
A single improvement score cannot say where a gain came from. rsi-bench is an open-source framework that scores a self-improving system on six axes: self-modification depth, improvement trajectories, operator discovery, meta-adaptation, safety, and autonomy.
Reading the axes separately shows which part of the process changed. An aggregate alone can hide a system that scores well on one axis and has nothing to report on another.
Scope: the methods note explains the published benchmark and its scoring. It does not report a new benchmark run or an independent replication.
Program synthesis generates candidate programs from a search space. In validation-gated synthesis, a candidate must pass explicit checks before it is accepted. The aim is to distinguish a useful change from an artifact of the tasks, search procedure, or evaluation.
rsi-metaforge-core explores meta-meta learning loops, analogical transfer, grammar-mediated expansion, and anti-cheat verification within a program synthesis runtime.
A rejected claim: a 15/15 adoption result was traced to the task generator and search engine having been co-designed. It was caught before publication. Read the standing gates and limitations; there is no separate postmortem for that result.
Reading the evidence
What does “validated” mean here?
On this site, “validated” means a result survives a stated counterfactual or held-out test linked from the project entry. It is not a claim of independent replication, peer review, or general applicability.
Check the task and scale. Bounded tasks and small search budgets impose limits on what a result can establish.
Inspect the control. Compare compute budgets and evaluation conditions before attributing a gain to the mechanism.
Read the rejected results. Retractions and limitations are part of the evidence, not footnotes to skip.
Follow the source. Repository documentation and run artifacts provide the detailed record behind the website summary.
Each run is bounded by the compute it can afford. In the latest experiment a run gives the improvement process about 50 decisions, where measuring its options would take roughly 400. Research funding would buy longer runs and more of them. Every result will be published, including the failures.