Research note / Experimental results

Does experience at improving carry over?

Two pre-registered experiments from gated-self-improvement, each run once on untouched seeds after its design was frozen, and the earlier headline that the project’s own audit retracted.

This note explains published project artifacts. It does not report a new run or an independent replication.

The question: does an improver get better at improving?

The gated-self-improvement project studies bounded improvement loops in deterministic, LLM-free Python systems. Its latest experiment asks whether what an improver learned while improving on earlier problems helps it improve on a new one. The RSI test checklist explains why the comparator matters as much as the number.

A run is a sequence of ten independent improvement problems. Each problem has its own task families, tasks and starting solver, and the improver gets five rounds on it. In each round it generates candidate edits to its solver, tests them against the incumbent on probe tasks, and adopts the winner or rolls back. Alongside, it keeps an improvement memory: a record of which kinds of edits paid off.

Three arms share everything else, including the tasks, the random streams and a cap of 4,976,000 program executions per run:

  • MEMORY_CARRY keeps its improvement memory from one problem to the next.
  • NO_CARRY runs the identical learner, but its memory is wiped at every problem boundary. Through the first problem the two arms are bit-identical.
  • FROZEN never learns; it uses a fixed, uninformed improvement policy.

Every arm is scored on a sealed final holdout: task families that no improver saw while it ran.

Tested once, after the design was frozen

  1. Development used seeds 3001–3100 only. Every development battery is recorded in an append-only, hash-chained ledger.
  2. A development check found that the first FROZEN control was weak. A stronger fixed prior was adopted for every arm, which cut the development margin over FROZEN from +6.9 to +2.1 tasks.
  3. A go/no-go criterion with five required conditions was written to the ledger before any of the reserved seeds 3201–3300 ran. The look returned GO, narrowly on its primary condition (p = 0.046).
  4. The pre-registration was committed, and a one-shot freeze hashed it together with the hyperparameters, every source file, the evaluator, the task library and the seed manifests.
  5. The confirmatory battery ran once on seeds 7001–7300, which had never been run: 5 arms × 300 seeds = 1,500 units, with 0 missing, 0 abandoned and 0 restarts.

In the public history, the freeze commit precedes the first confirmatory unit, and the ledger records no confirmatory seed before the freeze. The project’s verifier reports that the hash chain and every frozen hash match.

Four pre-registered hypotheses, 300 untouched seeds

Confirmatory results from the frozen report. Tasks are final-holdout tasks solved per run, about 64 of 240. Adjusted p-values use fixed-sequence gatekeeping, with Holm among H2–H4.
HypothesisMean95% CIWins / ties / lossesAdjusted pVerdict
H1 (primary) · memory carried − memory wiped+1.52[+0.55, +2.48]174 / 16 / 110p = 0.0011Supported
H2 · memory carried − frozen+1.78[+0.85, +2.73]176 / 14 / 110p = 0.0006Supported
H3 · memory wiped − frozen+0.25[−0.67, +1.15]132 / 15 / 153p = 0.60Null
H4 · growth of the carried advantage, problems 6–10 minus 1–5+0.05[−0.84, +0.95]136 / 17 / 147p = 0.60Null

MEMORY_CARRY solved 65.89 final-holdout tasks per run, NO_CARRY 64.36 and FROZEN 64.11. The primary effect, +1.52, is about 2.4% of the final metric. The means, win counts and intervals in this table were recomputed from the raw unit log for this note, and they match the frozen report.

Three readings follow. Carried experience improves later self-improvement on problems the improver has never seen (H1, H2). Learning inside a single five-round problem is not enough: the memory-wiped learner does not beat the frozen policy (H3). And the carried advantage appears once one problem of experience exists, then holds steady rather than growing (H4).

Download the reported results as CSV · Read the full report · Inspect the frozen report

What the carried memory does

Two further arms split the memory’s channels. MEMORY_GENONLY uses carried memory to decide which edits to generate and how to rank them, with round-level options left at their defaults. MEMORY_RANKONLY uses carried memory only to reorder a fixed generator’s candidates. These contrasts were pre-registered as exploratory and are descriptive, without multiplicity correction.

Exploratory contrasts on the final holdout, same 300 seeds.
StepEffectReading
Frozen → learner with memory wiped+0.25 (n.s.)Five rounds give too little experience to beat an uninformed policy.
Memory wiped → memory carried, used for generation+1.27 (p = 0.006)Carried experience about which kinds of edits to generate is the main channel.
Frozen → memory carried, used only for ranking+1.17 (p = 0.011)Carried experience helps even when it only reorders a fixed pool.
Ranking only → full process control+0.60 (n.s.)Control beyond ranking is not established on the final metric.
Generation → plus round-level options+0.26 (n.s.)Compute, probe, exploration and adoption options add nothing detectable.

What the memory learned shows in the adopted edits. It flattens the search prior less often: exploration edits were 16.4% of FROZEN’s adoptions and 6.1% of MEMORY_CARRY’s. It prunes less. And it favours mined macro bundles, 50.5% of MEMORY_CARRY’s adoptions against 35.9% of FROZEN’s. The report reads this as knowledge about types of edits, which needs more observations than one problem supplies. That is why only carried memory can use it, and why it saturates within a problem or two.

Earlier results in the same project

v2: recursion against a compute-matched single round

Also pre-registered and run once on 300 fresh seeds. Five recursive rounds beat one round given five times the compute, with the same tasks, the same attempt streams and the same 610,000-execution cap, by +0.40 of 24 external-holdout tasks (95% CI +0.10 to +0.70; Holm p = 0.0098). The effect is small. Its second hypothesis, that adapting the improvement policy within a run beats a frozen policy, was null (−0.08; CI −0.35 to +0.19). The v3 result is consistent with that null: learning inside one problem does not help, while carrying it across problems does. Read the v2 report.

v1: the retracted headline

This note used to feature a repaired five-round chain that beat a single round at five times the compute by +1.55 and +1.47 tasks per seed. An audit found two problems. Held-out evaluation seeded each arm’s search with the arm’s own name, so the arms were scored under different randomness. And the single-round control saw 5 distinct source tasks where the chain saw 25. Rerun on the same seeds with shared random streams, the repaired chain fell below the untrained baseline (−0.70, p = 0.0005), and a single round over the same 25 tasks beat it (−0.59, p = 0.002). The headline is retracted. The original report keeps its numbers under an audit correction, and the failure log lists the retraction.

What the evidence does not establish

  • A small effect in a synthetic domain. The gain is about +1.5 of 64 final-holdout tasks per run. It transfers to unseen task families of the same stack substrate, not to other domains.
  • No growth. A memory that kept improving the improver would show a widening gap. This one reaches its useful level after about one problem.
  • Process control beyond ranking is not established on the final metric (+0.60, n.s.). From the third problem on, the round-level options depart from their defaults in about half of all rounds, but they are data-starved: a run offers about 50 decisions, where measuring their effect would take roughly 400.
  • Controls. The first control was too weak and was replaced by a stronger one before the look. Other fixed policies that were not tested might be stronger still.
  • A narrow go/no-go. The look passed at p = 0.046 on its primary condition. The confirmatory battery is what carries the claim.
  • Scope. These are project-reported runs, not independent replications. They concern bounded tasks and do not establish open-ended recursive improvement or general intelligence.

For complementary measurement questions, see the RSI-Bench methods note.

Inspect the exact research artifact

git clone https://github.com/sunghunkwag/gated-self-improvement.git
cd gated-self-improvement
git checkout 8e69e9f3149db5310e7571b6f5ef5a2a422d69ae
python3 -m unittest discover -s tests -v
(cd src && python3 -m rsi_v3 verify)

The first command runs the project’s anti-cheat tests. The second checks the ledger’s hash chain and the frozen hashes, and counts missing or abandoned confirmatory units. The report documents how to recompute the frozen report and how to re-run any confirmatory unit from scratch.

The project is a Python standard-library implementation without API calls. This webpage publishes an explanatory note and a transcription of reported results; no new experiment was run to produce it.

Cite this research note

A stable reference for your work

Intelligence Research Project (2026). Gated Self-Improvement: Pre-Registered Results. https://sunghunkwag.github.io/research/gated-self-improvement/

These files cite this explanatory webpage. For software attribution and its license, consult the original repository. Link the exact revision when discussing a reported result.

Browse all research notes →

Support independent research

Help fund the next experiment.

Each run is bounded by the compute it can afford. In the latest experiment a run gives the improvement process about 50 decisions, where measuring its options would take roughly 400. Research funding would buy longer runs and more of them. Every result will be published, including the failures.