Research note / Benchmark methods

How do you measure recursive self-improvement?

Explore RSI-Bench’s six evaluation axes, scoring behavior, reproducibility checklist, and a verified external reference to the open-source benchmark.

This note explains published project artifacts. It does not report a new benchmark run or an independent replication.

Six questions before one score

RSI-Bench is an open-source measurement framework for systems that change their own methods. Its contribution is a common way to inspect several dimensions of that process. A benchmark framework is not itself evidence that a particular system can improve without bounds. For background on the concept, see the guide to recursive self-improvement.

The pinned project documentation describes six axes. Each produces a score between zero and one; publish the individual scores as well as the aggregate.

Measurement categories described by RSI-Bench, not achieved results.
AxisQuestionInspect
Self-modification depthWhich levels can the system change?Stable modification level
Improvement trajectoryHow does performance evolve?Trend and stagnation
Operator discoveryAre new operators useful beyond one task?Novelty and transfer
Meta-adaptationWhat happens after a distribution shift?Recovery and resilience
Safety and stabilityDo modifications preserve integrity?Crashes and violations
Goal generationAre generated objectives meaningful?Novelty and feasibility

What the aggregate can hide

The scoring implementation uses a weighted harmonic mean over recognized axes present in the input. It clamps very small scores to an epsilon and clips the output to the interval zero to one. A low score therefore has a larger effect than under an arithmetic mean.

Worked arithmetic example, not a benchmark result: five axis scores of 0.8 and one of 0.2 have an arithmetic mean of 0.7000, but an equally weighted harmonic mean of approximately 0.5333.

scores = [0.8, 0.8, 0.8, 0.8, 0.8, 0.2]
arithmetic = sum(scores) / len(scores)  # 0.7000
harmonic = len(scores) / sum(1 / s for s in scores)  # 0.5333

Missing is not zero. At the cited revision, the scorer skips absent axes. Submitting only the five 0.8 scores would return 0.8. Report which axes were supplied and compare complete, matched score vectors.

The framework also documents bootstrap intervals, convergence tracking, and Pareto analysis. Their availability does not make every run conclusive: publish the sample, protocol, and uncertainty for the experiment actually performed.

A reproducible evaluation record

The interface accepts functions to modify the system, evaluate it, and report its state, plus an optional reset function. The integration test supplies a concrete example. Select the cited revision before inspecting or running the code:

git clone https://github.com/sunghunkwag/rsi-bench.git
cd rsi-bench
git checkout cf41943a23840e3c5b4ad1cd4a56e12f37488106
python -m pip install -r requirements.txt
python -m pip install -e .
python -m pytest tests/ -v

These commands are a reproduction entry point, not a claim that a new run was performed for this note. The source reports 74 tests; software tests do not establish general intelligence or an independently replicated research outcome.

  1. Record the revision, dependencies, task split, seeds, and per-arm compute budget.
  2. Freeze the evaluator and separate held-out tasks from search.
  3. Export all six scores, weights, failures, and uncertainty alongside the composite.
  4. Compare with a control denied the mechanism under test. The matched-control note shows why the comparator changes the claim.

An external bibliographic reference

Mindverse’s August 23, 2026 article on recursive self-improvement lists the sunghunkwag/rsi-bench repository in its bibliography. Read the source article and bibliography (German). Reference checked September 23, 2026.

This is a bibliographic mention, not peer review, an endorsement, or an independent replication. The article discusses another benchmark, AI4AI-Bench; its reported scores must not be attributed to this project. The inspected bibliography contains the repository URL as text, so this note does not claim a clickable backlink.

Cite this research note

A stable reference for your work

Intelligence Research Project (2026). RSI-Bench: Measuring Recursive Self-Improvement. https://sunghunkwag.github.io/research/rsi-bench/

These files cite this explanatory webpage. For software attribution and its license, consult the original repository. Link the exact revision when discussing a reported result.

Browse all research notes →

Support independent research

Help fund the next experiment.

Each run is bounded by the compute it can afford. In the latest experiment a run gives the improvement process about 50 decisions, where measuring its options would take roughly 400. Research funding would buy longer runs and more of them. Every result will be published, including the failures.