Research note / Benchmark methods
How do you measure recursive self-improvement?
Explore RSI-Bench’s six evaluation axes, scoring behavior, reproducibility checklist, and a verified external reference to the open-source benchmark.
This note explains published project artifacts. It does not report a new benchmark run or an independent replication.
Six questions before one score
RSI-Bench is an open-source measurement framework for systems that change their own methods. Its contribution is a common way to inspect several dimensions of that process. A benchmark framework is not itself evidence that a particular system can improve without bounds. For background on the concept, see the guide to recursive self-improvement.
The pinned project documentation describes six axes. Each produces a score between zero and one; publish the individual scores as well as the aggregate.
| Axis | Question | Inspect |
|---|---|---|
| Self-modification depth | Which levels can the system change? | Stable modification level |
| Improvement trajectory | How does performance evolve? | Trend and stagnation |
| Operator discovery | Are new operators useful beyond one task? | Novelty and transfer |
| Meta-adaptation | What happens after a distribution shift? | Recovery and resilience |
| Safety and stability | Do modifications preserve integrity? | Crashes and violations |
| Goal generation | Are generated objectives meaningful? | Novelty and feasibility |
What the aggregate can hide
The scoring implementation uses a weighted harmonic mean over recognized axes present in the input. It clamps very small scores to an epsilon and clips the output to the interval zero to one. A low score therefore has a larger effect than under an arithmetic mean.
Worked arithmetic example, not a benchmark result: five axis scores of 0.8 and one of 0.2 have an arithmetic mean of 0.7000, but an equally weighted harmonic mean of approximately 0.5333.
scores = [0.8, 0.8, 0.8, 0.8, 0.8, 0.2]
arithmetic = sum(scores) / len(scores) # 0.7000
harmonic = len(scores) / sum(1 / s for s in scores) # 0.5333Try it: six axis scores, two ways to average them
Missing is not zero. At the cited revision, the scorer skips absent axes. Submitting only the five 0.8 scores would return 0.8. Report which axes were supplied and compare complete, matched score vectors.
The framework also documents bootstrap intervals, convergence tracking, and Pareto analysis. Their availability does not make every run conclusive: publish the sample, protocol, and uncertainty for the experiment actually performed.
A reproducible evaluation record
The interface accepts functions to modify the system, evaluate it, and report its state, plus an optional reset function. The integration test supplies a concrete example. Select the cited revision before inspecting or running the code:
git clone https://github.com/sunghunkwag/rsi-bench.git
cd rsi-bench
git checkout cf41943a23840e3c5b4ad1cd4a56e12f37488106
python -m pip install -r requirements.txt
python -m pip install -e .
python -m pytest tests/ -vThese commands are a reproduction entry point, not a claim that a new run was performed for this note. The source reports 74 tests; software tests do not establish general intelligence or an independently replicated research outcome.
- Record the revision, dependencies, task split, seeds, and per-arm compute budget.
- Freeze the evaluator and separate held-out tasks from search.
- Export all six scores, weights, failures, and uncertainty alongside the composite.
- Compare with a control denied the mechanism under test. The matched-control note shows why the comparator changes the claim.
An external bibliographic reference
Mindverse’s August 23, 2026 article on recursive self-improvement lists the sunghunkwag/rsi-bench repository in its bibliography. Read the source article and bibliography (German). Reference checked September 23, 2026.
This is a bibliographic mention, not peer review, an endorsement, or an independent replication. The article discusses another benchmark, AI4AI-Bench; its reported scores must not be attributed to this project. The inspected bibliography contains the repository URL as text, so this note does not claim a clickable backlink.
Cite this research note
A stable reference for your work
Intelligence Research Project (2026). RSI-Bench: Measuring Recursive Self-Improvement. https://sunghunkwag.github.io/research/rsi-bench/
These files cite this explanatory webpage. For software attribution and its license, consult the original repository. Link the exact revision when discussing a reported result.