Capability is a snapshot. Plasticity is a slope.
Most agent benchmarks ask what a model can do today. But agents increasingly outlive a single task: they keep notes, write tools and build up skills that later instances inherit. Two agents that start out equally strong can then diverge, one steadily getting better with experience and the other staying where it began.
We call this property agent plasticity: how efficiently an agent turns experience into lasting improvement. To measure it, we track each agent’s held-out score as it gains experience, and divide the gain by the cost of learning.
Learning curves: chess, Go and Hex (mean)
Psat (plasticity to saturation)
Each bar rises when its model reaches saturation (★)
How a frozen model learns
The model’s weights never change. What changes is an inventory of files that each new instance inherits: Python tools, written strategies and skills, and memory notes. We call each file an artifact, and the model plus its inventory after t rounds is checkpoint t (checkpoint 0 is the bare harness). Each round follows the same loop.
- Fresh every time. Each game starts in a new context. Nothing carries over except the inventory.
- Checked, not trusted. Proposed edits must pass structural and executable checks. These establish validity, not improvement, and rejected attempts still count towards the learning cost.
- Scored on unseen games. Held-out in-distribution (ID) opponents match the training difficulty. Held-out out-of-distribution (OOD) opponents are stronger.
We run 20 rounds in chess (two difficulty profiles, Hard and Easy), 5×5 Go and 6×6 Hex, and 40 rounds in NetHack. Opponents are engines: Stockfish 16, GNU Go and a native Hex engine. In chess we compare Claude Fable 5, Claude Opus 5, GPT-5.6 Sol, Claude Opus 4.8, Gemini 3.1 Pro, GPT-5.5, Claude Opus 4.6 and GPT-5.6 Luna; Go and Hex use a six-model subset. NetHack uses the six models available through Claude Code and Codex: Claude Opus 5.5, Claude Opus 5, GPT-5.6 Sol, Claude Opus 4.8, Claude Sonnet 5 and GPT-5.6 Luna.
Measuring plasticity
Learning cost is the API spend on training games and reflection, including rejected edits; held-out evaluation is not counted. The simplest measure of plasticity is gain per dollar: how many points the held-out score S rose from checkpoint 0 to checkpoint t, divided by the learning cost Ct spent up to that checkpoint (in thousands of dollars).
But this ratio keeps shrinking once a learner plateaus, because spending continues after the gains stop. So for our headline number, Psat, we fit a smooth S-shaped curve to each learning curve and stop the clock at saturation: the first checkpoint that reaches 90% of the fitted rise. For example, GPT-5.6 Sol and Claude Fable 5 each gain about 35 points before levelling off, but Sol gets there for about $120 and Fable for about $610, so Sol’s Psat is about five times higher (298 against 57). If the rise is too small (under 10 points) or not statistically supported, Psat is 0. If no checkpoint reaches saturation, we extrapolate the fit up to three times the final cost; if saturation lies even further out, we report plasticity to date instead.
Some agents get much better. Others barely move.
Everyone gets the same interface for feedback, reflection and artifact building. The resulting curves still come apart almost immediately. Claude Fable 5, Claude Opus 5 and GPT-5.6 Sol each add 30 to 40 points on held-out ID games, averaged over chess, Go and Hex. GPT-5.5 improves slowly, and Claude Opus 4.8 and GPT-5.6 Luna barely move.
Held-out ID gain at every checkpoint (chess, Go and Hex)
Chess shows the spread most clearly. In Chess (Hard), Claude Fable 5’s held-out ID score rises from 37.5% at checkpoint 0 to an average of 73.3% over checkpoints 16–20, and Claude Opus 5’s from 25.0% to 66.4%. GPT-5.6 Sol starts near zero but gains about as much (to 36.9%): it trails in absolute score, not in improvement. Several other models stay close to where they started. Gains against stronger OOD opponents are smaller or less consistent: getting better in the training regime and transferring beyond it are different outcomes.
Score improvement with checkpoints in Chess (Hard)
Same model, different inventory
What does improvement look like on the board? Each tab pairs two games by the same model against the same Stockfish level: an early checkpoint and a later one, after the agent has rebuilt its inventory from training games. The first three pairs are held-out games that reflection never saw; the fourth shows a training-game loss and the fix it prompted. We picked these four pairs as illustrations, not typical games, so their captions give average scores. Red and orange dots on the slider mark the agent’s blunders and mistakes as rated by Stockfish, and the bar beside the board tracks Stockfish’s evaluation from the agent’s side.
The last two tabs, Live: skill 8 and Live: skill 5, are new games we played for this post, with Claude Fable 5’s checkpoint-0 and checkpoint-20 inventories as Black, under the paper’s held-out settings. At checkpoint 0 the inventory is empty: against skill 5 the model reasons through each move itself (102,000 output tokens), and against skill 8 it first writes a chess engine from scratch (67,000), which is lost when the game ends. At checkpoint 20 it loads the engine it built during training and writes only 2,000 to 7,000. Across all six games we played (both colours for checkpoint 20), checkpoint 20 won three and drew one, while checkpoint 0 drew one and forfeited the other. The six games cost $54.52 at the paper’s prices.
It happens in other games, but the ranking shifts
In 5×5 Go, Claude Fable 5’s held-out ID score rises from 20% to 80%, and its OOD score from 0% to 50%. In 6×6 Hex, GPT-5.6 Sol climbs from 40% to 77.5% on held-out ID and from 28.1% to 78.1% on OOD. But the ranking is not stable: Claude Opus 4.8 loses ground in Go while improving substantially in Hex. Self-edits can also break things: Claude Opus 5 improves substantially in Go before a tool-making error at the final checkpoint makes its score drop sharply.
NetHack is the longest-horizon test: a partially observed dungeon crawl where a single bad decision can end the game. Here the agents start from native Claude Code and Codex, and may edit CLAUDE.md/AGENTS.md, skills, Python tools and memory. At every checkpoint, each lineage plays the same ten fresh, unseen games; they are scored first, then handed to reflection. We ran 40 checkpoints.
NetHack, checkpoints 0–40
Claude Opus 5.5 shows a clear, persistent gain: its mean score grows from about 2,000 over checkpoints 0–2 to about 6,000 over checkpoints 36–40. Its learning curve saturates at checkpoint 11, giving Psat = 61.75 normalized score points per $1,000 (NetHack scores are rescaled to 0–100 for this, so the value is not comparable with the board games). Over the same windows, Claude Opus 5 and Claude Sonnet 5 rise more modestly (about 1,500 to 1,900, and 680 to 800), Claude Opus 4.8 stays roughly flat, and GPT-5.6 Luna declines (about 210 to 120). So does GPT-5.6 Sol (about 570 to 420), the most plastic model on the board games. None of these other five shows a reliable rise.
One persistent fix makes this concrete. In NetHack, praying heals you only if you are in serious trouble and have not prayed recently. At checkpoint 1, Claude Opus 5.5 prays again too soon, by typing raw keystrokes, and is killed. Reflecting on that game, it adds a rule to CLAUDE.md: pray only through its own controller (nhctl), which checks whether prayer is safe. At checkpoint 10, in another low-health emergency, it waits until the controller allows a prayer (8 of 94 hit points) and is fully healed. The rule does not always hold: later agents still sometimes bypass the check and die while praying.
The strongest agent is not the most efficient learner
Checkpoint curves compare agents after the same number of update rounds, but those rounds cost very different amounts. Plotting score against money spent (the first figure) separates two questions: how good does the agent end up? and how quickly does experience pay off?
Psat in points of held-out score per $1,000 of learning cost, averaged over chess, Go and Hex.
The three leaders gain similar amounts before levelling off, about 32–36 points. What separates them is the bill: Claude Fable 5 ends highest, but reaches its saturation point only after about $610 of learning cost, roughly five times the $120 at which GPT-5.6 Sol levels off (at a lower score). GPT-5.5 is still improving when its runs end, so its Psat cannot be established. GPT-5.6 Sol also starts below Claude Opus 4.8, yet ends above it in both held-out performance and estimated plasticity. These estimates depend on provider prices and on the saturation fit, so we read plasticity as a measure tied to this protocol, not a universal ranking.
Where the loop breaks
Curves tell us whether an agent improves, not why not. So we trace every identified mistake back through the artifact lifecycle. Did the agent have an artifact that covered this situation? Did it actually use it (read the file or run the tool)? And if it did, did the mistake happen anyway?
Artifact reuse in Chess (Hard), held-out ID games
Low reuse accompanies weak improvement. Pooled over Hard and Easy chess, GPT-5.6 Luna makes about 98% of its relevant decisions without using any artifact, Gemini 3.1 Pro about 87%, Claude Opus 4.6 about 51% and GPT-5.5 about 28%, against 2–6% for Claude Fable 5, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Sol. Across 19 model–game cells in Chess (Hard), Go and Hex, held-out ID reuse correlates with the held-out ID gain. In one GPT-5.5 run, reflection keeps telling the agent to load its saved move-picking program right away, yet the agent still spends its first steps listing and reading files. Writing down an instruction is not the same as following it.
High reuse does not explain the remaining gap. Claude Fable 5, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Sol all reuse artifacts at 94–98% of relevant chess decisions (pooled over Hard and Easy), yet reach substantially different game scores. Neither the reuse rate nor the success rate when an artifact is used explains this gap.
Strong improvers still fail when using artifacts. These models make fewer mistakes overall (about 0.16 identified failures per relevant decision, against 0.38 for the other models). But 83–99% of the mistakes they still make happen while a relevant artifact is in use. The artifact may not generalize to the position, or the agent may apply it poorly: the bottleneck shifts from finding and reusing artifacts to building ones that generalize and applying them well.
What gets carried forward
One Claude Fable 5 run writes its own chess engine, a program that looks ahead to pick moves, without outside libraries. It then patches the engine after specific mistakes (not necessarily in this order):
Every later game starts with all of these changes. These examples show how artifacts and behaviour change over time; they do not show which single change caused an improvement.
What this means for building agents
Static evaluations measure where an agent stands. For agents meant to accumulate experience, how efficiently they improve matters as much as where they start. Our results suggest three practical lessons:
- Measure the slope, not just the endpoint. The best final score and the best learning efficiency can belong to different models.
- Check that artifacts are actually used. When relevant artifacts are ignored, it may help to make them harder to skip, for example by loading or invoking them automatically.
- Then look at artifact quality. When reuse is already frequent, better procedures and better application become the natural targets. Validation should test whether an artifact helps, not only whether it runs.
Limitations. Frozen weights and fresh contexts isolate persistent adaptation, but do not separate the benefit of training feedback from the extra computation spent developing artifacts. Our cost measure excludes deployment costs, the reuse diagnostics are observational, and plasticity estimates depend on pricing and on the saturation fit. Small evaluation sets add checkpoint noise; in NetHack, changes across checkpoints mix learning with game difficulty, because each checkpoint plays new games; separately trained runs do not establish cross-task transfer. Details are in the paper.
Citation
@article{singh2026agentplasticity,
title = {Agent Plasticity: Measuring Self-Improvement Through Experience},
author = {Singh, Harman and Bakhtin, Anton and Shao, Rulin and
Synnaeve, Gabriel and Kulikov, Ilia and Fergus, Rob and
Arora, Sanjeev and Keutzer, Kurt and Weston, Jason and
Mahajan, Anuj and Goyal, Anirudh},
journal = {arXiv preprint arXiv:2610.08902},
year = {2026}
}