Agent Plasticity

Measuring self-improvement through experience

Harman Singh1,2   Anton Bakhtin2   Rulin Shao2,3   Gabriel Synnaeve2   Ilia Kulikov2   Rob Fergus2   Sanjeev Arora4   Kurt Keutzer1   Jason Weston2   Anuj Mahajan2,†   Anirudh Goyal2,†
1UC Berkeley   2Meta Superintelligence Labs   3University of Washington   4Princeton University   †Joint supervision
TL;DR We freeze a model’s weights and let it improve only through persistent artifacts: tools, skills and notes that it writes and revises after each round of training games. We then score it on held-out games that reflection never sees. Agent plasticity measures how efficiently this works: held-out gain per dollar of learning cost. Given identical opportunities, frontier models follow very different learning curves, and the model that ends strongest (Claude Fable 5) is not the one that learns most per dollar (GPT-5.6 Sol). Models that improve little often skip the artifacts they wrote. The strongest improvers use their artifacts at almost every relevant chess move, yet most of their remaining mistakes happen while those artifacts are in use.

Capability is a snapshot. Plasticity is a slope.

Most agent benchmarks ask what a model can do today. But agents increasingly outlive a single task: they keep notes, write tools and build up skills that later instances inherit. Two agents that start out equally strong can then diverge, one steadily getting better with experience and the other staying where it began.

We call this property agent plasticity: how efficiently an agent turns experience into lasting improvement. To measure it, we track each agent’s held-out score as it gains experience, and divide the gain by the cost of learning.

Learning curves: chess, Go and Hex (mean)

Psat (plasticity to saturation)

Each bar rises when its model reaches saturation (★)

Frontier agents differ widely in how efficiently experience helps. Left: held-out in-distribution (ID) score against cumulative learning cost for the six models run in all three games, averaged equally over chess (Hard and Easy), Go and Hex. Faint markers are checkpoints, lines are fitted curves, and ★ marks where each curve levels off (for Claude Opus 4.8, extrapolated just beyond its last checkpoint). Right: Psat, the points gained up to that point per $1,000 (defined below). GPT-5.5 is still climbing when its runs end (☆; the lighter bar shows gain per $1,000 so far). GPT-5.6 Luna shows no reliable rise.
Setup

How a frozen model learns

The model’s weights never change. What changes is an inventory of files that each new instance inherits: Python tools, written strategies and skills, and memory notes. We call each file an artifact, and the model plus its inventory after t rounds is checkpoint t (checkpoint 0 is the bare harness). Each round follows the same loop.

the next checkpoint inherits the validated inventory Inventory It tools · skills memory notes Act training games, fresh context each Feedback outcomes, traces, move quality Reflect a fresh instance of the same model Validate structural and executable checks Held-out games (ID and OOD) score St recorded at every checkpoint Held-out trajectories and scores are never shown to reflection.
  • Fresh every time. Each game starts in a new context. Nothing carries over except the inventory.
  • Checked, not trusted. Proposed edits must pass structural and executable checks. These establish validity, not improvement, and rejected attempts still count towards the learning cost.
  • Scored on unseen games. Held-out in-distribution (ID) opponents match the training difficulty. Held-out out-of-distribution (OOD) opponents are stronger.

We run 20 rounds in chess (two difficulty profiles, Hard and Easy), 5×5 Go and 6×6 Hex, and 40 rounds in NetHack. Opponents are engines: Stockfish 16, GNU Go and a native Hex engine. In chess we compare Claude Fable 5, Claude Opus 5, GPT-5.6 Sol, Claude Opus 4.8, Gemini 3.1 Pro, GPT-5.5, Claude Opus 4.6 and GPT-5.6 Luna; Go and Hex use a six-model subset. NetHack uses the six models available through Claude Code and Codex: Claude Opus 5.5, Claude Opus 5, GPT-5.6 Sol, Claude Opus 4.8, Claude Sonnet 5 and GPT-5.6 Luna.

Measuring plasticity

Learning cost is the API spend on training games and reflection, including rejected edits; held-out evaluation is not counted. The simplest measure of plasticity is gain per dollar: how many points the held-out score S rose from checkpoint 0 to checkpoint t, divided by the learning cost Ct spent up to that checkpoint (in thousands of dollars).

P  =  St − S0Ct / $1,000

But this ratio keeps shrinking once a learner plateaus, because spending continues after the gains stop. So for our headline number, Psat, we fit a smooth S-shaped curve to each learning curve and stop the clock at saturation: the first checkpoint that reaches 90% of the fitted rise. For example, GPT-5.6 Sol and Claude Fable 5 each gain about 35 points before levelling off, but Sol gets there for about $120 and Fable for about $610, so Sol’s Psat is about five times higher (298 against 57). If the rise is too small (under 10 points) or not statistically supported, Psat is 0. If no checkpoint reaches saturation, we extrapolate the fit up to three times the final cost; if saturation lies even further out, we report plasticity to date instead.

Finding 1

Some agents get much better. Others barely move.

Everyone gets the same interface for feedback, reflection and artifact building. The resulting curves still come apart almost immediately. Claude Fable 5, Claude Opus 5 and GPT-5.6 Sol each add 30 to 40 points on held-out ID games, averaged over chess, Go and Hex. GPT-5.5 improves slowly, and Claude Opus 4.8 and GPT-5.6 Luna barely move.

Held-out ID gain at every checkpoint (chess, Go and Hex)

Held-out ID gain ΔS relative to checkpoint 0, averaged over chess (Hard and Easy), Go and Hex. Lines are trailing five-checkpoint means; faint markers are single checkpoints.

Chess shows the spread most clearly. In Chess (Hard), Claude Fable 5’s held-out ID score rises from 37.5% at checkpoint 0 to an average of 73.3% over checkpoints 16–20, and Claude Opus 5’s from 25.0% to 66.4%. GPT-5.6 Sol starts near zero but gains about as much (to 36.9%): it trails in absolute score, not in improvement. Several other models stay close to where they started. Gains against stronger OOD opponents are smaller or less consistent: getting better in the training regime and transferring beyond it are different outcomes.

Score improvement with checkpoints in Chess (Hard)

Scores are 100(W + 0.5D)/N against Stockfish 16 at 20,000 nodes per move: skills 1–8 in training, 3–8 for held-out ID and 12 for held-out OOD. Unfinished games count as losses. Shaded bands show ±1 s.d. across three runs; Claude Fable 5, Claude Opus 4.6 and GPT-5.6 Luna have one run each. Models that stay near their starting score are drawn lighter.
Replays

Same model, different inventory

What does improvement look like on the board? Each tab pairs two games by the same model against the same Stockfish level: an early checkpoint and a later one, after the agent has rebuilt its inventory from training games. The first three pairs are held-out games that reflection never saw; the fourth shows a training-game loss and the fix it prompted. We picked these four pairs as illustrations, not typical games, so their captions give average scores. Red and orange dots on the slider mark the agent’s blunders and mistakes as rated by Stockfish, and the bar beside the board tracks Stockfish’s evaluation from the agent’s side.

The last two tabs, Live: skill 8 and Live: skill 5, are new games we played for this post, with Claude Fable 5’s checkpoint-0 and checkpoint-20 inventories as Black, under the paper’s held-out settings. At checkpoint 0 the inventory is empty: against skill 5 the model reasons through each move itself (102,000 output tokens), and against skill 8 it first writes a chess engine from scratch (67,000), which is lost when the game ends. At checkpoint 20 it loads the engine it built during training and writes only 2,000 to 7,000. Across all six games we played (both colours for checkpoint 20), checkpoint 20 won three and drew one, while checkpoint 0 drew one and forfeited the other. The six games cost $54.52 at the paper’s prices.

Beyond chess

It happens in other games, but the ranking shifts

In 5×5 Go, Claude Fable 5’s held-out ID score rises from 20% to 80%, and its OOD score from 0% to 50%. In 6×6 Hex, GPT-5.6 Sol climbs from 40% to 77.5% on held-out ID and from 28.1% to 78.1% on OOD. But the ranking is not stable: Claude Opus 4.8 loses ground in Go while improving substantially in Hex. Self-edits can also break things: Claude Opus 5 improves substantially in Go before a tool-making error at the final checkpoint makes its score drop sharply.

NetHack is the longest-horizon test: a partially observed dungeon crawl where a single bad decision can end the game. Here the agents start from native Claude Code and Codex, and may edit CLAUDE.md/AGENTS.md, skills, Python tools and memory. At every checkpoint, each lineage plays the same ten fresh, unseen games; they are scored first, then handed to reflection. We ran 40 checkpoints.

NetHack, checkpoints 0–40

Only Claude Opus 5.5 improves reliably in NetHack. Median score (log scale) and mean deepest dungeon level over the ten games of each checkpoint. Lines are smoothed (Gaussian, σ = 3 checkpoints; log scale for score); faint markers are single checkpoints.

Claude Opus 5.5 shows a clear, persistent gain: its mean score grows from about 2,000 over checkpoints 0–2 to about 6,000 over checkpoints 36–40. Its learning curve saturates at checkpoint 11, giving Psat = 61.75 normalized score points per $1,000 (NetHack scores are rescaled to 0–100 for this, so the value is not comparable with the board games). Over the same windows, Claude Opus 5 and Claude Sonnet 5 rise more modestly (about 1,500 to 1,900, and 680 to 800), Claude Opus 4.8 stays roughly flat, and GPT-5.6 Luna declines (about 210 to 120). So does GPT-5.6 Sol (about 570 to 420), the most plastic model on the board games. None of these other five shows a reliable rise.

One persistent fix makes this concrete. In NetHack, praying heals you only if you are in serious trouble and have not prayed recently. At checkpoint 1, Claude Opus 5.5 prays again too soon, by typing raw keystrokes, and is killed. Reflecting on that game, it adds a rule to CLAUDE.md: pray only through its own controller (nhctl), which checks whether prayer is safe. At checkpoint 10, in another low-health emergency, it waits until the controller allows a prayer (8 of 94 hit points) and is fully healed. The rule does not always hold: later agents still sometimes bypass the check and die while praying.

Finding 1Persistent self-improvement is model- and task-dependent. The same interface for feedback, reflection and artifact construction produces sharply different learning curves in all four games, and transfer to stronger opponents is not uniform. This motivates evaluating plasticity across a distribution of agentic tasks, rather than treating it as a task-independent model property.
Finding 2

The strongest agent is not the most efficient learner

Checkpoint curves compare agents after the same number of update rounds, but those rounds cost very different amounts. Plotting score against money spent (the first figure) separates two questions: how good does the agent end up? and how quickly does experience pay off?

57Psat
Claude Fable 5: highest fitted endpoint, saturating at about $610
233Psat
Claude Opus 5, saturating at about $140
298Psat
GPT-5.6 Sol: highest Psat, saturating at about $120
14Psat
Claude Opus 4.8: saturation extrapolated beyond its runs

Psat in points of held-out score per $1,000 of learning cost, averaged over chess, Go and Hex.

The three leaders gain similar amounts before levelling off, about 32–36 points. What separates them is the bill: Claude Fable 5 ends highest, but reaches its saturation point only after about $610 of learning cost, roughly five times the $120 at which GPT-5.6 Sol levels off (at a lower score). GPT-5.5 is still improving when its runs end, so its Psat cannot be established. GPT-5.6 Sol also starts below Claude Opus 4.8, yet ends above it in both held-out performance and estimated plasticity. These estimates depend on provider prices and on the saturation fit, so we read plasticity as a measure tied to this protocol, not a universal ranking.

Finding 2Endpoint capability and acquisition efficiency answer different questions. Claude Fable 5 reaches the highest performance on the combined-game curve, while GPT-5.6 Sol has the largest estimated plasticity. GPT-5.6 Sol starts below Claude Opus 4.8 yet ultimately exceeds it in both held-out performance and estimated plasticity.
Finding 3

Where the loop breaks

Curves tell us whether an agent improves, not why not. So we trace every identified mistake back through the artifact lifecycle. Did the agent have an artifact that covered this situation? Did it actually use it (read the file or run the tool)? And if it did, did the mistake happen anyway?

Artifact reuse in Chess (Hard), held-out ID games

Low reuse accompanies weak improvement. Pooled over Hard and Easy chess, GPT-5.6 Luna makes about 98% of its relevant decisions without using any artifact, Gemini 3.1 Pro about 87%, Claude Opus 4.6 about 51% and GPT-5.5 about 28%, against 2–6% for Claude Fable 5, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Sol. Across 19 model–game cells in Chess (Hard), Go and Hex, held-out ID reuse correlates with the held-out ID gain. In one GPT-5.5 run, reflection keeps telling the agent to load its saved move-picking program right away, yet the agent still spends its first steps listing and reading files. Writing down an instruction is not the same as following it.

High reuse does not explain the remaining gap. Claude Fable 5, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Sol all reuse artifacts at 94–98% of relevant chess decisions (pooled over Hard and Easy), yet reach substantially different game scores. Neither the reuse rate nor the success rate when an artifact is used explains this gap.

Strong improvers still fail when using artifacts. These models make fewer mistakes overall (about 0.16 identified failures per relevant decision, against 0.38 for the other models). But 83–99% of the mistakes they still make happen while a relevant artifact is in use. The artifact may not generalize to the position, or the agent may apply it poorly: the bottleneck shifts from finding and reusing artifacts to building ones that generalize and applying them well.

What gets carried forward

One Claude Fable 5 run writes its own chess engine, a program that looks ahead to pick moves, without outside libraries. It then patches the engine after specific mistakes (not necessarily in this order):

After its king is left under attackadds a search for moves that get the king out of check.
After throwing away winning positionschanges how it handles repeated positions, since repetition can end the game in a draw.
After failing to win an endgame while aheadgives the engine more search time in endgames.
After a checkpoint-6 game is scored 0 at the 300-ply limitmakes the engine value draws more as the 300-ply limit nears, since an unfinished game scores 0 but a draw scores half a point (the “300-ply lesson” replay above).
To stop its tools timing outwrites a helper that plays several moves per call under a fixed total search budget.

Every later game starts with all of these changes. These examples show how artifacts and behaviour change over time; they do not show which single change caused an improvement.

Finding 3Artifact reuse is associated with improvement, but is not a proxy for capability. High reuse may still leave significant failures, hinting at problems with artifact quality, generality, or use. These diagnostics are observational, not causal.

What this means for building agents

Static evaluations measure where an agent stands. For agents meant to accumulate experience, how efficiently they improve matters as much as where they start. Our results suggest three practical lessons:

  • Measure the slope, not just the endpoint. The best final score and the best learning efficiency can belong to different models.
  • Check that artifacts are actually used. When relevant artifacts are ignored, it may help to make them harder to skip, for example by loading or invoking them automatically.
  • Then look at artifact quality. When reuse is already frequent, better procedures and better application become the natural targets. Validation should test whether an artifact helps, not only whether it runs.
NextEvaluate and hill-climb plasticity on a distribution of useful agentic tasks: building agents that are not only more capable, but better at becoming more capable.

Limitations. Frozen weights and fresh contexts isolate persistent adaptation, but do not separate the benefit of training feedback from the extra computation spent developing artifacts. Our cost measure excludes deployment costs, the reuse diagnostics are observational, and plasticity estimates depend on pricing and on the saturation fit. Small evaluation sets add checkpoint noise; in NetHack, changes across checkpoints mix learning with game difficulty, because each checkpoint plays new games; separately trained runs do not establish cross-task transfer. Details are in the paper.

Citation

@article{singh2026agentplasticity,
  title     = {Agent Plasticity: Measuring Self-Improvement Through Experience},
  author    = {Singh, Harman and Bakhtin, Anton and Shao, Rulin and
               Synnaeve, Gabriel and Kulikov, Ilia and Fergus, Rob and
               Arora, Sanjeev and Keutzer, Kurt and Weston, Jason and
               Mahajan, Anuj and Goyal, Anirudh},
  journal   = {arXiv preprint arXiv:2610.08902},
  year      = {2026}
}