Self-Harnessing Recursive Language Models

By William Stanford, Mohammed Akram Khan Lodi, Jaloliddin Boymakhammadov, Eliaz Calvar and Simon Coumes • 13 minutes read •


Authors
William Stanford, Mohammed Akram Khan Lodi, Jaloliddin Boymakhammadov, Eliaz Calvar, Simon Coumes
Venue
Submitted to the First Workshop on Meta Agents: Managing Agents that Manage Agents, NeurIPS 2026 (under review)
Program
Algoverse AI Research Program · Advisor: Dr. Simon Coumes
Timeline
July 2026 – present
Status
Infrastructure built · weakness mining running · proposal, validation and evaluation in progress

TL;DR

Recursive language models (RLMs) let a model with a bounded context window work over inputs far larger than that window: the model writes code that splits the input, calls fresh copies of itself on the pieces, and combines what comes back. In practice, plain RLMs fail in recurring, recognizable ways, and today those failures are fixed either by fine-tuning the model or by an expert hand-editing the harness.

We adapt the Self-Harness framework of Zhang et al. to RLMs, so that a fixed model improves its own recursive harness from evidence it produces while running:

  1. it mines recurring failure mechanisms from verifier-grounded execution traces,
  2. it proposes minimal edits to a small set of declared harness surfaces, and
  3. only edits that survive a non-regressive validation gate are promoted.

The optimized harness is then frozen and evaluated on inputs 8–32× longer than anything seen during optimization, and on a different environment it never saw. No weights change at any point.

Motivation

Frontier models degrade on long inputs well before they hit their context limit, a failure mode known as context rot. An input is in-distribution when its length and structure resemble what the model reliably handles, and out-of-distribution otherwise. Beyond that point, behavior degrades unpredictably.

Recursive language models scale inference-time compute instead of the context window. The model decomposes its own context and invokes copies of itself on the pieces, aiming to turn one out-of-distribution input into many sub-calls that are in-distribution. Whether that works depends on the harness that governs the recursion, not just the model. Known failures implicate the harness directly:

There are three ways people fix this today:

ApproachLimitation
Fine-tune the model to work inside the harnessAny gain is tied to one checkpoint
Hand-engineer the harnessNeeds repeated expert work, and a new model family can outpace it
A meta-agent edits the scaffold from its own trajectoriesNot yet tested on recursive harnesses

RLMs are a natural but untested setting for the third approach. Their failures are individually checkable, so a failure can be traced to a specific node in the call tree rather than to a proposer’s reading of a long trace. Recursion depth and sub-call decomposition are themselves things a meta-agent can modify.

Research question. Can a language model identify recurring failures in its own recursive behavior, convert them into targeted harness edits, and build a harness that keeps performing on inputs longer than those it was optimized on?

Background: the RLM turn loop

An RLM answers a query by driving a Python REPL turn by turn. On each turn the model writes one code block, sees only that block’s printed output, and works toward a final answer accumulated in a REPL variable. We call this sequence, from query to submitted answer, the turn loop.

Three invariants of the reference implementation stay fixed throughout:

Self-Harness edits the policies governing this loop, not the model. It acts through four runtime injection points (an injected metadata function, answer-protocol middleware, sub-call retry/validation, and enforceable batch caps), each of which defaults to the reference implementation’s behavior. The three invariants, the model’s weights, external tools, and the evaluator never change.

Method

Ten editable surfaces

We declare ten editable surfaces, one per phase of the turn loop, each implemented as a single builder function the optimizer may rewrite. The surfaces are derived from the loop’s own phase structure rather than from a catalogue of known RLM failures, so anything the optimizer discovers can be credited to the loop rather than to a designer who already knew what to fix.

SurfaceWhat it governs
S1REPL contractThe factual contract given to the root model: available names, answer protocol, per-turn REPL and stdout conventions, truncation behavior
S2Decomposition instructionTurns 1–2: how the root probes the input and whether/how it plans a decomposition
S3Execution instructionPer-turn discipline: what to print, when to offload to a sub-call vs. read directly, how to aggregate
S4Verification instructionWhat the root checks before marking an answer ready
S5Recovery instructionWhat the root does when a sub-call errors or returns something unusable
S6Runtime policyNumeric limits and switches: chars per prompt, batch width, call caps, recursion depth, retries, sub-output validation
S7Metadata functionWhat carries across turns: the harness’s memory of prior calls
S8REPL helpersHarness-local functions injected into the REPL (chunkers, batch wrappers, …)
S9Answer middlewareProgrammatic inspection of the detected final answer, with the ability to redirect instead of accepting it
S10SkillsNamed, reusable procedures. Only a name/description index enters the prompt; bodies load on demand via load_skill(name)

Optimization starts from a deliberately sparse initial harness H₀: most surfaces return nothing, a disabled policy, or a single generic line, and the skill library starts empty. Any orchestration strategy the tuned harness ends up carrying must therefore be recovered from the loop’s own execution traces, not inherited from H₀.

One surface is permanently excluded: the proposer may never truncate or pre-summarize the prompt variable into the root context. That would violate the prompt-as-variable invariant and turn the RLM back into a compaction agent.

One optimization round

One Self-Harness optimization round: the current harness is run for weakness mining, the same model proposes single-surface edits, validation promotes or rejects them, and the promoted harness becomes the next round's starting point.Harness Hₜ10 editable surfacesWeakness miningrun on mining data,cluster failures by ϕHarness proposalK minimal edits,one surface eachValidationmining + held-outno-regression gateHarness Hₜ₊₁merged, re-checkedall candidates rejected → keep Hₜpromoted harness becomes the next round's Hₜ
One optimization round. The same frozen model executes tasks, mines its own failures, and proposes edits; only edits that survive validation are promoted.

The same fixed model both executes tasks and proposes harness changes. Each proposal modifies exactly one of the ten surfaces. Each round has three stages.

1. Weakness mining. The current harness runs on short mining instances, recording the verifier outcome and full recursive trace of every run. Each sub-call is scored by the environment’s sub-verifier. Because the model writes its own sub-problems, a sub-call can be correct, incorrect, or uncheckable. Each failure becomes a structured record. The verifier-level cause and failing level come from the verifier and sub-verifier. The causal status and implicated mechanism are judged by the model from a compressed trace digest and checked against a closed mechanism vocabulary that maps one-to-one onto the editable surfaces. Failures are clustered by exact agreement on the signature

$$\phi(r_i) = \big(\text{verifier cause},\ \text{failing level},\ \text{causal status},\ \text{agent mechanism}\big)$$

and the resulting evidence bundle ranks each recurring pattern by instances affected and estimated actionability. It describes weaknesses without prescribing edits. Mining is also run with the sub-verifier withheld, so the two passes differ only in whether child-level evidence is available.

2. Harness proposal. The model receives the failure patterns, a record of passing behaviors to preserve, and a summary of previously attempted edits. It generates several distinct, minimal candidate edits. Each must target one mined pattern, modify one surface, state its predicted effect, and name possible regressions.

3. Proposal validation. The incumbent and each candidate are evaluated on the mining data and on a disjoint validation split the proposer never sees, with a preregistered number of repetitions per instance. Candidates that fail a structural check (stale base harness, zero or multiple surfaces changed, invariant violation) are rejected without evaluation. For the rest, with Δmine and Δval the change in pass count relative to the incumbent, a candidate is promoted only if

$$\Delta_{\text{mine}} \ge -\tau_{\text{reg}},\quad \Delta_{\text{val}} \ge -\tau_{\text{reg}},\quad \max(\Delta_{\text{mine}}, \Delta_{\text{val}}) > \tau_{\text{imp}}$$

and its mean sub-call count and cost fall inside a preregistered band. Both tolerances are fixed from pilot variance before optimization begins, so they can’t be tuned to favor a candidate. Setting them to zero recovers Zhang et al.’s strict no-regression rule. Evaluation limits are enforced outside the harness, so a candidate may tighten them but never exceed them.

Unlike the original Self-Harness, when several accepted edits touch disjoint surfaces we re-evaluate the merged harness before promotion. If the merge regresses, the round promotes nothing. Optimization stops after a fixed number of rounds or after several consecutive rounds without promotion.

Experimental design

The experiments run in two stages: (1) optimization, in which a frozen-weight RLM edits its own harness against mining and validation data, and (2) evaluation, in which the frozen harness is compared against fixed baselines on held-out data never touched during optimization. All conditions share a single frozen backbone (DeepSeek V4 Flash) with the same configuration for root and sub-calls. Only the harness differs between conditions.

Environments

Both come from the length-generalization suite of Zhang & Khattab. Both decompose into independently checkable units, which admits synthesized sub-verifiers, and both provide matched short/long instances (8–32× larger) under an unmodified deterministic verifier.

Data splits

SplitSourceSizeUsed for
Mining dataGraphWalks short24Weakness mining + proposal validation
Validation dataGraphWalks short40Proposal validation only, never shown to the proposer
Source-shortGraphWalks short40Improvement at the optimization length
Source-longGraphWalks long150Length generalization
Target-shortOOLONG-Pairs short40Cross-environment transfer
Target-longOOLONG-Pairs long150Transfer under a simultaneous length shift

Baselines

Together, H₀* and λ-RLM test Self-Harness against two different styles of human harness engineering: hand-tuned prompting within the same paradigm, and a structurally different typed alternative.

Metrics

The primary metric is verifier accuracy on all four held-out test sets, averaged over seeded runs with bootstrap confidence intervals. Secondary metrics are token counts, recursive-call count, recursion depth, and accuracy per million tokens. Trace analysis reports mined-failure-pattern frequency before and after optimization, whole-input sub-call collapse, and (via the sub-verifiers) the root/child share of failures. That last one shows whether optimization repaired errors or just relocated them.

Ablations and extensions

Status

Optimization and evaluation infrastructure is implemented for both environments, and weakness mining is operational. Harness proposal, validation, and the four-way evaluation are in progress. No results are reported yet; this page will be updated when they are.

My contributions

Key references

  1. A. L. Zhang, T. Kraska, O. Khattab. Recursive Language Models. 2025.
  2. H. Zhang et al. Self-Harness: Harnesses that Improve Themselves. 2026.
  3. A. L. Zhang, O. Khattab. Language Model Harnesses are Compositional Generalizers. 2026.
  4. A. Roy et al. The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus. 2026.
  5. J. Lin et al. Agentic Harness Engineering. 2026.
  6. S. Karten et al. Prime Agent: A Self-Improving RLM Harness. 2026.
  7. A. Bertsch et al. OOLONG: Evaluating Long Context Reasoning and Aggregation. 2025.
  8. S. Shao et al. Your Agent May Misevolve. 2025.