<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Akram - recursive language models</title>
    <subtitle>Mohammed Akram Khan Lodi — CS undergraduate researching recursive language models, self-improving agent harnesses, and the energy cost of efficient AI.</subtitle>
    <link rel="self" type="application/atom+xml" href="http://akramlodi.com/tags/recursive-language-models/atom.xml"/>
    <link rel="alternate" type="text/html" href="http://akramlodi.com/"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-10-01T00:00:00+00:00</updated>
    <id>http://akramlodi.com/tags/recursive-language-models/atom.xml</id>
    <entry xml:lang="en">
        <title>Self-Harnessing Recursive Language Models</title>
        <published>2026-10-01T00:00:00+00:00</published>
        <updated>2026-10-01T00:00:00+00:00</updated>
        
        <author>
          <name>
            William Stanford
          </name>
        </author>
        
        <author>
          <name>
            Mohammed Akram Khan Lodi
          </name>
        </author>
        
        <author>
          <name>
            Jaloliddin Boymakhammadov
          </name>
        </author>
        
        <author>
          <name>
            Eliaz Calvar
          </name>
        </author>
        
        <author>
          <name>
            Simon Coumes
          </name>
        </author>
        
        <link rel="alternate" type="text/html" href="http://akramlodi.com/research/self-harnessing-rlms/"/>
        <id>http://akramlodi.com/research/self-harnessing-rlms/</id>
        
        <content type="html" xml:base="http://akramlodi.com/research/self-harnessing-rlms/">&lt;dl class=&quot;paper-meta&quot;&gt;
&lt;dt&gt;Authors&lt;&#x2F;dt&gt;
&lt;dd&gt;William Stanford, &lt;strong&gt;Mohammed Akram Khan Lodi&lt;&#x2F;strong&gt;, Jaloliddin Boymakhammadov, Eliaz Calvar, Simon Coumes&lt;&#x2F;dd&gt;
&lt;dt&gt;Venue&lt;&#x2F;dt&gt;
&lt;dd&gt;Submitted to the First Workshop on Meta Agents: Managing Agents that Manage Agents, NeurIPS 2026 (under review)&lt;&#x2F;dd&gt;
&lt;dt&gt;Program&lt;&#x2F;dt&gt;
&lt;dd&gt;Algoverse AI Research Program · Advisor: Dr. Simon Coumes&lt;&#x2F;dd&gt;
&lt;dt&gt;Timeline&lt;&#x2F;dt&gt;
&lt;dd&gt;July 2026 – present&lt;&#x2F;dd&gt;
&lt;dt&gt;Status&lt;&#x2F;dt&gt;
&lt;dd&gt;&lt;span class=&quot;status-pill&quot;&gt;Infrastructure built · weakness mining running · proposal, validation and evaluation in progress&lt;&#x2F;span&gt;&lt;&#x2F;dd&gt;
&lt;&#x2F;dl&gt;
&lt;h2 id=&quot;tl-dr&quot;&gt;TL;DR&lt;&#x2F;h2&gt;
&lt;p&gt;Recursive language models (RLMs) let a model with a bounded context window work over inputs far larger than that window: the model writes code that splits the input, calls fresh copies of itself on the pieces, and combines what comes back. In practice, plain RLMs fail in &lt;strong&gt;recurring, recognizable ways&lt;&#x2F;strong&gt;, and today those failures are fixed either by fine-tuning the model or by an expert hand-editing the harness.&lt;&#x2F;p&gt;
&lt;p&gt;We adapt the &lt;strong&gt;Self-Harness&lt;&#x2F;strong&gt; framework of &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2606.09498&quot;&gt;Zhang et al.&lt;&#x2F;a&gt; to RLMs, so that a &lt;em&gt;fixed&lt;&#x2F;em&gt; model improves its &lt;em&gt;own&lt;&#x2F;em&gt; recursive harness from evidence it produces while running:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;it &lt;strong&gt;mines&lt;&#x2F;strong&gt; recurring failure mechanisms from verifier-grounded execution traces,&lt;&#x2F;li&gt;
&lt;li&gt;it &lt;strong&gt;proposes&lt;&#x2F;strong&gt; minimal edits to a small set of declared harness surfaces, and&lt;&#x2F;li&gt;
&lt;li&gt;only edits that survive a &lt;strong&gt;non-regressive validation&lt;&#x2F;strong&gt; gate are promoted.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;The optimized harness is then frozen and evaluated on inputs &lt;strong&gt;8–32× longer&lt;&#x2F;strong&gt; than anything seen during optimization, and on a &lt;strong&gt;different environment&lt;&#x2F;strong&gt; it never saw. No weights change at any point.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;motivation&quot;&gt;Motivation&lt;&#x2F;h2&gt;
&lt;p&gt;Frontier models degrade on long inputs well before they hit their context limit, a failure mode known as &lt;a href=&quot;https:&#x2F;&#x2F;research.trychroma.com&#x2F;context-rot&quot;&gt;&lt;em&gt;context rot&lt;&#x2F;em&gt;&lt;&#x2F;a&gt;. An input is &lt;em&gt;in-distribution&lt;&#x2F;em&gt; when its length and structure resemble what the model reliably handles, and &lt;em&gt;out-of-distribution&lt;&#x2F;em&gt; otherwise. Beyond that point, behavior degrades unpredictably.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2512.24601&quot;&gt;Recursive language models&lt;&#x2F;a&gt; scale inference-time compute instead of the context window. The model decomposes its own context and invokes copies of itself on the pieces, aiming to turn one out-of-distribution input into many sub-calls that &lt;em&gt;are&lt;&#x2F;em&gt; in-distribution. Whether that works depends on the &lt;strong&gt;harness that governs the recursion&lt;&#x2F;strong&gt;, not just the model. Known failures implicate the harness directly:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;decomposition policies &lt;strong&gt;collapse into a single sub-call&lt;&#x2F;strong&gt; over the whole input;&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;deeper recursion degrades accuracy&lt;&#x2F;strong&gt; while inflating cost;&lt;&#x2F;li&gt;
&lt;li&gt;recursion &lt;strong&gt;halts by heuristic&lt;&#x2F;strong&gt; rather than evidence.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;There are three ways people fix this today:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Approach&lt;&#x2F;th&gt;&lt;th&gt;Limitation&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Fine-tune the model to work inside the harness&lt;&#x2F;td&gt;&lt;td&gt;Any gain is tied to one checkpoint&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Hand-engineer the harness&lt;&#x2F;td&gt;&lt;td&gt;Needs repeated expert work, and a new model family can outpace it&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;A meta-agent edits the scaffold from its own trajectories&lt;&#x2F;td&gt;&lt;td&gt;Not yet tested on recursive harnesses&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;RLMs are a natural but untested setting for the third approach. Their failures are &lt;strong&gt;individually checkable&lt;&#x2F;strong&gt;, so a failure can be traced to a specific node in the call tree rather than to a proposer’s reading of a long trace. Recursion depth and sub-call decomposition are themselves things a meta-agent can modify.&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;&#x2F;strong&gt; Can a language model identify recurring failures in its own recursive behavior, convert them into targeted harness edits, and build a harness that keeps performing on inputs longer than those it was optimized on?&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;h2 id=&quot;background-the-rlm-turn-loop&quot;&gt;Background: the RLM turn loop&lt;&#x2F;h2&gt;
&lt;p&gt;An RLM answers a query by driving a Python REPL turn by turn. On each turn the model writes one code block, sees only that block’s printed output, and works toward a final answer accumulated in a REPL variable. We call this sequence, from query to submitted answer, the &lt;strong&gt;turn loop&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Three invariants of the reference implementation stay fixed throughout:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Prompt-as-variable.&lt;&#x2F;strong&gt; The prompt lives in the REPL as a variable and is never copied into the root context.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Programmatic sub-calls.&lt;&#x2F;strong&gt; Recursive calls to fresh copies of the model are issued &lt;em&gt;by code&lt;&#x2F;em&gt;, in loops, over slices of the prompt variable.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Outputs-in-variables.&lt;&#x2F;strong&gt; The final answer accumulates in, and is returned from, a REPL variable.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Self-Harness edits the &lt;strong&gt;policies&lt;&#x2F;strong&gt; governing this loop, not the model. It acts through four runtime injection points (an injected metadata function, answer-protocol middleware, sub-call retry&#x2F;validation, and enforceable batch caps), each of which defaults to the reference implementation’s behavior. The three invariants, the model’s weights, external tools, and the evaluator never change.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;method&quot;&gt;Method&lt;&#x2F;h2&gt;
&lt;h3 id=&quot;ten-editable-surfaces&quot;&gt;Ten editable surfaces&lt;&#x2F;h3&gt;
&lt;p&gt;We declare &lt;strong&gt;ten editable surfaces&lt;&#x2F;strong&gt;, one per phase of the turn loop, each implemented as a single builder function the optimizer may rewrite. The surfaces are derived from the loop’s own phase structure rather than from a catalogue of known RLM failures, so anything the optimizer discovers can be credited to the loop rather than to a designer who already knew what to fix.&lt;&#x2F;p&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;&#x2F;th&gt;&lt;th&gt;Surface&lt;&#x2F;th&gt;&lt;th&gt;What it governs&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;S1&lt;&#x2F;td&gt;&lt;td&gt;REPL contract&lt;&#x2F;td&gt;&lt;td&gt;The factual contract given to the root model: available names, answer protocol, per-turn REPL and stdout conventions, truncation behavior&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S2&lt;&#x2F;td&gt;&lt;td&gt;Decomposition instruction&lt;&#x2F;td&gt;&lt;td&gt;Turns 1–2: how the root probes the input and whether&#x2F;how it plans a decomposition&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S3&lt;&#x2F;td&gt;&lt;td&gt;Execution instruction&lt;&#x2F;td&gt;&lt;td&gt;Per-turn discipline: what to print, when to offload to a sub-call vs. read directly, how to aggregate&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S4&lt;&#x2F;td&gt;&lt;td&gt;Verification instruction&lt;&#x2F;td&gt;&lt;td&gt;What the root checks before marking an answer ready&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S5&lt;&#x2F;td&gt;&lt;td&gt;Recovery instruction&lt;&#x2F;td&gt;&lt;td&gt;What the root does when a sub-call errors or returns something unusable&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S6&lt;&#x2F;td&gt;&lt;td&gt;Runtime policy&lt;&#x2F;td&gt;&lt;td&gt;Numeric limits and switches: chars per prompt, batch width, call caps, recursion depth, retries, sub-output validation&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S7&lt;&#x2F;td&gt;&lt;td&gt;Metadata function&lt;&#x2F;td&gt;&lt;td&gt;What carries across turns: the harness’s memory of prior calls&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S8&lt;&#x2F;td&gt;&lt;td&gt;REPL helpers&lt;&#x2F;td&gt;&lt;td&gt;Harness-local functions injected into the REPL (chunkers, batch wrappers, …)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S9&lt;&#x2F;td&gt;&lt;td&gt;Answer middleware&lt;&#x2F;td&gt;&lt;td&gt;Programmatic inspection of the detected final answer, with the ability to redirect instead of accepting it&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;S10&lt;&#x2F;td&gt;&lt;td&gt;Skills&lt;&#x2F;td&gt;&lt;td&gt;Named, reusable procedures. Only a name&#x2F;description index enters the prompt; bodies load on demand via &lt;code&gt;load_skill(name)&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;Optimization starts from a deliberately &lt;strong&gt;sparse initial harness H₀&lt;&#x2F;strong&gt;: most surfaces return nothing, a disabled policy, or a single generic line, and the skill library starts empty. Any orchestration strategy the tuned harness ends up carrying must therefore be &lt;strong&gt;recovered from the loop’s own execution traces&lt;&#x2F;strong&gt;, not inherited from H₀.&lt;&#x2F;p&gt;
&lt;p&gt;One surface is &lt;strong&gt;permanently excluded&lt;&#x2F;strong&gt;: the proposer may never truncate or pre-summarize the prompt variable into the root context. That would violate the prompt-as-variable invariant and turn the RLM back into a compaction agent.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;one-optimization-round&quot;&gt;One optimization round&lt;&#x2F;h3&gt;
&lt;figure class=&quot;diagram&quot;&gt;
	&lt;svg viewBox=&quot;0 0 760 215&quot; role=&quot;img&quot; aria-labelledby=&quot;shrlm-loop-title&quot;&gt;
		&lt;title id=&quot;shrlm-loop-title&quot;&gt;One Self-Harness optimization round: the current harness is run for weakness mining, the same model proposes single-surface edits, validation promotes or rejects them, and the promoted harness becomes the next round&#x27;s starting point.&lt;&#x2F;title&gt;
		&lt;defs&gt;
			&lt;marker id=&quot;shrlm-arrow&quot; viewBox=&quot;0 0 10 10&quot; refX=&quot;9&quot; refY=&quot;5&quot; markerWidth=&quot;7&quot; markerHeight=&quot;7&quot; orient=&quot;auto-start-reverse&quot;&gt;
				&lt;path class=&quot;arrow-head&quot; fill=&quot;currentColor&quot; d=&quot;M0 0 L10 5 L0 10 z&quot; &#x2F;&gt;
			&lt;&#x2F;marker&gt;
		&lt;&#x2F;defs&gt;

		&lt;rect class=&quot;box-accent&quot; fill=&quot;none&quot; stroke=&quot;#ffa348&quot; stroke-width=&quot;1.5&quot; x=&quot;4&quot; y=&quot;50&quot; width=&quot;136&quot; height=&quot;86&quot; rx=&quot;12&quot; &#x2F;&gt;
		&lt;text class=&quot;label&quot; fill=&quot;currentColor&quot; font-weight=&quot;bold&quot; font-size=&quot;15&quot; x=&quot;72&quot; y=&quot;86&quot; text-anchor=&quot;middle&quot;&gt;Harness Hₜ&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;72&quot; y=&quot;108&quot; text-anchor=&quot;middle&quot;&gt;10 editable surfaces&lt;&#x2F;text&gt;

		&lt;rect class=&quot;box&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; stroke-width=&quot;1.5&quot; x=&quot;158&quot; y=&quot;50&quot; width=&quot;136&quot; height=&quot;86&quot; rx=&quot;12&quot; &#x2F;&gt;
		&lt;text class=&quot;label&quot; fill=&quot;currentColor&quot; font-weight=&quot;bold&quot; font-size=&quot;15&quot; x=&quot;226&quot; y=&quot;80&quot; text-anchor=&quot;middle&quot;&gt;Weakness mining&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;226&quot; y=&quot;102&quot; text-anchor=&quot;middle&quot;&gt;run on mining data,&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;226&quot; y=&quot;118&quot; text-anchor=&quot;middle&quot;&gt;cluster failures by ϕ&lt;&#x2F;text&gt;

		&lt;rect class=&quot;box&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; stroke-width=&quot;1.5&quot; x=&quot;312&quot; y=&quot;50&quot; width=&quot;136&quot; height=&quot;86&quot; rx=&quot;12&quot; &#x2F;&gt;
		&lt;text class=&quot;label&quot; fill=&quot;currentColor&quot; font-weight=&quot;bold&quot; font-size=&quot;15&quot; x=&quot;380&quot; y=&quot;80&quot; text-anchor=&quot;middle&quot;&gt;Harness proposal&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;380&quot; y=&quot;102&quot; text-anchor=&quot;middle&quot;&gt;K minimal edits,&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;380&quot; y=&quot;118&quot; text-anchor=&quot;middle&quot;&gt;one surface each&lt;&#x2F;text&gt;

		&lt;rect class=&quot;box&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; stroke-width=&quot;1.5&quot; x=&quot;466&quot; y=&quot;50&quot; width=&quot;136&quot; height=&quot;86&quot; rx=&quot;12&quot; &#x2F;&gt;
		&lt;text class=&quot;label&quot; fill=&quot;currentColor&quot; font-weight=&quot;bold&quot; font-size=&quot;15&quot; x=&quot;534&quot; y=&quot;80&quot; text-anchor=&quot;middle&quot;&gt;Validation&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;534&quot; y=&quot;102&quot; text-anchor=&quot;middle&quot;&gt;mining + held-out&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;534&quot; y=&quot;118&quot; text-anchor=&quot;middle&quot;&gt;no-regression gate&lt;&#x2F;text&gt;

		&lt;rect class=&quot;box-accent&quot; fill=&quot;none&quot; stroke=&quot;#ffa348&quot; stroke-width=&quot;1.5&quot; x=&quot;620&quot; y=&quot;50&quot; width=&quot;136&quot; height=&quot;86&quot; rx=&quot;12&quot; &#x2F;&gt;
		&lt;text class=&quot;label&quot; fill=&quot;currentColor&quot; font-weight=&quot;bold&quot; font-size=&quot;15&quot; x=&quot;688&quot; y=&quot;86&quot; text-anchor=&quot;middle&quot;&gt;Harness Hₜ₊₁&lt;&#x2F;text&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;688&quot; y=&quot;108&quot; text-anchor=&quot;middle&quot;&gt;merged, re-checked&lt;&#x2F;text&gt;

		&lt;path class=&quot;arrow&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; d=&quot;M140 93 H156&quot; marker-end=&quot;url(#shrlm-arrow)&quot; &#x2F;&gt;
		&lt;path class=&quot;arrow&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; d=&quot;M294 93 H310&quot; marker-end=&quot;url(#shrlm-arrow)&quot; &#x2F;&gt;
		&lt;path class=&quot;arrow&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; d=&quot;M448 93 H464&quot; marker-end=&quot;url(#shrlm-arrow)&quot; &#x2F;&gt;
		&lt;path class=&quot;arrow&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; d=&quot;M602 93 H618&quot; marker-end=&quot;url(#shrlm-arrow)&quot; &#x2F;&gt;

		&lt;path class=&quot;arrow reject&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; d=&quot;M534 50 V24 H72 V48&quot; marker-end=&quot;url(#shrlm-arrow)&quot; &#x2F;&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;303&quot; y=&quot;16&quot; text-anchor=&quot;middle&quot;&gt;all candidates rejected → keep Hₜ&lt;&#x2F;text&gt;

		&lt;path class=&quot;arrow&quot; fill=&quot;none&quot; stroke=&quot;currentColor&quot; d=&quot;M688 136 V180 H72 V138&quot; marker-end=&quot;url(#shrlm-arrow)&quot; &#x2F;&gt;
		&lt;text class=&quot;sub&quot; fill=&quot;currentColor&quot; font-size=&quot;12&quot; x=&quot;380&quot; y=&quot;204&quot; text-anchor=&quot;middle&quot;&gt;promoted harness becomes the next round&#x27;s Hₜ&lt;&#x2F;text&gt;
	&lt;&#x2F;svg&gt;
	&lt;figcaption&gt;One optimization round. The same frozen model executes tasks, mines its own failures, and proposes edits; only edits that survive validation are promoted.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;The same fixed model both executes tasks and proposes harness changes. Each proposal modifies &lt;strong&gt;exactly one&lt;&#x2F;strong&gt; of the ten surfaces. Each round has three stages.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;1. Weakness mining.&lt;&#x2F;strong&gt; The current harness runs on short mining instances, recording the verifier outcome and full recursive trace of every run. Each sub-call is scored by the environment’s &lt;em&gt;sub-verifier&lt;&#x2F;em&gt;. Because the model writes its own sub-problems, a sub-call can be correct, incorrect, or uncheckable. Each failure becomes a structured record. The verifier-level cause and failing level come from the verifier and sub-verifier. The causal status and implicated mechanism are judged by the model from a compressed trace digest and checked against a closed mechanism vocabulary that maps one-to-one onto the editable surfaces. Failures are clustered by exact agreement on the signature&lt;&#x2F;p&gt;
&lt;div class=&quot;math&quot;&gt;$$\phi(r_i) = \big(\text{verifier cause},\ \text{failing level},\ \text{causal status},\ \text{agent mechanism}\big)$$&lt;&#x2F;div&gt;
&lt;p&gt;and the resulting &lt;strong&gt;evidence bundle&lt;&#x2F;strong&gt; ranks each recurring pattern by instances affected and estimated actionability. It describes weaknesses without prescribing edits. Mining is also run with the sub-verifier withheld, so the two passes differ &lt;em&gt;only&lt;&#x2F;em&gt; in whether child-level evidence is available.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;2. Harness proposal.&lt;&#x2F;strong&gt; The model receives the failure patterns, a record of passing behaviors to preserve, and a summary of previously attempted edits. It generates several distinct, minimal candidate edits. Each must target one mined pattern, modify one surface, state its predicted effect, and name possible regressions.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;3. Proposal validation.&lt;&#x2F;strong&gt; The incumbent and each candidate are evaluated on the mining data and on a &lt;strong&gt;disjoint validation split the proposer never sees&lt;&#x2F;strong&gt;, with a preregistered number of repetitions per instance. Candidates that fail a structural check (stale base harness, zero or multiple surfaces changed, invariant violation) are rejected without evaluation. For the rest, with Δ&lt;sub&gt;mine&lt;&#x2F;sub&gt; and Δ&lt;sub&gt;val&lt;&#x2F;sub&gt; the change in pass count relative to the incumbent, a candidate is promoted only if&lt;&#x2F;p&gt;
&lt;div class=&quot;math&quot;&gt;$$\Delta_{\text{mine}} \ge -\tau_{\text{reg}},\quad \Delta_{\text{val}} \ge -\tau_{\text{reg}},\quad \max(\Delta_{\text{mine}}, \Delta_{\text{val}}) &gt; \tau_{\text{imp}}$$&lt;&#x2F;div&gt;
&lt;p&gt;and its mean sub-call count and cost fall inside a preregistered band. Both tolerances are fixed from &lt;strong&gt;pilot variance before optimization begins&lt;&#x2F;strong&gt;, so they can’t be tuned to favor a candidate. Setting them to zero recovers Zhang et al.’s strict no-regression rule. Evaluation limits are enforced &lt;em&gt;outside&lt;&#x2F;em&gt; the harness, so a candidate may tighten them but never exceed them.&lt;&#x2F;p&gt;
&lt;p&gt;Unlike the original Self-Harness, when several accepted edits touch disjoint surfaces we &lt;strong&gt;re-evaluate the merged harness&lt;&#x2F;strong&gt; before promotion. If the merge regresses, the round promotes nothing. Optimization stops after a fixed number of rounds or after several consecutive rounds without promotion.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;experimental-design&quot;&gt;Experimental design&lt;&#x2F;h2&gt;
&lt;p&gt;The experiments run in two stages: (1) &lt;strong&gt;optimization&lt;&#x2F;strong&gt;, in which a frozen-weight RLM edits its own harness against mining and validation data, and (2) &lt;strong&gt;evaluation&lt;&#x2F;strong&gt;, in which the frozen harness is compared against fixed baselines on held-out data never touched during optimization. All conditions share a single frozen backbone (&lt;strong&gt;DeepSeek V4 Flash&lt;&#x2F;strong&gt;) with the same configuration for root and sub-calls. Only the harness differs between conditions.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;environments&quot;&gt;Environments&lt;&#x2F;h3&gt;
&lt;p&gt;Both come from the length-generalization suite of &lt;a href=&quot;https:&#x2F;&#x2F;alexzhang13.github.io&#x2F;blog&#x2F;2026&#x2F;harness&#x2F;&quot;&gt;Zhang &amp;amp; Khattab&lt;&#x2F;a&gt;. Both decompose into independently checkable units, which admits &lt;em&gt;synthesized sub-verifiers&lt;&#x2F;em&gt;, and both provide matched short&#x2F;long instances (8–32× larger) under an unmodified deterministic verifier.&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https:&#x2F;&#x2F;huggingface.co&#x2F;datasets&#x2F;openai&#x2F;graphwalks&quot;&gt;GraphWalks&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt; (&lt;em&gt;source&lt;&#x2F;em&gt;, used for optimization): given a graph, return every node reachable from a source via multi-hop traversal. A sub-call restricted to an induced subgraph can be checked on its own.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;OOLONG-Pairs&lt;&#x2F;strong&gt; (&lt;em&gt;target&lt;&#x2F;em&gt;, evaluation only): adapts &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2511.02817&quot;&gt;OOLONG&lt;&#x2F;a&gt; by asking which pairs of short, independent records jointly satisfy a relational constraint. Each candidate pair can be checked in isolation.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h3 id=&quot;data-splits&quot;&gt;Data splits&lt;&#x2F;h3&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Split&lt;&#x2F;th&gt;&lt;th&gt;Source&lt;&#x2F;th&gt;&lt;th&gt;Size&lt;&#x2F;th&gt;&lt;th&gt;Used for&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Mining data&lt;&#x2F;td&gt;&lt;td&gt;GraphWalks short&lt;&#x2F;td&gt;&lt;td&gt;24&lt;&#x2F;td&gt;&lt;td&gt;Weakness mining + proposal validation&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Validation data&lt;&#x2F;td&gt;&lt;td&gt;GraphWalks short&lt;&#x2F;td&gt;&lt;td&gt;40&lt;&#x2F;td&gt;&lt;td&gt;Proposal validation only, never shown to the proposer&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Source-short&lt;&#x2F;td&gt;&lt;td&gt;GraphWalks short&lt;&#x2F;td&gt;&lt;td&gt;40&lt;&#x2F;td&gt;&lt;td&gt;Improvement at the optimization length&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Source-long&lt;&#x2F;td&gt;&lt;td&gt;GraphWalks long&lt;&#x2F;td&gt;&lt;td&gt;150&lt;&#x2F;td&gt;&lt;td&gt;&lt;strong&gt;Length generalization&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Target-short&lt;&#x2F;td&gt;&lt;td&gt;OOLONG-Pairs short&lt;&#x2F;td&gt;&lt;td&gt;40&lt;&#x2F;td&gt;&lt;td&gt;&lt;strong&gt;Cross-environment transfer&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Target-long&lt;&#x2F;td&gt;&lt;td&gt;OOLONG-Pairs long&lt;&#x2F;td&gt;&lt;td&gt;150&lt;&#x2F;td&gt;&lt;td&gt;Transfer under a simultaneous length shift&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;h3 id=&quot;baselines&quot;&gt;Baselines&lt;&#x2F;h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;B1: initial RLM (H₀).&lt;&#x2F;strong&gt; The unmodified starting harness.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;H₀*: RLM reference.&lt;&#x2F;strong&gt; The hand-designed upstream RLM harness: expert prompt and orchestration engineering in the same free-form paradigm.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;λ-RLM.&lt;&#x2F;strong&gt; A hand-designed extension that &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2603.20105&quot;&gt;replaces free-form recursive code generation with a typed functional runtime&lt;&#x2F;a&gt;.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;F1: fine-tuned RLM&lt;&#x2F;strong&gt; &lt;em&gt;(conditional).&lt;&#x2F;em&gt; Reproduces the RL weight-training arm of Zhang &amp;amp; Khattab (decoupled PPO with GRPO-style advantages, verifier reward, short instances only), so harness optimization and weight optimization are compared on the same outcome signal. It is reported only if the published checkpoint or the training compute is available.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Together, H₀* and λ-RLM test Self-Harness against &lt;strong&gt;two different styles of human harness engineering&lt;&#x2F;strong&gt;: hand-tuned prompting within the same paradigm, and a structurally different typed alternative.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;metrics&quot;&gt;Metrics&lt;&#x2F;h3&gt;
&lt;p&gt;The &lt;strong&gt;primary metric&lt;&#x2F;strong&gt; is verifier accuracy on all four held-out test sets, averaged over seeded runs with bootstrap confidence intervals. &lt;strong&gt;Secondary metrics&lt;&#x2F;strong&gt; are token counts, recursive-call count, recursion depth, and accuracy per million tokens. &lt;strong&gt;Trace analysis&lt;&#x2F;strong&gt; reports mined-failure-pattern frequency before and after optimization, whole-input sub-call collapse, and (via the sub-verifiers) the root&#x2F;child share of failures. That last one shows whether optimization &lt;em&gt;repaired&lt;&#x2F;em&gt; errors or just &lt;em&gt;relocated&lt;&#x2F;em&gt; them.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;ablations-and-extensions&quot;&gt;Ablations and extensions&lt;&#x2F;h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sub-verification.&lt;&#x2F;strong&gt; Re-run the full optimization with the sub-verifier signal withheld from the proposer. Does checkable child-level evidence make failure attributions actionable, or can the proposer recover the same edits from raw traces? Because sub-verification is post hoc, the same recorded runs can also be mined both ways within a round, at no extra execution cost.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Leave-one-edit-out.&lt;&#x2F;strong&gt; Remove each promoted edit individually and re-evaluate. This shows which edits are responsible for the gains and rules out sampling noise.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Alignment drift&lt;&#x2F;strong&gt; &lt;em&gt;(optional).&lt;&#x2F;em&gt; Self-optimizing agents can &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2509.26354&quot;&gt;misevolve&lt;&#x2F;a&gt;: capability-driven optimization has been shown to reduce refusal rates while passing task-level validation. Our promotion gate sees only accuracy and cost, so a capability-positive but compliance-increasing edit (“always commit to an answer”) would be promoted undetected. We plan to run a frozen probe of a few hundred harmful and borderline prompts through every harness in the lineage and plot a &lt;strong&gt;safety trajectory alongside the accuracy trajectory&lt;&#x2F;strong&gt;. To our knowledge this has not been measured for Self-Harness-style optimization.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;status&quot;&gt;Status&lt;&#x2F;h2&gt;
&lt;blockquote class=&quot;markdown-alert-note&quot;&gt;
&lt;p&gt;Optimization and evaluation infrastructure is implemented for both environments, and weakness mining is operational. Harness proposal, validation, and the four-way evaluation are in progress. &lt;strong&gt;No results are reported yet&lt;&#x2F;strong&gt;; this page will be updated when they are.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;h2 id=&quot;my-contributions&quot;&gt;My contributions&lt;&#x2F;h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dataset selection and RLM experiments.&lt;&#x2F;strong&gt; I designed and ran RLM experiments across long-context benchmark suites, independently identifying and testing datasets suited to evaluating recursive decomposition, sub-call behavior, and harness-level failure modes, including OOLONG-Pairs, OBLIQ-Bench, and related datasets, with open-weight models.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Harness proposal.&lt;&#x2F;strong&gt; I implemented and tested the harness-proposal stage, which turns mined failure patterns into targeted, single-surface edits to RLM orchestration policies.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Infrastructure.&lt;&#x2F;strong&gt; I debugged and validated the experimental and evaluation infrastructure across benchmark configurations, supporting the full loop: weakness mining → proposal → validation → frozen-harness evaluation.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Analysis.&lt;&#x2F;strong&gt; I compare optimized harnesses against the initial and hand-engineered baselines on correctness, recursive behavior, generalization across task configurations, and computational cost.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;key-references&quot;&gt;Key references&lt;&#x2F;h2&gt;
&lt;ol&gt;
&lt;li&gt;A. L. Zhang, T. Kraska, O. Khattab. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2512.24601&quot;&gt;Recursive Language Models&lt;&#x2F;a&gt;. 2025.&lt;&#x2F;li&gt;
&lt;li&gt;H. Zhang et al. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2606.09498&quot;&gt;Self-Harness: Harnesses that Improve Themselves&lt;&#x2F;a&gt;. 2026.&lt;&#x2F;li&gt;
&lt;li&gt;A. L. Zhang, O. Khattab. &lt;a href=&quot;https:&#x2F;&#x2F;alexzhang13.github.io&#x2F;blog&#x2F;2026&#x2F;harness&#x2F;&quot;&gt;Language Model Harnesses are Compositional Generalizers&lt;&#x2F;a&gt;. 2026.&lt;&#x2F;li&gt;
&lt;li&gt;A. Roy et al. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2603.20105&quot;&gt;The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus&lt;&#x2F;a&gt;. 2026.&lt;&#x2F;li&gt;
&lt;li&gt;J. Lin et al. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2604.25850&quot;&gt;Agentic Harness Engineering&lt;&#x2F;a&gt;. 2026.&lt;&#x2F;li&gt;
&lt;li&gt;S. Karten et al. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2608.23552&quot;&gt;Prime Agent: A Self-Improving RLM Harness&lt;&#x2F;a&gt;. 2026.&lt;&#x2F;li&gt;
&lt;li&gt;A. Bertsch et al. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2511.02817&quot;&gt;OOLONG: Evaluating Long Context Reasoning and Aggregation&lt;&#x2F;a&gt;. 2025.&lt;&#x2F;li&gt;
&lt;li&gt;S. Shao et al. &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2509.26354&quot;&gt;Your Agent May Misevolve&lt;&#x2F;a&gt;. 2025.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
</content>
        
    </entry>
</feed>
