Harness LearningEnables Generalizable Test-Time Adaptation

Alvin Zhang*†Xuecheng Liu*Zixuan Wang*Fahim TajwarDaman AroraRuslan SalakhutdinovDaniel KhashabiYuda Song‡Andrea Zanette‡

* Equal contribution. † Project lead. ‡ Equal advising.

Learning to revise the executable program around a language model,
and applying that skill to new tasks with model weights fixed.

Johns Hopkins UniversityCarnegie Mellon University

Harness learning in 39 seconds

Paper overview ↗

Learn to revise code. Adapt to new tasks.

Learn from execution outcomes

During training, task performance rewards the proposer for revising harnesses. At test time, both models stay fixed while the proposer edits code. The example gives the answering call retrieved passages instead of summaries. Training raises mean revision score on 21 unseen reasoning families from 0.32 to 0.62. A QA proposer trained on HotpotQA reaches 0.27 mean MuSiQue exact match after ten revision rounds.

During training, revised harnesses run the frozen solver on task questions. Their performance supplies the reward for updating the proposer.

The QA code is a simplified example from the paper. The final charts show aggregate experimental results. Explore the full results ↓

Key results on unseen tasks

0.32→0.62

Better individual revisions

Mean revision score on 21 unseen Reasoning Gym families, from the base to the trained proposer.

0.15→0.27

Further gains through revision

MuSiQue exact match: an independent revision versus ten rounds, using a proposer trained on HotpotQA.

Abstract

A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver’s harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.

How harness learning works

The harness determines how a solver calls tools, processes outputs, and continues or stops execution. We train a proposer to edit this executable program, conditioned on a task description, the current harness, and an execution report.

This gives harness learning a meta-learning interpretation: the harness is the adapted object, and the proposer is the learned adaptation rule.

During training
Task performance supplies the reward for reinforcement learning. The reasoning proposer begins with supervised fine-tuning on teacher revisions; the QA proposer trains directly with RL.
At test time
The proposer applies the learned revision skill to a new task. Repeated executions provide feedback for further edits. The harness changes, while both models’ parameters remain fixed.

Results

We evaluate two settings with separately trained proposers. Reasoning Gym tests transfer across diverse synthetic task families from a minimal seed harness. Multi-hop QA tests adaptation of an existing retrieval workflow.

Better individual revisions on unseen families

On 21 unseen task families, training raises the mean single-revision score from 0.32 to 0.62. The gain appears in the mean over proposals, so it does not depend on selecting the best candidate. Under the same revision protocol, the trained 4B proposer also exceeds its 35B teacher on average, although the teacher achieves a higher oracle best-of-eight score.

Three panels compare single-step revision scores on seen tasks, unseen tasks, and the average over 21 unseen families. Training improves the mean score from 0.32 to 0.62.
Single-step revision on format-varied questions. Bars average eight proposals per family; the right panel averages 21 unseen families. Circles show oracle best-of-eight scores. View full size ↗

Over successive rounds, proposers trained on individual revisions can continue improving harnesses. On format-varied questions, Single-step RL obtains most of its five-round gain in the first revision, while SFT improves more gradually.

What do the learned revisions change?

More reliable candidates

RL reduces the fraction of failed or zero-scoring proposals in both settings. The trained proposer supplies more viable candidates to selection, which may help sustain progress across rounds.

The answering call receives retrieved passages directly through a new red path, while summaries still guide retrieval.
A learned QA revision. The new evidence path is shown in red.

Changes to computation and information flow

Reasoning harnesses commonly introduce interpreter loops, letting the solver express a computation as code and delegate its execution to the runtime.

In QA, a common revision keeps summaries to guide the next retrieval, while passing the original retrieved passages directly to the answering call. This separates the evidence needed to formulate a search from the evidence needed to answer.

These edits change how the harness organizes the solver’s work. Their transfer to unseen tasks suggests that execution outcomes can teach a reusable adaptation skill.

Repeated adaptation from single-revision training

Proposers trained on individual revisions remain useful after earlier edits change the harness and its execution reports. Training on revision sequences does not consistently improve on training on individual revisions. This leaves open whether different intermediate training examples or credit assignment across rounds could make longer revision chains more effective.

Cite this work

Download .bib
@misc{zhang2026harnesslearningenablesgeneralizable,
  title={Harness Learning Enables Generalizable Test-Time Adaptation},
  author={Alvin Zhang and Xuecheng Liu and Zixuan Wang and Fahim Tajwar and Daman Arora and Ruslan Salakhutdinov and Daniel Khashabi and Yuda Song and Andrea Zanette},
  year={2026},
  eprint={2609.35738},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2609.35738},
}