Harness Engineering · Episode 03

Three Small Changes, Dramatic Outcomes

Add evidence. Bound outputs. Narrow choices. Verify outcomes.

Arc · The patterns Episode · 03 of 08 Next · Episode 04 — Five Paradoxes Every PM Must Hold
After this you will know
  • RetryWhy another attempt helps only when something changes.
  • SchemaWhat a shape contract can prove, and what it cannot.
  • ToolsHow to reduce tool confusion without removing useful capability.
  • LearningHow one failed run becomes a test that protects the next release.

01Where Episode 02 left us

At 9:10 on Monday, the reconciliation team opens a spreadsheet comparing four models.

One column is green. The production failure list is unchanged.

The agent still repeats malformed calls, returns fields the downstream system cannot use, and chooses the wrong operation when tool names overlap. In the most serious failure, it processes 847 of 2,347 eligible invoices and stops with 1,500 still waiting.

Episode 01 taught the team to find the first broken contract. Episode 02 showed where that contract lives: Identity, Memory Policy, Orchestration, Interception, or Observability and Evals.

The team can now draw the machine. The next question is practical: what can it change this sprint without replacing the model?

02The sprint that changed no model

On Monday, the team freezes the model.

Validation errors will return to the next attempt. Outputs must follow an agreed shape. The model will see only the tools needed for the current step. The workflow will verify the ledger before it accepts completion.

By Friday, the lesson is not that four tricks worked. It is why they worked.

Each change removed a decision the model had been forced to improvise.

The title says three changes and the body teaches four patterns. That is deliberate. The four patterns are the controls the team ships. Section 10 reduces them to three deeper design moves: map the rules, preserve the handoff, and separate proposal from authority and proof.

Reliability improves when the system replaces a vague choice with a specific signal, boundary, or check.

03The whole pattern in one table

Step What happens What the harness decides
Prepare Gather current state and relevant tools What the model may see and choose
Propose The model returns an answer or tool request Whether the proposal has an acceptable shape
Act A tool performs the approved operation Whether the action is permitted
Verify The system checks the real effect Whether the step moved the task forward
Continue The run advances, retries, escalates, or stops Whether external completion has passed
Learn The trace is stored and judged Whether this failure becomes a regression test

The model proposes. Reliability comes from the decisions before and after that proposal.

Figure 01 · Concept
Four ambiguity reducers
VAGUE CHOICE SPECIFIC SIGNAL 1 · RETRY “Try again.” Return the exact violation Field, observed value, rule, correction scope 2 · OUTPUT Any text the model can express One versioned shape Required fields, typed values, agreed labels 3 · TOOLS Forty names, several overlapping Only this step’s tools Exact names, stated boundaries, hidden specialists 4 · COMPLETION The model says it is done External evidence passes Coverage, exceptions, quality, authority Add evidence. Bound the shape. Narrow the choices. Verify the outcome. Each door removes a decision the model kept improvising.
Read it as Add evidence. Bound the shape. Narrow the choices. Verify the outcome. Each door removes a decision the model kept improvising.

For business, this means defining the accepted outcome. For product, it means deciding where ambiguity is useful and where it is dangerous. Engineering implements the constraints. Security enforces authority. Operations keeps the trace and the recovery path.

04Pattern 1: retry with evidence

A naive retry says, “Try again.”

The model receives the same task, context, and uncertainty. A different answer may appear, but the second attempt has no better reason to succeed.

A useful retry explains what failed.

Weak feedback Useful feedback
“Validation failed” amount_cents must be a whole number
“Try again” You returned "$42.50"; return 4250
“Wrong format” Change this field only; preserve the other values

The useful version supplies four things: the field, observed value, expected rule, and correction scope.

A retry must change at least one condition:

If none changes, the retry is repetition.

Three failures are often blurred together:

Failure What it means Invoice example What repairs it
Syntax The output breaks an agreed structure Invalid JSON, missing field, unsupported status Provider-enforced structure and boundary validation
Semantic The shape is valid, but the business value is wrong The amount is a whole number, but belongs to another invoice Business rules, source reconciliation, tests, or an eval
Goal Individual outputs are valid, but the job is unfinished 847 invoices pass while 1,500 are never processed External completion tied to the real workflow

One validator cannot do all three jobs.

PM requirementA retry ticket should state the failure it can repair, the new evidence supplied, the maximum attempts, the progress signal, the escalation path, and the recovery metric.

The useful metric is not retry count. It is the share of failed first attempts that recover without producing a wrong business outcome.

05Pattern 2: constrain the output space

Teams often treat structured output as a developer convenience. Its product value is a smaller failure space.

Free text lets the model return anything it can express. A schema limits the response to fields and values the next system knows how to handle.

Consider the invoice output:

Field Weak contract Stronger contract
Vendor Any text Non-empty text
Amount Any text Whole number in cents, zero or more
Due date Any text YYYY-MM-DD
Status Any text Paid, pending, overdue, or disputed
Extra fields Allowed silently Rejected unless added to the versioned contract

An instruction such as “return JSON” may still produce:

The response looks sensible to a person. The next system may fail or silently discard information.

Strict output proves Strict output does not prove
Required fields exist The amount belongs to the right invoice
Values use supported types The status matches the ledger
Finite business states use agreed labels The source data is current
Unagreed fields are rejected The whole job is finished
A schema checks shape. An eval checks meaning. A completion rule checks the job.

OpenAI's Structured Outputs documentation distinguishes JSON mode, which produces valid JSON, from strict schema adherence, which constrains supported outputs to a supplied schema.[1] Your application must still handle refusals, empty results, timeouts, unsupported schema features, and semantically wrong values.

Treat the schema as an API

A schema is a contract with every downstream consumer, not only the model. When it changes:

  1. Version the new shape.
  2. Identify affected consumers.
  3. Test old and new cases.
  4. Plan compatibility or migration.
  5. Monitor empty, refused, and semantically invalid outputs.

A field added casually today becomes a silent mismatch three systems later.

PM requirementA structured-output ticket should name the business consumer, required fields, allowed values, semantic checks after parsing, the experience when no valid result exists, the migration rule, and primary and counter-metrics.

Track schema validity and semantic correctness separately. A perfect shape can carry a wrong answer.

06Pattern 3: narrow the decision surface

Give an agent forty tools and it must distinguish forty operations before every action.

Some tools are irrelevant. Some overlap. Others use broad names such as query_data or update_record. The model must infer what each name means and whether the consequence is read, draft, or execute.

The goal is not always fewer total capabilities. It is fewer visible choices for this step.

Make the jobs explicit
Vague tool Better tool Decision removed
query_data get_invoice_by_id Which dataset and query shape?
run_analysis list_overdue_invoices Which analysis and output?
update_record draft_payment_adjustment Draft or execute? Which record?
refund draft_refund and issue_refund Proposal or irreversible action?

A precise name is part of the control surface.

Three operations improve the catalogue:

  1. Delete tools with no distinct job.
  2. Rename vague tools to exact actions.
  3. Reveal specialist tools only when the task needs them.

The third is progressive disclosure. The system keeps broad capability but presents a small, relevant set now. A reconciliation skill may expose invoice lookup, payment lookup, exception drafting, and completion checks. A supplier-onboarding skill exposes a different set.

Progressive disclosure creates a less obvious trade-off. Prompt caching rewards a stable request prefix, while changing the visible tool list can invalidate that prefix. CacheRouter, a preprint submitted on 24 August 2026, proposes separating tool discovery from the main model's fixed tool prefix. In its prototype, tested on 55 functional queries and one 30-turn dialogue, token-level cache hit rates reached 90.99% and 95.2%.[7]

Those figures come from the authors' prototype, not an independent production benchmark. The transferable design question is narrower: can you reduce the visible decision surface without rebuilding the entire prompt on every turn?

Do not cut by frequency alone

A rare escalation tool may be essential. Review each tool on four dimensions:

Dimension Question
Usage How often is it selected?
Contribution Does it help complete the task?
Overlap Does another tool perform the same job?
Consequence What happens if the model selects it incorrectly?

The decision may be keep, rename, tighten, hide behind a skill, require approval, merge, or remove. The goal is not the shortest list. It is the clearest set that covers the workflow.

Near-neighbour tools fail at their boundary. Descriptions should therefore say when not to use them:

Use get_invoice_by_id when an exact invoice ID is known. Do not use it for vendor search or status lists.
PM requirementA tool audit should produce a named job, side-effect class, negative boundary, approval policy, selection metric, and owner for every tool.

Track correct-tool selection and steps per validated completion. Fewer calls help only when completion remains correct.

07Pattern 4: verify before exit

The invoice failure in Episode 01 did not involve malformed output or the wrong tool. The run stopped before the workflow ended.

Replace model confidence with an external checklist.

Completion check Required evidence
Coverage Processed count equals 2,347 eligible source records
Exceptions Every unmatched record has an assigned state
Quality Reconciliation checks pass
Authority No prohibited action occurred

The model may propose completion. The harness checks the business state.

Output feedback says, “This field is invalid.” Completion feedback says, “This outcome is unfinished.” The first repairs a response. The second keeps the job open.

Anthropic's work on long-running agents uses feature inventories, progress files, startup scripts, and version history so a fresh session can reconstruct what remains.[3]

The same pattern now appears in evaluation and product tooling. Braintrust's Harbor integration runs each task in an isolated container and records the verifier that inspects the container after the agent stops.[6] Cursor's /goal, released on 19 August 2026, holds a long-lived objective until it is fully complete.[8]

The products differ. The rule is the same: done is a state of the world, not a sentence.

Continuation needs limits

A completion gate can create an infinite loop. Define:

“Continue until done” is not a policy unless both done and stop are defined.

08The measured receipt

LangChain held the model fixed and changed the system prompt, tools, and middleware, its term for hooks around model and tool calls.

On Terminal Bench 2.0, an 89-task coding benchmark, its reported score rose from 52.8% to 66.5%, a gain of 13.7 percentage points. The work included trace-based error analysis, self-verification, pre-completion checks, environment context, and loop detection.[2]

The result does not prove that every harness change helps every workflow. It supports a narrower claim: changing the control system can materially change measured behaviour while the model stays fixed.

A second receipt concerns cost. In its 26 August 2026 newsletter, LangChain reported that Deep Agents v0.7 simplified the base harness and used 65% fewer base input tokens at comparable performance.[5]

Both results are company-reported measurements on LangChain's systems. Neither predicts your workflow. Together, they identify the lever: with the model held constant, the harness changed both measured quality and input cost.

The reusable method matters more than either number:

  1. Run a defined test set.
  2. Read the failed traces.
  3. Group failures by behaviour.
  4. Change one system surface.
  5. Keep the change only if evaluation improves.
FalsifierIf the same model, tasks, and evaluation produce no material change after several isolated harness interventions, then the harness is not the present bottleneck. Investigate model capability, data quality, or the evaluation itself.

09The feedback loop

A production trace records what happened. It becomes an eval only after the team states what should have happened.

Stage Invoice example
Trace Agent used list_invoices_by_status despite an exact ID
Failure label Wrong tool selected when invoice_id is explicit
Eval Exact-ID cases must choose get_invoice_by_id
Change Add a negative boundary to the list tool
Holdout Test new exact-ID and ambiguous cases not used to design the fix
Ship Release if selection improves without harming ambiguous search

A holdout set contains test cases the team does not use while designing the change. It checks whether the improvement extends beyond the examples that inspired it.

Twenty well-labelled cases from real failures may be worth more than a thousand synthetic cases no user has produced.

Figure 02 · Practice
Trace to eval, and back
TRACE TO EVAL, AND BACK 01 · TRACE What the run actually did 02 · LABEL What should have happened 03 · EVAL The rule, written as a test case 04 · CHANGE One surface: prompt, tool, hook 05 · HOLDOUT Cases the fix never saw 06 · SHIP Release only if evaluation improves THE NEXT RELEASE INHERITS THE TEST A failure compounds only after it becomes a test.
Read it as A trace is evidence, not learning. Learning starts at the label and is banked only when the holdout passes.

10Three changes, restated

Strip the four patterns back to their mechanism and three changes remain. Each removes a different kind of guesswork.

Figure 03 · Framework
Three changes, three kinds of guesswork
THREE CHANGES · THREE KINDS OF GUESSWORK REMOVED None of them change the model. 01 MAP, NOT MANUAL A short entry point that says where authoritative rules live. Load detail on demand. Removes guesswork about the rules 02 CHECKED HANDOFF Plan, evidence, open questions and next action stored outside the conversation. Removes guesswork about progress 03 PROPOSE / AUTHORISE / VERIFY Model proposes, policy authorises, environment evidence closes the job. Removes guesswork about authority Read left to right. Rules, progress, and authority are three things an agent should never have to infer.
Read left to right Rules, progress, and authority are three things an agent should never have to infer.

Change 1: give the agent a map, not a manual

A large instruction file feels like control. Over time, it becomes the opposite.

Old rules remain beside new ones. Everything looks equally important. The file consumes the context the agent needs for the job, and nobody can tell which instruction still earns its place.

Keep a short entry point that tells the agent where authoritative information lives. Store architecture, workflows, policies, examples, and tool guidance in smaller versioned sources. Load the detail the current job requires.

OpenAI reports moving in this direction after a central AGENTS.md became too large and stale. The transferable lesson is progressive disclosure: start with a map, then load the relevant source of truth.[9]

TestCan a fresh agent session and a new colleague both find the current rule without reading the full history?
Bound the job before the model begins

A vague job invites the agent to invent work.

A preregistered coding-agent study by Sarel Weinberger and Amir Hozez, revised on 24 August 2026, tested an instruction many teams add by reflex: develop several approaches, compare them, and implement the best. Across six open-weight models and 4,644 runs, the instruction increased reasoning tokens by 2.4 to 7.4 times without improving success. Among discordant task pairs, the plain instruction won 26 to 13.[4]

The result is specific to the study's models, tasks, and harness. It does not prove that planning is useless. It proves that extra reasoning steps must earn their place through end-to-end success, not through the appearance of diligence.

A useful job card names:

Keep diagnosis and a final independent check. Do not invite open-ended exploration or repeated self-review unless evidence shows that they improve the complete job.

Change 2: make the handoff a checked artifact

A context window is temporary attention. It is not durable project state.

For work that outlives one call or session, store the plan, completed work, open questions, decisions, evidence, pending approvals, and next action outside the conversation.

A new run should begin by reading the handoff and checking its important claims against the environment. It should end by updating the artifact.

A weak handoff says:

Finished step three. Continue with step four.

A useful handoff says:

Updated the policy index to version 14. Retrieval passed for 28 of 30 cases. Two German-language cases still return version 13. Do not release. Evidence: retrieval-report-14. Next action: inspect locale routing.

The second handoff carries progress, uncertainty, evidence, and the next safe step.

TestCan the job recover after the model process, worker, or machine stops completely?

Change 3: let the model propose, let policy authorise, let evidence close

Models are useful for proposing the next action. They should not be the final authority for actions that change important state.

Before execution, enforced policy should check:

  1. which agent is acting,
  2. who or what it represents,
  3. the requested operation,
  4. the resource being changed,
  5. time and value limits,
  6. whether approval exists.

The tool then runs inside network, file, credential, time, and spending limits.

After execution, inspect the real result. Read the record back. Run the test. Open the page. Check the transaction state.

Do not close the job because the model wrote “success.”

One-week experiment

Measure Before After
End-to-end completion
Cost per accepted job
Failed resumptions
Human investigation minutes
Unauthorised proposals blocked

11In practice: one sprint

Pick one high-volume model decision or one failure with serious consequences.

Move Action Primary metric Counter-metric
Retry Return the exact violation and correction scope Recovery after first failure Wrong outcome after retry
Output Enforce one versioned shape Schema-valid output Semantic correctness
Tools Show relevant tools; rename one vague operation Correct-tool selection Valid completion rate
Completion Add one external business check before exit Valid completed workflows Loops and escalations
Learning Turn the next failed trace into an eval Time to regression coverage Holdout regressions

Do not ship all five moves at once if you cannot attribute the result. Start with the failure that appears most often or carries the largest consequence.

For the selected move, complete this record:

Field Record
Evidence Trace, validator result, tool call, ledger state, or missing artifact
Immediate containment What reduces harm before the full repair ships
Long-term repair The smallest durable change to the harness
Owner Person or team able to make and maintain the change
Regression case The test that proves the failure remains fixed
Release decision Ship, hold, narrow, or revert

12Connecting the dots

The four patterns share one design principle: let the model reason over the uncertain part, but do not make it improvise the contract around the uncertain part.

A precise error reduces wasted retries. A schema protects downstream systems. A narrow tool surface reduces wrong actions. A completion check lets the organisation trust longer work.

Each constraint can also add latency, cost, maintenance, or blocked edge cases. A stable tool prefix may save cache cost while exposing too many choices. A completion gate may prevent premature exit while creating a loop. A strict schema may protect a downstream system while rejecting a legitimate new case.

That tension is the bridge to Episode 04. The next question is not whether constraints work. It is when they buy enough reliability, when they limit useful autonomy, and when a stronger model makes them unnecessary.

You now hold four ambiguity reducers: evidence in the retry, a versioned output shape, a narrowed tool surface, and an external completion check. You also hold the loop that turns one failed trace into a regression test. Episode 04 asks when each constraint buys reliability, and when it quietly costs autonomy.

You now hold
Patterns 03 Four ambiguity reducers: evidence in the retry, a versioned output shape, a narrowed tool surface, and an external completion check. You also hold the loop that turns one failed trace into a regression test.
The next question
When does a constraint buy reliability, and when does it quietly cost autonomy, latency, or a legitimate edge case?
Continue
Harness 04 Five Paradoxes Every PM Must Hold → — when a constraint buys reliability, and when it quietly costs autonomy.
Read alongside
Environment 02 Construct the Task World → — where the completion evidence in Pattern 4 comes from.
Sources
  1. OpenAI — “Structured model outputs,” JSON mode, strict schemas, refusals, and unsupported features.
    developers.openai.com/api/docs/guides/structured-outputs
  2. LangChain — “Improving Deep Agents with harness engineering,” fixed model, changed prompt, tools, and middleware; Terminal Bench 2.0 score from 52.8% to 66.5%.
    langchain.com/blog/improving-deep-agents-with-harness-engineering
  3. Anthropic — “Effective harnesses for long-running agents,” feature inventories, progress files, startup scripts, and version history.
    anthropic.com/engineering/effective-harnesses-for-long-running-agents
  4. Sarel Weinberger and Amir Hozez — “Prompt-Induced Waste in Coding Agents: Reasoning, Effort, Harness Design, and End-to-End Cost,” arXiv 2608.01347, version 5, 24 August 2026.
    arxiv.org/abs/2608.01347
  5. LangChain — “August 2026: LangChain Newsletter,” 26 August 2026; Deep Agents v0.7 reported 65% fewer base input tokens at comparable performance.
    langchain.com/blog/august-2026-langchain-newsletter
  6. Ornella Altunyan, Braintrust — “Sandboxed agent evals with Harbor,” 24 August 2026.
    braintrust.dev/blog/harbor-agent-evals
  7. Donghui Zha, Lingwei Xu, Linxiao Wu, Yixue Dong, and Haochen Li — “CacheRouter,” arXiv 2608.22708, 24 August 2026.
    arxiv.org/abs/2608.22708
  8. Cursor — “Cloud Agents and Cursor Harness Improvements,” 19 August 2026.
    cursor.com/changelog/08-19-26
  9. OpenAI — “Harness engineering: leveraging Codex in an agent-first world.”
    openai.com/index/harness-engineering