- RetryWhy another attempt helps only when something changes.
- SchemaWhat a shape contract can prove, and what it cannot.
- ToolsHow to reduce tool confusion without removing useful capability.
- LearningHow one failed run becomes a test that protects the next release.
01Where Episode 02 left us
At 9:10 on Monday, the reconciliation team opens a spreadsheet comparing four models.
One column is green. The production failure list is unchanged.
The agent still repeats malformed calls, returns fields the downstream system cannot use, and chooses the wrong operation when tool names overlap. In the most serious failure, it processes 847 of 2,347 eligible invoices and stops with 1,500 still waiting.
Episode 01 taught the team to find the first broken contract. Episode 02 showed where that contract lives: Identity, Memory Policy, Orchestration, Interception, or Observability and Evals.
The team can now draw the machine. The next question is practical: what can it change this sprint without replacing the model?
02The sprint that changed no model
On Monday, the team freezes the model.
Validation errors will return to the next attempt. Outputs must follow an agreed shape. The model will see only the tools needed for the current step. The workflow will verify the ledger before it accepts completion.
By Friday, the lesson is not that four tricks worked. It is why they worked.
Each change removed a decision the model had been forced to improvise.
The title says three changes and the body teaches four patterns. That is deliberate. The four patterns are the controls the team ships. Section 10 reduces them to three deeper design moves: map the rules, preserve the handoff, and separate proposal from authority and proof.
Reliability improves when the system replaces a vague choice with a specific signal, boundary, or check.
03The whole pattern in one table
| Step | What happens | What the harness decides |
|---|---|---|
| Prepare | Gather current state and relevant tools | What the model may see and choose |
| Propose | The model returns an answer or tool request | Whether the proposal has an acceptable shape |
| Act | A tool performs the approved operation | Whether the action is permitted |
| Verify | The system checks the real effect | Whether the step moved the task forward |
| Continue | The run advances, retries, escalates, or stops | Whether external completion has passed |
| Learn | The trace is stored and judged | Whether this failure becomes a regression test |
The model proposes. Reliability comes from the decisions before and after that proposal.
For business, this means defining the accepted outcome. For product, it means deciding where ambiguity is useful and where it is dangerous. Engineering implements the constraints. Security enforces authority. Operations keeps the trace and the recovery path.
04Pattern 1: retry with evidence
A naive retry says, “Try again.”
The model receives the same task, context, and uncertainty. A different answer may appear, but the second attempt has no better reason to succeed.
A useful retry explains what failed.
| Weak feedback | Useful feedback |
|---|---|
| “Validation failed” | amount_cents must be a whole number |
| “Try again” | You returned "$42.50"; return 4250 |
| “Wrong format” | Change this field only; preserve the other values |
The useful version supplies four things: the field, observed value, expected rule, and correction scope.
A retry must change at least one condition:
- InformationReturn the exact validation error.
- StrategyRequire a different method after repeated failure.
- ContextRestart with a clean view and durable progress.
- AuthorityEscalate to another model, tool, or human.
If none changes, the retry is repetition.
Three failures are often blurred together:
| Failure | What it means | Invoice example | What repairs it |
|---|---|---|---|
| Syntax | The output breaks an agreed structure | Invalid JSON, missing field, unsupported status | Provider-enforced structure and boundary validation |
| Semantic | The shape is valid, but the business value is wrong | The amount is a whole number, but belongs to another invoice | Business rules, source reconciliation, tests, or an eval |
| Goal | Individual outputs are valid, but the job is unfinished | 847 invoices pass while 1,500 are never processed | External completion tied to the real workflow |
One validator cannot do all three jobs.
The useful metric is not retry count. It is the share of failed first attempts that recover without producing a wrong business outcome.
05Pattern 2: constrain the output space
Teams often treat structured output as a developer convenience. Its product value is a smaller failure space.
Free text lets the model return anything it can express. A schema limits the response to fields and values the next system knows how to handle.
Consider the invoice output:
| Field | Weak contract | Stronger contract |
|---|---|---|
| Vendor | Any text | Non-empty text |
| Amount | Any text | Whole number in cents, zero or more |
| Due date | Any text | YYYY-MM-DD |
| Status | Any text | Paid, pending, overdue, or disputed |
| Extra fields | Allowed silently | Rejected unless added to the versioned contract |
An instruction such as “return JSON” may still produce:
- Amount
"$42.50" - Due date
"end of August" - Status
"awaiting payment" - Extra fieldAn invented explanation the caller ignores
The response looks sensible to a person. The next system may fail or silently discard information.
| Strict output proves | Strict output does not prove |
|---|---|
| Required fields exist | The amount belongs to the right invoice |
| Values use supported types | The status matches the ledger |
| Finite business states use agreed labels | The source data is current |
| Unagreed fields are rejected | The whole job is finished |
A schema checks shape. An eval checks meaning. A completion rule checks the job.
OpenAI's Structured Outputs documentation distinguishes JSON mode, which produces valid JSON, from strict schema adherence, which constrains supported outputs to a supplied schema.[1] Your application must still handle refusals, empty results, timeouts, unsupported schema features, and semantically wrong values.
A schema is a contract with every downstream consumer, not only the model. When it changes:
- Version the new shape.
- Identify affected consumers.
- Test old and new cases.
- Plan compatibility or migration.
- Monitor empty, refused, and semantically invalid outputs.
A field added casually today becomes a silent mismatch three systems later.
Track schema validity and semantic correctness separately. A perfect shape can carry a wrong answer.
06Pattern 3: narrow the decision surface
Give an agent forty tools and it must distinguish forty operations before every action.
Some tools are irrelevant. Some overlap. Others use broad names such as
query_data or update_record. The model must infer what each
name means and whether the consequence is read, draft, or execute.
The goal is not always fewer total capabilities. It is fewer visible choices for this step.
| Vague tool | Better tool | Decision removed |
|---|---|---|
query_data |
get_invoice_by_id |
Which dataset and query shape? |
run_analysis |
list_overdue_invoices |
Which analysis and output? |
update_record |
draft_payment_adjustment |
Draft or execute? Which record? |
refund |
draft_refund and issue_refund |
Proposal or irreversible action? |
A precise name is part of the control surface.
Three operations improve the catalogue:
- Delete tools with no distinct job.
- Rename vague tools to exact actions.
- Reveal specialist tools only when the task needs them.
The third is progressive disclosure. The system keeps broad capability but presents a small, relevant set now. A reconciliation skill may expose invoice lookup, payment lookup, exception drafting, and completion checks. A supplier-onboarding skill exposes a different set.
Progressive disclosure creates a less obvious trade-off. Prompt caching rewards a stable request prefix, while changing the visible tool list can invalidate that prefix. CacheRouter, a preprint submitted on 24 August 2026, proposes separating tool discovery from the main model's fixed tool prefix. In its prototype, tested on 55 functional queries and one 30-turn dialogue, token-level cache hit rates reached 90.99% and 95.2%.[7]
Those figures come from the authors' prototype, not an independent production benchmark. The transferable design question is narrower: can you reduce the visible decision surface without rebuilding the entire prompt on every turn?
A rare escalation tool may be essential. Review each tool on four dimensions:
| Dimension | Question |
|---|---|
| Usage | How often is it selected? |
| Contribution | Does it help complete the task? |
| Overlap | Does another tool perform the same job? |
| Consequence | What happens if the model selects it incorrectly? |
The decision may be keep, rename, tighten, hide behind a skill, require approval, merge, or remove. The goal is not the shortest list. It is the clearest set that covers the workflow.
Near-neighbour tools fail at their boundary. Descriptions should therefore say when not to use them:
Use get_invoice_by_id when an exact invoice ID is known. Do not use it for
vendor search or status lists.
Track correct-tool selection and steps per validated completion. Fewer calls help only when completion remains correct.
07Pattern 4: verify before exit
The invoice failure in Episode 01 did not involve malformed output or the wrong tool. The run stopped before the workflow ended.
Replace model confidence with an external checklist.
| Completion check | Required evidence |
|---|---|
| Coverage | Processed count equals 2,347 eligible source records |
| Exceptions | Every unmatched record has an assigned state |
| Quality | Reconciliation checks pass |
| Authority | No prohibited action occurred |
The model may propose completion. The harness checks the business state.
Output feedback says, “This field is invalid.” Completion feedback says, “This outcome is unfinished.” The first repairs a response. The second keeps the job open.
Anthropic's work on long-running agents uses feature inventories, progress files, startup scripts, and version history so a fresh session can reconstruct what remains.[3]
The same pattern now appears in evaluation and product tooling. Braintrust's Harbor
integration runs each task in an isolated container and records the verifier that
inspects the container after the agent stops.[6] Cursor's /goal,
released on 19 August 2026, holds a long-lived objective until it is fully
complete.[8]
The products differ. The rule is the same: done is a state of the world, not a sentence.
A completion gate can create an infinite loop. Define:
- BudgetThe time and cost ceiling.
- AttemptsThe maximum attempts without progress.
- RecordThe durable progress artifact.
- EscalationWhere the run goes when it stalls.
- StopThe condition under which the task cannot complete.
“Continue until done” is not a policy unless both done and stop are defined.
08The measured receipt
LangChain held the model fixed and changed the system prompt, tools, and middleware, its term for hooks around model and tool calls.
On Terminal Bench 2.0, an 89-task coding benchmark, its reported score rose from 52.8% to 66.5%, a gain of 13.7 percentage points. The work included trace-based error analysis, self-verification, pre-completion checks, environment context, and loop detection.[2]
The result does not prove that every harness change helps every workflow. It supports a narrower claim: changing the control system can materially change measured behaviour while the model stays fixed.
A second receipt concerns cost. In its 26 August 2026 newsletter, LangChain reported that Deep Agents v0.7 simplified the base harness and used 65% fewer base input tokens at comparable performance.[5]
Both results are company-reported measurements on LangChain's systems. Neither predicts your workflow. Together, they identify the lever: with the model held constant, the harness changed both measured quality and input cost.
The reusable method matters more than either number:
- Run a defined test set.
- Read the failed traces.
- Group failures by behaviour.
- Change one system surface.
- Keep the change only if evaluation improves.
09The feedback loop
A production trace records what happened. It becomes an eval only after the team states what should have happened.
| Stage | Invoice example |
|---|---|
| Trace | Agent used list_invoices_by_status despite an exact ID |
| Failure label | Wrong tool selected when invoice_id is explicit |
| Eval | Exact-ID cases must choose get_invoice_by_id |
| Change | Add a negative boundary to the list tool |
| Holdout | Test new exact-ID and ambiguous cases not used to design the fix |
| Ship | Release if selection improves without harming ambiguous search |
A holdout set contains test cases the team does not use while designing the change. It checks whether the improvement extends beyond the examples that inspired it.
Twenty well-labelled cases from real failures may be worth more than a thousand synthetic cases no user has produced.
10Three changes, restated
Strip the four patterns back to their mechanism and three changes remain. Each removes a different kind of guesswork.
Change 1: give the agent a map, not a manual
A large instruction file feels like control. Over time, it becomes the opposite.
Old rules remain beside new ones. Everything looks equally important. The file consumes the context the agent needs for the job, and nobody can tell which instruction still earns its place.
Keep a short entry point that tells the agent where authoritative information lives. Store architecture, workflows, policies, examples, and tool guidance in smaller versioned sources. Load the detail the current job requires.
OpenAI reports moving in this direction after a central AGENTS.md became too
large and stale. The transferable lesson is progressive disclosure: start with a map, then
load the relevant source of truth.[9]
A vague job invites the agent to invent work.
A preregistered coding-agent study by Sarel Weinberger and Amir Hozez, revised on 24 August 2026, tested an instruction many teams add by reflex: develop several approaches, compare them, and implement the best. Across six open-weight models and 4,644 runs, the instruction increased reasoning tokens by 2.4 to 7.4 times without improving success. Among discordant task pairs, the plain instruction won 26 to 13.[4]
The result is specific to the study's models, tasks, and harness. It does not prove that planning is useless. It proves that extra reasoning steps must earn their place through end-to-end success, not through the appearance of diligence.
A useful job card names:
- ScopeWhat is included and excluded?
- AcceptanceWhat evidence proves completion?
- StopWhen must the agent stop, even if it could continue exploring?
Keep diagnosis and a final independent check. Do not invite open-ended exploration or repeated self-review unless evidence shows that they improve the complete job.
Change 2: make the handoff a checked artifact
A context window is temporary attention. It is not durable project state.
For work that outlives one call or session, store the plan, completed work, open questions, decisions, evidence, pending approvals, and next action outside the conversation.
A new run should begin by reading the handoff and checking its important claims against the environment. It should end by updating the artifact.
A weak handoff says:
Finished step three. Continue with step four.
A useful handoff says:
Updated the policy index to version 14. Retrieval passed for 28 of 30 cases. Two
German-language cases still return version 13. Do not release. Evidence:
retrieval-report-14. Next action: inspect locale routing.
The second handoff carries progress, uncertainty, evidence, and the next safe step.
Change 3: let the model propose, let policy authorise, let evidence close
Models are useful for proposing the next action. They should not be the final authority for actions that change important state.
Before execution, enforced policy should check:
- which agent is acting,
- who or what it represents,
- the requested operation,
- the resource being changed,
- time and value limits,
- whether approval exists.
The tool then runs inside network, file, credential, time, and spending limits.
After execution, inspect the real result. Read the record back. Run the test. Open the page. Check the transaction state.
Do not close the job because the model wrote “success.”
One-week experiment
| Measure | Before | After |
|---|---|---|
| End-to-end completion | ||
| Cost per accepted job | ||
| Failed resumptions | ||
| Human investigation minutes | ||
| Unauthorised proposals blocked |
11In practice: one sprint
Pick one high-volume model decision or one failure with serious consequences.
| Move | Action | Primary metric | Counter-metric |
|---|---|---|---|
| Retry | Return the exact violation and correction scope | Recovery after first failure | Wrong outcome after retry |
| Output | Enforce one versioned shape | Schema-valid output | Semantic correctness |
| Tools | Show relevant tools; rename one vague operation | Correct-tool selection | Valid completion rate |
| Completion | Add one external business check before exit | Valid completed workflows | Loops and escalations |
| Learning | Turn the next failed trace into an eval | Time to regression coverage | Holdout regressions |
Do not ship all five moves at once if you cannot attribute the result. Start with the failure that appears most often or carries the largest consequence.
For the selected move, complete this record:
| Field | Record |
|---|---|
| Evidence | Trace, validator result, tool call, ledger state, or missing artifact |
| Immediate containment | What reduces harm before the full repair ships |
| Long-term repair | The smallest durable change to the harness |
| Owner | Person or team able to make and maintain the change |
| Regression case | The test that proves the failure remains fixed |
| Release decision | Ship, hold, narrow, or revert |
12Connecting the dots
The four patterns share one design principle: let the model reason over the uncertain part, but do not make it improvise the contract around the uncertain part.
A precise error reduces wasted retries. A schema protects downstream systems. A narrow tool surface reduces wrong actions. A completion check lets the organisation trust longer work.
Each constraint can also add latency, cost, maintenance, or blocked edge cases. A stable tool prefix may save cache cost while exposing too many choices. A completion gate may prevent premature exit while creating a loop. A strict schema may protect a downstream system while rejecting a legitimate new case.
That tension is the bridge to Episode 04. The next question is not whether constraints work. It is when they buy enough reliability, when they limit useful autonomy, and when a stronger model makes them unnecessary.
You now hold four ambiguity reducers: evidence in the retry, a versioned output shape, a narrowed tool surface, and an external completion check. You also hold the loop that turns one failed trace into a regression test. Episode 04 asks when each constraint buys reliability, and when it quietly costs autonomy.
-
OpenAI — “Structured model outputs,” JSON mode, strict schemas,
refusals, and unsupported features.
developers.openai.com/api/docs/guides/structured-outputs -
LangChain — “Improving Deep Agents with harness engineering,” fixed
model, changed prompt, tools, and middleware; Terminal Bench 2.0 score from 52.8% to
66.5%.
langchain.com/blog/improving-deep-agents-with-harness-engineering -
Anthropic — “Effective harnesses for long-running agents,” feature
inventories, progress files, startup scripts, and version history.
anthropic.com/engineering/effective-harnesses-for-long-running-agents -
Sarel Weinberger and Amir Hozez — “Prompt-Induced Waste in Coding Agents:
Reasoning, Effort, Harness Design, and End-to-End Cost,” arXiv 2608.01347,
version 5, 24 August 2026.
arxiv.org/abs/2608.01347 -
LangChain — “August 2026: LangChain Newsletter,” 26 August 2026;
Deep Agents v0.7 reported 65% fewer base input tokens at comparable performance.
langchain.com/blog/august-2026-langchain-newsletter -
Ornella Altunyan, Braintrust — “Sandboxed agent evals with Harbor,”
24 August 2026.
braintrust.dev/blog/harbor-agent-evals -
Donghui Zha, Lingwei Xu, Linxiao Wu, Yixue Dong, and Haochen Li —
“CacheRouter,” arXiv 2608.22708, 24 August 2026.
arxiv.org/abs/2608.22708 -
Cursor — “Cloud Agents and Cursor Harness Improvements,” 19 August
2026.
cursor.com/changelog/08-19-26 -
OpenAI — “Harness engineering: leveraging Codex in an agent-first
world.”
openai.com/index/harness-engineering