Harness Engineering · Episode 08

When the Harness Becomes the Habit

Capabilities move. Product obligations remain.

Arc · The disciplineEpisode · 8 of 8Next · Environment Engineering — Operate
After this you will know
  • LocateWhich capabilities move into models and vendor products.
  • KeepWhich product obligations remain local.
  • OperateHow traces become evals, changes, and releases.
  • RetireHow old controls leave without reopening old failures.
  • ReviewFour questions to ask after every meaningful failure.

Episode 8 of 8: The discipline

01After this you will know

02The dashboard changed

A year after the 2 AM invoice failure, the pager is quiet.

The reconciliation agent has completed two hundred nightly runs without leaving invoices behind. The team watches a different set of numbers now: validated completion, recovery time, review load, and cost per completed workflow.

Half the controls built during the first urgent quarter are gone. Newer models no longer need them.

What remains matters more because it encodes the company's workflow, evidence standard, permissions, and definition of a completed reconciliation.

The team did not preserve the first harness. It preserved the habit of testing what still deserves to exist.

Capabilities move. Product obligations remain.

That sentence closes the season.

03The path from failure to discipline

The eight episodes have followed one team and one changing question.

EpisodeQuestionWhat changed
1. Why Your Agent FailsWhere did the run first diverge?"The model failed" became a layer diagnosis
2. Inside a Production HarnessWhich decision lives where?The team drew four responsibilities and five harness jobs
3. Three Small ChangesWhat can improve this sprint?Retry, output, tools, and completion became explicit
4. Five ParadoxesWhen is each pattern right?Every change gained a condition and counter-metric
5. Costs and ReturnsWhat is the work worth?Cost moved from tokens to acceptable completed workflows
6. Monday Morning KitWhat do we do first?Nine days of diagnosis produced four tickets
7. Organizations RebuiltWho may decide?Components gained owners and behavior gained an accountable release authority
8. The HabitWhat survives model change?The team keeps the loop and retires dead scaffolding

The series began with attribution. It ends with operating cadence.

04Job versus implementation

Figure 01 · Concept
Capabilities move. Obligations remain.
Implementations move into models and vendor harnesses while five local product obligations remain
Read it as Delete local machinery when the platform earns it. Never delete the duty it performed.

A control can disappear while the job it performed remains.

The implementation moves into the model, provider API, or managed platform. The product team still owes the user an acceptable outcome.

JobOlder implementationNewer implementationObligation that remains
Return machine-readable outputValidator and repair loopNative strict schemaCheck semantic correctness
Select and call toolsHand-built ReAct loopNative tool calling or vendor harnessAuthorize, verify, and record effects
Work with large knowledge setsHeavy chunking and retrievalLonger context plus retrievalFreshness, access, provenance, cost
Reason across stepsVisible scratchpad scaffoldNative reasoningCheck consequential boundaries
Filter common unsafe contentCustom generic classifiersModel and platform defaultsEnforce domain policy
Orient a coding agentLarge default promptLeaner defaults plus dynamic skillsPreserve local operating rules

The capability has not vanished. Its implementation moved.

This distinction prevents two common mistakes:

  1. Keeping obsolete machinery because it once fixed a real problem.
  2. Deleting the local obligation because a vendor absorbed one implementation.

05Two destinations, two economics

A capability can move in two directions.

Into the model or provider API

Structured schema enforcement is an example. The team can remove some local repair logic and keep its architecture.

This is a clean simplification, provided semantic checks remain.

Into the vendor's harness

A coding product may absorb planning, tool routing, compaction, skills, parallel workers, and completion checks.

The team sheds engineering work and takes on vendor coupling at the same time.

One is closer to a gift. The other is closer to a lease.

Neither is automatically good or bad. The product question is which local judgment remains visible, exportable, and replaceable after the move.

06Five permanent residents

Some obligations remain local because they depend on facts, incentives, and duties no general model can own for you.

071. Your definition of good

A model provider can train for general helpfulness, accuracy, and safety.

It cannot decide that a partial payment must become an exception under your finance policy, that a support response violates your brand, or that an audit requires a specific evidence chain.

Those judgments belong in your evals and rubrics.

If a vendor becomes the only place your definition of good exists, the vendor owns more than infrastructure. It owns the product standard.

082. Your workflow rules

The model does not know your approval hierarchy, exception routing, source systems, change-freeze periods, or which action another team must sign.

These are facts about the organization.

They may be expressed through prompts, tools, policy engines, or workflow code. Their implementation can change. The organization still owns the rule.

093. Your evidence and audit

A production system needs a durable record of who requested the work, which context and policy applied, which actions were proposed, which authority approved them, what changed, and how completion was verified.

The record must remain inspectable across model and vendor changes.

A model can generate an explanation. It cannot be the sole keeper of evidence the organization may need to defend.

104. Your cost controls

The provider earns more when the system uses more compute. Your product earns more when enough compute produces an acceptable outcome.

The incentives differ.

Per-workflow budgets, retry ceilings, model routing, review capacity, and cost per validated completion stay in your control plane.

115. User and tenant context

Permissions, consent, history, locale, private state, and customer-specific policy must remain scoped to the correct user and request.

A general model should not carry that state across customers. The harness decides what enters this turn and what survives afterward.

12The feedback loop

Figure 02 · Practice
The weekly harness habit
A six-step weekly loop from thread evidence through eval, change, holdout, ship, and retirement
Read it as A production failure is useful only when it becomes a test, a controlled change, and a decision about what to remove.

The five residents stay useful only if they change with the product.

The feedback loop is release practice, not a metaphor.

1. Source evidence

Collect:

2. Design the test

Define the expected behavior, metric, and trade-off.

Split cases into:

3. Change one thing

Change one prompt rule, tool description, schema, routing policy, memory rule, completion check, or approval gate.

One change keeps attribution possible.

4. Review and ship

Check the holdout and the counter-metrics. Review effects the eval may miss. Record version, reason, rollout, and rollback.

After release, new production behavior creates the evidence for the next cycle.

13The measured example

LangChain published one same-model example of this method.

It kept gpt-5.2-codex fixed and changed system prompt, tools, and middleware, its term for hooks around model and tool calls. Trace analysis produced targeted changes including self-verification, pre-completion checks, context onboarding, and loop detection.

Terminal Bench 2.0 performance moved from 52.8% to 66.5%.

The benchmark does not prove every change helps every workflow. It demonstrates the mechanism: read failed runs, group the behavior, change one control surface, and keep the change only if evaluation improves.

The invoice team runs the same method on a different distribution. Its benchmark is not Terminal Bench. It is the cases real finance users produce.

14Evaluate at the level where failure lives

A mature suite uses three altitudes.

LevelQuestionInvoice example
RunDid one step behave correctly?Was the correct tool selected?
TraceDid the complete turn produce the right effect?Did the invoice state match the payment evidence?
ThreadDid the multi-turn or multi-session goal finish?Was the full queue reconciled across the night?

A run-level check can pass while the thread fails. A final answer can look right after the path crossed an unacceptable boundary.

The eval should match the failure.

Deep methodology belongs in AI Evals. The operating rule here is simple: preserve enough evidence to judge at the level where the user's outcome exists.

15From traces to an asset

Production traffic alone is not a moat.

A pile of traces is exhaust.

The advantage appears when a team can:

That loop improves on data competitors do not have: your users, workflows, exceptions, and consequences.

Most teams stop earlier. LangChain's survey found that observability adoption outpaced eval adoption. Teams can see what agents do without turning those observations into release criteria.

The gap between watching and learning is an operating gap, not a tooling gap.

16Four small edits show what mature work looks like

LangChain documented several targeted instructions produced through evaluation work:

The edits are short. The evidence behind them is not.

A one-line change may represent weeks of trace review, failure grouping, and controlled validation.

Mature harness work often looks small because the diagnosis is precise.

17Treat the harness as a versioned product

A harness that changes without versions is not a product. It is a rumour. Prompts, tools, policies, checks, and thresholds should each carry a version, an owner, a reason for the change, and a way to go back.

Three habits make that real:

The same discipline governs a harness that improves itself. When the system proposes changes to its own prompts, tools, or checks, those proposals enter the ladder below at step three — never at step five.

Figure 01 · Framework
The promotion ladder
THE PROMOTION LADDER FOR REPEATED JUDGEMENTJudgement that repeats is a candidate for a check. Judgement that varies is not.1OBSERVE
A judgement is made by a person, repeatedly, the same way.
2RECORD
The decision, its inputs and its outcome are captured as evidence.
3PROPOSE
A rule or check is written and run in shadow beside the human.
4COMPARE
Agreement is measured on real cases, including the hard ones.
5PROMOTE
The check runs automatically, with the exception path still open.
6REVIEW
The rule keeps a version, an owner, and a rollback trigger.
Judgement moves down the ladder, never skips a rung. A rule that reaches step five without steps three and four is an untested policy running in production.

18The retirement loop

Improvement is incomplete until deletion becomes normal.

For every temporary control, record:

FieldExample
Built forModel and version
Failure addressedPremature stopping near context limit
EvidenceRegression cases and production traces
Review triggerModel upgrade or quarterly review
Retirement testEnable-versus-disable eval on representative and holdout cases

Anthropic reported one context-reset rule that fixed a real behavior in Sonnet 4.5 and became dead weight with Opus 4.5.

The lesson is not to avoid temporary controls. It is to give them a measured exit.

A system that only adds controls becomes harder to reason about. A system that deletes without evals repeats old failures. The discipline does both.

19The dissolve test

For each meaningful component, ask three questions.

Is the job generic or specific to us?

Generic work is more likely to move into models and platforms. Local approval rules and domain ground truth are not.

Does vendor scale improve the implementation?

Model providers can train tool selection across many tasks. They cannot generalize the exact exception path between your ERP and finance team without your context.

If the implementation moves, which obligation remains?

Native JSON removes some repair logic. Semantic correctness remains. Managed sessions remove runtime plumbing. Session portability and audit remain. Better tool choice removes a router. Authority and effect verification remain.

The third question prevents deletion from erasing the product duty.

20Four instincts that survive every implementation

Probabilistic components need deterministic contract surfaces

The model may reason flexibly. Tool arguments, permission, identity, and completion still need enforceable boundaries.

Trust is a set of properties

Replace "we trust the agent" with authorization, verification, rollback, traceability, and escalation.

Cost and reliability share one loop

Retries, context, review, and verification affect both the invoice and the incident. Optimize cost per acceptable completed workflow.

The demo-to-production gap is a design concern

A demo proves possibility. Production requires coverage of ambiguity, failure, concurrency, and consequence.

The specific patterns will change. These four instincts will keep producing useful questions.

21The two boundaries beside the harness

The series focused on the control plane that decides what happens next.

Two bonus essays complete the operating picture.

22The Tool Is the Contract

The tool is where a model proposal becomes an operation.

A production tool contract includes purpose, schema, identity, authorization, side effects, reversibility, idempotency, timeout, cost, and verification.

Read [The Tool Is the Contract (Harness Engineering Bonus)] when you need to design the boundary between reasoning and action.

23The Environment Is the Product Boundary

The environment is where an approved action gains real reach.

Process, filesystem, network, identity, and persistence define what the agent can observe, affect, and leave behind.

Read [The Environment Is the Product Boundary (Harness Engineering Bonus)] when you need to design the boundary between action and consequence.

Together:

That is the full operating picture.

24The case-study trail

The companion case studies show the same principles at different levels.

Read [Harness Engineering in Practice (Harness Engineering Case Studies)] after the season if you want architecture and control-loop examples rather than another conceptual layer.

25In practice: the weekly habit

Run one sixty-minute review each week.

TimeActivityOutput
15 minReview one failed or borderline threadFirst divergence and failure label
15 minCheck whether an eval covers itNew or revised eval candidate
15 minPropose one targeted changeOwner, metric, counter-metric
10 minReview one temporary controlKeep, revise, or run retirement test
5 minRecord the decisionVersion, rationale, next check

Do not wait for a major incident. Borderline and recovered cases often reveal the next failure class before it reaches a user.

26The four questions after every failure

The old review asked, "Which model failed?"

The new review asks:

  1. Where did the run first diverge from the intended outcome?
  2. Which control should have caught it?
  3. Which eval makes the failure reproducible?
  4. Which existing control no longer earns its cost?

Those questions contain the whole season.

The first comes from Episode 1. The second needs Episode 2's anatomy. The third uses Episodes 3 and 6. The fourth holds Episodes 4 and 5. Episode 7 supplies the decision owner.

27The final return to 2 AM

One year ago, the dashboard was green while 1,500 invoices sat untouched.

Today, completion compares processed and eligible counts. The session preserves progress. Retries receive specific evidence. The tool set is narrow. High-consequence actions have separate authority. Failed traces enter the eval suite. Temporary controls carry retirement conditions. One owner can accept or reject a behavior change.

The user never sees this machinery.

They see that the queue is finished in the morning.

That is the point.

28Closing

No agent change ships without an eval appropriate to its risk.

No consequential action proceeds without an authority boundary.

No production failure closes without a replayable trace or an explicit evidence gap.

No temporary workaround survives without a retirement condition.

When those expectations become ordinary release practice, the team stops debating whether agents require a different discipline.

The harness becomes less visible as a category because its obligations have become the way the organization builds software.

That is when the harness becomes the habit.

29Where the curriculum continues

This season covered the control plane. The surrounding curriculum continues from four seats:

You now hold
Harness 08 · FinaleA retirement loop, three eval altitudes, five permanent obligations, and a weekly operating habit.
Go deeper
Environment EngineeringThe World Around the Agent →— move from the control plane to the boundary where approved action gains reach and consequence.

30Sources