- LocateWhich capabilities move into models and vendor products.
- KeepWhich product obligations remain local.
- OperateHow traces become evals, changes, and releases.
- RetireHow old controls leave without reopening old failures.
- ReviewFour questions to ask after every meaningful failure.
Episode 8 of 8: The discipline
01After this you will know
- Which agent capabilities are moving into models and vendor products.
- Which product obligations remain local to your organization.
- How traces become evals, evals become changes, and old controls get removed.
- Why the feedback loop matters more than any one pattern in this series.
- The four questions your team should ask after every meaningful failure.
02The dashboard changed
A year after the 2 AM invoice failure, the pager is quiet.
The reconciliation agent has completed two hundred nightly runs without leaving invoices behind. The team watches a different set of numbers now: validated completion, recovery time, review load, and cost per completed workflow.
Half the controls built during the first urgent quarter are gone. Newer models no longer need them.
What remains matters more because it encodes the company's workflow, evidence standard, permissions, and definition of a completed reconciliation.
The team did not preserve the first harness. It preserved the habit of testing what still deserves to exist.
Capabilities move. Product obligations remain.
That sentence closes the season.
03The path from failure to discipline
The eight episodes have followed one team and one changing question.
| Episode | Question | What changed |
|---|---|---|
| 1. Why Your Agent Fails | Where did the run first diverge? | "The model failed" became a layer diagnosis |
| 2. Inside a Production Harness | Which decision lives where? | The team drew four responsibilities and five harness jobs |
| 3. Three Small Changes | What can improve this sprint? | Retry, output, tools, and completion became explicit |
| 4. Five Paradoxes | When is each pattern right? | Every change gained a condition and counter-metric |
| 5. Costs and Returns | What is the work worth? | Cost moved from tokens to acceptable completed workflows |
| 6. Monday Morning Kit | What do we do first? | Nine days of diagnosis produced four tickets |
| 7. Organizations Rebuilt | Who may decide? | Components gained owners and behavior gained an accountable release authority |
| 8. The Habit | What survives model change? | The team keeps the loop and retires dead scaffolding |
The series began with attribution. It ends with operating cadence.
04Job versus implementation
A control can disappear while the job it performed remains.
The implementation moves into the model, provider API, or managed platform. The product team still owes the user an acceptable outcome.
| Job | Older implementation | Newer implementation | Obligation that remains |
|---|---|---|---|
| Return machine-readable output | Validator and repair loop | Native strict schema | Check semantic correctness |
| Select and call tools | Hand-built ReAct loop | Native tool calling or vendor harness | Authorize, verify, and record effects |
| Work with large knowledge sets | Heavy chunking and retrieval | Longer context plus retrieval | Freshness, access, provenance, cost |
| Reason across steps | Visible scratchpad scaffold | Native reasoning | Check consequential boundaries |
| Filter common unsafe content | Custom generic classifiers | Model and platform defaults | Enforce domain policy |
| Orient a coding agent | Large default prompt | Leaner defaults plus dynamic skills | Preserve local operating rules |
The capability has not vanished. Its implementation moved.
This distinction prevents two common mistakes:
- Keeping obsolete machinery because it once fixed a real problem.
- Deleting the local obligation because a vendor absorbed one implementation.
05Two destinations, two economics
A capability can move in two directions.
Into the model or provider API
Structured schema enforcement is an example. The team can remove some local repair logic and keep its architecture.
This is a clean simplification, provided semantic checks remain.
Into the vendor's harness
A coding product may absorb planning, tool routing, compaction, skills, parallel workers, and completion checks.
The team sheds engineering work and takes on vendor coupling at the same time.
One is closer to a gift. The other is closer to a lease.
Neither is automatically good or bad. The product question is which local judgment remains visible, exportable, and replaceable after the move.
06Five permanent residents
Some obligations remain local because they depend on facts, incentives, and duties no general model can own for you.
071. Your definition of good
A model provider can train for general helpfulness, accuracy, and safety.
It cannot decide that a partial payment must become an exception under your finance policy, that a support response violates your brand, or that an audit requires a specific evidence chain.
Those judgments belong in your evals and rubrics.
If a vendor becomes the only place your definition of good exists, the vendor owns more than infrastructure. It owns the product standard.
082. Your workflow rules
The model does not know your approval hierarchy, exception routing, source systems, change-freeze periods, or which action another team must sign.
These are facts about the organization.
They may be expressed through prompts, tools, policy engines, or workflow code. Their implementation can change. The organization still owns the rule.
093. Your evidence and audit
A production system needs a durable record of who requested the work, which context and policy applied, which actions were proposed, which authority approved them, what changed, and how completion was verified.
The record must remain inspectable across model and vendor changes.
A model can generate an explanation. It cannot be the sole keeper of evidence the organization may need to defend.
104. Your cost controls
The provider earns more when the system uses more compute. Your product earns more when enough compute produces an acceptable outcome.
The incentives differ.
Per-workflow budgets, retry ceilings, model routing, review capacity, and cost per validated completion stay in your control plane.
115. User and tenant context
Permissions, consent, history, locale, private state, and customer-specific policy must remain scoped to the correct user and request.
A general model should not carry that state across customers. The harness decides what enters this turn and what survives afterward.
12The feedback loop
The five residents stay useful only if they change with the product.
The feedback loop is release practice, not a metaphor.
1. Source evidence
Collect:
- Labeled production failures
- Representative successful runs
- Hand-written high-consequence cases
- Cases from new tools, policies, users, and models
2. Design the test
Define the expected behavior, metric, and trade-off.
Split cases into:
- Optimization set: cases the team uses while designing the change.
- Holdout set: cases the team does not use until validation. The holdout tests whether the improvement generalizes.
3. Change one thing
Change one prompt rule, tool description, schema, routing policy, memory rule, completion check, or approval gate.
One change keeps attribution possible.
4. Review and ship
Check the holdout and the counter-metrics. Review effects the eval may miss. Record version, reason, rollout, and rollback.
After release, new production behavior creates the evidence for the next cycle.
13The measured example
LangChain published one same-model example of this method.
It kept gpt-5.2-codex fixed and changed system prompt, tools, and middleware, its term for hooks around model and tool calls. Trace analysis produced targeted changes including self-verification, pre-completion checks, context onboarding, and loop detection.
Terminal Bench 2.0 performance moved from 52.8% to 66.5%.
The benchmark does not prove every change helps every workflow. It demonstrates the mechanism: read failed runs, group the behavior, change one control surface, and keep the change only if evaluation improves.
The invoice team runs the same method on a different distribution. Its benchmark is not Terminal Bench. It is the cases real finance users produce.
14Evaluate at the level where failure lives
A mature suite uses three altitudes.
| Level | Question | Invoice example |
|---|---|---|
| Run | Did one step behave correctly? | Was the correct tool selected? |
| Trace | Did the complete turn produce the right effect? | Did the invoice state match the payment evidence? |
| Thread | Did the multi-turn or multi-session goal finish? | Was the full queue reconciled across the night? |
A run-level check can pass while the thread fails. A final answer can look right after the path crossed an unacceptable boundary.
The eval should match the failure.
Deep methodology belongs in AI Evals. The operating rule here is simple: preserve enough evidence to judge at the level where the user's outcome exists.
15From traces to an asset
Production traffic alone is not a moat.
A pile of traces is exhaust.
The advantage appears when a team can:
- Capture representative behavior
- Label failures that matter
- Convert them into stable tests
- Make targeted changes
- Avoid overfitting through holdouts
- Ship faster than the failure distribution changes
That loop improves on data competitors do not have: your users, workflows, exceptions, and consequences.
Most teams stop earlier. LangChain's survey found that observability adoption outpaced eval adoption. Teams can see what agents do without turning those observations into release criteria.
The gap between watching and learning is an operating gap, not a tooling gap.
16Four small edits show what mature work looks like
LangChain documented several targeted instructions produced through evaluation work:
- Use reasonable defaults when the request implies them.
- Do not ask for information the user already supplied.
- Stop near-duplicate searches when enough evidence exists.
- Ask domain-defining questions before implementation questions.
The edits are short. The evidence behind them is not.
A one-line change may represent weeks of trace review, failure grouping, and controlled validation.
Mature harness work often looks small because the diagnosis is precise.
17Treat the harness as a versioned product
A harness that changes without versions is not a product. It is a rumour. Prompts, tools, policies, checks, and thresholds should each carry a version, an owner, a reason for the change, and a way to go back.
Three habits make that real:
- Named releasesA change to the definition of done is a release, not a tweak.
- Shadow firstNew checks run beside the current behaviour before they gate it.
- ReversibleEvery promotion has a stated trigger that reverses it.
The same discipline governs a harness that improves itself. When the system proposes changes to its own prompts, tools, or checks, those proposals enter the ladder below at step three — never at step five.
18The retirement loop
Improvement is incomplete until deletion becomes normal.
For every temporary control, record:
| Field | Example |
|---|---|
| Built for | Model and version |
| Failure addressed | Premature stopping near context limit |
| Evidence | Regression cases and production traces |
| Review trigger | Model upgrade or quarterly review |
| Retirement test | Enable-versus-disable eval on representative and holdout cases |
Anthropic reported one context-reset rule that fixed a real behavior in Sonnet 4.5 and became dead weight with Opus 4.5.
The lesson is not to avoid temporary controls. It is to give them a measured exit.
A system that only adds controls becomes harder to reason about. A system that deletes without evals repeats old failures. The discipline does both.
19The dissolve test
For each meaningful component, ask three questions.
Is the job generic or specific to us?
Generic work is more likely to move into models and platforms. Local approval rules and domain ground truth are not.
Does vendor scale improve the implementation?
Model providers can train tool selection across many tasks. They cannot generalize the exact exception path between your ERP and finance team without your context.
If the implementation moves, which obligation remains?
Native JSON removes some repair logic. Semantic correctness remains. Managed sessions remove runtime plumbing. Session portability and audit remain. Better tool choice removes a router. Authority and effect verification remain.
The third question prevents deletion from erasing the product duty.
20Four instincts that survive every implementation
Probabilistic components need deterministic contract surfaces
The model may reason flexibly. Tool arguments, permission, identity, and completion still need enforceable boundaries.
Trust is a set of properties
Replace "we trust the agent" with authorization, verification, rollback, traceability, and escalation.
Cost and reliability share one loop
Retries, context, review, and verification affect both the invoice and the incident. Optimize cost per acceptable completed workflow.
The demo-to-production gap is a design concern
A demo proves possibility. Production requires coverage of ambiguity, failure, concurrency, and consequence.
The specific patterns will change. These four instincts will keep producing useful questions.
21The two boundaries beside the harness
The series focused on the control plane that decides what happens next.
Two bonus essays complete the operating picture.
22The Tool Is the Contract
The tool is where a model proposal becomes an operation.
A production tool contract includes purpose, schema, identity, authorization, side effects, reversibility, idempotency, timeout, cost, and verification.
Read [The Tool Is the Contract (Harness Engineering Bonus)] when you need to design the boundary between reasoning and action.
23The Environment Is the Product Boundary
The environment is where an approved action gains real reach.
Process, filesystem, network, identity, and persistence define what the agent can observe, affect, and leave behind.
Read [The Environment Is the Product Boundary (Harness Engineering Bonus)] when you need to design the boundary between action and consequence.
Together:
- The model proposes.
- The harness decides.
- The tool performs.
- The environment contains.
- Evals judge the outcome.
That is the full operating picture.
24The case-study trail
The companion case studies show the same principles at different levels.
- OpenAI Codex and Symphony: repository as memory, issue-driven orchestration, executable policy, and controlled review.
- Anthropic long-running agents: initializer, durable feature state, incremental sessions, and independent evidence gates.
- Cursor repository-native controls: plans, scoped rules, skills, hooks, worktrees, and review inside ordinary development.
Read [Harness Engineering in Practice (Harness Engineering Case Studies)] after the season if you want architecture and control-loop examples rather than another conceptual layer.
25In practice: the weekly habit
Run one sixty-minute review each week.
| Time | Activity | Output |
|---|---|---|
| 15 min | Review one failed or borderline thread | First divergence and failure label |
| 15 min | Check whether an eval covers it | New or revised eval candidate |
| 15 min | Propose one targeted change | Owner, metric, counter-metric |
| 10 min | Review one temporary control | Keep, revise, or run retirement test |
| 5 min | Record the decision | Version, rationale, next check |
Do not wait for a major incident. Borderline and recovered cases often reveal the next failure class before it reaches a user.
26The four questions after every failure
The old review asked, "Which model failed?"
The new review asks:
- Where did the run first diverge from the intended outcome?
- Which control should have caught it?
- Which eval makes the failure reproducible?
- Which existing control no longer earns its cost?
Those questions contain the whole season.
The first comes from Episode 1. The second needs Episode 2's anatomy. The third uses Episodes 3 and 6. The fourth holds Episodes 4 and 5. Episode 7 supplies the decision owner.
27The final return to 2 AM
One year ago, the dashboard was green while 1,500 invoices sat untouched.
Today, completion compares processed and eligible counts. The session preserves progress. Retries receive specific evidence. The tool set is narrow. High-consequence actions have separate authority. Failed traces enter the eval suite. Temporary controls carry retirement conditions. One owner can accept or reject a behavior change.
The user never sees this machinery.
They see that the queue is finished in the morning.
That is the point.
28Closing
No agent change ships without an eval appropriate to its risk.
No consequential action proceeds without an authority boundary.
No production failure closes without a replayable trace or an explicit evidence gap.
No temporary workaround survives without a retirement condition.
When those expectations become ordinary release practice, the team stops debating whether agents require a different discipline.
The harness becomes less visible as a category because its obligations have become the way the organization builds software.
That is when the harness becomes the habit.
29Where the curriculum continues
This season covered the control plane. The surrounding curriculum continues from four seats:
- Agentic Stack: what the model sees and how context is assembled.
- Environment Engineering: what an action can reach and how the organization recovers.
- AI Evals: how acceptable behavior is defined and measured.
- AI PM OS: how capability becomes product economics, adoption, and authority.