Harness Engineering · Episode 06

Your Monday Morning Harness Kit

Nine days to diagnose. Twelve weeks to ship.

Arc · The practice Episode · 06 of 08 Next · Episode 07 — Organizations That Rebuilt Around Agents
After this you will know
  • InspectHow to examine the agent system already running.
  • SampleWhich production sessions to read and what to record.
  • PlaceHow to diagnose the workflow’s current level of operating maturity.
  • FileThe four tickets that begin a measurable improvement loop.
  • TranslateHow to explain the same plan to engineering, design, leadership, finance, and security.

01The first five episodes, compressed

At 9:04 on Monday morning, the reconciliation team opens the same dashboard that failed in Episode 01.

It still shows a green run. This time, the team knows that green is not evidence.

It asks for the source count, processed count, exception states, authority decisions, trace, reviewer time, and cost per accepted completion.

The first five episodes changed how the team sees the system:

  1. Episode 01: Diagnose. Find the first point where the run departed from the user’s intended outcome.
  2. Episode 02: Locate. Place the failure in the Model, Harness, Tool, Runtime, or Environment, then identify the harness decision involved.
  3. Episode 03: Intervene. Add evidence to retries, constrain output, narrow visible tools, and verify before exit.
  4. Episode 04: Decide. State what must be true before moving toward more capability, autonomy, reuse, or production exposure.
  5. Episode 05: Price. Track Build, Run, Review, and Failure cost, then measure cost per accepted result rather than cost per call.

Episode 06 turns those ideas into a working plan.

02Monday, 9:04 AM

The team agrees that the current system needs work.

Nobody has approved a platform migration, a new department, or a six-month transformation programme.

Good. The first move is evidence, not architecture.

For nine working days, the team will inspect one production workflow, read how it behaves, and turn the findings into four owned tickets. The next twelve weeks will ship those changes one decision at a time.

Days 1 to 9 diagnose. The next twelve weeks ship.

Choose one workflow before beginning. Pick recurring operational pain, not the most impressive demonstration.

The invoice workflow qualifies because it has:

If the workflow has no clear definition of completion, defining it becomes Day 1.

Figure 01 · Concept
Nine days diagnose. Twelve weeks ship.
EVIDENCE BEFORE ARCHITECTURE DAYS 1–9 · DIAGNOSE Read what is true Map · sessions · interviews Tools · eval gap · maturity Tickets · calibration · ownership 7 ARTIFACTS WEEKS 1–12 · SHIP One decision at a time Identity · memory · tools · observability Orchestration · interception · portability Feedback · second model · discipline → WHAT CROSSES THE LINE 01 System map with owners 02 Failure tally 03 Operator workarounds 04 Tool decision table 05 Eval-debt list 06 Maturity diagnosis 07 Four owned tickets: the only output that reaches the backlog
Read it as If you are holding a platform proposal instead, you moved too early. Nine days of evidence earn the right to spend the next twelve weeks.

03What this kit produces

At the end of Day 9, the team should hold seven artifacts:

ArtifactWhat it answersFormat
System mapWhere do decisions and state live?One page
Failure tallyHow does the workflow fail?Tagged session list
Operator workaroundsWhere is hidden labour already being paid?Two interview notes plus overlap
Tool decision tableWhich actions are useful, confusing, or dangerous?One row per tool
Eval-debt listWhich promised behaviours can break silently?Behaviour-to-check table
Maturity diagnosisWhat improvement can this workflow absorb next?One evidence-based paragraph
Four owned ticketsWhat will the team change first?Real backlog tickets

The kit is not trying to produce a complete target architecture.

It is trying to replace four common guesses:

Each statement may eventually be correct. The nine days produce the evidence required to know.

What not to do

Do not change the model while diagnosing unless an incident requires containment. Do not redesign the whole tool layer. Do not create a new maturity score for the company. Do not ask every stakeholder for a wish list.

Freeze the comparison point. Study one workflow. Record what is true.

04Days 1 to 3: map, read, interview

Day 1: Map what production runs

QuestionWhat system is handling this workflow today?

Most teams already have a harness. They do not call it one.

The system prompt lives in one repository. Retrieval is configured in a vendor console. Retry logic sits in a utility. Tool permissions come from an API gateway. Traces land in an observability product. A spreadsheet holds the eval cases. An operator knows which run to restart manually.

Three people own fragments. Nobody owns the behaviour produced by the whole.

Draw one page with four columns: Model, Harness, Tools, Environment. Put Runtime underneath.

Inside Harness, add five rows: Identity, Memory Policy, Orchestration, Interception, Observability and Evals.

For each row, record:

FieldQuestion
ImplementationWhere does this decision happen today?
OwnerWho can change it?
EvidenceWhat proves it works?
FailureWhat does a defect look like?
VersionCan we identify what ran?

Write missing when no implementation, owner, or evidence exists. Write the vendor’s name when the decision lives outside your code.

Do not fix anything today.

Artifact: One system map with owners, dependencies, and missing rows.

Connection: Episode 02 supplied the architecture. Day 1 applies it to the system you actually ship.

Day 2: Read complete production sessions

QuestionHow does the workflow fail in real use?

Read sessions, not final responses.

A complete session includes the request, context sources, model calls, tool calls, retries, state changes, approval decisions, final outcome, reviewer actions, and completion evidence.

The final response may look correct while the path contains wasted loops, unsafe calls, stale context, or unverified effects.

How many sessions should you read?

Use two samples for two different jobs.

Discovery sample: Read 30 to 50 sessions chosen to expose variety. Include successful sessions, user-corrected sessions, tool failures, escalations, long or expensive runs, and sessions that operators remember.

Rate sample: Draw a random or properly stratified sample if you want to estimate how often each failure occurs. A hand-picked failure sample discovers patterns but cannot estimate production rates.

If the product has fewer than 50 sessions, read all of them. If it has millions, begin with 50 for discovery, then validate the pattern against a representative sample.

Tag the six failure shapes from Episode 01:

Failure shapeWhat to look for
Premature stopWork remains after a clean exit
Infinite loopRepeated action without new evidence or progress
Silent driftPolished work against stale intent or policy
Vocabulary mismatchThe same term carries different business meanings
Unauthorised actionThe action exceeds delegated authority
Worker divergenceParallel work conflicts or is lost during merge

For every failed session, record:

Artifact: A failure tally plus three representative sessions that business, product, engineering, and security can inspect together.

Connection: Episode 01 supplied the diagnostic ladder. Episode 05 supplied the cost buckets. Day 2 joins them in one record.

Day 3: Interview the people who repair the system

QuestionWhat does the organisation quietly work around?

Interview the staff engineer and the person who responds when the workflow fails. Speak with them separately before comparing answers.

Ask:

  1. What do you repair manually without filing anymore?
  2. Which alert do you distrust or ignore?
  3. Which customer or data pattern makes you nervous?
  4. If feature work stopped for one sprint, what would you fix first?
  5. What evidence do you wish the trace contained during an incident?

Listen for manual restarts, hand-corrected state, secret spreadsheets, production prompt edits, repeated tool overrides, reviewers who always reverse one decision, and checks that happen outside the product.

The overlap between both interviews is often the first high-confidence ticket.

Episode 05 showed why this matters economically. McKinsey estimated human oversight at 70% to 75% of variable run cost in one banking customer-service example, compared with 20% to 25% for tokens.[7] Those are example-specific estimates, but they show why hidden operator labour belongs in the audit.

Artifact: Two short interview records, an overlap section, and the estimated Build, Run, Review, or Failure cost of each workaround.

05Days 4 to 6: tools, evals, maturity

Day 4: Audit the tool surface

QuestionWhich tools create useful action, and which create confusion or risk?

Export the tool registry the agent can actually access in production. Do not use the architecture document if runtime configuration differs.

For every tool, record:

FieldQuestion
JobWhat distinct user or workflow job does it perform?
VisibilityIn which steps can the model see it?
InputsAre fields, formats, and boundaries explicit?
Side effectRead, draft, reversible write, or irreversible action?
IdentityWhich principal and agent identity call it?
ApprovalWhich calls require permission or human confirmation?
VerificationHow is the real effect checked?
OwnerWho changes the contract and permission?
UsageHow often is it selected?
ContributionHow often does it lead to an accepted completion?

Then choose one action per tool: keep visible, rename, tighten, hide behind a skill, require approval, merge, or remove.

Do not delete a tool because it is rare. A rare escalation operation may be essential. Do not keep a tool because it is popular. A vague general-purpose tool may be popular because the model sees no better choice.

Artifact: Tool decision table with one owner and one decision per tool.

Connection: Episode 03 narrowed the decision surface. Day 4 tests whether the production surface is clear and bounded.

Day 5: Find the eval gap

QuestionWhich promised behaviour can break without a test failing?

List what the product claims to do. For invoice reconciliation:

For each claim, ask:

If this behaviour broke tomorrow, which automated check would fail?

A blank answer is eval debt: promised behaviour with no repeatable test.

Put completion first when it is missing. A system can pass format, tool-selection, and tone checks while leaving 1,500 invoices untouched.

Evaluate at the right level

Evaluation tools use different names. Keep the meaning clear:

LevelPlain meaningInvoice example
Step or runOne model or tool operationDid the model choose get_invoice_by_id?
TraceOne complete agent executionDid this invoice reconciliation action succeed?
Thread or journeyThe user’s goal across turns or sessionsDid all 2,347 invoices finish across restarts?

A final-answer check can miss a dangerous path. A step-level check can pass while the complete job fails.

For each behaviour, prefer the cheapest valid check:

Artifact: Behaviour-to-eval table, with blank rows highlighted and the verification method named.

Day 6: Write the maturity diagnosis

QuestionWhat improvement is this workflow equipped to absorb next?

Use five rungs as a diagnostic, not a maturity badge.

RungEvidence
1: PromptOne model call; little durable state or traceability
2: RetryParsing and retries exist; behavioural measurement is thin
3: Eval suiteGround truth and regression checks run before release
4: HarnessAll five control-plane jobs have implementations and owners
5: DisciplineVersioning, trace-to-eval learning, migration, and retirement run on a cadence

Do not average the company into one score. Place the selected workflow.

Write one paragraph:

We are at Rung [N] because [evidence]. The next rung requires [capability]. The smallest move is [ticket]. We will know it worked when [metric], while [counter-metric] remains acceptable.

For the reconciliation workflow:

We are at Rung 2. Model calls are traced and invalid output is retried, but there is no ground-truth test for full-queue completion. The next rung requires a workflow eval in CI. The smallest move is Ticket 2. It works when the 847-of-2,347 pagination failure blocks release without increasing false escalation beyond the agreed limit.

Artifact: One evidence-based paragraph the team can read aloud without explanation.

06Days 7 to 9: tickets, calibration, ownership

Day 7: Write four tickets

QuestionWhich changes are small enough to own and specific enough to measure?

Compress the audit into four backlog tickets:

  1. Enforce a structured contract on the highest-volume model output.
  2. Add a ground-truth eval for the highest-consequence journey.
  3. Trace the call the on-call team cannot explain.
  4. Complete one trace-to-eval-to-change cycle.

Every ticket needs:

Required fieldWhy it matters
Failure and evidencePrevents a generic solution looking for a problem
One workflow and boundaryKeeps the change attributable
OwnerCreates a person or team able to make the decision
Baseline and targetShows whether behaviour improved
Counter-metricDetects the harm introduced by the fix
Economic variableConnects the ticket to Build, Run, Review, or Failure cost
Rollout and rollbackMakes the change reversible
Regression caseKeeps the failure fixed after the sprint

Artifact: Four tickets in the real backlog, not a strategy deck.

Day 8: Calibrate evaluation

QuestionCan the team turn production evidence into a release decision?

LangChain’s 2026 survey of 1,340 respondents reported that 89% had some agent observability, while 52.4% ran offline evals and 37.3% ran online evals.[1] Many teams can watch a run without turning what they observe into a release test.

Start small:

  1. Review 20 to 50 real traces with a domain expert.
  2. Define success for one task.
  3. Write five eval cases from Day 5’s gaps.
  4. Run them against the current system and record the baseline.
  5. Put one regression case in CI.
  6. Attach the judgement to the trace or journey it evaluates.
  7. Name an owner for the evaluator and rubric.

A trace records what happened. Feedback records what that behaviour meant.

Do not treat a model judge as ground truth

Use code when code can decide. Use a model judge when meaning requires interpretation. Use a qualified person when the decision is consequential or the judge is not calibrated.

MobileJudgeBench, published on 11 August 2026, evaluated 30 judge variants on 931 human-labelled mobile-agent trajectories across six benchmarks, four agent models, and 68 Android apps. Accuracy ranged from 76.4% to 90.9%, depending on judge method and model backbone.[5]

The dataset itself required judgement: nine graduate annotators labelled each trajectory, each case received two to four labels, pairwise agreement was 88.4%, and disagreements were resolved through discussion.[5]

The study does not establish a universal 9% to 24% judge error rate for every agent product. It establishes a more useful operating rule: judge quality varies materially, and more elaborate judge pipelines do not automatically perform better.

The failure pattern matters as much as the average score. In the study, some backends were conservative and missed successful work. Another was permissive and accepted incomplete or constraint-violating trajectories. For a release gate, false acceptance may matter more than false rejection. For a customer-experience score, the trade-off may reverse.

Calibrate the judge against labelled examples from your workflow. Track false acceptance and false rejection separately. Recalibrate when the agent, rubric, judge model, or traffic changes.

Artifact: Five evals, one baseline, one CI gate, a labelled calibration set, and a named eval owner.

Day 9: Audit ownership, portability, and measurement

QuestionWhich product judgement remains available if a vendor changes, a model retires, or the measurement stack is wrong?

Run the ownership question across the five harness clusters:

ClusterOwnership question
IdentityCan we export and version the role, instructions, and operating rules?
Memory PolicyCan we inspect retrieval, context assembly, caching, and compaction?
OrchestrationCan we change models or runtimes without rewriting the workflow?
InterceptionCan we control validation, approval, and completion checks?
Observability and EvalsCan we export traces, datasets, rubrics, evaluator versions, and results?

Score each row:

Add four continuity checks:

  1. Has the fallback model passed the same workflow eval?
  2. Can unfinished work be exported and resumed?
  3. Can the team reproduce cost from provider usage records?
  4. Is the measurement stack itself tested?

The fourth check is not theoretical. Arize’s August 2026 release notes include fixes for session-average scores that had counted the same evaluation once per span, and for cached-token costs that had not used the correct cache rate.[6]

The lesson is not that one observability vendor is unreliable. It is that telemetry is software. Version it, test it against hand-calculated cases, and investigate discontinuities after upgrades.

Artifact: Portability and measurement table with one accepted dependency, one migration risk, and one instrumentation test.

Connection: Episode 05’s rule was to rent operations while owning meaning. Day 9 tests whether your evidence and economics remain portable.

07The eight decisions

The kit is not a documentation exercise. It writes down eight decisions every agent system already makes.

Figure 02 · Framework
The eight decisions
EIGHT DECISIONS · WRITE THEM DOWN OR INHERIT THEM Every system already answers these. Most answer them by accident. 01 JOB What exactly is this agent responsible for, and where does it stop? 02 ACCEPTANCE What evidence proves a result is acceptable? 03 CONTEXT What is the smallest complete context for each step? 04 TOOLS Which capabilities are exposed, and which change state? 05 AUTHORITY Which identity acts, under what limits, with what approvals? 06 RECOVERY What survives a restart, and what expires? 07 VERIFICATION Which checks are automatic, and which need a person? 08 CHANGE How does the harness change safely, and how is it reversed? One page, eight answers. If a decision has no named owner, the default belongs to whoever changed the system last. THE FIRST MONDAY Morning · job/stop/evidence  ·  Midday · tools  ·  Afternoon · authority  ·  EOD · first ticket
Use it One page, eight answers. If a decision has no named owner, the default belongs to whoever changed the system last.

The first Monday:

TimeWorkOutput
MorningWrite the job, stop condition, and completion evidenceOne agreed paragraph
MiddayList tools and mark which change stateTool table with owners
AfternoonName identity, limits, and approvalsAuthority line for each consequential tool
End of daySelect the first verification to automateOne ticket with baseline and counter-metric

08The four tickets

Ticket 1: Enforce a structured contract

Failure: A high-volume model output repeatedly reaches another system in an unusable or ambiguous shape.

Change: Define required fields and allowed values, reject unagreed fields, add semantic checks, version the contract, define refusal behaviour, and release behind a flag.

Done when: Shape failures fall, semantic correctness holds, and empty or refused outcomes remain within the agreed limit.

Economic variable: Accepted completion and recovery cost.

Counter-metric: Empty-result and escalation rate.

Ticket 2: Add a ground-truth workflow eval

Failure: The team discovers important regressions through users or on-call.

Change: Select successful, borderline, and failed cases; record approved outcomes and evidence; include a full-workflow completion check; establish a baseline; add the regression to CI.

Done when: A known production failure reliably blocks release.

Economic variable: Consequential failure rate and expected loss.

Counter-metric: Acceptable work incorrectly blocked.

Ticket 3: Trace the unexplained call

Failure: On-call receives an incident without enough evidence to locate the first divergence.

Change: Record trace ID, redacted context with provenance, model and harness versions, tool arguments, authority decisions, retries, state changes, latency, cost, and completion evidence.

Done when: On-call identifies the first divergence within an agreed time and closes one real incident using the trace.

Economic variable: Diagnosis and recovery cost.

Counter-metric: Storage cost and sensitive-data exposure.

Ticket 4: Complete one feedback cycle

Failure: Traces and evals exist, but production evidence does not change the released system.

Change: Label one failure, add representative cases, reserve holdouts, change one harness surface, test the change and counter-metrics, release with version and rollback, and record what changed.

Done when: One production failure has moved from trace to regression case to validated release.

Economic variable: Time from failure to prevention.

Counter-metric: Regression in another behaviour class.

Figure 03 · Practice
Four tickets on the map
EVERY TICKET NAMES A FAILURE, A LAYER, A METRIC, AND AN OWNER TICKET HARNESS CLUSTER ECONOMIC VARIABLE OWNER 01 Structured contract Highest-volume output boundary Interception Accepted completion, recovery cost Counter: empty-result rate Product and engineering 02 Ground-truth eval Highest-consequence journey Observability and evals Expected failure loss Counter: false blocks Domain and eval owner 03 Trace coverage The call on-call cannot explain Memory policy and runtime Diagnosis and recovery cost Counter: storage and exposure Platform 04 Feedback cycle Most useful failure from the audit Relevant cluster Time from failure to prevention Counter: regression elsewhere Cross-functional owner
Read it as See it, judge it, prove it, compound it. That is the order. A ticket without a failure is a generic improvement. A ticket without an owner will not change behaviour. A ticket without a counter-metric may create the next incident.

09Why these four come first

TicketSeason connectionGap it closes
Structured contractEpisode 03Narrows invalid output at a high-volume boundary
Ground-truth evalEpisodes 01 and 05Defines good and prices consequential failure
Trace coverageEpisode 02Makes the lifecycle visible and replayable
Feedback cycleEpisodes 03 and 08Makes learning repeatable

The order is logical rather than rigid.

A team cannot improve a failure it cannot see. It cannot judge a trace without a definition of good. It cannot prove a repair without an eval. It cannot retain the improvement until the case joins the release process.

If the audit reveals a critical authority gap, containment moves ahead of this order. Safety and incident response outrank the teaching sequence.

10The twelve-week runway

The nine days produce the diagnosis. The next twelve weeks spread implementation so changes remain attributable.

WeekFocusEvidence produced
1IdentityVersioned role, objective, stop rule, and operating policy
2Memory PolicyMeasurement of what enters, leaves, persists, and expires
3Tools and IdentityRenamed, narrowed, and classified tool contracts with authority
4ObservabilityFirst complete traces, baseline, and CI eval
5OrchestrationRouting or retry rule tested against the workflow suite
6InterceptionOne enforced approval, validation, or completion checkpoint
7Tools and IdentityProgressive disclosure with stable authority boundaries
8PortabilityOwnership audit, state export test, and fallback-model eval
9FeedbackOne production failure promoted into the eval set
10IdentityOne vague rule replaced by a versioned boundary or example
11ModelSame workflow suite run on a second model
12DisciplineFull feedback cycle reviewed, released, and recorded

Do not run the calendar mechanically. If Day 2 finds an uncontained authority failure, Week 6 becomes Week 1. If Day 6 finds no usable traces, observability moves ahead of model comparison.

The runway is a default order, not a substitute for diagnosis.

11The stakeholder translations

The plan remains the same. Only the reason for caring changes.

Engineering

The harness already exists across prompts, services, vendor settings, APIs, and manual workarounds. The audit makes it visible and gives on-call enough evidence to stop diagnosing every incident from scratch.

Design

Users experience the harness as consistency. State, retries, completion, refusal, and escalation determine whether the same request feels dependable or arbitrary.

Leadership and finance

The plan targets accepted completion, review load, failure loss, recovery time, and the evidence required to delegate more volume. It begins with four bounded tickets rather than a platform programme.

Security and compliance

The audit records context sources, identities, requested actions, authority decisions, actual effects, and retained evidence. It separates prompt instructions from controls that remain enforceable when the model is wrong.

The translation must not change the project. If each stakeholder hears a different plan, alignment will disappear during delivery.

12Connecting the dots

This kit is product discovery for a system whose behaviour varies from run to run.

Traditional discovery asks what users need and whether the product creates value. Harness discovery adds a production question: which hidden decision prevents that value from surviving real tools, state, failure, and scale?

The nine days follow the same logic as the season:

  1. Observe the real business outcome.
  2. Locate the first divergence.
  3. Name the responsible layer and decision.
  4. Convert the failure into a repeatable test.
  5. Change the smallest attributable surface.
  6. Price the result at the workflow level.
  7. Assign an owner who can keep it working.

That is why the kit begins with traces rather than a technology choice.

The audit will expose an organisational problem. Several rows will have implementations but no clear decision owner. Episode 07 is about that gap.

13In practice: the one-page output

At the end of Day 9, publish one page:

SectionContent
WorkflowOne sentence defining the business outcome and stop condition
Current rungEvidence-based maturity diagnosis for this workflow
Top failuresThree shapes, prevalence estimate if available, representative traces
Hidden costBuild, Run, Review, and Failure cost discovered during the audit
Missing decisionsClusters or boundaries with no implementation, evidence, or owner
Four ticketsOwner, baseline, target, counter-metric, sprint, rollback
Economic gateEvidence required to expand delegated volume or autonomy
Portability riskOne accepted dependency and one risk to reduce
Measurement checkOne hand-calculated case proving the dashboard is correct

Take that page to sprint planning. Do not take a 40-slide transformation deck.

The kit produces owners as often as it produces code. The prompt may be edited by product, retry logic owned by engineering, permissions set by platform, and eval thresholds maintained in a spreadsheet no team officially owns.

You now hold
Practice 06 A nine-day diagnostic, seven artifacts, five maturity rungs, four owned tickets, eight written decisions, and a twelve-week runway that keeps changes attributable.
The next question
Who has decision rights over the behaviour produced by all four layers?
Continue
Harness 07 Organizations That Rebuilt Around Agents → — the ownership map, three operating models, and the Harness PM’s decision rights.
Read alongside
Environment 00 The World Around the Agent → — the operate-side audit Day 9 deliberately leaves out.
Sources
  1. LangChain, “State of Agent Engineering,” 12 June 2026; survey of 1,340 respondents conducted from 18 November to 2 December 2025.
    langchain.com/state-of-agent-engineering
  2. LangChain, “Agent Evaluation Readiness Checklist.”
    langchain.com/blog/agent-evaluation-readiness-checklist
  3. LangChain, “Evaluating AI agents at the run, trace, and thread level.”
    langchain.com/resources/agent-evals
  4. LangChain, “Agent observability needs feedback to power learning.”
    langchain.com/blog/agent-observability-needs-feedback-to-power-learning
  5. Ziqiang Wang et al., “Benchmarking LLM Judges for Mobile Agent Evaluation,” MobileJudgeBench, arXiv 2608.11434v1, 11 August 2026.
    arxiv.org/html/2608.11434v1
  6. Arize AX, release notes, 20 to 26 August 2026; session-average and cached-token cost corrections.
    arize.com/docs/ax/release-notes
  7. Chandana Asif, Dieter Kiewell, Lari Hämäläinen, Tom Kolaja, and Tunde Olanrewaju, McKinsey/QuantumBlack, “Where AI agents pay off,” 24 August 2026.
    mckinsey.com/.../where-ai-agents-pay-off