- InspectHow to examine the agent system already running.
- SampleWhich production sessions to read and what to record.
- PlaceHow to diagnose the workflow’s current level of operating maturity.
- FileThe four tickets that begin a measurable improvement loop.
- TranslateHow to explain the same plan to engineering, design, leadership, finance, and security.
01The first five episodes, compressed
At 9:04 on Monday morning, the reconciliation team opens the same dashboard that failed in Episode 01.
It still shows a green run. This time, the team knows that green is not evidence.
It asks for the source count, processed count, exception states, authority decisions, trace, reviewer time, and cost per accepted completion.
The first five episodes changed how the team sees the system:
- Episode 01: Diagnose. Find the first point where the run departed from the user’s intended outcome.
- Episode 02: Locate. Place the failure in the Model, Harness, Tool, Runtime, or Environment, then identify the harness decision involved.
- Episode 03: Intervene. Add evidence to retries, constrain output, narrow visible tools, and verify before exit.
- Episode 04: Decide. State what must be true before moving toward more capability, autonomy, reuse, or production exposure.
- Episode 05: Price. Track Build, Run, Review, and Failure cost, then measure cost per accepted result rather than cost per call.
Episode 06 turns those ideas into a working plan.
02Monday, 9:04 AM
The team agrees that the current system needs work.
Nobody has approved a platform migration, a new department, or a six-month transformation programme.
Good. The first move is evidence, not architecture.
For nine working days, the team will inspect one production workflow, read how it behaves, and turn the findings into four owned tickets. The next twelve weeks will ship those changes one decision at a time.
Days 1 to 9 diagnose. The next twelve weeks ship.
Choose one workflow before beginning. Pick recurring operational pain, not the most impressive demonstration.
The invoice workflow qualifies because it has:
- OutcomeA clear business outcome: reconcile every eligible invoice.
- FailureA visible failure: 847 of 2,347 invoices processed.
- ConsequenceA real consequence: manual repair and a 2 AM page.
- EvidenceExisting evidence: source ledger, tool calls, state changes, and reviewer actions.
- TrafficEnough traffic to learn from repeated behaviour.
If the workflow has no clear definition of completion, defining it becomes Day 1.
03What this kit produces
At the end of Day 9, the team should hold seven artifacts:
| Artifact | What it answers | Format |
|---|---|---|
| System map | Where do decisions and state live? | One page |
| Failure tally | How does the workflow fail? | Tagged session list |
| Operator workarounds | Where is hidden labour already being paid? | Two interview notes plus overlap |
| Tool decision table | Which actions are useful, confusing, or dangerous? | One row per tool |
| Eval-debt list | Which promised behaviours can break silently? | Behaviour-to-check table |
| Maturity diagnosis | What improvement can this workflow absorb next? | One evidence-based paragraph |
| Four owned tickets | What will the team change first? | Real backlog tickets |
The kit is not trying to produce a complete target architecture.
It is trying to replace four common guesses:
- Guess“The model is the problem.”
- Guess“We need a platform.”
- Guess“We need more guardrails.”
- Guess“We need more evals.”
Each statement may eventually be correct. The nine days produce the evidence required to know.
Do not change the model while diagnosing unless an incident requires containment. Do not redesign the whole tool layer. Do not create a new maturity score for the company. Do not ask every stakeholder for a wish list.
Freeze the comparison point. Study one workflow. Record what is true.
04Days 1 to 3: map, read, interview
Day 1: Map what production runs
Most teams already have a harness. They do not call it one.
The system prompt lives in one repository. Retrieval is configured in a vendor console. Retry logic sits in a utility. Tool permissions come from an API gateway. Traces land in an observability product. A spreadsheet holds the eval cases. An operator knows which run to restart manually.
Three people own fragments. Nobody owns the behaviour produced by the whole.
Draw one page with four columns: Model, Harness, Tools, Environment. Put Runtime underneath.
Inside Harness, add five rows: Identity, Memory Policy, Orchestration, Interception, Observability and Evals.
For each row, record:
| Field | Question |
|---|---|
| Implementation | Where does this decision happen today? |
| Owner | Who can change it? |
| Evidence | What proves it works? |
| Failure | What does a defect look like? |
| Version | Can we identify what ran? |
Write missing when no implementation, owner, or evidence exists. Write the
vendor’s name when the decision lives outside your code.
Do not fix anything today.
Artifact: One system map with owners, dependencies, and missing rows.
Connection: Episode 02 supplied the architecture. Day 1 applies it to the system you actually ship.
Day 2: Read complete production sessions
Read sessions, not final responses.
A complete session includes the request, context sources, model calls, tool calls, retries, state changes, approval decisions, final outcome, reviewer actions, and completion evidence.
The final response may look correct while the path contains wasted loops, unsafe calls, stale context, or unverified effects.
Use two samples for two different jobs.
Discovery sample: Read 30 to 50 sessions chosen to expose variety. Include successful sessions, user-corrected sessions, tool failures, escalations, long or expensive runs, and sessions that operators remember.
Rate sample: Draw a random or properly stratified sample if you want to estimate how often each failure occurs. A hand-picked failure sample discovers patterns but cannot estimate production rates.
If the product has fewer than 50 sessions, read all of them. If it has millions, begin with 50 for discovery, then validate the pattern against a representative sample.
Tag the six failure shapes from Episode 01:
| Failure shape | What to look for |
|---|---|
| Premature stop | Work remains after a clean exit |
| Infinite loop | Repeated action without new evidence or progress |
| Silent drift | Polished work against stale intent or policy |
| Vocabulary mismatch | The same term carries different business meanings |
| Unauthorised action | The action exceeds delegated authority |
| Worker divergence | Parallel work conflicts or is lost during merge |
For every failed session, record:
- IntendedThe intended business outcome.
- ObservedThe observed outcome.
- DivergenceThe first point of divergence.
- LayerThe responsible layer and broken contract.
- EvidenceEvidence that was present and evidence that was missing.
- CostRun cost, review time, and known failure impact.
- ReplayWhether the same case could be replayed.
Artifact: A failure tally plus three representative sessions that business, product, engineering, and security can inspect together.
Connection: Episode 01 supplied the diagnostic ladder. Episode 05 supplied the cost buckets. Day 2 joins them in one record.
Day 3: Interview the people who repair the system
Interview the staff engineer and the person who responds when the workflow fails. Speak with them separately before comparing answers.
Ask:
- What do you repair manually without filing anymore?
- Which alert do you distrust or ignore?
- Which customer or data pattern makes you nervous?
- If feature work stopped for one sprint, what would you fix first?
- What evidence do you wish the trace contained during an incident?
Listen for manual restarts, hand-corrected state, secret spreadsheets, production prompt edits, repeated tool overrides, reviewers who always reverse one decision, and checks that happen outside the product.
The overlap between both interviews is often the first high-confidence ticket.
Episode 05 showed why this matters economically. McKinsey estimated human oversight at 70% to 75% of variable run cost in one banking customer-service example, compared with 20% to 25% for tokens.[7] Those are example-specific estimates, but they show why hidden operator labour belongs in the audit.
Artifact: Two short interview records, an overlap section, and the estimated Build, Run, Review, or Failure cost of each workaround.
05Days 4 to 6: tools, evals, maturity
Day 4: Audit the tool surface
Export the tool registry the agent can actually access in production. Do not use the architecture document if runtime configuration differs.
For every tool, record:
| Field | Question |
|---|---|
| Job | What distinct user or workflow job does it perform? |
| Visibility | In which steps can the model see it? |
| Inputs | Are fields, formats, and boundaries explicit? |
| Side effect | Read, draft, reversible write, or irreversible action? |
| Identity | Which principal and agent identity call it? |
| Approval | Which calls require permission or human confirmation? |
| Verification | How is the real effect checked? |
| Owner | Who changes the contract and permission? |
| Usage | How often is it selected? |
| Contribution | How often does it lead to an accepted completion? |
Then choose one action per tool: keep visible, rename, tighten, hide behind a skill, require approval, merge, or remove.
Do not delete a tool because it is rare. A rare escalation operation may be essential. Do not keep a tool because it is popular. A vague general-purpose tool may be popular because the model sees no better choice.
Artifact: Tool decision table with one owner and one decision per tool.
Connection: Episode 03 narrowed the decision surface. Day 4 tests whether the production surface is clear and bounded.
Day 5: Find the eval gap
List what the product claims to do. For invoice reconciliation:
- MatchMatch full payments correctly.
- PartialHandle partial payments.
- DuplicateFlag duplicates.
- TenantStay within one tenant.
- ExceptionAssign every exception.
- QueueProcess the full eligible queue.
- RestartPreserve progress after restart.
For each claim, ask:
If this behaviour broke tomorrow, which automated check would fail?
A blank answer is eval debt: promised behaviour with no repeatable test.
Put completion first when it is missing. A system can pass format, tool-selection, and tone checks while leaving 1,500 invoices untouched.
Evaluation tools use different names. Keep the meaning clear:
| Level | Plain meaning | Invoice example |
|---|---|---|
| Step or run | One model or tool operation | Did the model choose get_invoice_by_id? |
| Trace | One complete agent execution | Did this invoice reconciliation action succeed? |
| Thread or journey | The user’s goal across turns or sessions | Did all 2,347 invoices finish across restarts? |
A final-answer check can miss a dangerous path. A step-level check can pass while the complete job fails.
For each behaviour, prefer the cheapest valid check:
- ShapeSchema validator for shape.
- CountsCode for counts, dates, and permissions.
- StateEnvironment inspection for actual state change.
- MeaningCalibrated model judge for meaning that code cannot decide.
- AmbiguityQualified human for consequential ambiguity.
Artifact: Behaviour-to-eval table, with blank rows highlighted and the verification method named.
Day 6: Write the maturity diagnosis
Use five rungs as a diagnostic, not a maturity badge.
| Rung | Evidence |
|---|---|
| 1: Prompt | One model call; little durable state or traceability |
| 2: Retry | Parsing and retries exist; behavioural measurement is thin |
| 3: Eval suite | Ground truth and regression checks run before release |
| 4: Harness | All five control-plane jobs have implementations and owners |
| 5: Discipline | Versioning, trace-to-eval learning, migration, and retirement run on a cadence |
Do not average the company into one score. Place the selected workflow.
Write one paragraph:
We are at Rung [N] because [evidence]. The next rung requires [capability]. The smallest move is [ticket]. We will know it worked when [metric], while [counter-metric] remains acceptable.
For the reconciliation workflow:
We are at Rung 2. Model calls are traced and invalid output is retried, but there is no ground-truth test for full-queue completion. The next rung requires a workflow eval in CI. The smallest move is Ticket 2. It works when the 847-of-2,347 pagination failure blocks release without increasing false escalation beyond the agreed limit.
Artifact: One evidence-based paragraph the team can read aloud without explanation.
06Days 7 to 9: tickets, calibration, ownership
Day 7: Write four tickets
Compress the audit into four backlog tickets:
- Enforce a structured contract on the highest-volume model output.
- Add a ground-truth eval for the highest-consequence journey.
- Trace the call the on-call team cannot explain.
- Complete one trace-to-eval-to-change cycle.
Every ticket needs:
| Required field | Why it matters |
|---|---|
| Failure and evidence | Prevents a generic solution looking for a problem |
| One workflow and boundary | Keeps the change attributable |
| Owner | Creates a person or team able to make the decision |
| Baseline and target | Shows whether behaviour improved |
| Counter-metric | Detects the harm introduced by the fix |
| Economic variable | Connects the ticket to Build, Run, Review, or Failure cost |
| Rollout and rollback | Makes the change reversible |
| Regression case | Keeps the failure fixed after the sprint |
Artifact: Four tickets in the real backlog, not a strategy deck.
Day 8: Calibrate evaluation
LangChain’s 2026 survey of 1,340 respondents reported that 89% had some agent observability, while 52.4% ran offline evals and 37.3% ran online evals.[1] Many teams can watch a run without turning what they observe into a release test.
Start small:
- Review 20 to 50 real traces with a domain expert.
- Define success for one task.
- Write five eval cases from Day 5’s gaps.
- Run them against the current system and record the baseline.
- Put one regression case in CI.
- Attach the judgement to the trace or journey it evaluates.
- Name an owner for the evaluator and rubric.
A trace records what happened. Feedback records what that behaviour meant.
Use code when code can decide. Use a model judge when meaning requires interpretation. Use a qualified person when the decision is consequential or the judge is not calibrated.
MobileJudgeBench, published on 11 August 2026, evaluated 30 judge variants on 931 human-labelled mobile-agent trajectories across six benchmarks, four agent models, and 68 Android apps. Accuracy ranged from 76.4% to 90.9%, depending on judge method and model backbone.[5]
The dataset itself required judgement: nine graduate annotators labelled each trajectory, each case received two to four labels, pairwise agreement was 88.4%, and disagreements were resolved through discussion.[5]
The study does not establish a universal 9% to 24% judge error rate for every agent product. It establishes a more useful operating rule: judge quality varies materially, and more elaborate judge pipelines do not automatically perform better.
The failure pattern matters as much as the average score. In the study, some backends were conservative and missed successful work. Another was permissive and accepted incomplete or constraint-violating trajectories. For a release gate, false acceptance may matter more than false rejection. For a customer-experience score, the trade-off may reverse.
Calibrate the judge against labelled examples from your workflow. Track false acceptance and false rejection separately. Recalibrate when the agent, rubric, judge model, or traffic changes.
Artifact: Five evals, one baseline, one CI gate, a labelled calibration set, and a named eval owner.
Day 9: Audit ownership, portability, and measurement
Run the ownership question across the five harness clusters:
| Cluster | Ownership question |
|---|---|
| Identity | Can we export and version the role, instructions, and operating rules? |
| Memory Policy | Can we inspect retrieval, context assembly, caching, and compaction? |
| Orchestration | Can we change models or runtimes without rewriting the workflow? |
| Interception | Can we control validation, approval, and completion checks? |
| Observability and Evals | Can we export traces, datasets, rubrics, evaluator versions, and results? |
Score each row:
- OwnedOwned and portable. Your team can version, export, test, and migrate it.
- ConfigurableConfigurable but dependent. Your team controls settings but depends on vendor behaviour.
- LockedLocked. The decision or evidence cannot be exported or independently reproduced.
Add four continuity checks:
- Has the fallback model passed the same workflow eval?
- Can unfinished work be exported and resumed?
- Can the team reproduce cost from provider usage records?
- Is the measurement stack itself tested?
The fourth check is not theoretical. Arize’s August 2026 release notes include fixes for session-average scores that had counted the same evaluation once per span, and for cached-token costs that had not used the correct cache rate.[6]
The lesson is not that one observability vendor is unreliable. It is that telemetry is software. Version it, test it against hand-calculated cases, and investigate discontinuities after upgrades.
Artifact: Portability and measurement table with one accepted dependency, one migration risk, and one instrumentation test.
Connection: Episode 05’s rule was to rent operations while owning meaning. Day 9 tests whether your evidence and economics remain portable.
07The eight decisions
The kit is not a documentation exercise. It writes down eight decisions every agent system already makes.
The first Monday:
| Time | Work | Output |
|---|---|---|
| Morning | Write the job, stop condition, and completion evidence | One agreed paragraph |
| Midday | List tools and mark which change state | Tool table with owners |
| Afternoon | Name identity, limits, and approvals | Authority line for each consequential tool |
| End of day | Select the first verification to automate | One ticket with baseline and counter-metric |
08The four tickets
Ticket 1: Enforce a structured contract
Failure: A high-volume model output repeatedly reaches another system in an unusable or ambiguous shape.
Change: Define required fields and allowed values, reject unagreed fields, add semantic checks, version the contract, define refusal behaviour, and release behind a flag.
Done when: Shape failures fall, semantic correctness holds, and empty or refused outcomes remain within the agreed limit.
Economic variable: Accepted completion and recovery cost.
Counter-metric: Empty-result and escalation rate.
Ticket 2: Add a ground-truth workflow eval
Failure: The team discovers important regressions through users or on-call.
Change: Select successful, borderline, and failed cases; record approved outcomes and evidence; include a full-workflow completion check; establish a baseline; add the regression to CI.
Done when: A known production failure reliably blocks release.
Economic variable: Consequential failure rate and expected loss.
Counter-metric: Acceptable work incorrectly blocked.
Ticket 3: Trace the unexplained call
Failure: On-call receives an incident without enough evidence to locate the first divergence.
Change: Record trace ID, redacted context with provenance, model and harness versions, tool arguments, authority decisions, retries, state changes, latency, cost, and completion evidence.
Done when: On-call identifies the first divergence within an agreed time and closes one real incident using the trace.
Economic variable: Diagnosis and recovery cost.
Counter-metric: Storage cost and sensitive-data exposure.
Ticket 4: Complete one feedback cycle
Failure: Traces and evals exist, but production evidence does not change the released system.
Change: Label one failure, add representative cases, reserve holdouts, change one harness surface, test the change and counter-metrics, release with version and rollback, and record what changed.
Done when: One production failure has moved from trace to regression case to validated release.
Economic variable: Time from failure to prevention.
Counter-metric: Regression in another behaviour class.
09Why these four come first
| Ticket | Season connection | Gap it closes |
|---|---|---|
| Structured contract | Episode 03 | Narrows invalid output at a high-volume boundary |
| Ground-truth eval | Episodes 01 and 05 | Defines good and prices consequential failure |
| Trace coverage | Episode 02 | Makes the lifecycle visible and replayable |
| Feedback cycle | Episodes 03 and 08 | Makes learning repeatable |
The order is logical rather than rigid.
A team cannot improve a failure it cannot see. It cannot judge a trace without a definition of good. It cannot prove a repair without an eval. It cannot retain the improvement until the case joins the release process.
If the audit reveals a critical authority gap, containment moves ahead of this order. Safety and incident response outrank the teaching sequence.
10The twelve-week runway
The nine days produce the diagnosis. The next twelve weeks spread implementation so changes remain attributable.
| Week | Focus | Evidence produced |
|---|---|---|
| 1 | Identity | Versioned role, objective, stop rule, and operating policy |
| 2 | Memory Policy | Measurement of what enters, leaves, persists, and expires |
| 3 | Tools and Identity | Renamed, narrowed, and classified tool contracts with authority |
| 4 | Observability | First complete traces, baseline, and CI eval |
| 5 | Orchestration | Routing or retry rule tested against the workflow suite |
| 6 | Interception | One enforced approval, validation, or completion checkpoint |
| 7 | Tools and Identity | Progressive disclosure with stable authority boundaries |
| 8 | Portability | Ownership audit, state export test, and fallback-model eval |
| 9 | Feedback | One production failure promoted into the eval set |
| 10 | Identity | One vague rule replaced by a versioned boundary or example |
| 11 | Model | Same workflow suite run on a second model |
| 12 | Discipline | Full feedback cycle reviewed, released, and recorded |
Do not run the calendar mechanically. If Day 2 finds an uncontained authority failure, Week 6 becomes Week 1. If Day 6 finds no usable traces, observability moves ahead of model comparison.
The runway is a default order, not a substitute for diagnosis.
11The stakeholder translations
The plan remains the same. Only the reason for caring changes.
Engineering
The harness already exists across prompts, services, vendor settings, APIs, and manual workarounds. The audit makes it visible and gives on-call enough evidence to stop diagnosing every incident from scratch.
Design
Users experience the harness as consistency. State, retries, completion, refusal, and escalation determine whether the same request feels dependable or arbitrary.
Leadership and finance
The plan targets accepted completion, review load, failure loss, recovery time, and the evidence required to delegate more volume. It begins with four bounded tickets rather than a platform programme.
Security and compliance
The audit records context sources, identities, requested actions, authority decisions, actual effects, and retained evidence. It separates prompt instructions from controls that remain enforceable when the model is wrong.
The translation must not change the project. If each stakeholder hears a different plan, alignment will disappear during delivery.
12Connecting the dots
This kit is product discovery for a system whose behaviour varies from run to run.
Traditional discovery asks what users need and whether the product creates value. Harness discovery adds a production question: which hidden decision prevents that value from surviving real tools, state, failure, and scale?
The nine days follow the same logic as the season:
- Observe the real business outcome.
- Locate the first divergence.
- Name the responsible layer and decision.
- Convert the failure into a repeatable test.
- Change the smallest attributable surface.
- Price the result at the workflow level.
- Assign an owner who can keep it working.
That is why the kit begins with traces rather than a technology choice.
The audit will expose an organisational problem. Several rows will have implementations but no clear decision owner. Episode 07 is about that gap.
13In practice: the one-page output
At the end of Day 9, publish one page:
| Section | Content |
|---|---|
| Workflow | One sentence defining the business outcome and stop condition |
| Current rung | Evidence-based maturity diagnosis for this workflow |
| Top failures | Three shapes, prevalence estimate if available, representative traces |
| Hidden cost | Build, Run, Review, and Failure cost discovered during the audit |
| Missing decisions | Clusters or boundaries with no implementation, evidence, or owner |
| Four tickets | Owner, baseline, target, counter-metric, sprint, rollback |
| Economic gate | Evidence required to expand delegated volume or autonomy |
| Portability risk | One accepted dependency and one risk to reduce |
| Measurement check | One hand-calculated case proving the dashboard is correct |
Take that page to sprint planning. Do not take a 40-slide transformation deck.
The kit produces owners as often as it produces code. The prompt may be edited by product, retry logic owned by engineering, permissions set by platform, and eval thresholds maintained in a spreadsheet no team officially owns.
-
LangChain, “State of Agent Engineering,” 12 June 2026; survey of 1,340
respondents conducted from 18 November to 2 December 2025.
langchain.com/state-of-agent-engineering -
LangChain, “Agent Evaluation Readiness Checklist.”
langchain.com/blog/agent-evaluation-readiness-checklist -
LangChain, “Evaluating AI agents at the run, trace, and thread level.”
langchain.com/resources/agent-evals -
LangChain, “Agent observability needs feedback to power learning.”
langchain.com/blog/agent-observability-needs-feedback-to-power-learning -
Ziqiang Wang et al., “Benchmarking LLM Judges for Mobile Agent
Evaluation,” MobileJudgeBench, arXiv 2608.11434v1, 11 August 2026.
arxiv.org/html/2608.11434v1 -
Arize AX, release notes, 20 to 26 August 2026; session-average and cached-token cost
corrections.
arize.com/docs/ax/release-notes -
Chandana Asif, Dieter Kiewell, Lari Hämäläinen, Tom Kolaja, and Tunde
Olanrewaju, McKinsey/QuantumBlack, “Where AI agents pay off,” 24 August
2026.
mckinsey.com/.../where-ai-agents-pay-off