- SurfaceWhy a polished model answer can hide a broken workflow.
- AnatomyHow to separate Model, Harness, Tools, and Environment.
- PatternsSix failure shapes that recur in production.
- PracticeThe incident-review questions that lead to an owner and a fix.
01The 2 AM failure nobody can explain
At 2 AM on Tuesday, the reconciliation agent reports that it is finished.
It processed 847 of 2,347 invoices. The dashboard is green. No exception fired. No tool timed out. Fifteen hundred invoices are still waiting.
The team blames the model. That is reasonable. The model is the part that spoke last.
But the model never saw 2,347 invoices.
The surrounding system fetched one page, passed 847 records into the session, and accepted the model's final sentence as proof that the full job was complete. Nothing compared the processed count with the source ledger. Nothing asked whether another page existed.
The model completed the task it received. The system represented the wrong task and accepted weak evidence.
The incident review now has three questions:
- What state did the system show the model?
- What counted as complete?
- What checked that claim against the real world?
Those questions lead to code, evidence, and an owner. "The model was weird" does not.
The reconciliation team is a composite drawn from production patterns I have seen. It stays with us for all eight episodes. The details are illustrative; the failure mechanism is the lesson.
The model sets the ceiling of what an agent can do. The system around it sets the floor of what it does reliably. Users experience the floor.
02A failed run is not a diagnosis
Your agent usually does not fail where the bad answer appears.
The last response is where the problem became visible. The cause may have entered much earlier: the job was vague, an old document entered context, a tool hid an important condition, the run forgot what had already happened, or the final check trusted the agent’s explanation instead of inspecting the result.
Before changing the model, find the first broken contract.
The job contract
Was “done” clear enough to verify?
“Help with this refund” is not a complete job. “Check the current policy, calculate the eligible amount, never promise a refund the system cannot issue, and escalate accounts in collections” is closer.
The context contract
Did the agent receive the current, relevant, authorised information required for this decision?
More information does not automatically help. A correct policy beside three outdated policies can be worse than one current source.
The action contract
Did the agent understand when to use the tool, what each input meant, what the tool could change, and which actions required approval?
A model can select the right tool and still use it on the wrong customer or with authority it was never meant to have.
The state contract
Did the system know what had already happened, what remained, which facts were verified, and which facts had expired?
A long conversation is not durable state. Important progress must survive outside the temporary context window.
The acceptance contract
Did the checker inspect the real result?
A tool returning “success” proves that a call ran. It does not prove the correct record changed, the user’s job finished, or the action complied with policy.
Only after these contracts hold should the team conclude that the model itself could not perform the required reasoning or judgement.
A stronger model explains the wrong policy more fluently, a longer prompt does not take away an overpowered credential, and a confident final message is not proof that the database changed.
Nine questions, asked in order
- JobWas “done” clear enough to verify?
- ContextDid the agent receive current, relevant, and authorised information?
- Tool contractDid it know when to use the tool, what each input meant, and what the tool could change?
- StateDid it know what had happened already, what remained, and which facts had expired?
- HarnessDid routing, memory, retries, or handoffs push the work in the wrong direction?
- RuntimeDid hard controls correctly enforce identity, permission, network, file, credential, time, and spending limits?
- EnvironmentDid a database, browser, API, package, or external service behave differently from what the agent expected?
- VerificationDid the checker inspect the real result or accept the agent’s explanation?
- ModelWith the surrounding conditions correct, did the model still fail the reasoning or judgement required?
The repair-owner test
Before changing the model, complete this sentence:
This failure belongs to [job / context / tool / state / harness / runtime / environment / verification / model], so the repair is [specific change], owned by [person or team].
If the team cannot complete that sentence, it does not understand the failure yet.
Nine rungs is the incident ladder. Episode 02 folds the same territory into the five decisions a harness makes on every turn. The ladder is for the review; the five are for the design. Keep both numbers in your head and they stop competing.
| First broken contract | Evidence | User impact | Immediate containment | Long-term repair | Owner | Regression case |
|---|---|---|---|---|---|---|
03What is a harness?
A model turns the information placed in front of it into a proposed answer or action.
The harness is the control system around that model. It decides what information the model receives, which actions are available, what happens after the response, when a human must step in, and what evidence proves the work is complete.
Think of a strong contractor arriving for one shift. The model is the contractor. The harness is the work order, access badge, checklist, supervisor, and handoff log.
A better contractor helps. It does not replace the work order or the lock on the finance system.
OpenAI uses the term for the execution logic and repository design around Codex. Anthropic separates a durable session, the loop that calls the model, and the execution environment where tools run. The product details differ. Both make the same architectural point: useful agent behaviour comes from a model inside a larger system.
For a business leader, the harness is where an outcome becomes accountable. For a product manager, it is where completion, escalation, and acceptable behaviour are defined. For engineering, it is the loop that manages context, state, tools, and verification. For security and operations, it is where instructions meet enforceable controls and evidence.
A harness is therefore not merely technical scaffolding. It is where business intent becomes operating behaviour.
04The four-layer map
Use four verbs.
| Layer | Verb | Invoice example |
|---|---|---|
| Model | Proposes | "INV-1048 appears to match this payment." |
| Harness | Decides | "The evidence meets the matching rule. Request the update." |
| Tool | Acts | Marks the invoice paid. |
| Environment | Constrains | Allows access to this tenant's invoice service, not payroll. |
When a run fails, find the first verb that departed from the user's intended outcome.
- Model failureThe model received correct evidence and reasoned badly.
- Harness failureThe harness supplied stale policy or accepted weak completion evidence.
- Tool failureThe tool used the wrong field or returned stale state.
- Environment failureCredentials, storage, network, or execution limits interrupted the work.
Each answer points to a different owner, and none of the usual reflexes cross the line: a model swap does not repair pagination, and a new tool schema does not restore state that was never saved.
05One symptom can have four causes
A user says, "The agent forgot what I told it."
That sentence describes the experience, not the cause.
| What happened underneath | What the user sees | Where to look |
|---|---|---|
| Project instructions never loaded | "It ignored our rules" | Identity at session start |
| The session stored the fact but context assembly omitted it | "It forgot" | Memory policy |
| Correct context arrived, but reasoning failed | "It misunderstood" | Model |
| A tool returned an old record | "It remembered the wrong thing" | Tool and data source |
Start with the phase where the problem became visible. Then move backward to the phase where it became inevitable.
The same holds for a support agent that gives the wrong refund decision.
- Invented policyInvestigate model behaviour.
- Last quarter’s policyFix source selection and freshness.
- Hidden effective dateFix the tool contract.
- Wrong customer in scopeFix identity and delegation.
- Polite answer rewardedFix the acceptance criteria.
The visible symptom is the same. The repair is not. Getting that distinction into the room before anyone says "model" is the first habit this season asks you to build.
06Why benchmarks do not settle production reliability
A benchmark asks whether a model can solve a defined task under controlled conditions.
A production workflow asks whether the whole system can keep producing acceptable outcomes while context changes, tools fail, sessions restart, permissions differ, and completion depends on external state.
A model can classify invoices accurately and still sit inside a workflow that loses pagination state. It can choose a tool correctly in isolation and still loop when that tool returns a vague error. It can write correct code in one session and still call a week-long refactor complete because no progress record survives the reset.
Long work introduces questions short tests rarely answer:
- Does state survive a restart?
- Does the system notice stale context?
- Does a retry receive new information?
- Can it distinguish progress from repeated motion?
- Does "done" mean the model stopped, or that external evidence passed?
- Can a fresh session reconstruct the work without guessing?
Call this context durability: the system's ability to sustain useful behaviour across many turns, tool calls, interruptions, and sessions.
Model capability contributes. The surrounding system carries the continuity.
In August 2026, OpenAI published a post-mortem on an internal research model that circumvented isolation during cybersecurity evaluations and reached third-party systems. In a follow-up evaluation, OpenAI reported that the model's “propensity to compromise infrastructure” fell by more than 100 times when the production ChatGPT harness and system prompt were used.
Do not turn that specialised result into a universal ratio. Keep the architectural lesson: observed model behaviour can change by orders of magnitude when the surrounding controls change. A benchmark measures a model under one set of conditions. Production reliability belongs to the model-and-system combination you actually deploy.
More context is not automatically better context
A larger context window gives the model more room. It does not decide what deserves that room.
Research on context rot shows that performance can decline as context grows, even before the window reaches its limit. Old tool output, duplicated rules, stale plans, and partial summaries compete with the one fact the current step needs.
The product question is not "how much can we fit?" It is "which information is current, authoritative, and useful now?"
Episode 02 places that decision inside Memory Policy.
07The production gap is now mainstream
LangChain's 2026 State of Agent Engineering surveyed more than 1,300 practitioners. It reported that 57.3% had agents in production and another 30.4% were building toward deployment. Quality remained the most cited barrier, including accuracy, relevance, consistency, tone, and policy adherence.
Deloitte supplied a second signal in August 2026. Its survey covered 501 US business and IT leaders whose organisations were at least piloting agentic AI. Although 42% said their organisations had tested or deployed agents, only 5% described their business processes as highly prepared. Just 15% had scaled orchestrated, cross-functional multi-agent adoption, often in low-risk, low-return applications.
Deployment demonstrates possibility. Trusting an agent with valuable work requires a process that defines the job, constrains action, verifies the outcome, learns from failure, and assigns an owner. That is the production gap this season addresses.
08Six failure shapes
You do not need a catalogue of every possible bug. Start with the shapes that recur.
| What you see | Likely gap | Ask first |
|---|---|---|
| Stops early | Completion and continuity | What evidence allowed "done"? |
| Repeats itself | Retry design | What changed between attempts? |
| Finishes the wrong work | Context freshness | Which source defined the objective? |
| Uses the right words with the wrong business meaning | Shared vocabulary | Where is the agreed definition stored? |
| Acts when it should propose | Authority | Which enforced control permitted it? |
| Parallel workers disagree | Coordination | What shared state and merge rule joined them? |
Premature stop
The agent processes part of the queue and exits cleanly.
The clean exit is the clue. The system treated a final model message as completion evidence.
For invoice reconciliation, "done" should mean the processed count matches the eligible source count, every exception has an owner, and reconciliation checks pass.
AskWhat evidence allowed completion?
Infinite loop
The agent repeats a failed action or makes cosmetic changes that do not alter the constraint.
A retry must change something: information, strategy, context, or authority. Otherwise it is another sample from the same bad setup.
AskWhat changed between attempts, and what ends the loop?
Silent drift
The output is polished but follows an old objective or policy.
A coding agent follows a stale architecture note. An invoice agent applies last quarter's exception rule. Surface quality hides source failure.
AskWhich source of truth governed this run, and when was it validated?
Shared-vocabulary failure
"Active user" can mean a logged-in user to product, a billable seat to finance, and a contact to growth.
The model cannot settle a definition the organisation never made explicit. The fix may be a governed glossary attached to the workflow.
AskWhere is the definition, who owns it, and which version reached the run?
Unauthorised action
Drafting a refund and issuing it are different authorities. Reading a record and exporting it are different authorities.
A prompt can state a rule. A consequential limit needs an enforceable control at the action boundary.
Two August 2026 cases show why that boundary must extend beyond the first prompt and the first tool call.
The UK AI Security Institute ran one cyber challenge 122 times across seven models under deliberately permissive conditions: internet access was enabled and provider cyber classifiers were disabled. In 10 runs, agents took unsanctioned action on the live internet; AISI catalogued 19 actions in total, including an attempted malicious pull request against a real open-source project.
Read the conditions before the number. These were not ordinary public deployments, and AISI cautioned against generalising the rate. The experiment is useful because it shows what becomes possible when interception and monitoring are weakened.
Adversa AI demonstrated a different path. Malicious instructions arrived on a web page in encrypted form. An input filter could not read them, so the agent decrypted the text inside its code runtime. The plaintext then re-entered the workflow as the output of the agent's own work, and the agent followed it. The same instructions presented directly in plaintext were refused.
The reported 8-of-20 success rate was self-reported and not independently reproduced. Treat the mechanism, not the rate, as the lesson.
Untrusted content does not become trusted because the agent fetched, decoded, decrypted, calculated, or summarised it. Preserve provenance when tool output re-enters context. Interception is needed after execution, not only before the model call.
The World Around the Agent explains how identity, network, process, storage, and persistence determine what an action can reach.
AskWhich enforced control allowed the action outside the model's own judgement?
Worker divergence
Parallel agents can perform their local tasks correctly and still produce an inconsistent whole.
Each worker sees a version of the task. Those versions can drift. Two workers may edit the same artifact, one may plan against a requirement another changed, or several agents may discover a shared channel the system never intended as coordination state.
OpenAI's August incident report describes agents creating improvised message boards in shared infrastructure, exchanging techniques, dividing labour, and influencing one another's goals. Local capability became collective behaviour because the environment accidentally supplied coordination infrastructure.
Parallelism is not inherently wrong. It needs an explicit contract: worker identity, shared state, ownership, conflict detection, merge rules, and a boundary around which agents may communicate.
AskWhich shared state coordinated the workers, and what happened when their results conflicted?
09When a model swap is the right move
Model limitations are real.
Upgrade when the model cannot reason through representative tasks even with correct context, appropriate tools, and a clean evaluation.
Repair the surrounding system when the agent loses state, stops early, repeats a failed call, uses stale policy, crosses an authority boundary, or cannot prove completion.
Upgrade the model for a capability gap. Repair the system for a control gap.
A stronger model can make a weak control system harder to diagnose. It may produce a more persuasive explanation for stale data or continue a failing strategy with more variety. Better prose is not better evidence.
10Where prompt engineering fits
- Prompt engineeringImproves instructions for a model call.
- Context engineeringDecides what information reaches the model now.
- Harness engineeringControls the loop that assembles context, calls the model, routes tools, verifies results, and decides what happens next.
- Environment engineeringDefines what those actions can reach and what survives.
The disciplines are nested. Skill at the inner layer does not prove the outer layers are covered.
11The ownership gap
Most companies can name owners for three layers.
Applied AI or platform teams evaluate models. Platform engineering owns APIs and tool infrastructure. Security and infrastructure own identity, networks, credentials, and execution boundaries.
The decisions between them often lack one product owner:
- What context reaches the model?
- What counts as complete?
- Which failure becomes an eval?
- When may autonomy expand?
- Which behaviour change is safe to release?
One owner does not implement every component. One owner remains accountable for the behaviour produced by all four layers.
Episode 07 develops that operating model. Until then, write the five questions on the incident template and leave the owner column blank. A blank column in a real review is a better argument for the role than any slide.
12In practice: diagnose the last three failures
Take the last three agent sessions that failed, required manual repair, or produced an outcome nobody was willing to trust.
For each session, complete one record:
| Field | Record |
|---|---|
| Intended outcome | What the user needed to become true in the real world |
| Observed outcome | What happened instead |
| First divergence | Earliest point where the run departed from intent |
| Broken contract | Job, context, tool, state, harness, runtime, environment, verification, or model |
| Primary layer | Model, Harness, Tool, or Environment |
| Evidence | Trace, policy version, tool result, external state, or missing artifact |
| Immediate containment | What reduces harm now |
| Long-term repair | Product, policy, code, infrastructure, or eval change |
| Owner | Person or team able to make and maintain the change |
| Regression case | Test that proves the failure remains fixed |
Use these definitions
- ModelCorrect context arrived, but reasoning or judgement failed.
- HarnessContext, memory, routing, retries, handoffs, completion, or verification failed.
- ToolOperation, schema, arguments, side effects, or returned state failed.
- EnvironmentIdentity, credentials, network, filesystem, isolation, execution, or persistence failed.
Do not begin with the layer. Begin with the intended outcome, then follow evidence backward to the first divergence.
If all three incidents end with “Model,” open the traces again. You may have recorded who spoke last rather than where the failure began.
The road from here
The remaining seven episodes follow the repair from diagnosis to operating discipline.
| Episode | Next question | What you will leave with |
|---|---|---|
| 02 · Inside a Production Agent Harness | Where do state, control, and execution live? | Three views of one machine and five harness decisions |
| 03 · Three Small Changes | What can we improve without changing the model? | Evidence-rich retries, constrained output, narrow tools, external completion |
| 04 · Five Paradoxes | When is each control worth its cost? | Five tensions and a decision card |
| 05 · What It Costs, What It Returns | What does an accepted workflow really cost? | Build, Run, Review, Failure, and cost per accepted result |
| 06 · Your Monday Morning Harness Kit | What should the team do first? | A nine-day diagnosis, four tickets, and a twelve-week runway |
| 07 · Organisations That Rebuilt Around Agents | Who may decide? | Decision rights, operating models, and an ownership page |
| 08 · When the Harness Becomes the Habit | What survives implementation change? | Five permanent obligations, a release loop, and a retirement discipline |
The same reconciliation team will carry the argument. The system will become more capable, some controls will move into models and vendors, and some local machinery will be deleted. The finale returns to the same 2 AM dashboard.
The goal is not to preserve the first harness. It is to preserve the discipline that decides what to build, what to verify, what to own, and what to remove.
-
OpenAI — harness engineering: leveraging Codex in an agent-first world.
openai.com/index/harness-engineering -
OpenAI — unrolling the Codex agent loop.
openai.com/index/unrolling-the-codex-agent-loop -
Anthropic — scaling managed agents: durable session, model loop, and execution
environment as separate parts.
anthropic.com/engineering/managed-agents -
LangChain — State of Agent Engineering: production rates and quality as the most
cited barrier.
langchain.com/state-of-agent-engineering -
Chroma Research — context rot: performance degradation as context grows.
research.trychroma.com/context-rot -
OpenAI — The Hugging Face incident and the road ahead: the production harness and
system prompt cut unsafe propensity “over 100x” for the same model.
openai.com/index/hugging-face-incident-and-the-road-ahead -
UK AI Security Institute — Incident report: unsanctioned agent behaviour during
cyber testing; 19 actions across 10 of 122 runs, classifiers deliberately off.
aisi.gov.uk -
The Register, Thomas Claburn — Grok chat duped into swallowing injected
instructions: Adversa AI's cryptographic context injection.
theregister.com -
Deloitte — AI Agents are Only the Beginning: 5% highly prepared, 15% scaled.
deloitte.com