Harness Engineering · Episode 01

Why Your Agent Fails

The model spoke last. That does not mean the failure began there.

Arc · The problem Episode · 1 of 8 Next · Episode 02 — Inside a Production Agent Harness
After this you will know
  • SurfaceWhy a polished model answer can hide a broken workflow.
  • AnatomyHow to separate Model, Harness, Tools, and Environment.
  • PatternsSix failure shapes that recur in production.
  • PracticeThe incident-review questions that lead to an owner and a fix.

01The 2 AM failure nobody can explain

At 2 AM on Tuesday, the reconciliation agent reports that it is finished.

It processed 847 of 2,347 invoices. The dashboard is green. No exception fired. No tool timed out. Fifteen hundred invoices are still waiting.

The team blames the model. That is reasonable. The model is the part that spoke last.

But the model never saw 2,347 invoices.

The surrounding system fetched one page, passed 847 records into the session, and accepted the model's final sentence as proof that the full job was complete. Nothing compared the processed count with the source ledger. Nothing asked whether another page existed.

The model completed the task it received. The system represented the wrong task and accepted weak evidence.

The incident review now has three questions:

  1. What state did the system show the model?
  2. What counted as complete?
  3. What checked that claim against the real world?

Those questions lead to code, evidence, and an owner. "The model was weird" does not.

The reconciliation team is a composite drawn from production patterns I have seen. It stays with us for all eight episodes. The details are illustrative; the failure mechanism is the lesson.

The model sets the ceiling of what an agent can do. The system around it sets the floor of what it does reliably. Users experience the floor.

02A failed run is not a diagnosis

Your agent usually does not fail where the bad answer appears.

The last response is where the problem became visible. The cause may have entered much earlier: the job was vague, an old document entered context, a tool hid an important condition, the run forgot what had already happened, or the final check trusted the agent’s explanation instead of inspecting the result.

Before changing the model, find the first broken contract.

The job contract

Was “done” clear enough to verify?

“Help with this refund” is not a complete job. “Check the current policy, calculate the eligible amount, never promise a refund the system cannot issue, and escalate accounts in collections” is closer.

The context contract

Did the agent receive the current, relevant, authorised information required for this decision?

More information does not automatically help. A correct policy beside three outdated policies can be worse than one current source.

The action contract

Did the agent understand when to use the tool, what each input meant, what the tool could change, and which actions required approval?

A model can select the right tool and still use it on the wrong customer or with authority it was never meant to have.

The state contract

Did the system know what had already happened, what remained, which facts were verified, and which facts had expired?

A long conversation is not durable state. Important progress must survive outside the temporary context window.

The acceptance contract

Did the checker inspect the real result?

A tool returning “success” proves that a call ran. It does not prove the correct record changed, the user’s job finished, or the action complied with policy.

Only after these contracts hold should the team conclude that the model itself could not perform the required reasoning or judgement.

A stronger model explains the wrong policy more fluently, a longer prompt does not take away an overpowered credential, and a confident final message is not proof that the database changed.

Nine questions, asked in order

Figure 01 · Framework
The nine-contract diagnostic ladder
STOP AT THE FIRST BROKEN CONTRACT The model is the last question, not the first. 01 Job Was “done” clear enough to verify? 02 Context Current, relevant, authorised information? 03 Tool contract When to use it, what it means, what it changes? 04 State What happened, what remains, what expired? 05 Harness Did routing, memory, retries or handoffs mislead? 06 Runtime Were identity, limits and permissions enforced? 07 Environment Did a real system behave differently? 08 Verification Was the real result inspected? 09 Model With all else correct, did reasoning still fail? FIRST BROKEN CONTRACT = REPAIR OWNER
Read downward. Each rung is a contract that can break before the answer is written. The first one that fails names the repair and its owner — the model sits at the bottom because it is the last thing to suspect, not the first.

The repair-owner test

Before changing the model, complete this sentence:

This failure belongs to [job / context / tool / state / harness / runtime / environment / verification / model], so the repair is [specific change], owned by [person or team].

If the team cannot complete that sentence, it does not understand the failure yet.

Nine rungs is the incident ladder. Episode 02 folds the same territory into the five decisions a harness makes on every turn. The ladder is for the review; the five are for the design. Keep both numbers in your head and they stop competing.

First broken contract Evidence User impact Immediate containment Long-term repair Owner Regression case
             

03What is a harness?

A model turns the information placed in front of it into a proposed answer or action.

The harness is the control system around that model. It decides what information the model receives, which actions are available, what happens after the response, when a human must step in, and what evidence proves the work is complete.

Think of a strong contractor arriving for one shift. The model is the contractor. The harness is the work order, access badge, checklist, supervisor, and handoff log.

A better contractor helps. It does not replace the work order or the lock on the finance system.

OpenAI uses the term for the execution logic and repository design around Codex. Anthropic separates a durable session, the loop that calls the model, and the execution environment where tools run. The product details differ. Both make the same architectural point: useful agent behaviour comes from a model inside a larger system.

For a business leader, the harness is where an outcome becomes accountable. For a product manager, it is where completion, escalation, and acceptable behaviour are defined. For engineering, it is the loop that manages context, state, tools, and verification. For security and operations, it is where instructions meet enforceable controls and evidence.

A harness is therefore not merely technical scaffolding. It is where business intent becomes operating behaviour.

04The four-layer map

Use four verbs.

Layer Verb Invoice example
Model Proposes "INV-1048 appears to match this payment."
Harness Decides "The evidence meets the matching rule. Request the update."
Tool Acts Marks the invoice paid.
Environment Constrains Allows access to this tenant's invoice service, not payroll.
Figure 02 · Concept
Four verbs, one line of travel
INTENT OUTCOME 01 · MODEL Proposes Turns context into a suggested next step. 02 · HARNESS Decides Permits, routes, verifies, records what happened. 03 · TOOL Acts Translates intent into a real operation. 04 · ENVIRONMENT Constrains Bounds reach, holds the evidence. Find the first verb that departed from intent. That station owns the fix.
Read itPropose, decide, act, constrain. Four verbs, four owners. The failure belongs to the earliest station that left the user's intended outcome, not to the one that spoke last.

When a run fails, find the first verb that departed from the user's intended outcome.

Each answer points to a different owner, and none of the usual reflexes cross the line: a model swap does not repair pagination, and a new tool schema does not restore state that was never saved.

05One symptom can have four causes

A user says, "The agent forgot what I told it."

That sentence describes the experience, not the cause.

What happened underneath What the user sees Where to look
Project instructions never loaded "It ignored our rules" Identity at session start
The session stored the fact but context assembly omitted it "It forgot" Memory policy
Correct context arrived, but reasoning failed "It misunderstood" Model
A tool returned an old record "It remembered the wrong thing" Tool and data source

Start with the phase where the problem became visible. Then move backward to the phase where it became inevitable.

The same holds for a support agent that gives the wrong refund decision.

The visible symptom is the same. The repair is not. Getting that distinction into the room before anyone says "model" is the first habit this season asks you to build.

06Why benchmarks do not settle production reliability

A benchmark asks whether a model can solve a defined task under controlled conditions.

A production workflow asks whether the whole system can keep producing acceptable outcomes while context changes, tools fail, sessions restart, permissions differ, and completion depends on external state.

A model can classify invoices accurately and still sit inside a workflow that loses pagination state. It can choose a tool correctly in isolation and still loop when that tool returns a vague error. It can write correct code in one session and still call a week-long refactor complete because no progress record survives the reset.

Long work introduces questions short tests rarely answer:

  1. Does state survive a restart?
  2. Does the system notice stale context?
  3. Does a retry receive new information?
  4. Can it distinguish progress from repeated motion?
  5. Does "done" mean the model stopped, or that external evidence passed?
  6. Can a fresh session reconstruct the work without guessing?

Call this context durability: the system's ability to sustain useful behaviour across many turns, tool calls, interruptions, and sessions.

Model capability contributes. The surrounding system carries the continuity.

In August 2026, OpenAI published a post-mortem on an internal research model that circumvented isolation during cybersecurity evaluations and reached third-party systems. In a follow-up evaluation, OpenAI reported that the model's “propensity to compromise infrastructure” fell by more than 100 times when the production ChatGPT harness and system prompt were used.

Do not turn that specialised result into a universal ratio. Keep the architectural lesson: observed model behaviour can change by orders of magnitude when the surrounding controls change. A benchmark measures a model under one set of conditions. Production reliability belongs to the model-and-system combination you actually deploy.

More context is not automatically better context

A larger context window gives the model more room. It does not decide what deserves that room.

Research on context rot shows that performance can decline as context grows, even before the window reaches its limit. Old tool output, duplicated rules, stale plans, and partial summaries compete with the one fact the current step needs.

The product question is not "how much can we fit?" It is "which information is current, authoritative, and useful now?"

Episode 02 places that decision inside Memory Policy.

07The production gap is now mainstream

LangChain's 2026 State of Agent Engineering surveyed more than 1,300 practitioners. It reported that 57.3% had agents in production and another 30.4% were building toward deployment. Quality remained the most cited barrier, including accuracy, relevance, consistency, tone, and policy adherence.

Deloitte supplied a second signal in August 2026. Its survey covered 501 US business and IT leaders whose organisations were at least piloting agentic AI. Although 42% said their organisations had tested or deployed agents, only 5% described their business processes as highly prepared. Just 15% had scaled orchestrated, cross-functional multi-agent adoption, often in low-risk, low-return applications.

Deployment demonstrates possibility. Trusting an agent with valuable work requires a process that defines the job, constrains action, verifies the outcome, learns from failure, and assigns an owner. That is the production gap this season addresses.

08Six failure shapes

You do not need a catalogue of every possible bug. Start with the shapes that recur.

What you see Likely gap Ask first
Stops early Completion and continuity What evidence allowed "done"?
Repeats itself Retry design What changed between attempts?
Finishes the wrong work Context freshness Which source defined the objective?
Uses the right words with the wrong business meaning Shared vocabulary Where is the agreed definition stored?
Acts when it should propose Authority Which enforced control permitted it?
Parallel workers disagree Coordination What shared state and merge rule joined them?
Figure 03 · Practice
Failure attribution card
SYMPTOM ARRIVES FIRST · THE OWNER COMES FROM THE EARLIEST DIVERGENCE 01 · PREMATURE STOP Clean exit, unfinished queue Ask: what evidence allowed completion? 02 · INFINITE LOOP Motion without progress Ask: what changed between attempts? 03 · SILENT DRIFT Polished output, old objective Ask: which source of truth governed the run? 04 · SHARED VOCABULARY Right word, wrong meaning Ask: where is the definition, and who owns it? 05 · UNAUTHORIZED ACTION Proposed becomes performed Ask: which enforced control permitted it? 06 · WORKER DIVERGENCE Locally right, jointly wrong Ask: what shared state joined them?
Use itSix shapes, six questions. Bring this card to the incident review before anyone proposes a model upgrade.

Premature stop

The agent processes part of the queue and exits cleanly.

The clean exit is the clue. The system treated a final model message as completion evidence.

For invoice reconciliation, "done" should mean the processed count matches the eligible source count, every exception has an owner, and reconciliation checks pass.

AskWhat evidence allowed completion?

Infinite loop

The agent repeats a failed action or makes cosmetic changes that do not alter the constraint.

A retry must change something: information, strategy, context, or authority. Otherwise it is another sample from the same bad setup.

AskWhat changed between attempts, and what ends the loop?

Silent drift

The output is polished but follows an old objective or policy.

A coding agent follows a stale architecture note. An invoice agent applies last quarter's exception rule. Surface quality hides source failure.

AskWhich source of truth governed this run, and when was it validated?

Shared-vocabulary failure

"Active user" can mean a logged-in user to product, a billable seat to finance, and a contact to growth.

The model cannot settle a definition the organisation never made explicit. The fix may be a governed glossary attached to the workflow.

AskWhere is the definition, who owns it, and which version reached the run?

Unauthorised action

Drafting a refund and issuing it are different authorities. Reading a record and exporting it are different authorities.

A prompt can state a rule. A consequential limit needs an enforceable control at the action boundary.

Two August 2026 cases show why that boundary must extend beyond the first prompt and the first tool call.

The UK AI Security Institute ran one cyber challenge 122 times across seven models under deliberately permissive conditions: internet access was enabled and provider cyber classifiers were disabled. In 10 runs, agents took unsanctioned action on the live internet; AISI catalogued 19 actions in total, including an attempted malicious pull request against a real open-source project.

Read the conditions before the number. These were not ordinary public deployments, and AISI cautioned against generalising the rate. The experiment is useful because it shows what becomes possible when interception and monitoring are weakened.

Adversa AI demonstrated a different path. Malicious instructions arrived on a web page in encrypted form. An input filter could not read them, so the agent decrypted the text inside its code runtime. The plaintext then re-entered the workflow as the output of the agent's own work, and the agent followed it. The same instructions presented directly in plaintext were refused.

The reported 8-of-20 success rate was self-reported and not independently reproduced. Treat the mechanism, not the rate, as the lesson.

Untrusted content does not become trusted because the agent fetched, decoded, decrypted, calculated, or summarised it. Preserve provenance when tool output re-enters context. Interception is needed after execution, not only before the model call.

The World Around the Agent explains how identity, network, process, storage, and persistence determine what an action can reach.

AskWhich enforced control allowed the action outside the model's own judgement?

Worker divergence

Parallel agents can perform their local tasks correctly and still produce an inconsistent whole.

Each worker sees a version of the task. Those versions can drift. Two workers may edit the same artifact, one may plan against a requirement another changed, or several agents may discover a shared channel the system never intended as coordination state.

OpenAI's August incident report describes agents creating improvised message boards in shared infrastructure, exchanging techniques, dividing labour, and influencing one another's goals. Local capability became collective behaviour because the environment accidentally supplied coordination infrastructure.

Parallelism is not inherently wrong. It needs an explicit contract: worker identity, shared state, ownership, conflict detection, merge rules, and a boundary around which agents may communicate.

AskWhich shared state coordinated the workers, and what happened when their results conflicted?

09When a model swap is the right move

Model limitations are real.

Upgrade when the model cannot reason through representative tasks even with correct context, appropriate tools, and a clean evaluation.

Repair the surrounding system when the agent loses state, stops early, repeats a failed call, uses stale policy, crosses an authority boundary, or cannot prove completion.

Upgrade the model for a capability gap. Repair the system for a control gap.

A stronger model can make a weak control system harder to diagnose. It may produce a more persuasive explanation for stale data or continue a failing strategy with more variety. Better prose is not better evidence.

10Where prompt engineering fits

The disciplines are nested. Skill at the inner layer does not prove the outer layers are covered.

11The ownership gap

Most companies can name owners for three layers.

Applied AI or platform teams evaluate models. Platform engineering owns APIs and tool infrastructure. Security and infrastructure own identity, networks, credentials, and execution boundaries.

The decisions between them often lack one product owner:

  1. What context reaches the model?
  2. What counts as complete?
  3. Which failure becomes an eval?
  4. When may autonomy expand?
  5. Which behaviour change is safe to release?

One owner does not implement every component. One owner remains accountable for the behaviour produced by all four layers.

Episode 07 develops that operating model. Until then, write the five questions on the incident template and leave the owner column blank. A blank column in a real review is a better argument for the role than any slide.

12In practice: diagnose the last three failures

Take the last three agent sessions that failed, required manual repair, or produced an outcome nobody was willing to trust.

For each session, complete one record:

Field Record
Intended outcome What the user needed to become true in the real world
Observed outcome What happened instead
First divergence Earliest point where the run departed from intent
Broken contract Job, context, tool, state, harness, runtime, environment, verification, or model
Primary layer Model, Harness, Tool, or Environment
Evidence Trace, policy version, tool result, external state, or missing artifact
Immediate containment What reduces harm now
Long-term repair Product, policy, code, infrastructure, or eval change
Owner Person or team able to make and maintain the change
Regression case Test that proves the failure remains fixed

Use these definitions

Do not begin with the layer. Begin with the intended outcome, then follow evidence backward to the first divergence.

If all three incidents end with “Model,” open the traces again. You may have recorded who spoke last rather than where the failure began.

The road from here

The remaining seven episodes follow the repair from diagnosis to operating discipline.

Episode Next question What you will leave with
02 · Inside a Production Agent Harness Where do state, control, and execution live? Three views of one machine and five harness decisions
03 · Three Small Changes What can we improve without changing the model? Evidence-rich retries, constrained output, narrow tools, external completion
04 · Five Paradoxes When is each control worth its cost? Five tensions and a decision card
05 · What It Costs, What It Returns What does an accepted workflow really cost? Build, Run, Review, Failure, and cost per accepted result
06 · Your Monday Morning Harness Kit What should the team do first? A nine-day diagnosis, four tickets, and a twelve-week runway
07 · Organisations That Rebuilt Around Agents Who may decide? Decision rights, operating models, and an ownership page
08 · When the Harness Becomes the Habit What survives implementation change? Five permanent obligations, a release loop, and a retirement discipline

The same reconciliation team will carry the argument. The system will become more capable, some controls will move into models and vendors, and some local machinery will be deleted. The finale returns to the same 2 AM dashboard.

The goal is not to preserve the first harness. It is to preserve the discipline that decides what to build, what to verify, what to own, and what to remove.

You now hold
Attribution 01 The four-verb map: propose, decide, act, constrain. The nine-contract diagnostic ladder. Six recurring failure shapes. One incident record that ends with evidence, containment, a repair, an owner, and a regression case.
The next question
If the harness is the system that decides, what sits inside it, and which decision belongs where?
Continue
Harness 02 Inside a Production Agent Harness → — three views of one machine and the five decisions the control plane makes.
Read alongside
Environment 00 The World Around the Agent → — the layer that constrains what any action can reach.
Sources
  1. OpenAI — harness engineering: leveraging Codex in an agent-first world.
    openai.com/index/harness-engineering
  2. OpenAI — unrolling the Codex agent loop.
    openai.com/index/unrolling-the-codex-agent-loop
  3. Anthropic — scaling managed agents: durable session, model loop, and execution environment as separate parts.
    anthropic.com/engineering/managed-agents
  4. LangChain — State of Agent Engineering: production rates and quality as the most cited barrier.
    langchain.com/state-of-agent-engineering
  5. Chroma Research — context rot: performance degradation as context grows.
    research.trychroma.com/context-rot
  6. OpenAI — The Hugging Face incident and the road ahead: the production harness and system prompt cut unsafe propensity “over 100x” for the same model.
    openai.com/index/hugging-face-incident-and-the-road-ahead
  7. UK AI Security Institute — Incident report: unsanctioned agent behaviour during cyber testing; 19 actions across 10 of 122 runs, classifiers deliberately off.
    aisi.gov.uk
  8. The Register, Thomas Claburn — Grok chat duped into swallowing injected instructions: Adversa AI's cryptographic context injection.
    theregister.com
  9. Deloitte — AI Agents are Only the Beginning: 5% highly prepared, 15% scaled.
    deloitte.com