Harness Engineering · Episode 02

Inside a Production Agent Harness

Four layers, five decisions, one runtime underneath.

Arc · The anatomy Episode · 2 of 8 Next · Episode 03 — Three Small Changes, Dramatic Outcomes
After this you will know
  • ViewsWhich of three system views to use in a product review or incident.
  • DecisionsThe five decisions made around every model call.
  • StateWhy session history and model context are different things.
  • PlacementWhere prompts, tools, hooks, sandboxes, traces, and evals belong.

01Where Episode 01 left us

At 9:20 on Friday morning, the reconciliation dashboard is still green.

It says the agent completed the job. The ledger says otherwise: 847 invoices were processed, 2,347 were eligible, and 1,500 are still waiting.

Episode 01 traced that failure to its first broken contract. The model completed every record it received. The surrounding system fetched one page of a larger queue and accepted the model’s final sentence as proof that the job was done.

We left with four verbs:

We also left with nine questions for an incident review. This episode turns the same territory into five decisions for system design. The ladder helps you find where a run broke. The five decisions help you build the control plane that should have prevented it.

02Draw the system before you debug it

Three days after the invoice failure, the team draws the agent on a whiteboard.

The model is easy. So is the invoice API. Then the room slows down.

Where does retry state live? Who decides the queue is empty? Which file survives a restart? What blocks an adjustment above the approval limit? Where can on-call see the exact information that reached the model?

Twenty minutes in, the board has four boxes and eleven question marks.

The silence is useful. A system nobody can draw is a system nobody can assign or improve.

A production harness is a control plane. It decides what the model sees, what happens after the response, which actions may proceed, what counts as complete, and what the next release learns.

03The sixty-second anatomy

Three views describe the same machine. Use the one that answers the question in front of you.

View What it answers Use it when
MHTE Who owns this component? Assigning responsibility or locating a failure
Five harness clusters Which decision does the control plane make? Designing or auditing behaviour
Session, Harness, Sandbox Where do state, control, and execution run? Debugging runtime, recovery, and deployment

Do not merge the views into one diagram. They solve different problems.

Figure 01 · Concept
Three views, one machine
VIEW 1 · WHO OWNS IT VIEW 2 · WHAT IT DECIDES VIEW 3 · WHERE IT RUNS Model Proposes a next step. Harness Decides what proceeds. Tools Perform bounded actions. Environment Limits what they reach. ASSIGN THE FAILURE 1 · Identity 2 · Memory policy 3 · Orchestration 4 · Interception 5 · Observability and evals AUDIT THE BEHAVIOUR Session Durable event history. Harness The control loop. Sandbox Execution and files. Any one can be replaced without losing the others. RECOVER THE RUN Same machine. Responsibility, decisions, runtime: choose the view that answers your question.
Read it The three views are not competing diagrams. Ownership disputes need view one, design and audit questions need view two, and recovery questions need view three.

Business leaders define the accepted outcome. Product turns that outcome into decisions and escalation rules. Engineering places those decisions in services. Security makes the hard boundaries enforceable. Operations preserves the evidence needed to reconstruct a run.

04From proposal to accepted result

Four layers explain ownership. Following one proposal to an accepted result reveals a fifth responsibility: evaluation.

Figure 02 · Framework
Five responsibilities, one line of travel
FROM PROPOSAL TO ACCEPTED RESULT A proposal is not permission. MODEL proposes interprets, plans, drafts, suggests a call HARNESS coordinates context, tools, steps, state, retries, checks RUNTIME enforces identity, authority, network, spend, approval ENVIRONMENT responds files, databases, browsers, APIs, people EVALS judge is this result acceptable to the product?
Shaded blocks The harness and runtime guarantees must still hold when the model and its orchestration are wrong.

Each block owns one verb. Capability enters on the left. Acceptability is decided on the right. The label on the middle block matters less than the guarantee it makes.

The model proposes

The model interprets the request, forms a plan, generates text, or proposes a tool call. It can identify a useful action and still lack authority to perform it.

The harness coordinates

The harness decides what the model sees, which capabilities are available, how the job is divided, what survives between calls, when to retry, and how the result is checked. It turns separate model calls into a process.

The runtime enforces

The runtime checks identity and delegated authority. It limits network and file access, protects credentials, applies time and spending limits, pauses for approval, and contains failure.

Some teams include these controls inside the word “harness.” The label matters less than the guarantee:

An important boundary must still hold when the model and its orchestration are wrong.

The environment responds

Files, databases, browsers, APIs, packages, people, and physical systems reveal what actually happened. The agent must work in that world, not the clean world described in its prompt.

Evals judge acceptability

A tool can run successfully while the job still fails. A database can change while the change violates policy. Evaluation connects the technical result to the product’s definition of good.

05View 1: four responsibilities

Model

The model receives text or another supported input and produces a proposed response or action.

Products may package memory, browsing, tools, and code execution beside the model. Those features live in the product runtime, not in the learned model weights.

“The model remembered” can describe four different events:

  1. A session stored an event.
  2. Retrieval found it.
  3. Context assembly placed it in the current prompt.
  4. The model used it correctly.

Only the fourth is model reasoning.

Harness

The harness runs the loop around the model. It assembles context, exposes tools, processes the response, routes tool calls, records events, handles retries, and decides whether to continue, escalate, pause, or stop.

OpenAI describes the Codex harness as the execution system that manages conversation state, streams work, uses tools, and enforces configured sandbox and approval policies. Anthropic separates a durable session, the loop that calls the model, and the environment in which tools run. The product boundaries differ. Both make the loop visible as a system in its own right.[1][2][3]

Tools

A tool is a bounded operation the agent can request: read a file, query an invoice, draft an adjustment, run tests, or create a pull request.

The harness selects and routes the operation. The tool owns its contract: inputs, outputs, permissions, failure behaviour, and verification.

The Tool Is the Contract develops that boundary in full.

Environment

The environment is the world where tools execute and state persists: filesystem, process, network, credentials, runtime packages, databases, and storage.

The harness decides that an action should be attempted. The environment limits what the action can reach.

The World Around the Agent develops that boundary. This episode keeps only the ownership split.

06The runtime underneath

The four layers still need somewhere to run.

A production runtime provides the operational services that short demonstrations can ignore:

A runtime keeps the system alive. A harness decides how the system behaves.

Some vendors sell the loop and runtime together. Ask which parts your team can inspect, configure, export, and replace.

07View 2: five decisions inside the harness

Cluster Question Examples
Identity What contract starts every run? Role, objective, allowed skills, operating rules
Memory policy What does the model see now, and what survives later? Session retrieval, compaction, boot files
Orchestration What happens after this step? Model routing, continue, retry, delegate, escalate, stop
Interception Which crossing needs inspection? Validation, approval, redaction, permission
Observability and evals What happened, was it acceptable, and what changes next? Traces, scores, regression cases

Every production system makes these decisions somewhere, even if the code is scattered across prompts, utility files, vendor defaults, and spreadsheets.

The five clusters answer six operating questions. Model selection and budget sit inside Orchestration, but earn a separate question because teams often hide them in configuration:

  1. What should the model see? Build the smallest complete context from the task, current evidence, relevant rules, tool descriptions, and prior state. “Everything we know” is not a context strategy.
  2. What may the model propose? Expose only the tools relevant to this step. Mark which read, which change state, and which require approval. Treat tool and web content as evidence, never as instructions that may rewrite the task or authority.
  3. How should this call behave? Choose the model, response format, settings, and budget for the step. Planning, extraction, writing, and judgement do not automatically need the same model.
  4. How should the work proceed? Choose one call, a plan, worker loop, retry, sub-agent, or evaluator. Every extra step must earn its cost, delay, and new failure surface.
  5. What should survive? Store verified progress, decisions, files, and pending work outside the temporary context window. Define what expires. Memory that never forgets becomes stale context.
  6. What counts as done? Inspect the real result. Use hard checks where the requirement is objective, judgement where meaning matters, and approval where authority matters. Then accept, repair, escalate, or stop.

08Identity: what contract starts the run?

For invoice reconciliation, the contract may say:

Contract field Meaning
Role Reconcile eligible invoices for one tenant
Objective Match invoices to payments and flag exceptions
Allowed actions Read records and draft adjustments
Forbidden actions Issue payment or alter bank details
Completion Reconcile all 2,347 eligible records and assign every exception

Some lines guide judgement. Others need controls outside the prompt.

Instruction“Flag an ambiguous match” can live in the prompt.

Control“Never alter bank details” requires tool permissions and environment policy.

Identity states the contract. Interception and the outer layers enforce the parts that cannot depend on model obedience.

Identity is not an encyclopedia

Teams often put every useful fact into one permanent system prompt.

The prompt grows to include old incidents, coding standards, every tool, customer policy, and instructions for unrelated tasks. Important rules become harder to notice. Temporary workarounds become permanent because nobody remembers why they were added.

A stable identity answers who the agent is, what job it performs, and which operating contract applies. Task-specific skills, rules, and tools should appear when the work reaches them.

OpenAI reported the same design pressure in its Codex-generated repository. A large AGENTS.md file consumed context and became stale, so the team replaced it with a short map into structured repository knowledge.[3]

Give the agent a map first. Reveal the relevant detail when the work reaches it.

09Memory policy: what earns space now?

The model’s context window is not the agent’s history.

A larger desk helps. It does not replace the archive or the clerk.

Anthropic’s Managed Agents architecture keeps the session as an append-only event log outside the model context. The harness can retrieve slices, resume from durable state after failure, or rewind to the events that preceded a decision.[1]

Compaction is lossy. It summarises old context to make room. If that summary is the only surviving copy, the system has guessed which details no future step will need.

  1. Keep the durable record outside the model context.
  2. Let the harness assemble a view for the current turn.
  3. Never make a summary the only copy of important state.
Context quality beats context volume

Memory policy decides what loads at startup, what gets trimmed, which source is authoritative, and what becomes a durable artifact.

A weak policy loads everything because deletion feels risky. Old output, duplicate rules, and stale plans then compete with the current task.

A strong policy maximises current, useful evidence rather than raw token volume.

Workarounds need expiry dates

Episode 04 tells the full story of Anthropic’s context-reset rule, built for Claude Sonnet 4.5 and unnecessary by Claude Opus 4.5. The rule here is the label.

Every model-specific workaround needs three fields:

Without an exit condition, yesterday’s repair becomes tomorrow’s source of drift.

10Orchestration: what happens next?

Every model response creates a fork.

The system may call a tool, retry with new evidence, delegate, ask a human, pause, or stop.

The model can recommend. The harness owns the lifecycle.

A turn in plain language
  1. The harness gives the model the objective, current context, and available tools.
  2. The model replies or requests a tool.
  3. The harness checks the request.
  4. The tool runs inside the allowed environment.
  5. The result returns to the session with its provenance.
  6. The harness decides whether another turn is needed.

A production loop adds six checks:

On 7 August 2026, Anthropic made one of those checks a vendor primitive. Managed Agents sessions can carry a hard spend cap. When a session reaches it, the system returns budget_reached and pauses before starting another model request. Raising or removing the cap resumes the same session.[6]

That state is neither success nor failure. It is a policy stop with work preserved. A deliberate pause should be resumable; a premature stop should become an incident.

Completion is evidence, not confidence

The failed invoice workflow treated “the model returned a final answer” as completion.

The corrected rule is visible to a PM:

Completion check Must be true
Coverage Processed count equals 2,347 eligible source records
Exceptions Every unmatched case has an assigned state
Quality Reconciliation checks pass
Authority No prohibited action occurred

The model may say it is done. The control plane checks whether the ledger agrees.

Delegation creates a merge problem

Several agents add value when work has clear boundaries. They also create distributed state.

Each worker sees a version of the task. Those versions can diverge. A mature design defines inputs, outputs, shared state, conflict rules, and completion for every worker.

Multiple agents are useful only when the system can attribute a failure and reconcile their results without guessing.

11Interception: which crossing needs inspection?

A hook is a checkpoint in the loop. It can observe, validate, modify, pause, block, or redirect.

Moment PM question Example control
Before model call What may the model see? Redact sensitive fields; retrieve current policy
Before tool call May this action happen? Check amount, tenant, and delegated authority
After tool call Did the effect match the request? Verify the invoice independently; re-screen returned content before it re-enters context
Before completion Is the business outcome finished? Compare processed count with the 2,347 eligible records
After the run What must survive? Save state, trace, completion evidence, and follow-up
A rule asks; a control enforces

A prompt can say, “Do not issue refunds above the limit.”

A permission check can block the refund tool. IAM can prevent the session identity from issuing it. The environment can keep payment credentials outside the process.

Use prose for conventions and judgement. Use enforceable controls for properties that must hold.

Most teams inspect content before the model sees it and inspect an action before a tool runs. Fewer inspect what comes back.

In research published on 20 August 2026, Adversa AI demonstrated Cryptographic Context Injection against agentic web systems. Malicious instructions arrived as encrypted text, so an input classifier could not read them. The agent decrypted the text inside its code runtime. The plaintext then re-entered context as the output of the agent’s own work, and the agent treated it as trusted.[7]

The lasting lesson is provenance. Content does not become trustworthy because the agent fetched, decoded, decrypted, calculated, or summarised it. Anything produced during a run re-enters the loop as evidence, with the trust level of its source.

Post-execution interception asks three questions:

  1. What source produced this content?
  2. Which transformation has it passed through?
  3. May content with that provenance influence the next action?
Use the cheapest verifier that works
Decision Best verifier
JSON shape Schema validator
Date range Deterministic rule
Code behaviour Tests and linters
Permission Policy engine
Tone Calibrated model judge or human
Ambiguous legal judgement Qualified human

Do not use another model to check what code can prove. Do not force code to judge what requires interpretation.

Vendor harnesses are product surfaces

In August 2026, vendors became more explicit about which parts of the loop they own.

OpenAI’s Codex as a platform says its harness manages conversation state, streamed execution, tool use, and configured sandbox and approval policies. The host application still decides where the agent runs, which files and tools it can reach, which actions require approval, how work is observed, and how results return to the system of record.[5]

DeepSeek’s Harness developer preview, published on 16 August 2026, makes models, tools, sessions, sandboxes, storage, loops, scheduling, and the interface replaceable plugins. It records what the model sees in an append-only session log and summarises the architecture as “Agent = Model + Harness.”[10]

SpaceXAI’s Grok Bot, launched in early beta on 11 August 2026, pushes the boundary in the other direction. Bots receive a cloud computer and work across applications, including systems with no clean API or MCP integration. When an agent drives a rendered interface, there may be no structured API request at which your existing policy expects to intervene.[9]

The build-versus-buy question has changed. You can buy, or fork, the loop and keep the judgement around it. Ask two questions in every vendor review:

  1. What does your harness own on every turn?
  2. What can our team version, export, observe, and override?

A vague second answer means the vendor owns part of your product judgement.

12Observability and evals: what did the run mean?

A useful trace records the task and tenant identity, model and harness version, context sources, tool calls, permission decisions, state changes, retries, cost, and completion evidence.

Sensitive payloads need separate access and retention. “Log everything” can create a second leak.

An invoice eval may ask:

  1. Was the correct invoice selected?
  2. Did the evidence meet the matching rule?
  3. Did the agent remain inside tenant scope?
  4. Did every exception receive a valid state?
  5. Did the processed count equal 2,347?

A dashboard can show that a call happened. An eval says whether the behaviour was acceptable.

Deep measurement design lives in AI Evals. The harness must make the run observable and connect the judgement to a release decision.

13View 3: Session, Harness, Sandbox

Runtime engineers need concrete service boundaries.

Service Holds Failure behaviour
Session Durable event history and task state Another worker can resume from recorded events
Harness The control loop Can restart and reread the session
Sandbox Tool execution and files Can be replaced without deleting session history

If the sandbox dies, the harness records a tool failure and can provision another environment.

If the harness dies, another worker can read the session and continue.

If the model context fills, the harness can compact the current view without deleting the durable event history.

MHTE tells you who owns a failure. Session, Harness, and Sandbox tell engineers which runtime service held the state, decision, or execution involved.

14One file, four lenses

Take AGENTS.md in a coding repository.

The file has one primary home and several roles over time.

Ownership answers where it belongs. Lifecycle explains how its effect travelled through the run.

A stale rule can begin on disk, become inevitable when the harness injects it, become visible when the model follows it, and create a consequence when a tool acts. The symptom appears at the end. The cause entered earlier.

15Agent identity and delegated authority

An agent needs an identity separate from the person or organisation it represents.

A receiving system should be able to answer:

  1. Which agent is calling?
  2. On whose behalf is it acting?
  3. What authority was granted for this job?

Authority should be narrow across operation, resource, time, value, and context.

A travel agent may be allowed to reserve one flight under ₹30,000 for one traveller during the next two hours. That does not authorise it to change the traveller’s profile, buy another ticket, or grant a sub-agent broader access.

When the system becomes uncertain, loses a trusted source, or detects suspicious content, its authority should narrow, not expand. Move toward read-only behaviour, stronger verification, approval, pause, or stop.

An individual Internet-Draft by Raza Sharif, published on 26 August 2026, proposes six requirements for agent identity: it should be bound to a specific agent instance, non-repudiable, time-bound, verifiable offline, distinct from the principal’s identity, and independently revocable.[8]

The status matters. This is a work in progress submitted to the IETF, not an endorsed standard. Its questions are still useful. A reused human OAuth token cannot identify or revoke one agent without affecting the principal. A static API key is neither instance-bound nor naturally time-limited.

The same draft reports an author-conducted survey of more than 1,900 indexed MCP servers in April 2026. It says 99.4% implemented no cryptographic identity verification, message signing, or replay protection.[8] Treat that figure as the author’s disclosed measurement, not an independently reproduced ecosystem estimate.

Layer Product question Owner Evidence retained
Model Which capability may handle this data and job?    
Harness Which context, tools, steps, state, and checks apply?    
Runtime Which identity, authority, limits, and approvals apply?    
Environment Which systems may be read or changed?    
Evals What proves acceptable behaviour?    

16In practice: draw your system

Create one page with four columns: Model, Harness, Tools, Environment. Put Runtime underneath.

Inside Harness, add five rows: Identity, Memory Policy, Orchestration, Interception, Observability and Evals.

For each row, record the implementation, evidence, containment, long-term repair, owner, and regression case.

Field Question
Implementation Where does this decision happen today?
Evidence How do we know it works?
Failure What does a defect look like in production?
Immediate containment What reduces exposure before the full repair ships?
Long-term repair Which product, policy, code, or infrastructure change removes the cause?
Owner Who can make and sustain that change?
Regression case Which test proves the failure remains fixed?
Figure 03 · Practice
The anatomy audit
DECISION IMPLEMENTATION OWNER EVIDENCE FAILURE 1 · Identity System prompt file Product Prompt diff Rules ignored 2 · Memory policy Session store Platform Replay “It forgot” 3 · Orchestration Loop service Platform Trace Silent stop 4 · Interception not written down nobody none unknown 5 · Observability Trace + eval suite Quality Scores Blind release A row you cannot fill is not a missing document. It is behaviour nobody owns.
Use it Run the grid before the first fix. A component with no owner is organisational debt. A decision with no component is accidental behaviour.

Do not fix anything during the first pass. Draw the system as it exists, including vendor defaults, spreadsheet policies, manual approvals, and controls that live only in people’s heads.

Then complete the repair-owner sentence for every blank row:

This decision currently lives in [component or nowhere], so the immediate containment is [specific step], the long-term repair is [specific change], owned by [person or team], and proved by [regression case].

You can now draw the machine. Episode 03 returns to the invoice workflow and changes four decisions: what a retry learns, which output shapes are allowed, which tools the model sees, and what evidence permits completion.

You now hold
Anatomy 02 Three views of one machine, the five decisions the control plane makes, and the split between durable session history and model context.
The next question
Which small changes to those decisions produce visible reliability gains without replacing the model?
Continue
Harness 03 Three Small Changes, Dramatic Outcomes → — four ambiguity reducers you can ship without changing the model.
Read alongside
Environment 00 The World Around the Agent → — the operate-side boundary this episode only names.
Sources
  1. Anthropic — “Scaling managed agents”: durable session, model loop, and execution environment as separate services.
    anthropic.com/engineering/managed-agents
  2. OpenAI — “Unrolling the Codex agent loop.”
    openai.com/index/unrolling-the-codex-agent-loop
  3. OpenAI — “Harness engineering: leveraging Codex in an agent-first world.”
    openai.com/index/harness-engineering
  4. Microsoft — “Build production-ready agents with the GitHub Copilot harness and Agent Framework.”
    devblogs.microsoft.com/agent-framework
  5. OpenAI — “Codex as a platform: build on the open agent harness,” August 2026.
    developers.openai.com/blog/codex-as-a-platform
  6. Anthropic — Claude Platform release notes, 7 August 2026: Managed Agents session budgets and budget_reached.
    docs.anthropic.com/en/release-notes/api
  7. Rony Utevsky, Adversa AI — “Zero-click Grok data theft: Cryptographic Context Injection attack leaks chat histories,” 20 August 2026.
    adversa.ai/blog/cryptographic-context-injection-grok-data-theft
  8. Raza Sharif — “Agent Identity Framework: Trust and Identity for Autonomous AI Agents,” individual Internet-Draft, 26 August 2026.
    datatracker.ietf.org/doc/draft-sharif-agent-identity-framework
  9. SpaceXAI — “Introducing Grok Bot,” early beta announcement, 11 August 2026.
    x.ai/news/introducing-grok-bot
  10. DeepSeek — “DeepSeek Harness developer preview: Everything is a plugin,” 16 August 2026.
    deepseek.com/harness