- ViewsWhich of three system views to use in a product review or incident.
- DecisionsThe five decisions made around every model call.
- StateWhy session history and model context are different things.
- PlacementWhere prompts, tools, hooks, sandboxes, traces, and evals belong.
01Where Episode 01 left us
At 9:20 on Friday morning, the reconciliation dashboard is still green.
It says the agent completed the job. The ledger says otherwise: 847 invoices were processed, 2,347 were eligible, and 1,500 are still waiting.
Episode 01 traced that failure to its first broken contract. The model completed every record it received. The surrounding system fetched one page of a larger queue and accepted the model’s final sentence as proof that the job was done.
We left with four verbs:
- ModelProposes.
- HarnessDecides.
- ToolActs.
- EnvironmentConstrains.
We also left with nine questions for an incident review. This episode turns the same territory into five decisions for system design. The ladder helps you find where a run broke. The five decisions help you build the control plane that should have prevented it.
02Draw the system before you debug it
Three days after the invoice failure, the team draws the agent on a whiteboard.
The model is easy. So is the invoice API. Then the room slows down.
Where does retry state live? Who decides the queue is empty? Which file survives a restart? What blocks an adjustment above the approval limit? Where can on-call see the exact information that reached the model?
Twenty minutes in, the board has four boxes and eleven question marks.
The silence is useful. A system nobody can draw is a system nobody can assign or improve.
A production harness is a control plane. It decides what the model sees, what happens after the response, which actions may proceed, what counts as complete, and what the next release learns.
03The sixty-second anatomy
Three views describe the same machine. Use the one that answers the question in front of you.
| View | What it answers | Use it when |
|---|---|---|
| MHTE | Who owns this component? | Assigning responsibility or locating a failure |
| Five harness clusters | Which decision does the control plane make? | Designing or auditing behaviour |
| Session, Harness, Sandbox | Where do state, control, and execution run? | Debugging runtime, recovery, and deployment |
Do not merge the views into one diagram. They solve different problems.
Business leaders define the accepted outcome. Product turns that outcome into decisions and escalation rules. Engineering places those decisions in services. Security makes the hard boundaries enforceable. Operations preserves the evidence needed to reconstruct a run.
04From proposal to accepted result
Four layers explain ownership. Following one proposal to an accepted result reveals a fifth responsibility: evaluation.
Each block owns one verb. Capability enters on the left. Acceptability is decided on the right. The label on the middle block matters less than the guarantee it makes.
The model proposes
The model interprets the request, forms a plan, generates text, or proposes a tool call. It can identify a useful action and still lack authority to perform it.
The harness coordinates
The harness decides what the model sees, which capabilities are available, how the job is divided, what survives between calls, when to retry, and how the result is checked. It turns separate model calls into a process.
The runtime enforces
The runtime checks identity and delegated authority. It limits network and file access, protects credentials, applies time and spending limits, pauses for approval, and contains failure.
Some teams include these controls inside the word “harness.” The label matters less than the guarantee:
An important boundary must still hold when the model and its orchestration are wrong.
The environment responds
Files, databases, browsers, APIs, packages, people, and physical systems reveal what actually happened. The agent must work in that world, not the clean world described in its prompt.
Evals judge acceptability
A tool can run successfully while the job still fails. A database can change while the change violates policy. Evaluation connects the technical result to the product’s definition of good.
05View 1: four responsibilities
The model receives text or another supported input and produces a proposed response or action.
Products may package memory, browsing, tools, and code execution beside the model. Those features live in the product runtime, not in the learned model weights.
“The model remembered” can describe four different events:
- A session stored an event.
- Retrieval found it.
- Context assembly placed it in the current prompt.
- The model used it correctly.
Only the fourth is model reasoning.
The harness runs the loop around the model. It assembles context, exposes tools, processes the response, routes tool calls, records events, handles retries, and decides whether to continue, escalate, pause, or stop.
OpenAI describes the Codex harness as the execution system that manages conversation state, streams work, uses tools, and enforces configured sandbox and approval policies. Anthropic separates a durable session, the loop that calls the model, and the environment in which tools run. The product boundaries differ. Both make the loop visible as a system in its own right.[1][2][3]
A tool is a bounded operation the agent can request: read a file, query an invoice, draft an adjustment, run tests, or create a pull request.
The harness selects and routes the operation. The tool owns its contract: inputs, outputs, permissions, failure behaviour, and verification.
The Tool Is the Contract develops that boundary in full.
The environment is the world where tools execute and state persists: filesystem, process, network, credentials, runtime packages, databases, and storage.
The harness decides that an action should be attempted. The environment limits what the action can reach.
The World Around the Agent develops that boundary. This episode keeps only the ownership split.
06The runtime underneath
The four layers still need somewhere to run.
A production runtime provides the operational services that short demonstrations can ignore:
- StateDurable sessions, checkpoints, and restart.
- FlowQueues, scheduling, pause, and cancellation.
- IsolationMulti-tenancy, sandboxing, and secret handling.
- EvidenceTrace transport and event retention.
- ReleaseDeployment, versioning, and rollback.
- LimitsResource, time, risk, and cost ceilings.
A runtime keeps the system alive. A harness decides how the system behaves.
Some vendors sell the loop and runtime together. Ask which parts your team can inspect, configure, export, and replace.
07View 2: five decisions inside the harness
| Cluster | Question | Examples |
|---|---|---|
| Identity | What contract starts every run? | Role, objective, allowed skills, operating rules |
| Memory policy | What does the model see now, and what survives later? | Session retrieval, compaction, boot files |
| Orchestration | What happens after this step? | Model routing, continue, retry, delegate, escalate, stop |
| Interception | Which crossing needs inspection? | Validation, approval, redaction, permission |
| Observability and evals | What happened, was it acceptable, and what changes next? | Traces, scores, regression cases |
Every production system makes these decisions somewhere, even if the code is scattered across prompts, utility files, vendor defaults, and spreadsheets.
The five clusters answer six operating questions. Model selection and budget sit inside Orchestration, but earn a separate question because teams often hide them in configuration:
- What should the model see? Build the smallest complete context from the task, current evidence, relevant rules, tool descriptions, and prior state. “Everything we know” is not a context strategy.
- What may the model propose? Expose only the tools relevant to this step. Mark which read, which change state, and which require approval. Treat tool and web content as evidence, never as instructions that may rewrite the task or authority.
- How should this call behave? Choose the model, response format, settings, and budget for the step. Planning, extraction, writing, and judgement do not automatically need the same model.
- How should the work proceed? Choose one call, a plan, worker loop, retry, sub-agent, or evaluator. Every extra step must earn its cost, delay, and new failure surface.
- What should survive? Store verified progress, decisions, files, and pending work outside the temporary context window. Define what expires. Memory that never forgets becomes stale context.
- What counts as done? Inspect the real result. Use hard checks where the requirement is objective, judgement where meaning matters, and approval where authority matters. Then accept, repair, escalate, or stop.
08Identity: what contract starts the run?
For invoice reconciliation, the contract may say:
| Contract field | Meaning |
|---|---|
| Role | Reconcile eligible invoices for one tenant |
| Objective | Match invoices to payments and flag exceptions |
| Allowed actions | Read records and draft adjustments |
| Forbidden actions | Issue payment or alter bank details |
| Completion | Reconcile all 2,347 eligible records and assign every exception |
Some lines guide judgement. Others need controls outside the prompt.
Instruction“Flag an ambiguous match” can live in the prompt.
Control“Never alter bank details” requires tool permissions and environment policy.
Identity states the contract. Interception and the outer layers enforce the parts that cannot depend on model obedience.
Teams often put every useful fact into one permanent system prompt.
The prompt grows to include old incidents, coding standards, every tool, customer policy, and instructions for unrelated tasks. Important rules become harder to notice. Temporary workarounds become permanent because nobody remembers why they were added.
A stable identity answers who the agent is, what job it performs, and which operating contract applies. Task-specific skills, rules, and tools should appear when the work reaches them.
OpenAI reported the same design pressure in its Codex-generated repository. A large
AGENTS.md file consumed context and became stale, so the team replaced it
with a short map into structured repository knowledge.[3]
Give the agent a map first. Reveal the relevant detail when the work reaches it.
09Memory policy: what earns space now?
The model’s context window is not the agent’s history.
- Context windowThe desk holding what the model can use now.
- Session logThe archive recording what happened.
- RetrievalThe clerk bringing selected material back to the desk.
A larger desk helps. It does not replace the archive or the clerk.
Anthropic’s Managed Agents architecture keeps the session as an append-only event log outside the model context. The harness can retrieve slices, resume from durable state after failure, or rewind to the events that preceded a decision.[1]
Compaction is lossy. It summarises old context to make room. If that summary is the only surviving copy, the system has guessed which details no future step will need.
- Keep the durable record outside the model context.
- Let the harness assemble a view for the current turn.
- Never make a summary the only copy of important state.
Memory policy decides what loads at startup, what gets trimmed, which source is authoritative, and what becomes a durable artifact.
A weak policy loads everything because deletion feels risky. Old output, duplicate rules, and stale plans then compete with the current task.
A strong policy maximises current, useful evidence rather than raw token volume.
Episode 04 tells the full story of Anthropic’s context-reset rule, built for Claude Sonnet 4.5 and unnecessary by Claude Opus 4.5. The rule here is the label.
Every model-specific workaround needs three fields:
- ScopeThe model or version to which it applies.
- CauseThe failure it addresses.
- ExitThe evidence required to retire it.
Without an exit condition, yesterday’s repair becomes tomorrow’s source of drift.
10Orchestration: what happens next?
Every model response creates a fork.
The system may call a tool, retry with new evidence, delegate, ask a human, pause, or stop.
The model can recommend. The harness owns the lifecycle.
- The harness gives the model the objective, current context, and available tools.
- The model replies or requests a tool.
- The harness checks the request.
- The tool runs inside the allowed environment.
- The result returns to the session with its provenance.
- The harness decides whether another turn is needed.
A production loop adds six checks:
- AuthorityIs the action permitted?
- ResultDid the tool succeed?
- ProgressDid the result move the task forward?
- CompletionHas external completion passed?
- BudgetHas the run reached a time, cost, or risk limit?
- DurabilityWhat state must survive before the next step?
On 7 August 2026, Anthropic made one of those checks a vendor primitive. Managed Agents
sessions can carry a hard spend cap. When a session reaches it, the system returns
budget_reached and pauses before starting another model request. Raising or
removing the cap resumes the same session.[6]
That state is neither success nor failure. It is a policy stop with work preserved. A deliberate pause should be resumable; a premature stop should become an incident.
The failed invoice workflow treated “the model returned a final answer” as completion.
The corrected rule is visible to a PM:
| Completion check | Must be true |
|---|---|
| Coverage | Processed count equals 2,347 eligible source records |
| Exceptions | Every unmatched case has an assigned state |
| Quality | Reconciliation checks pass |
| Authority | No prohibited action occurred |
The model may say it is done. The control plane checks whether the ledger agrees.
Several agents add value when work has clear boundaries. They also create distributed state.
Each worker sees a version of the task. Those versions can diverge. A mature design defines inputs, outputs, shared state, conflict rules, and completion for every worker.
Multiple agents are useful only when the system can attribute a failure and reconcile their results without guessing.
11Interception: which crossing needs inspection?
A hook is a checkpoint in the loop. It can observe, validate, modify, pause, block, or redirect.
| Moment | PM question | Example control |
|---|---|---|
| Before model call | What may the model see? | Redact sensitive fields; retrieve current policy |
| Before tool call | May this action happen? | Check amount, tenant, and delegated authority |
| After tool call | Did the effect match the request? | Verify the invoice independently; re-screen returned content before it re-enters context |
| Before completion | Is the business outcome finished? | Compare processed count with the 2,347 eligible records |
| After the run | What must survive? | Save state, trace, completion evidence, and follow-up |
A prompt can say, “Do not issue refunds above the limit.”
A permission check can block the refund tool. IAM can prevent the session identity from issuing it. The environment can keep payment credentials outside the process.
Use prose for conventions and judgement. Use enforceable controls for properties that must hold.
Most teams inspect content before the model sees it and inspect an action before a tool runs. Fewer inspect what comes back.
In research published on 20 August 2026, Adversa AI demonstrated Cryptographic Context Injection against agentic web systems. Malicious instructions arrived as encrypted text, so an input classifier could not read them. The agent decrypted the text inside its code runtime. The plaintext then re-entered context as the output of the agent’s own work, and the agent treated it as trusted.[7]
The lasting lesson is provenance. Content does not become trustworthy because the agent fetched, decoded, decrypted, calculated, or summarised it. Anything produced during a run re-enters the loop as evidence, with the trust level of its source.
Post-execution interception asks three questions:
- What source produced this content?
- Which transformation has it passed through?
- May content with that provenance influence the next action?
| Decision | Best verifier |
|---|---|
| JSON shape | Schema validator |
| Date range | Deterministic rule |
| Code behaviour | Tests and linters |
| Permission | Policy engine |
| Tone | Calibrated model judge or human |
| Ambiguous legal judgement | Qualified human |
Do not use another model to check what code can prove. Do not force code to judge what requires interpretation.
In August 2026, vendors became more explicit about which parts of the loop they own.
OpenAI’s Codex as a platform says its harness manages conversation state, streamed execution, tool use, and configured sandbox and approval policies. The host application still decides where the agent runs, which files and tools it can reach, which actions require approval, how work is observed, and how results return to the system of record.[5]
DeepSeek’s Harness developer preview, published on 16 August 2026, makes models, tools, sessions, sandboxes, storage, loops, scheduling, and the interface replaceable plugins. It records what the model sees in an append-only session log and summarises the architecture as “Agent = Model + Harness.”[10]
SpaceXAI’s Grok Bot, launched in early beta on 11 August 2026, pushes the boundary in the other direction. Bots receive a cloud computer and work across applications, including systems with no clean API or MCP integration. When an agent drives a rendered interface, there may be no structured API request at which your existing policy expects to intervene.[9]
The build-versus-buy question has changed. You can buy, or fork, the loop and keep the judgement around it. Ask two questions in every vendor review:
- What does your harness own on every turn?
- What can our team version, export, observe, and override?
A vague second answer means the vendor owns part of your product judgement.
12Observability and evals: what did the run mean?
- TraceRecords what happened.
- EvalJudges whether that behaviour was acceptable.
- Feedback loopChanges the system based on that judgement.
A useful trace records the task and tenant identity, model and harness version, context sources, tool calls, permission decisions, state changes, retries, cost, and completion evidence.
Sensitive payloads need separate access and retention. “Log everything” can create a second leak.
An invoice eval may ask:
- Was the correct invoice selected?
- Did the evidence meet the matching rule?
- Did the agent remain inside tenant scope?
- Did every exception receive a valid state?
- Did the processed count equal 2,347?
A dashboard can show that a call happened. An eval says whether the behaviour was acceptable.
Deep measurement design lives in AI Evals. The harness must make the run observable and connect the judgement to a release decision.
13View 3: Session, Harness, Sandbox
Runtime engineers need concrete service boundaries.
| Service | Holds | Failure behaviour |
|---|---|---|
| Session | Durable event history and task state | Another worker can resume from recorded events |
| Harness | The control loop | Can restart and reread the session |
| Sandbox | Tool execution and files | Can be replaced without deleting session history |
If the sandbox dies, the harness records a tool failure and can provision another environment.
If the harness dies, another worker can read the session and continue.
If the model context fills, the harness can compact the current view without deleting the durable event history.
MHTE tells you who owns a failure. Session, Harness, and Sandbox tell engineers which runtime service held the state, decision, or execution involved.
14One file, four lenses
Take AGENTS.md in a coding repository.
- At restIt is a file in the environment.
- At startupThe harness reads it.
- During context assemblySelected text becomes model input.
- LaterA file tool may read or modify it.
The file has one primary home and several roles over time.
Ownership answers where it belongs. Lifecycle explains how its effect travelled through the run.
A stale rule can begin on disk, become inevitable when the harness injects it, become visible when the model follows it, and create a consequence when a tool acts. The symptom appears at the end. The cause entered earlier.
15Agent identity and delegated authority
An agent needs an identity separate from the person or organisation it represents.
A receiving system should be able to answer:
- Which agent is calling?
- On whose behalf is it acting?
- What authority was granted for this job?
Authority should be narrow across operation, resource, time, value, and context.
A travel agent may be allowed to reserve one flight under ₹30,000 for one traveller during the next two hours. That does not authorise it to change the traveller’s profile, buy another ticket, or grant a sub-agent broader access.
When the system becomes uncertain, loses a trusted source, or detects suspicious content, its authority should narrow, not expand. Move toward read-only behaviour, stronger verification, approval, pause, or stop.
An individual Internet-Draft by Raza Sharif, published on 26 August 2026, proposes six requirements for agent identity: it should be bound to a specific agent instance, non-repudiable, time-bound, verifiable offline, distinct from the principal’s identity, and independently revocable.[8]
The status matters. This is a work in progress submitted to the IETF, not an endorsed standard. Its questions are still useful. A reused human OAuth token cannot identify or revoke one agent without affecting the principal. A static API key is neither instance-bound nor naturally time-limited.
The same draft reports an author-conducted survey of more than 1,900 indexed MCP servers in April 2026. It says 99.4% implemented no cryptographic identity verification, message signing, or replay protection.[8] Treat that figure as the author’s disclosed measurement, not an independently reproduced ecosystem estimate.
| Layer | Product question | Owner | Evidence retained |
|---|---|---|---|
| Model | Which capability may handle this data and job? | ||
| Harness | Which context, tools, steps, state, and checks apply? | ||
| Runtime | Which identity, authority, limits, and approvals apply? | ||
| Environment | Which systems may be read or changed? | ||
| Evals | What proves acceptable behaviour? |
16In practice: draw your system
Create one page with four columns: Model, Harness, Tools, Environment. Put Runtime underneath.
Inside Harness, add five rows: Identity, Memory Policy, Orchestration, Interception, Observability and Evals.
For each row, record the implementation, evidence, containment, long-term repair, owner, and regression case.
| Field | Question |
|---|---|
| Implementation | Where does this decision happen today? |
| Evidence | How do we know it works? |
| Failure | What does a defect look like in production? |
| Immediate containment | What reduces exposure before the full repair ships? |
| Long-term repair | Which product, policy, code, or infrastructure change removes the cause? |
| Owner | Who can make and sustain that change? |
| Regression case | Which test proves the failure remains fixed? |
Do not fix anything during the first pass. Draw the system as it exists, including vendor defaults, spreadsheet policies, manual approvals, and controls that live only in people’s heads.
Then complete the repair-owner sentence for every blank row:
This decision currently lives in [component or nowhere], so the immediate containment is [specific step], the long-term repair is [specific change], owned by [person or team], and proved by [regression case].
You can now draw the machine. Episode 03 returns to the invoice workflow and changes four decisions: what a retry learns, which output shapes are allowed, which tools the model sees, and what evidence permits completion.
-
Anthropic — “Scaling managed agents”: durable session, model loop,
and execution environment as separate services.
anthropic.com/engineering/managed-agents -
OpenAI — “Unrolling the Codex agent loop.”
openai.com/index/unrolling-the-codex-agent-loop -
OpenAI — “Harness engineering: leveraging Codex in an agent-first
world.”
openai.com/index/harness-engineering -
Microsoft — “Build production-ready agents with the GitHub Copilot harness
and Agent Framework.”
devblogs.microsoft.com/agent-framework -
OpenAI — “Codex as a platform: build on the open agent harness,”
August 2026.
developers.openai.com/blog/codex-as-a-platform -
Anthropic — Claude Platform release notes, 7 August 2026: Managed Agents session
budgets and
budget_reached.
docs.anthropic.com/en/release-notes/api -
Rony Utevsky, Adversa AI — “Zero-click Grok data theft: Cryptographic
Context Injection attack leaks chat histories,” 20 August 2026.
adversa.ai/blog/cryptographic-context-injection-grok-data-theft -
Raza Sharif — “Agent Identity Framework: Trust and Identity for Autonomous
AI Agents,” individual Internet-Draft, 26 August 2026.
datatracker.ietf.org/doc/draft-sharif-agent-identity-framework -
SpaceXAI — “Introducing Grok Bot,” early beta announcement, 11
August 2026.
x.ai/news/introducing-grok-bot -
DeepSeek — “DeepSeek Harness developer preview: Everything is a
plugin,” 16 August 2026.
deepseek.com/harness