A model swap does not repair pagination. A new tool schema does not restore state that was never saved. Before touching the model, name the failing layer — and its owner.
The model spoke last.
That doesn't mean the failure began there.
This season is about what did.
The model proposes what could be done.
The harness decides what may be attempted.
The tool translates intent into an operation.
The environment supplies the world, absorbs the effect, and preserves the evidence.The season, in four sentences
Pick the one that matches this week's problem.
The one-sitting autopsy. Six failure shapes, one repair-owner test, and the question that ends every blame-the-model review.
Start here → I need the systemThree views of one machine, five decisions the control plane makes on every turn, and the runtime underneath. The whiteboard your team will actually redraw.
Open the anatomy → I need to apply itRetry with evidence. Constrain the output space. Narrow the decision surface. Verify before exit. One sprint, measured before and after — no model change.
Run the patterns →A model swap does not repair pagination. A new tool schema does not restore state that was never saved. Before touching the model, name the failing layer — and its owner.
LangChain's 2026 State of Agent Engineering survey: the most-cited barrier isn't capability — it's quality. And quality is not a model property. It's a harness property.
Retry with evidence. Constrain the output space. Narrow the decision surface. Verify before exit. The sprint that changed no model — and changed the outcome.
When failure is contained and proved, the organisation delegates more useful work. That increase — not the token bill — is the return on the harness.
By the end: diagnose your last three failures, sketch the harness in three views, hold the five paradoxes under review, price the reliability dividend, run the nine-day audit, and name the owner of the behaviour between the layers.
Three layers already have owners. The harness has none. That's the opening.
The boundary test: the model proposes; the harness permits, executes, verifies, records, and recovers. If a component isn't doing one of those five verbs, it isn't in the harness.
Episode 02 mapped the harness. These two companions stop at the boundaries that hide the most production risk: the tools the agent can call, and the world those calls run in.
Function calling, MCP, A2A, permissions, verification, the registry pattern — and the eight failure modes that eat production agents.
Why the same agent action can be harmless in one world and catastrophic in another. Isolation, sandboxes, identity, state, blast radius, reversibility.
Three deep case studies of the systems that make coding agents reliable — architecture, failure modes, and control loops.
Three systems that show what the model cannot do alone. Comparison matrix, shared design principles, and an adoption order for teams starting now.
How a small team ran an approximately one-million-line agent-generated codebase — repository as memory, Symphony orchestration, agent-to-agent review, and continuous cleanup.
How persistent files, incremental sessions, evidence gates, and a fresh-context evaluator let coding agents continue reliable work across many context windows.
How plans, scoped rules, dynamic skills, hooks, worktrees, and review surfaces turn harness engineering into ordinary repository engineering.
The harness is the contract you make with the model. And the only thing you actually ship.