Series 2 of 5 · 8 episodes

HARNESS
ENGINEERING

The model spoke last.
That doesn't mean the failure began there.
This season is about what did.

The model proposes what could be done.
The harness decides what may be attempted.
The tool translates intent into an operation.
The environment supplies the world, absorbs the effect, and preserves the evidence.

The season, in four sentences
Start Episode 01 · Why Your Agent Fails → See the full map ↓
Runtime, not prompt 57.3% ship agents — quality is still the barrier
Enter here

Three ways into the season

Pick the one that matches this week's problem.

I need the idea
Episode 01Why Your Agent Fails
The autopsy

The one-sitting autopsy. Six failure shapes, one repair-owner test, and the question that ends every blame-the-model review.

Start here →
I need the system
Episode 02Inside a Production Agent Harness
The full anatomy

Three views of one machine, five decisions the control plane makes on every turn, and the runtime underneath. The whiteboard your team will actually redraw.

Open the anatomy →
I need to apply it
Episode 03Three Small Changes
Four patterns · One sprint

Retry with evidence. Constrain the output space. Narrow the decision surface. Verify before exit. One sprint, measured before and after — no model change.

Run the patterns →

The complete season map ↓
Evidence · Selected

Four receipts. One season.

01
Not the model
The repair-owner test.

A model swap does not repair pagination. A new tool schema does not restore state that was never saved. Before touching the model, name the failing layer — and its owner.

02
57.3%
of 1,300+ practitioners run agents in production.

LangChain's 2026 State of Agent Engineering survey: the most-cited barrier isn't capability — it's quality. And quality is not a model property. It's a harness property.

03
Four patterns, one sprint
Measured before and after.

Retry with evidence. Constrain the output space. Narrow the decision surface. Verify before exit. The sprint that changed no model — and changed the outcome.

04
The reliability dividend
Bounded downside, visible evidence.

When failure is contained and proved, the organisation delegates more useful work. That increase — not the token bill — is the return on the harness.

The Thesis
  1. The model sets the capability ceiling.
  2. The harness establishes the reliable floor.
  3. The environment bounds the blast radius.
  4. The organization decides the authority.
  5. Evals provide the evidence.
  6. The improvement loop decides whether the system compounds or decays.
Ceiling · Floor/ Radius · Authority/ Evidence · Compounding
SEASON 018 episodes

Act 01What the Harness Owns

By the end: diagnose your last three failures, sketch the harness in three views, hold the five paradoxes under review, price the reliability dividend, run the nine-day audit, and name the owner of the behaviour between the layers.

The MHTE stack · the harness is the moat
01Modelrented, commodifying
02 · yoursHarnessorchestration · retries · guardrails
03Toolsowned by IT
04Environmentowned by data

Three layers already have owners. The harness has none. That's the opening.

The boundary test: the model proposes; the harness permits, executes, verifies, records, and recovers. If a component isn't doing one of those five verbs, it isn't in the harness.

T01 Why Your Agent Fails The model spoke last. That does not mean the failure began there. Six failure shapes, one repair-owner test, and the end of blame-the-model reviews. If you only read one — start here Ep01 → 6 failure shapes Read episode → T02 Inside a Production Agent Harness Three views of one machine, five decisions the control plane makes on every turn, and the runtime underneath — including why guardrails live in IAM, not the system prompt. Ep02 → 5 decisions Read episode → T03 Three Small Changes, Dramatic Outcomes Add evidence. Bound outputs. Narrow choices. Verify outcomes. Four patterns, one sprint, a measured before-and-after — and no model change. Ep03 → 4 patterns Read episode → T04 Five Paradoxes Every PM Must Hold Useful goals pull against each other. Five paradoxes — plus a sixth inherited from Episode 02 — each closed by the review question that makes the tension testable. Ep04 → 5 paradoxes Read episode → T05 What It Costs, What It Returns The model call is one line on the bill. Five cost centres it hides, the reliability dividend, and a monthly review that ends expand, hold, or narrow. CFO-ready Ep05 → 5 cost centres Read episode → T06 Your Monday Morning Harness Kit Nine days to diagnose. Twelve weeks to ship. Seven artifacts, five maturity rungs, and four owned tickets that keep every change attributable. Ep06 → 9-day audit Read episode → T07 Organisations That Rebuilt Around Agents Three layers have natural owners. The behaviour between them still needs one. An ownership map, three operating models, and the Harness PM's decision rights. Ep07 → 3 org models Read episode → T08 When the Harness Becomes the Habit Capabilities move. Product obligations remain. Five capabilities the model already absorbed; five obligations it never will — and how to tell scaffolding from the forever work. Season finale Ep08 → 5+5 patterns Read episode →
In the margins · Bonus deep dives

The two most misunderstood MHTE boundaries

Episode 02 mapped the harness. These two companions stop at the boundaries that hide the most production risk: the tools the agent can call, and the world those calls run in.

2 essays · Episode 02 companions
Companion to Episode 02 · Tools

The Tool Is the Contract

Field guide

Function calling, MCP, A2A, permissions, verification, the registry pattern — and the eight failure modes that eat production agents.

CompanionRead →
Companion to Episode 02 · Environment

The Environment Is the Product Boundary

Outline · body copy in progress

Why the same agent action can be harmless in one world and catastrophic in another. Isolation, sandboxes, identity, state, blast radius, reversibility.

CompanionRead outline →
In practice · Companion collection

Case Studies in Practice

Three deep case studies of the systems that make coding agents reliable — architecture, failure modes, and control loops.

3 systems
Collection · Start here

Harness Engineering in Practice

Three systems that show what the model cannot do alone. Comparison matrix, shared design principles, and an adoption order for teams starting now.

Open the collection→
Case 01 · OpenAI

Codex & Symphony

How a small team ran an approximately one-million-line agent-generated codebase — repository as memory, Symphony orchestration, agent-to-agent review, and continuous cleanup.

Case 01Read →
Case 02 · Anthropic

Long-Running Agent Harness

How persistent files, incremental sessions, evidence gates, and a fresh-context evaluator let coding agents continue reliable work across many context windows.

Case 02Read →
Case 03 · Cursor

Repository-Native Harness

How plans, scoped rules, dynamic skills, hooks, worktrees, and review surfaces turn harness engineering into ordinary repository engineering.

Case 03Read →

The harness is the contract you make with the model. And the only thing you actually ship.

Start Episode 01 · Why Your Agent Fails → See the full map ↓
The syllabus. · Three next-hops — pick your angle.
Agentic Stack · Ep 01 Prompt vs. Context The iceberg. The user's message is less than 1% of what the model actually sees. Read → AI Evals · Ep 01 Benchmarks ≠ Evals The needle. Benchmarks measure the model. Evals measure what you shipped. Read → AI PM OS · Ep 01 Why AI PM ≠ SaaS PM The compass. The job changed underneath the title. Here's the new one. Read →