Series 2 of 5 · 8 episodes · 1 season

HARNESS
ENGINEERING

The model is 10%.
The harness is 90%.
Design the system around it.

The model proposes what could be done.
The harness decides what may be attempted.
The tool translates intent into an operation.
The environment supplies the world, absorbs the effect, and preserves the evidence.

The season, in four sentences
Start Episode 01 · Why Your Agent Fails → See the full map ↓
Runtime, not prompt 0.95⁷ = 70% compound reliability
Enter here

Three ways into the season

Pick the one that matches this week's problem.

I need the idea
Episode 01Why Your Agent Fails
The autopsy

The one-sitting read that reframes every agent failure you've investigated — and names what the industry keeps missing.

Start here →
I need the system
Episode 02Inside the Harness
The full anatomy

Four layers, five clusters, one runtime. The whiteboard your team will actually reference in the sprint review.

Open the anatomy →
I need to apply it
Episode 03Three Small Changes
Four patterns · One sprint

Retry with evidence, one versioned output shape, a narrower tool surface, and an external completion check — the sprint that changes no model.

Run the patterns →

The complete season map ↓
Evidence · Selected

Four receipts. One season.

01
Not the model
Every post-mortem.

The hallucination wasn't a hallucination. The window was stale. The tools were wrong. The structure was missing. That's harness.

02
0.95⁷ = 70%
Compound reliability.

A 95%-good step chained seven times succeeds 70% of the time. No prompt undoes that compounding. The harness interrupts it — validating intermediate outputs, retrying recoverable failures, checkpointing progress, escalating uncertainty.

03
31 → 9
Fewer moving parts. Half the errors. One week.

A team spent three months hunting a smarter model. It didn't help. Then a new hire made three small fixes to the plumbing around the model — nothing to the model itself. Errors dropped by half in five days. Same model, same team. The leverage was never in the model. It was in what they built around it.

04
The boundary was the product
Containment is not a prompt.

A sandbox with an allowlisted package proxy still failed when the agent found a path the harness did not revalidate. Containment is not a prompt. It is environment design under adversarial goal-seeking.

The Thesis
  1. The model sets the capability ceiling.
  2. The harness establishes the reliable floor.
  3. The environment bounds the blast radius.
  4. The organization decides the authority.
  5. Evals provide the evidence.
  6. The improvement loop decides whether the system compounds or decays.
Ceiling · Floor/ Radius · Authority/ Evidence · Compounding
SEASON 018 episodes

Act 01What the Harness Owns

By the end: sketch your own harness, hold the five paradoxes under pressure, price the program for a skeptical CFO, run the nine-day audit, and tell scaffolding apart from the forever work.

The MHTE stack · the harness is the moat
01Modelrented, commodifying
02 · yoursHarnessorchestration · retries · guardrails
03Toolsowned by IT
04Environmentowned by data

Three layers already have owners. The harness has none. That's the opening.

The boundary test: the model proposes; the harness permits, executes, verifies, records, and recovers. If a component isn't doing one of those five verbs, it isn't in the harness.

T01 Why Your Agent Fails 2 AM. The most capable model money buys just declared itself done at invoice 847 of 2,347. The autopsy says model. It isn't. If you only read one — start here Ep01 4 signatures Read episode → T02 Inside the Harness You bought a harness. Nobody in the room can name what's inside it. Four layers, five clusters, one runtime — and why guardrails belong in IAM, not the system prompt. Ep02 3 views Read episode → T03 Three Small Changes A quarter chasing a better model. Then three fixes to the plumbing around it. Errors dropped by half in five days, same model. Ep03 4 patterns Read episode → T04 Five Paradoxes Every PM Must Hold Useful goals pull against each other. Five tensions — plus the sixth inherited from Episode 02 — each with the review question that makes it testable. Ep04 5 paradoxes Read episode → T05 What It Costs, What It Returns Month four. The CFO opens the spend report. Five cost centers the token bill hides, the runtime line, the reliability dividend, and a six-line monthly review that ends in expand, hold, or narrow. CFO-ready Ep05 5 cost centers Read episode → T06 Your Monday Morning Harness Kit Nine days to diagnose, twelve weeks to ship. Seven artifacts, a five-rung maturity ladder, and four tickets with owners, metrics, and counter-metrics. Ep06 9-day audit Read episode → T07 Organizations That Rebuilt Around Agents Model, tools, and environment already have owners. The behavior between them does not. An ownership map, three operating models, and the Harness PM's decision rights. Ep07 3 org models Read episode → T08 When the Harness Becomes the Habit Every pattern here has a shelf life. That isn't the bug — it's the engine. Five capabilities the model absorbed; five it never will. Season finale Ep08 5+5 patterns Read episode →
In the margins · Bonus deep dives

The two most misunderstood MHTE boundaries

Episode 02 mapped the harness. These two companions stop at the boundaries that hide the most production risk: the tools the agent can call, and the world those calls run in.

2 essays · Episode 02 companions
Companion to Episode 02 · Tools

The Tool Is the Contract

Field guide

Function calling, MCP, A2A, permissions, verification, the registry pattern — and the eight failure modes that eat production agents.

CompanionRead →
Companion to Episode 02 · Environment

The Environment Is the Product Boundary

Outline · body copy in progress

Why the same agent action can be harmless in one world and catastrophic in another. Isolation, sandboxes, identity, state, blast radius, reversibility.

CompanionRead outline →
In practice · Companion collection

Case Studies in Practice

Three deep case studies of the systems that make coding agents reliable — architecture, failure modes, and control loops.

3 systems
Collection · Start here

Harness Engineering in Practice

Three systems that show what the model cannot do alone. Comparison matrix, shared design principles, and an adoption order for teams starting now.

Open the collection
Case 01 · OpenAI

Codex & Symphony

How a small team ran an approximately one-million-line agent-generated codebase — repository as memory, Symphony orchestration, agent-to-agent review, and continuous cleanup.

Case 01Read →
Case 02 · Anthropic

Long-Running Agent Harness

How persistent files, incremental sessions, evidence gates, and a fresh-context evaluator let coding agents continue reliable work across many context windows.

Case 02Read →
Case 03 · Cursor

Repository-Native Harness

How plans, scoped rules, dynamic skills, hooks, worktrees, and review surfaces turn harness engineering into ordinary repository engineering.

Case 03Read →

The harness is the contract you make with the model. And the only thing you actually ship.

Start Episode 01 · Why Your Agent Fails → See the full map ↓
The syllabus. · Three next-hops — pick your angle.
Agentic Stack · Ep 01 Prompt vs. Context The iceberg. The user's message is less than 1% of what the model actually sees. Read → AI Evals · Ep 01 Benchmarks ≠ Evals The needle. Benchmarks measure the model. Evals measure what you shipped. Read → AI PM OS · Ep 01 Why AI PM ≠ SaaS PM The compass. The job changed underneath the title. Here's the new one. Read →