Harness Engineering · Episode 04

Five Paradoxes Every PM Must Hold

Useful goals pull against each other. The PM makes the tension testable.

Arc · The tensions Episode · 04 of 08 Next · Episode 05 — What It Costs, What It Returns
After this you will know
  • ScorecardsWhy capability and reliability need separate measures.
  • AutonomyWhen a constraint expands useful autonomy instead of limiting it.
  • FundingWhich controls deserve permanent funding and which need an expiry date.
  • BoundaryWhat to centralise and what to keep close to the workflow.
  • LaunchWhat a demo proves and what production still needs.

01Where Episode 03 left us

At 3 PM on Thursday, five teams enter the same review with five reasonable requests.

Engineering wants a stronger model. Design wants fewer restrictions. Platform wants one shared harness. Marketing wants the capability in next month’s demo. Finance wants to stop funding controls the next model may absorb.

No request is obviously wrong. Approving any one without conditions could still damage the product.

Episode 01 taught us to find the first broken contract. Episode 02 placed those contracts inside a production system. Episode 03 changed four decisions without changing the model: retry with evidence, constrain the output, narrow the tool surface, and verify before exit.

Each pattern improved control. Each also introduced a cost.

A strict output contract may reject a legitimate new case. A narrow tool surface may strand an edge case. A completion gate may keep a run alive after another attempt is no longer worth its cost. A workaround built for today’s model may become tomorrow’s dead weight.

Patterns tell a team what can work. Judgement decides when to use them.

02The meeting where every request made sense

The five requests pull the product in different directions:

The disagreement is not caused by one team understanding the system and another missing the point. Each team is optimising a different legitimate goal.

A weak decision picks the loudest goal. A stronger decision names the evidence required to move toward either side.

The PM’s job is not to choose one pole. It is to state what must be true before the system moves toward either side.

03The sixty-second harness recap

A model proposes an answer or action.

The harness decides what context the model receives, which tools it may request, what gets checked, what survives, and when the workflow may stop.

The runtime enforces identity, authority, limits, and containment. The environment reveals what actually happened. Evals decide whether that outcome was acceptable.

Every paradox in this episode concerns a choice across those responsibilities. Intelligence changes what the model can propose. Constraints, workflow design, verification, and ownership determine what the organisation can trust.

For a business leader, each paradox is a delegation decision. For product, it is a rule and a counter-metric. For engineering, it is a boundary between components. For security, it is an enforceable limit. For operations, it is evidence that the limit held.

04The five tensions

Tension Goal on the left Goal on the right Review question
Intelligence and reliability Solve harder tasks Behave acceptably across real conditions How wide is the best-to-worst gap?
Constraints and autonomy Bound consequence Allow useful judgement and action Does the control expand trusted work?
Scaffolding and permanence Fix today’s model Preserve local obligations Would we need it if capability doubled?
Specificity and generality Fit the workflow Reuse infrastructure What benefits from scale, and what loses meaning?
Demo and production Prove possibility Cover the failure distribution Which conditions did the demo exclude?

The goal is not permanent balance. A payment action may sit close to constraint. A drafting task may sit close to autonomy. The marker can move by workflow and release. The review question must remain visible.

Figure 01 · Concept
Five tensions, one movable marker each
WHERE DOES THIS WORKFLOW SIT TODAY? INTELLIGENCE Solve harder tasks RELIABILITY Behave acceptably across real conditions How wide is the best-to-worst gap? CONSTRAINTS Bound consequence AUTONOMY Allow useful judgement and action Does the control expand trusted work? SCAFFOLDING Fix today’s model PERMANENCE Preserve local obligations Would we need it if capability doubled? SPECIFICITY Fit the workflow GENERALITY Reuse infrastructure What benefits from scale, and what loses meaning? DEMO Prove possibility PRODUCTION Cover the failure distribution Which conditions did the demo exclude? Do not choose a pole before naming the condition.
Read it as The marker moves per workflow and release. What must not move is the requirement to answer the question underneath it.

05Paradox 1: intelligence and reliability

The reasonable instinct

When reliability drops, upgrade the model.

Stronger models solve harder tasks, follow more complex instructions, and recover from mistakes weaker models cannot understand.

The missing distinction

Capability and reliability are different variables.

A model upgrade can raise capability while leaving the system failure untouched.

Failure Would a stronger model help?
Cannot reason through partial versus duplicate payment Possibly; test it
Lost pagination state No; repair state handling
Stale tool result No; repair source selection and freshness
Weak completion rule No; verify against external evidence

A more capable model can also make unsupported output harder to notice. Fluency can make a stale source sound authoritative.

Anthropic’s 1 September 2026 announcement of Claude Fable 5.1 is a useful test case. The company presents the model as stronger on long-running problem solving, and a customer testimonial describes a 38-hour unattended run that diagnosed a result, launched parallel experiments, and returned with evidence and next steps. Those are meaningful capability claims, but they do not remove the need for a reliability scorecard.[4]

Anthropic’s benchmark footnote makes the boundary unusually clear: its Terminal-Bench-Science results were measured with the Claude Code harness. The published number is already a model-plus-harness result.[4]

For the reconciliation team, a model may become better at recovering from a failed step. The ledger must still prove that all 2,347 eligible invoices were processed. Recovery is capability. Reconciliation is reliability.

The two scorecards

Capability Reliability
Hardest representative task solved Full workflow completion rate
Remaining task classes outside ability Performance after tool failure or restart
Cost and latency of the gain Spread across common and edge cases
Model or task benchmark Failures reaching users

The distance between the columns is the product problem.

Decision rule

Approve a model change when a named capability gap improves on representative tasks and the full workflow regression suite still passes. Run an experiment when the source of failure is uncertain. Reject the upgrade as a fix when the defect sits in state, tools, authority, or completion.

Review questionWhich failure are we asking the model upgrade to remove, and what evidence says that failure comes from reasoning?

06Paradox 2: constraints and autonomy

The reasonable instinct

Too many guardrails remove what makes an agent useful. Enough approvals can turn an adaptive system into a slow workflow engine.

The missing mechanism

An unbounded agent often receives less real authority.

Security blocks the launch. Operations routes only low-consequence work. Users check every result. The product is technically autonomous and organisationally powerless.

A bounded agent can receive more work because the organisation can price the downside.

Action Autonomy posture
Read one customer’s order Autonomous with audit
Draft a refund Autonomous
Issue a small eligible refund Autonomous with policy and rollback
Issue a large refund Human approval
Change payment destination Separate authority or prohibited

The boundaries do not remove autonomy. They allocate it by consequence.

When constraints become harmful

  1. Reviewers approve routine actions without reading.
  2. Safe and reversible work waits behind an irreversible-action gate.
  3. Policy grows without reducing actual loss.
  4. Users cannot understand why the system refuses.

A control earns its place when it reduces expected loss or expands the work the organisation is willing to delegate.

Decision rule

Approve a constraint when consequence, verification, and rollback justify it. Revise it when it blocks safe work or creates approval fatigue. Remove it when another control covers the same failure and evals confirm that the workflow remains safe.

Review questionWhich consequence are we containing, and does this control expand trusted work?

07Paradox 3: scaffolding and permanence

The reasonable instinct

A reliability fix that works in production should become durable infrastructure.

The missing distinction

Some controls compensate for a model weakness. Others encode facts about your organisation.

Compensation Institution
JSON repair loop Financial approval policy
Context reset for a model quirk Tenant isolation
Prompt scaffold for planning Audit retention
Tool-selection workaround Domain completion rule

A compensation has a timer. An institution has an owner.

The context-reset case

Anthropic reported that Claude Sonnet 4.5 sometimes wrapped up tasks as its context limit approached. The Managed Agents team added context resets. The behaviour disappeared with Claude Opus 4.5, so the same resets became dead weight.[1]

The original repair was correct. Keeping it after its assumption expired would have been wrong.

A second compatibility lesson arrived on 1 September 2026. Anthropic’s release notes state that thinking blocks produced by Claude Fable 5.1 can be read only by the same model or a newer one. If a harness routes the conversation to an earlier model, the API drops the replayed block.[7]

This is not evidence that the earlier context-reset repair returned. It is evidence that a model upgrade can change assumptions inside a running harness. A migration can remove scaffolding, invalidate scaffolding, or create a new compatibility boundary. All three require tests.

Put an expiry label on compensation

Field Record
Built for Model and version
Failure addressed Observable behaviour
Retire when Eval no longer regresses without it
Review date Next model migration or quarterly review
Owner Team responsible for testing and removal

Build compensation thin, measurable, and easy to disable.

Fund institutions as durable product infrastructure. Financial authority, tenant isolation, audit retention, and domain completion rules do not disappear because a general model becomes smarter.

Decision rule

Approve compensation with an expiry condition. Retire it after an enabled-versus-disabled evaluation shows no loss. Treat local policy, audit, identity, and workflow rules as permanent unless the organisation itself changes.

Review questionIf model capability doubled tomorrow, would we still need this component?

08Paradox 4: specificity and generality

The reasonable instinct

One shared harness should serve the enterprise. Common infrastructure lowers cost and prevents every product team from rebuilding tracing, deployment, identity, and model access.

The missing boundary

General infrastructure can serve many workflows. Differentiating judgement usually cannot.

Improves when shared: model-provider integration, identity and secrets, trace storage, deployment and rollback, cost accounting, tool standards, and eval execution.

Loses meaning when generalised: definition of done, authoritative sources, approval thresholds, exception routing, domain evals, and local latency or cost tolerance.

A compliance review and a support conversation can share trace storage. They should not share one generic completion rule.

The layered design

Shared platform Workflow harness
Model access Context policy
Deployment Completion rule
Identity Domain tools
Traces Escalation
Cost controls Workflow evals
Tool standards Local autonomy limits

Microsoft’s GitHub Copilot integration illustrates the split. Copilot supplies a specialist coding loop. Agent Framework supplies a general surface for tools, middleware, observability, sessions, and approval.[3]

The architecture is not platform versus product. It is an explicit interface between reusable infrastructure and local judgement.

Multi-agent systems obey the same rule

Add agents only when:

  1. Each has a narrow job.
  2. Inputs and outputs are explicit.
  3. Each job has an eval.
  4. Shared state is visible.
  5. Failure can be attributed.
  6. Merge and conflict rules exist.

More workers without those conditions create more ambiguity.

Decision rule

Centralise a component when scale makes it cheaper, safer, or more reliable without stripping workflow meaning. Keep it local when the decision depends on domain judgement or must evolve with the team accountable for the outcome.

Review questionWhich part benefits from scale, and which loses meaning when generalised?

09Paradox 5: demo and production

The reasonable instinct

A successful demo shows that the architecture works.

It proves something useful: the system can create value on at least one path.

The missing distribution

A demo selects a favourable case. Production samples ambiguity, stale context, tool failure, concurrency, retries, long sessions, changing policy, and rare costly errors.

Confusing the questions creates reckless launches or endless pilots.

The hardening contract

For every demonstrated capability, name:

Field What to record
Traffic Expected input distribution
Failure modes Named cases, not “edge cases”
Completion External evidence required
Recovery Retry, rollback, or handoff
Autonomy Actions allowed without review
Evals Cases required before release
Rollout Staged exposure and rollback
Owner Person accountable for the accepted behaviour

LangChain’s 2026 survey reported substantial production adoption while quality remained the leading barrier.[2] Production access is no longer sufficient proof. Sustained acceptable behaviour is the scaling test.

Reuters supplied a sharper organisational example on 26 August 2026. It reported that Meta’s Project OT considered cuts of up to 60% in some teams while reorganising work around AI-enabled pods. Meta scaled back the effort after its agents and coding tools produced reliability and security problems and less productivity improvement than expected. Meta described the effort as scenario planning and said it did not affect every team.[6]

The report is second-hand evidence based on Reuters’ internal documents, recordings, and interviews, not an independently audited agent evaluation. Its lesson should remain equally bounded: the organisation changed its assumptions about staffing before agent behaviour was reliable enough to support them.

A demo proved possibility. Production exposed the missing conditions at the level of the org chart.

Decision rule

Approve a demo when it answers a real value question. Approve production rollout when the failure inventory, evals, recovery, and ownership exist for the first traffic slice. Pause or narrow when the team cannot name the distribution it has not tested.

Review questionWhich conditions did the demo exclude, and how will we test them before launch?

10A companion lens: segregation and lifecycle

The five paradoxes are choices between legitimate goals. Segregation and lifecycle is not a sixth paradox. It is a lens for investigating what happens after one of those choices enters the system.

Episode 02 separated four responsibilities and followed one AGENTS.md file through a run.

At rest, the file belongs to the environment. The harness reads it. Selected text becomes model context. A tool may later modify it.

Both views matter:

A stale file can begin in storage, become inevitable when the harness injects it, become visible when the model follows it, and become consequential when a tool acts.

In an incident, ask:

  1. Where did the failure become visible?
  2. Where did it become inevitable?
  3. Which layer owned the earlier preventable condition?
Review questionWhere did the failure become visible, and where did it become inevitable?

11The same tensions in the sprint

Sections 05 to 09 stated five design tensions. The cards below are not five additional paradoxes. They are five ways the same tensions appear during ordinary product work.

Figure 02 · Framework
Five paradoxes and their PM rules
FIVE TENSIONS A PM HOLDS AT THE SAME TIME Each pair is a setting to tune, not a problem to solve once. More context → Less understanding PM RULE Smallest complete, current, authorised context for the next decision. More agents → Less reliable work PM RULE Each handoff adds instructions, delay and failure surface. Make it earn its place. More autonomy → Controls below the agent PM RULE Flexible about the path. Strict about the boundary. Stronger model → Less harness, more opportunity PM RULE After a model change, test what can be deleted before adding machinery. Faster automation → Scarcer human judgement PM RULE People for boundaries, exceptions, ambiguity and new patterns.
One card per tension. The left term is what teams add first. The right term is the cost or obligation that appears next. The rule underneath names the decision the PM still owns.

More context can create less understanding

The agent needs enough information to act well. Giving it everything makes important facts harder to find, increases cost, and carries stale information forward.

PM ruleAsk what the next step needs, where that fact comes from, who may see it, and how freshness is checked.

More agents can produce less reliable work

Specialisation can improve a difficult subtask. Every additional worker also adds an instruction set, context boundary, handoff, delay, and failure surface.

In a simplified chain where three independent components must all succeed and each succeeds 90% of the time, end-to-end success is about 73%. Five such components produce about 59%. Real systems include dependencies, retries, optional steps, and recovery, so this arithmetic is an illustration, not a forecast.

PM ruleAdd an agent only when the complete job improves enough to justify the additional coordination, cost, and failure risk.

More autonomy requires controls below the agent

An agent needs freedom to choose useful steps. Important permissions cannot live only inside the same model-controlled process making those choices.

PM ruleBe flexible about the path and strict about the boundary. Test whether the boundary survives a wrong model decision.

A stronger model can need less harness, and create more harness opportunity

Better models can make old scaffolding unnecessary. The same improvement can make longer and more consequential work possible, creating a new boundary where planning, recovery, and external verification matter more.

OpenAI’s 26 August 2026 post-mortem on an internal cybersecurity incident offers unusually strong evidence for the second half. OpenAI reported that a model’s “propensity to compromise infrastructure” fell by more than 100 times when tested with the production ChatGPT harness and system prompt.[5]

The conditions were specialised, so the ratio should not be treated as a universal estimate. The architectural lesson is broader: after a model change, test what the harness can delete and what the more capable system now requires.

A benchmark outside the context, tools, runtime, and checks you intend to ship is not the product score.

PM ruleAfter a model change, test what can be deleted before adding machinery. Include both ordinary performance and the failures you would rather not discover in production.

Faster automation makes human judgement more valuable

As agents produce more work, qualified human attention becomes scarce.

Reviewing everything is impossible. Removing people entirely is unsafe. Place human judgement where it changes the decision:

If the review queue contains everything, it hides the important work inside volume.

PM ruleUse people for boundaries, exceptions, ambiguity, and new patterns. Use automated checks for stable, objective requirements.

Paradox worksheet

Current choice Benefit Cost or risk Evidence to keep it Evidence to reverse it
         

Governance of a harness that rewrites itself belongs in Episode 08, where the promotion loop and rollback rules are set out in full.

12In practice: the decision card

Decision Evidence required Outcome choices
Upgrade model Capability eval plus full-workflow regression Approve, experiment, reject as fix
Remove approval Consequence, reversibility, verifier, rollback Approve, revise, retain
Keep workaround Enabled-versus-disabled eval Retain, time-box, retire
Centralise component Reuse benefit and local-meaning cost Platform, workflow, shared interface
Launch capability Failure inventory, evals, recovery, owner Demo only, staged release, hold
Assign incident Earliest preventable divergence Owner and action item

The card does not choose for the PM. It forces the condition into the room before politics chooses by default.

Figure 03 · Practice
The decision card
DECISION EVIDENCE OUTCOME CHOICES No decision moves without the evidence in its row. Upgrade model Capability eval plus full workflow regression Approve · experiment · reject as fix Remove approval Consequence, reversibility, verifier, rollback Approve · revise · retain Keep workaround Enabled-versus-disabled eval Retain · time-box · retire Centralise component Reuse benefit and local-meaning cost Platform · workflow · shared interface Launch capability Failure inventory, evals, recovery, owner Demo only · staged release · hold Assign incident Earliest preventable divergence Owner · action item Approve, test, revise, retire. Every decision needs evidence.
Read it as The outcome column always contains more than yes and no. That keeps the tension testable instead of political.

For the decision currently blocking your team, complete one record:

Field Record
Decision The choice being proposed
Evidence Trace, eval, cost, incident, or missing artifact
Immediate containment What reduces exposure while the decision is tested
Long-term change What will be built, revised, retained, or removed
Owner Person or team accountable for the change
Regression case Test that must remain green after the decision
Reversal condition Evidence that would make the team undo the choice

13Connecting the dots

Return to the sprint meeting.

Engineering can test a stronger model without calling it the fix. Design can remove a control after its covered failure remains caught elsewhere. Platform can centralise primitives while workflow teams keep domain judgement. Marketing can run the demo with a hardening contract. Finance can stop funding compensation after evidence says it has expired.

No department lost. None received an unconditional yes.

Episodes 01 to 04 now form one operating sequence:

  1. Find the first broken contract.
  2. Locate the responsibility and control.
  3. Replace vague choices with specific signals.
  4. Test what each new control gains and what it costs.

The fourth step takes us from system design to economics. Every condition has a price: evals, review time, runtime cost, duplicated workflow work, or delayed release. Episode 05 asks whether those costs reduce expected failure loss and expand useful delegation.

You now hold
Tensions 04 Five paradoxes with one review question each, a companion lifecycle lens, and a decision card that turns a sprint argument into an evidence requirement.
The next question
Every condition has a price. Do those costs reduce expected failure loss enough to expand the work the organisation is willing to delegate?
Continue
Harness 05 What It Costs, What It Returns → — the five cost centres the token bill hides, and the reliability dividend that decides how much work you may delegate.
Read alongside
Environment 06 Decide What Runs Without You → — the autonomy ladder that Paradox 2 is arguing about.
Sources
  1. Anthropic — “Scaling managed agents,” context resets, model-version behaviour, and controls that expire with the model.
    anthropic.com/engineering/managed-agents
  2. LangChain — “State of Agent Engineering,” production adoption alongside quality as the leading barrier.
    langchain.com/state-of-agent-engineering
  3. Microsoft — “Build production-ready agents with the GitHub Copilot harness and Agent Framework,” specialist loop and general platform surface.
    devblogs.microsoft.com/agent-framework/build-production-ready-agents-with-the-github-copilot-harness-and-agent-framework
  4. Anthropic — “Introducing Claude Fable 5.1 and Claude Mythos 5.1,” 1 September 2026.
    anthropic.com/claude-fable-and-mythos-5-1
  5. OpenAI — “The Hugging Face incident and the road ahead,” 26 August 2026.
    openai.com/index/hugging-face-incident-and-the-road-ahead
  6. Greg Bensinger, Reuters — “How Meta’s AI layoff plans went kaput,” 26 August 2026.
    reuters.com/technology/artificial-intelligence/how-metas-ai-workforce-transformation-plans-went-kaput-2026-08-26
  7. Anthropic — Claude Platform release notes, 1 September 2026, thinking-block compatibility for Claude Fable 5.1 and Claude Mythos 5.1.
    docs.anthropic.com/en/release-notes/api