Harness Engineering · Episode 07

Organisations That Rebuilt Around Agents

Three layers have natural owners. The behaviour between them still needs one.

Arc · The operating model Episode · 07 of 08 Next · Episode 08 — When the Harness Becomes the Habit
After this you will know
  • LocateWhy agent ownership falls between existing functions.
  • CompareThree internal operating models and the failure each creates when misapplied.
  • DefineThe decision rights a Harness PM or Agent Reliability PM needs.
  • OperateHow to turn human review into a service rather than a reassuring phrase.
  • KeepWhat judgement must stay in-house even when the infrastructure is rented.

01Where Episode 06 left us

At 9:30 on Monday morning, the reconciliation team opens the one-page output from Episode 06.

The page names four tickets. Every ticket has an owner. The system still does not.

Product owns the instructions. Engineering owns memory and retries. Platform owns tool contracts. Security owns credentials and runtime limits. Finance experts define correct invoice treatment. The eval threshold lives in a spreadsheet maintained by one analyst.

Each component has a contributor. No one can answer the full product question:

Is this agent’s behaviour acceptable enough to release, expand, or keep running?

The first six episodes brought the team here:

  1. Diagnose the first broken contract.
  2. Locate the responsible component and control.
  3. Replace vague choices with specific signals.
  4. Test what each control gains and costs.
  5. Price the completed workflow, not the model call.
  6. Turn evidence into four owned tickets.

Episode 07 addresses the organisational gap the audit exposed. The problem is not that nobody works on the agent. The problem is that contribution is distributed while end-to-end behaviour remains unowned.

02The org chart built for different software

Two quarters after the 2 AM invoice failure, the company has three agents in production and a fourth in testing.

The COO asks:

If agents are now part of how work gets done, who owns their behaviour?

The existing org chart has natural homes for important parts:

The unresolved decisions sit between them:

  1. What context reaches the model?
  2. Which tools appear for this workflow?
  3. What counts as complete?
  4. Which production failures become evals?
  5. When must a person review the case?
  6. When may autonomy expand or contract?
  7. Which behaviour change is safe to release?

Those decisions form the product layer of the harness.

The first organisational change is not a new department or title. It is a named decision owner with authority to accept, reject, narrow, or stop the behaviour.

Shared implementation is normal. Shared accountability without a decider is not an operating model.

03The ownership map

Figure 01 · Concept
Three owners and a hole
ACCOUNTABILITY NEEDS AUTHORITY OVER BEHAVIOUR MODEL Applied AI / platform Selection · routing · providers Cannot settle: is the workflow ok? TOOLS Platform engineering Contract · availability · build Cannot settle: allow this action now? ENVIRONMENT Security & infrastructure Identity · network · limits Cannot settle: did it meet intent? HARNESS · NAMED PRODUCT OWNER + ENGINEERING LEAD Decision rights, not a question mark • Context policy • Definition of complete • Eval coverage • Autonomy envelope • Release readiness for behaviour changes • Coordination across model, tool, environment owners Several teams can change the behaviour. One role accepts or rejects it as a product.
Read it as The three owned layers push their limits downward. What lands in the harness band is exactly what no existing function can settle alone.
Layer Natural owner Decision it can make Decision it cannot settle alone
Model Applied AI or platform Model selection, routing, provider integration Whether the full workflow is acceptable
Harness Named product owner plus engineering lead Context policy, completion, evals, autonomy, release Model, tool, and environment implementation alone
Tools Platform engineering API contract, availability, implementation Whether the agent should receive the action now
Environment Security and infrastructure Identity, network, storage, execution limits Whether the workflow outcome meets user intent

The harness owner is accountable for end-to-end behaviour. The role does not absorb every adjacent function.

Use a four-part test:

  1. Decision: What may this role approve or reject?
  2. Evidence: What must it inspect before deciding?
  3. Control: Which team can implement the decision?
  4. Escalation: Who resolves conflict or accepts exceptional risk?

Assigning everything to one person creates a fictional super-role. Assigning nothing creates shared responsibility without decision rights.

04Four decisions most organisations leave implicit

1. Who defines good?

A finance expert decides how a partial payment should be treated. Product turns that judgement into a versioned requirement. Engineering implements the check. The eval owner ensures the case blocks a bad release.

No one role can perform every step. One role must ensure the chain exists.

2. Who decides what the agent may do?

Product decides which action belongs in the user journey. Security defines the authority policy. Platform implements the permission. Operations defines recovery and rollback.

“Product and security” is not a complete answer until the organisation says who decides when they disagree.

3. Who decides the work is complete?

Engineering can implement processed_count == eligible_count. The workflow owner defines why that condition, the exception states, and the reconciliation checks constitute done.

The code owns enforcement. The business owns meaning.

4. Who may expand autonomy?

Product proposes the value. Engineering provides performance evidence. Security and operations assess consequence and recovery. The accountable product owner accepts, narrows, or rejects the change within a pre-agreed risk boundary.

An executive or risk committee should be required only when the change crosses that boundary.

If every decision needs a new committee, the organisation has participants but no operating model. If one team can expand autonomy without evidence or challenge, it has authority without governance.

05Shared platform, product domain

Most ownership arguments dissolve once a team separates the parts of the harness that should be shared from the parts that must stay with the product.

Figure 02 · Framework
The platform / domain split
WHERE HARNESS WORK LIVES SHARED PLATFORM Identity and authority Runtime limits Tool registry State and memory stores Tracing and evidence Deployment and rollback GUARANTEES → PRODUCT DOMAIN Definition of done Domain evals Escalation policy Tone and copy Customer-facing recovery Acceptable cost per result ← EVIDENCE Platform owns the boundary. Product owns the definition of good.
Two owners, one contract The platform side promises guarantees; the product side returns evidence about whether results were acceptable.

The frontier labs have not published the org charts behind these systems, so this episode does not infer them. Their product interfaces still reveal a useful ownership split.

OpenAI’s August 2026 Codex as a platform article says the Codex harness manages conversation state, tool use, streaming, the configured sandbox, and approval enforcement. The host application decides where the agent runs, what it can reach, which actions need approval, how work is observed, and how results return to the system of record.[3]

Anthropic’s Managed Agents architecture separates the durable session, the harness loop, and the sandbox so each can fail or be replaced independently. It also keeps credentials outside the sandbox and leaves context management in the harness while the session preserves durable history.[2]

The implementations differ. Both expose a stable infrastructure layer while leaving workflow meaning and application policy to the product using it.

Test If a product team cannot change its definition of done without a platform release, the line is in the wrong place.

06Three internal models

The useful variable is where workflow judgement sits.

Model 1 · Platform-owned primitives

A central team owns shared infrastructure:

Product teams own domain rules, completion, escalation, and workflow evals.

How it fails The platform team starts owning workflow decisions as well as infrastructure. Every product waits for the same backlog, local teams build shadow loops and tool wrappers to move faster, and the organisation pays for the central platform plus the duplication it was meant to remove.

When it works Several teams use genuinely common primitives, and the central group can operate them more reliably and cheaply than each team can.

Operating signal Product teams can change workflow context, completion, and evals without requesting a platform roadmap slot. If they cannot, the platform owns too much.

Model 2 · Product-owned harness

Each product team owns its workflow behaviour, orchestration, and evals. Shared infrastructure stays thin.

How it fails Six teams create six retry policies, six eval styles, and six failure taxonomies. A lesson found by one team does not travel, and provider outages, policy changes, or common tool defects require several separate repairs.

When it works The organisation has a small number of different workflows. Customer support, contract review, code assistance, and incident response need different state, tools, latency, and approval rules, and local ownership keeps iteration close to the user outcome.

Operating signal Listen for repeated conversations beginning with “How are you handling…”. When several teams independently solve the same runtime or eval-infrastructure problem, a shared primitive is ready to be extracted.

Model 3 · Harness as operating model

The agent and its control system are the product.

Product managers write behaviour and autonomy specifications. Engineers build inside the control plane. Designers own interaction, refusal, progress, and review surfaces. Evals act as release criteria. Operations watches completed workflows rather than model uptime alone.

How it fails An incumbent renames teams without changing code, incentives, release evidence, or decision rights. A transformation office appears. The operating model remains unchanged beneath new labels.

When it works AI-native companies with one dominant product shape can design this way from the beginning.

Operating signal If eval review, autonomy review, and failure triage are not normal release rituals, the company has adopted the language, not the operating model.

07Choosing among them

Do not use fixed employee or team counts as universal thresholds. Choose using three questions:

Question Favors shared platform Favors product ownership
How similar are the workflows? Similar tools and controls Different data, outcomes, and risk
Where does iteration happen? Infrastructure changes dominate Domain behaviour changes dominate
What repeats across teams? Runtime, identity, traces, eval execution Completion, escalation, domain judgement

Most mature organisations land on a layered version: platform-owned primitives with product-owned workflow harnesses. The interface between them matters more than the label.

08The Harness PM

Titles vary: Harness PM, Agent Reliability PM, Agent Product Lead, or Applied AI Lead. The title matters less than the decision rights.

The role in one sentence

The Harness PM is accountable for the quality, authority, economics, and release readiness of the complete agent workflow.

The role does not personally implement every layer. It makes sure the layers produce one acceptable product behaviour.

What the role owns

What the role does not own alone

Decision table

Decision Decider Required evidence Implementers and required partners Escalation
Definition of done Workflow product owner Business rule, examples, completion evidence Domain expert, engineering, eval owner Business owner
Eval coverage and release gate Harness PM with engineering lead Failure inventory, baseline, holdout, counter-metrics Domain expert, quality, Applied AI Product leadership
Tool authority Product and security within agreed policy Consequence, identity, reversibility, verification Platform, legal, operations Risk owner
Model routing Applied AI or platform Workflow quality, latency, cost, fallback result Harness PM, finance Engineering leadership
Autonomy expansion Harness PM within approved risk envelope Completion, failure loss, review load, rollback Security, operations, domain owner Executive or risk committee if boundary changes
Incident command SRE or security Live incident evidence Harness PM, business owner, vendor Existing incident policy

The table prevents two bad outcomes: a PM held accountable for controls they cannot influence, and a platform team operating business behaviour it cannot define.

09What good ownership looks like

Good ownership is visible in decisions, not titles.

Avoid a vague “eval coverage percentage” unless the denominator is defined. Coverage of what: tool selection, high-value journeys, known failure classes, traffic, or regulatory obligations?

Titles are spreading faster than decision rights

A 2026 JAPAN AI listing for an Agent Harness Engineer covers execution engines, orchestration, session management, policy guardrails, model routing, and related platform concerns. Evaluation ownership is not explicit in the published scope.[6]

OpenAI’s Forward Deployed Engineer, Legal role uses a different title but includes the work this episode cares about: deploying AI into consequential legal workflows, working across product, engineering, research, legal, privacy, and security, and turning operational experience into reusable systems.[7]

Neither posting proves a universal organisation design. Together they make the useful point: the work may appear under engineering, product, reliability, or forward-deployed titles.

Test the role with three questions:

  1. Can it accept or reject a behaviour change?
  2. Can it require evidence before release or autonomy expansion?
  3. Can it assign the repair across product, platform, security, and operations?

If the answer is no, the title adds an attendee, not an owner.

10Human in the harness

“Human in the loop” usually describes one synchronous checkpoint: the agent proposes, a person approves, and the action proceeds.

That pattern is useful for consequential actions. It is not a complete operating model for human judgement.

At volume, human attention becomes a service with demand, capacity, routing, quality, and delay.

Suppose 20,000 actions run each day. If 4% require review and each review takes three minutes, the queue needs 40 person-hours a day before breaks, training, disagreement, and management.

Episode 05 provided the formula and economics. McKinsey estimated human oversight at 70% to 75% of variable run cost in one banking customer-service example, with tokens at 20% to 25%.[5] Those figures are example-specific, but the organisational lesson is general: review capacity must be designed, funded, and owned.

A human-review service needs:

Measure Question
Escalation rate What share of work reaches people?
Queue length and age Is demand exceeding capacity?
Review time What does judgement cost, and how long does the user wait?
Service level How long may a case remain undecided?
Override rate Are routed cases ones where people change the outcome?
Reviewer agreement Is the policy clear enough to apply consistently?
Return path Does the decision update the workflow, trace, and future evals?

Do not route every case to people because the agent is untrusted. That converts automation into data entry. Do not remove people because review is expensive. That hides unresolved ambiguity inside production.

Route people where judgement changes the decision: new failure patterns, policy ambiguity, high consequence, weak verification, or a boundary change.

Call this human in the harness. People are part of escalation, correction, learning, and governance, not a ceremonial approval box.

11Skills need product ownership

A skill packages task-specific instructions, examples, and tool access that load when relevant.

Skills solve the progressive-disclosure problem from Episode 02. They also create a new lifecycle:

  1. Who writes the skill?
  2. Which eval proves it works?
  3. Who approves new tool access?
  4. How is it versioned?
  5. When is it retired?

Treat a skill as a product artifact, not a personal prompt file.

A small organisation can assign ownership inside each product team. A larger one may use a rotating guild to define templates, review standards, and deprecation policy. The guild should not own domain correctness. It owns the quality of the skill-making process.

12What to keep, what to buy

Own what depends on your circumstances. Buy what improves mainly through vendor scale.
Keep in-house Buy or reuse
Definition of good Model APIs
Workflow and escalation rules Generic observability plumbing
Domain failure taxonomy Runtime primitives
Autonomy decisions Common tool libraries
Product-specific memory policy Protocol implementations
Release and retirement judgement Infrastructure that does not differentiate the workflow

The boundary changes as vendors absorb common work. Review it instead of defending a permanent build-versus-buy ideology.

OpenAI’s Codex harness is Apache-2.0, so the choice now includes fork as well as build or buy.[3]

A fork creates a product obligation. Name who owns it, who reviews upstream changes, which security patches must be merged, which workflow eval gates an upgrade, and how the organisation returns to the upstream project or another runtime. A fork without an owner is custom infrastructure disguised as optionality.

Episode 05 called the preferred posture “rent operations, own meaning.” This is the organisational version of that choice.

Ownership also arrives from outside the org chart. From 2 August 2026, the European Commission and national authorities began enforcing AI Act requirements, including transparency obligations for interactive AI systems and generated content.[8] Product, legal, security, and platform therefore need a named owner for disclosure, evidence, complaints, and vendor-change review. A technical build-versus-buy decision does not transfer every downstream obligation.

13Two companion boundaries

The eight-episode series focuses on the harness, the layer that decides.

Two adjacent boundaries deserve their own field guides.

The Tool Is the Contract

A tool is where model intent becomes an operation. Its contract includes more than input fields. It includes caller identity, authority, side effects, reversibility, idempotency, failure behaviour, cost, and verification. Read the tool companion when your ownership problem centers on what the agent may request and how the result is proved.

The Environment Is the Product Boundary

The environment is where an approved action gains reach. Process, filesystem, network, identity, and persistence determine the blast radius. Read the environment companion, and the full Environment Engineering series, when the question is what execution can touch, what survives, and who owns recovery.

The bonus essays do not add two more harness clusters. They slow down at the two boundaries where decision becomes action and action becomes consequence.

14In practice: the ownership page

Take the one-page output from Episode 06 and add five sections.

Figure 03 · Practice
The ownership page
WHO DECIDES · WHO IMPLEMENTS · WHO VERIFIES · WHO APPROVES LAYER OWNER DECISION RIGHT EVIDENCE ESCALATION Model Applied AI Routing and provider Cost + quality run Product, finance Harness Harness PM + eng lead Completion, evals, autonomy Eval run + traces Business owner Tools Platform eng Contract and availability Verified result Security, legal Environment Security, SRE Identity and limits Containment log Incident command • Shared primitives — platform • Local judgement — workflow team • Human review — owner, capacity, SLA TONIGHT TEST If a production failure happened tonight, can you name the four people? “Shared” is not an answer.
Read it as One page, four columns. Every row must resolve to a person or a role before the next autonomy expansion is approved.
Section What to record
Layer owners Model, Harness, Tools, Environment
Decision rights What each owner may approve without committee
Shared primitives What platform owns for all workflows
Local judgement What the workflow team must control
Human-review service Owner, capacity, routing, service level
If a production failure happened tonight, could the team name who decides the fix, who implements it, who verifies it, and who approves wider rollout?

A “shared” answer is incomplete. Name people or roles.

15Connecting the dots

The org chart should follow the behaviour the product can control.

Reuters reported on 26 August 2026 that Meta had considered workforce reductions of up to 60% in some teams while planning AI-enabled pods, then scaled the programme back after internal agents and coding tools produced reliability, security, and productivity problems.[4]

Reuters based the report on internal materials, recordings, and interviews. It is not an audited agent evaluation, and Meta described the work as scenario planning. The bounded lesson is still valuable: changing the organisation before the behaviour is controllable turns technical uncertainty into workforce risk.

The sequence from Episodes 01 to 07 is now complete:

  1. Diagnose the first broken contract.
  2. Locate the responsible component.
  3. Change the smallest decision surface.
  4. Hold the resulting trade-off.
  5. Price the complete workflow.
  6. Turn evidence into an operating plan.
  7. Give that plan decision rights and owners.

Deterministic software concentrates much of its product judgement before release. Agent behaviour continues to change after release with context, tools, models, traffic, and policies. Product judgement must therefore continue through trace review, eval updates, autonomy decisions, incident learning, and retirement of old controls.

That is why this role sits between product management and reliability engineering. It does not exist because prompts need a manager. It exists because changing behaviour needs accountable release authority.

Episode 08 opens that release loop. The finale asks what remains valuable when models absorb more of today’s scaffolding.

You now hold
Operating model 07 An ownership map across four layers, three internal models with their failure modes, a Harness PM defined by decision rights, and a human-review service with metrics.
The next question
What stays valuable when the model absorbs more of today’s harness work?
Continue
Harness 08 When the Harness Becomes the Habit → — the season finale on keeping the operating loop while retiring dead scaffolding.
Read alongside
Environment 00 The World Around the Agent → — the boundary your security and infrastructure owners actually operate.
Sources
  1. LangChain — “State of Agent Engineering,” 12 June 2026.
    langchain.com/state-of-agent-engineering
  2. Lance Martin, Gabe Cemaj, and Michael Cohen, Anthropic — “Scaling Managed Agents: Decoupling the brain from the hands,” 2026.
    anthropic.com/engineering/managed-agents
  3. OpenAI — “Codex as a platform: build on the open agent harness,” August 2026.
    developers.openai.com/blog/codex-as-a-platform
  4. Greg Bensinger, Reuters — “How Meta’s AI layoff plans went kaput,” 26 August 2026.
    reuters.com/technology/artificial-intelligence/how-metas-ai-workforce-transformation-plans-went-kaput-2026-08-26
  5. Chandana Asif, Dieter Kiewell, Lari Hämäläinen, Tom Kolaja, and Tunde Olanrewaju, McKinsey/QuantumBlack — “Where AI agents pay off,” 24 August 2026.
    mckinsey.com/capabilities/quantumblack/our-insights/where-ai-agents-pay-off-a-practical-guide-to-the-economics-of-agentic-workflows
  6. JAPAN AI — “Agent Harness Engineer,” role listing, accessed September 2026.
    hrmos.co/pages/geniee/jobs/2237020752155283473
  7. OpenAI — “Forward Deployed Engineer, Legal,” role listing, accessed September 2026.
    openai.com/careers/forward-deployed-engineer-(fde)-legal-sf-new-york-city
  8. European Commission — “Commission starts enforcing AI Act rules and new transparency requirements,” 31 July 2026, effective 2 August 2026.
    digital-strategy.ec.europa.eu/en/news/commission-starts-enforcing-ai-act-rules-and-new-transparency-requirements-2-august