- LocateWhy agent ownership falls between existing functions.
- CompareThree internal operating models and the failure each creates when misapplied.
- DefineThe decision rights a Harness PM or Agent Reliability PM needs.
- OperateHow to turn human review into a service rather than a reassuring phrase.
- KeepWhat judgement must stay in-house even when the infrastructure is rented.
01Where Episode 06 left us
At 9:30 on Monday morning, the reconciliation team opens the one-page output from Episode 06.
The page names four tickets. Every ticket has an owner. The system still does not.
Product owns the instructions. Engineering owns memory and retries. Platform owns tool contracts. Security owns credentials and runtime limits. Finance experts define correct invoice treatment. The eval threshold lives in a spreadsheet maintained by one analyst.
Each component has a contributor. No one can answer the full product question:
Is this agent’s behaviour acceptable enough to release, expand, or keep running?
The first six episodes brought the team here:
- Diagnose the first broken contract.
- Locate the responsible component and control.
- Replace vague choices with specific signals.
- Test what each control gains and costs.
- Price the completed workflow, not the model call.
- Turn evidence into four owned tickets.
Episode 07 addresses the organisational gap the audit exposed. The problem is not that nobody works on the agent. The problem is that contribution is distributed while end-to-end behaviour remains unowned.
02The org chart built for different software
Two quarters after the 2 AM invoice failure, the company has three agents in production and a fourth in testing.
The COO asks:
If agents are now part of how work gets done, who owns their behaviour?
The existing org chart has natural homes for important parts:
- Applied AIEvaluates and routes models.
- Platform engineeringOwns APIs, tool infrastructure, and shared services.
- Security & infraOwns identity, credentials, networks, and execution boundaries.
- Business teamsOwn domain policy and operating outcomes.
The unresolved decisions sit between them:
- What context reaches the model?
- Which tools appear for this workflow?
- What counts as complete?
- Which production failures become evals?
- When must a person review the case?
- When may autonomy expand or contract?
- Which behaviour change is safe to release?
Those decisions form the product layer of the harness.
The first organisational change is not a new department or title. It is a named decision owner with authority to accept, reject, narrow, or stop the behaviour.
Shared implementation is normal. Shared accountability without a decider is not an operating model.
03The ownership map
| Layer | Natural owner | Decision it can make | Decision it cannot settle alone |
|---|---|---|---|
| Model | Applied AI or platform | Model selection, routing, provider integration | Whether the full workflow is acceptable |
| Harness | Named product owner plus engineering lead | Context policy, completion, evals, autonomy, release | Model, tool, and environment implementation alone |
| Tools | Platform engineering | API contract, availability, implementation | Whether the agent should receive the action now |
| Environment | Security and infrastructure | Identity, network, storage, execution limits | Whether the workflow outcome meets user intent |
The harness owner is accountable for end-to-end behaviour. The role does not absorb every adjacent function.
Use a four-part test:
- Decision: What may this role approve or reject?
- Evidence: What must it inspect before deciding?
- Control: Which team can implement the decision?
- Escalation: Who resolves conflict or accepts exceptional risk?
Assigning everything to one person creates a fictional super-role. Assigning nothing creates shared responsibility without decision rights.
04Four decisions most organisations leave implicit
1. Who defines good?
A finance expert decides how a partial payment should be treated. Product turns that judgement into a versioned requirement. Engineering implements the check. The eval owner ensures the case blocks a bad release.
No one role can perform every step. One role must ensure the chain exists.
2. Who decides what the agent may do?
Product decides which action belongs in the user journey. Security defines the authority policy. Platform implements the permission. Operations defines recovery and rollback.
“Product and security” is not a complete answer until the organisation says who decides when they disagree.
3. Who decides the work is complete?
Engineering can implement processed_count == eligible_count. The workflow
owner defines why that condition, the exception states, and the reconciliation checks
constitute done.
The code owns enforcement. The business owns meaning.
4. Who may expand autonomy?
Product proposes the value. Engineering provides performance evidence. Security and operations assess consequence and recovery. The accountable product owner accepts, narrows, or rejects the change within a pre-agreed risk boundary.
An executive or risk committee should be required only when the change crosses that boundary.
If every decision needs a new committee, the organisation has participants but no operating model. If one team can expand autonomy without evidence or challenge, it has authority without governance.
05Shared platform, product domain
Most ownership arguments dissolve once a team separates the parts of the harness that should be shared from the parts that must stay with the product.
The frontier labs have not published the org charts behind these systems, so this episode does not infer them. Their product interfaces still reveal a useful ownership split.
OpenAI’s August 2026 Codex as a platform article says the Codex harness manages conversation state, tool use, streaming, the configured sandbox, and approval enforcement. The host application decides where the agent runs, what it can reach, which actions need approval, how work is observed, and how results return to the system of record.[3]
Anthropic’s Managed Agents architecture separates the durable session, the harness loop, and the sandbox so each can fail or be replaced independently. It also keeps credentials outside the sandbox and leaves context management in the harness while the session preserves durable history.[2]
The implementations differ. Both expose a stable infrastructure layer while leaving workflow meaning and application policy to the product using it.
Test If a product team cannot change its definition of done without a platform release, the line is in the wrong place.
06Three internal models
The useful variable is where workflow judgement sits.
Model 1 · Platform-owned primitives
A central team owns shared infrastructure:
- AccessModel providers and identity integration.
- RuntimeDeployment and trace transport and storage.
- ControlsCommon permission framework and tool standards.
- MeasurementEval execution and cost accounting.
Product teams own domain rules, completion, escalation, and workflow evals.
How it fails The platform team starts owning workflow decisions as well as infrastructure. Every product waits for the same backlog, local teams build shadow loops and tool wrappers to move faster, and the organisation pays for the central platform plus the duplication it was meant to remove.
When it works Several teams use genuinely common primitives, and the central group can operate them more reliably and cheaply than each team can.
Operating signal Product teams can change workflow context, completion, and evals without requesting a platform roadmap slot. If they cannot, the platform owns too much.
Model 2 · Product-owned harness
Each product team owns its workflow behaviour, orchestration, and evals. Shared infrastructure stays thin.
How it fails Six teams create six retry policies, six eval styles, and six failure taxonomies. A lesson found by one team does not travel, and provider outages, policy changes, or common tool defects require several separate repairs.
When it works The organisation has a small number of different workflows. Customer support, contract review, code assistance, and incident response need different state, tools, latency, and approval rules, and local ownership keeps iteration close to the user outcome.
Operating signal Listen for repeated conversations beginning with “How are you handling…”. When several teams independently solve the same runtime or eval-infrastructure problem, a shared primitive is ready to be extracted.
Model 3 · Harness as operating model
The agent and its control system are the product.
Product managers write behaviour and autonomy specifications. Engineers build inside the control plane. Designers own interaction, refusal, progress, and review surfaces. Evals act as release criteria. Operations watches completed workflows rather than model uptime alone.
How it fails An incumbent renames teams without changing code, incentives, release evidence, or decision rights. A transformation office appears. The operating model remains unchanged beneath new labels.
When it works AI-native companies with one dominant product shape can design this way from the beginning.
Operating signal If eval review, autonomy review, and failure triage are not normal release rituals, the company has adopted the language, not the operating model.
07Choosing among them
Do not use fixed employee or team counts as universal thresholds. Choose using three questions:
| Question | Favors shared platform | Favors product ownership |
|---|---|---|
| How similar are the workflows? | Similar tools and controls | Different data, outcomes, and risk |
| Where does iteration happen? | Infrastructure changes dominate | Domain behaviour changes dominate |
| What repeats across teams? | Runtime, identity, traces, eval execution | Completion, escalation, domain judgement |
Most mature organisations land on a layered version: platform-owned primitives with product-owned workflow harnesses. The interface between them matters more than the label.
08The Harness PM
Titles vary: Harness PM, Agent Reliability PM, Agent Product Lead, or Applied AI Lead. The title matters less than the decision rights.
The role in one sentence
The Harness PM is accountable for the quality, authority, economics, and release readiness of the complete agent workflow.
The role does not personally implement every layer. It makes sure the layers produce one acceptable product behaviour.
What the role owns
- DefinitionThe workflow’s definition of good and done.
- EvalsEval coverage and the prioritised failure backlog.
- AutonomyThe autonomy boundary and evidence required to change it.
- RoadmapThe harness roadmap, including retirement of temporary controls.
- ReleaseRelease readiness for behaviour changes.
- EconomicsCost per accepted result and the evidence required to expand delegated volume.
- CoordinationAcross model, tool, runtime, environment, and domain owners.
What the role does not own alone
- ModelTraining or every provider integration.
- ToolsEvery tool API.
- RuntimeIAM, network, sandbox, or runtime implementation.
- IncidentsIncident command.
- LegalLegal interpretation.
- DomainDomain judgement without the domain expert.
Decision table
| Decision | Decider | Required evidence | Implementers and required partners | Escalation |
|---|---|---|---|---|
| Definition of done | Workflow product owner | Business rule, examples, completion evidence | Domain expert, engineering, eval owner | Business owner |
| Eval coverage and release gate | Harness PM with engineering lead | Failure inventory, baseline, holdout, counter-metrics | Domain expert, quality, Applied AI | Product leadership |
| Tool authority | Product and security within agreed policy | Consequence, identity, reversibility, verification | Platform, legal, operations | Risk owner |
| Model routing | Applied AI or platform | Workflow quality, latency, cost, fallback result | Harness PM, finance | Engineering leadership |
| Autonomy expansion | Harness PM within approved risk envelope | Completion, failure loss, review load, rollback | Security, operations, domain owner | Executive or risk committee if boundary changes |
| Incident command | SRE or security | Live incident evidence | Harness PM, business owner, vendor | Existing incident policy |
The table prevents two bad outcomes: a PM held accountable for controls they cannot influence, and a platform team operating business behaviour it cannot define.
09What good ownership looks like
Good ownership is visible in decisions, not titles.
- FailuresThe top uncovered failure modes are known.
- RegressionRecent production failures become regression cases.
- AutonomyAutonomy changes cite evidence, limits, and rollback.
- VersioningPrompts, context policy, tools, permissions, and completion rules have owners and versions.
- Temporary controlsHave review and retirement conditions.
- Stop authorityProduct, platform, security, and operations know who can stop or narrow the workflow.
- EconomicsCost per accepted result and human-review capacity have owners.
- IncidentsA production incident can be assigned without creating a new committee.
Avoid a vague “eval coverage percentage” unless the denominator is defined. Coverage of what: tool selection, high-value journeys, known failure classes, traffic, or regulatory obligations?
Titles are spreading faster than decision rights
A 2026 JAPAN AI listing for an Agent Harness Engineer covers execution engines, orchestration, session management, policy guardrails, model routing, and related platform concerns. Evaluation ownership is not explicit in the published scope.[6]
OpenAI’s Forward Deployed Engineer, Legal role uses a different title but includes the work this episode cares about: deploying AI into consequential legal workflows, working across product, engineering, research, legal, privacy, and security, and turning operational experience into reusable systems.[7]
Neither posting proves a universal organisation design. Together they make the useful point: the work may appear under engineering, product, reliability, or forward-deployed titles.
Test the role with three questions:
- Can it accept or reject a behaviour change?
- Can it require evidence before release or autonomy expansion?
- Can it assign the repair across product, platform, security, and operations?
If the answer is no, the title adds an attendee, not an owner.
10Human in the harness
“Human in the loop” usually describes one synchronous checkpoint: the agent proposes, a person approves, and the action proceeds.
That pattern is useful for consequential actions. It is not a complete operating model for human judgement.
At volume, human attention becomes a service with demand, capacity, routing, quality, and delay.
Suppose 20,000 actions run each day. If 4% require review and each review takes three minutes, the queue needs 40 person-hours a day before breaks, training, disagreement, and management.
Episode 05 provided the formula and economics. McKinsey estimated human oversight at 70% to 75% of variable run cost in one banking customer-service example, with tokens at 20% to 25%.[5] Those figures are example-specific, but the organisational lesson is general: review capacity must be designed, funded, and owned.
A human-review service needs:
| Measure | Question |
|---|---|
| Escalation rate | What share of work reaches people? |
| Queue length and age | Is demand exceeding capacity? |
| Review time | What does judgement cost, and how long does the user wait? |
| Service level | How long may a case remain undecided? |
| Override rate | Are routed cases ones where people change the outcome? |
| Reviewer agreement | Is the policy clear enough to apply consistently? |
| Return path | Does the decision update the workflow, trace, and future evals? |
Do not route every case to people because the agent is untrusted. That converts automation into data entry. Do not remove people because review is expensive. That hides unresolved ambiguity inside production.
Route people where judgement changes the decision: new failure patterns, policy ambiguity, high consequence, weak verification, or a boundary change.
Call this human in the harness. People are part of escalation, correction, learning, and governance, not a ceremonial approval box.
11Skills need product ownership
A skill packages task-specific instructions, examples, and tool access that load when relevant.
Skills solve the progressive-disclosure problem from Episode 02. They also create a new lifecycle:
- Who writes the skill?
- Which eval proves it works?
- Who approves new tool access?
- How is it versioned?
- When is it retired?
Treat a skill as a product artifact, not a personal prompt file.
A small organisation can assign ownership inside each product team. A larger one may use a rotating guild to define templates, review standards, and deprecation policy. The guild should not own domain correctness. It owns the quality of the skill-making process.
12What to keep, what to buy
Own what depends on your circumstances. Buy what improves mainly through vendor scale.
| Keep in-house | Buy or reuse |
|---|---|
| Definition of good | Model APIs |
| Workflow and escalation rules | Generic observability plumbing |
| Domain failure taxonomy | Runtime primitives |
| Autonomy decisions | Common tool libraries |
| Product-specific memory policy | Protocol implementations |
| Release and retirement judgement | Infrastructure that does not differentiate the workflow |
The boundary changes as vendors absorb common work. Review it instead of defending a permanent build-versus-buy ideology.
OpenAI’s Codex harness is Apache-2.0, so the choice now includes fork as well as build or buy.[3]
A fork creates a product obligation. Name who owns it, who reviews upstream changes, which security patches must be merged, which workflow eval gates an upgrade, and how the organisation returns to the upstream project or another runtime. A fork without an owner is custom infrastructure disguised as optionality.
Episode 05 called the preferred posture “rent operations, own meaning.” This is the organisational version of that choice.
Ownership also arrives from outside the org chart. From 2 August 2026, the European Commission and national authorities began enforcing AI Act requirements, including transparency obligations for interactive AI systems and generated content.[8] Product, legal, security, and platform therefore need a named owner for disclosure, evidence, complaints, and vendor-change review. A technical build-versus-buy decision does not transfer every downstream obligation.
13Two companion boundaries
The eight-episode series focuses on the harness, the layer that decides.
Two adjacent boundaries deserve their own field guides.
The Tool Is the Contract
A tool is where model intent becomes an operation. Its contract includes more than input fields. It includes caller identity, authority, side effects, reversibility, idempotency, failure behaviour, cost, and verification. Read the tool companion when your ownership problem centers on what the agent may request and how the result is proved.
The Environment Is the Product Boundary
The environment is where an approved action gains reach. Process, filesystem, network, identity, and persistence determine the blast radius. Read the environment companion, and the full Environment Engineering series, when the question is what execution can touch, what survives, and who owns recovery.
The bonus essays do not add two more harness clusters. They slow down at the two boundaries where decision becomes action and action becomes consequence.
14In practice: the ownership page
Take the one-page output from Episode 06 and add five sections.
| Section | What to record |
|---|---|
| Layer owners | Model, Harness, Tools, Environment |
| Decision rights | What each owner may approve without committee |
| Shared primitives | What platform owns for all workflows |
| Local judgement | What the workflow team must control |
| Human-review service | Owner, capacity, routing, service level |
If a production failure happened tonight, could the team name who decides the fix, who implements it, who verifies it, and who approves wider rollout?
A “shared” answer is incomplete. Name people or roles.
15Connecting the dots
The org chart should follow the behaviour the product can control.
Reuters reported on 26 August 2026 that Meta had considered workforce reductions of up to 60% in some teams while planning AI-enabled pods, then scaled the programme back after internal agents and coding tools produced reliability, security, and productivity problems.[4]
Reuters based the report on internal materials, recordings, and interviews. It is not an audited agent evaluation, and Meta described the work as scenario planning. The bounded lesson is still valuable: changing the organisation before the behaviour is controllable turns technical uncertainty into workforce risk.
The sequence from Episodes 01 to 07 is now complete:
- Diagnose the first broken contract.
- Locate the responsible component.
- Change the smallest decision surface.
- Hold the resulting trade-off.
- Price the complete workflow.
- Turn evidence into an operating plan.
- Give that plan decision rights and owners.
Deterministic software concentrates much of its product judgement before release. Agent behaviour continues to change after release with context, tools, models, traffic, and policies. Product judgement must therefore continue through trace review, eval updates, autonomy decisions, incident learning, and retirement of old controls.
That is why this role sits between product management and reliability engineering. It does not exist because prompts need a manager. It exists because changing behaviour needs accountable release authority.
Episode 08 opens that release loop. The finale asks what remains valuable when models absorb more of today’s scaffolding.
-
LangChain — “State of Agent Engineering,” 12 June 2026.
langchain.com/state-of-agent-engineering -
Lance Martin, Gabe Cemaj, and Michael Cohen, Anthropic — “Scaling
Managed Agents: Decoupling the brain from the hands,” 2026.
anthropic.com/engineering/managed-agents -
OpenAI — “Codex as a platform: build on the open agent harness,”
August 2026.
developers.openai.com/blog/codex-as-a-platform -
Greg Bensinger, Reuters — “How Meta’s AI layoff plans went
kaput,” 26 August 2026.
reuters.com/technology/artificial-intelligence/how-metas-ai-workforce-transformation-plans-went-kaput-2026-08-26 -
Chandana Asif, Dieter Kiewell, Lari Hämäläinen, Tom Kolaja, and Tunde
Olanrewaju, McKinsey/QuantumBlack — “Where AI agents pay off,” 24
August 2026.
mckinsey.com/capabilities/quantumblack/our-insights/where-ai-agents-pay-off-a-practical-guide-to-the-economics-of-agentic-workflows -
JAPAN AI — “Agent Harness Engineer,” role listing, accessed
September 2026.
hrmos.co/pages/geniee/jobs/2237020752155283473 -
OpenAI — “Forward Deployed Engineer, Legal,” role listing,
accessed September 2026.
openai.com/careers/forward-deployed-engineer-(fde)-legal-sf-new-york-city -
European Commission — “Commission starts enforcing AI Act rules and new
transparency requirements,” 31 July 2026, effective 2 August 2026.
digital-strategy.ec.europa.eu/en/news/commission-starts-enforcing-ai-act-rules-and-new-transparency-requirements-2-august