Harness Engineering · Episode 05

What It Costs, What It Returns

The model call is one line on the bill. The product is the completed workflow.

Arc · The economics Episode · 05 of 8 Next · Episode 06 — Your Monday Morning Harness Kit
After this you will know
  • BillWhy the provider invoice is not the total cost of an agent product.
  • BucketsHow to separate build, run, review, and failure costs.
  • UnitWhy cost per accepted result is more useful than cost per call.
  • ValueHow reliability can reduce losses and increase delegated volume.
  • DecisionHow to end a monthly review with expand, hold, or narrow.

01Where Episode 04 left us

At 10 AM on Monday, the CFO opens the quarterly spend report.

Model prices fell. The AI line rose.

The dashboard explains tokens in detail. It does not show the finance experts who labelled difficult invoice cases, the engineer who made failures replayable, the reviewers clearing exceptions, or the runtime preserving state between attempts.

The number is accurate. The picture is incomplete.

Episodes 01 to 04 built one chain of reasoning:

  1. Episode 01 found the first broken contract.
  2. Episode 02 located the responsible component and control.
  3. Episode 03 replaced vague choices with specific signals.
  4. Episode 04 asked what each control gained and what it cost.

This episode puts money around that system.

The economics are not separate from reliability. They are reliability viewed from finance.

02The budget meeting

The CFO asks:

If intelligence became cheaper, why did the agent become more expensive?

The answer begins with a distinction.

A model produces an answer or proposes an action. An agent product must also retrieve evidence, call tools, preserve state, enforce authority, verify the result, recover from failure, and send uncertain cases to a person.

Those activities create different bills.

For the invoice agent, a completed workflow means more than “the model responded.” It means:

The business is not buying tokens. It is buying that completed outcome.

This gives us the central rule for the episode:

Price the work from request to accepted business result.

The reader map

Remember Meaning
Four cost buckets Build, Run, Review, Failure
Three unit measures Cost per attempt, cost per accepted result, net workflow value
Two reliability returns Lower failure cost and more delegated volume
One PM rule Price the completed outcome, not the model call

The rest of the episode explains each line slowly, then puts them together in one business case.

03Four cost buckets

Use four ordinary words whenever the bill becomes confusing.

1. Build

What did it cost to create and change the system?

This includes product discovery, engineering, domain rules, eval cases, tool integration, security review, migration, and release work.

Some build cost happens once. Much of it returns whenever the model, tool, policy, or workflow changes.

2. Run

What did it cost to attempt the work?

This includes input and output tokens, cache writes and reads, web search, file search, code execution, API calls, queues, databases, containers, storage, and observability.

Run cost grows with traffic and with the number of steps inside each workflow.

3. Review

What did it cost people to inspect or decide cases?

This includes routine approval, expert review, escalation, reviewer tooling, calibration, and service-level coverage.

Review cost usually grows with volume, ambiguity, and consequence.

4. Failure

What did wrong or unfinished work cost?

This includes retries, investigation, rework, reversals, customer support, incidents, regulatory exposure, and work that the organisation refuses to delegate after trust falls.

These four buckets are deliberately plain:

Bucket Question Invoice example
Build What did we create or change? Define partial-payment rules and add regression tests
Run What did the system consume? Model calls, invoice API, state store, traces
Review What did people inspect or decide? Finance specialist reviews an ambiguous match
Failure What did wrong work cost? Reverse a bad update and investigate the incident

Model tokens sit inside Run. They are not the whole run cost, and run cost is not the whole business cost.

04What a real API bill contains

Provider bills differ because providers meter different resources.

Provider pricing at a glance

The providers meter different resources. This table uses public list prices available on 2 September 2026. Use it to understand the bill’s structure, not to select a vendor from this article alone.

Provider example New input Cached input Output Other separately billed items
OpenAI GPT-5.6 Terra, short context $2.00 / 1M tokens $0.20 / 1M cache reads; $2.50 / 1M cache writes $12.00 / 1M tokens Web search, file search, file storage, hosted containers
Anthropic Claude Fable 5.1 $10.00 / 1M tokens $0.25 / 1M cache reads; $12.50 / 1M five-minute cache writes $50.00 / 1M tokens Longer cache duration, web search, code execution, managed runtime
Google Gemini 3.7 Flash, promotional standard pricing $0.75 / 1M tokens $0.075 / 1M cached tokens plus $0.50 / 1M tokens per cache-storage hour $3.75 / 1M tokens Search grounding, maps grounding, file search, live APIs, higher service tiers

Sources: OpenAI API pricing, Anthropic prompt-caching documentation, and Google Gemini API pricing.[9][10][11]

Three details matter more than the headline rates:

  1. The same text can have three prices. New input, a cache write, and a cache read may be billed differently.
  2. Tools sit outside token pricing. Search, retrieval, code execution, containers, and storage can appear as separate line items.
  3. Storage duration can matter. Google charges for cached-context storage time. Anthropic offers different cache lifetimes. A cache is an economic choice, not merely a latency feature.

The rows are not quality-equivalent. This table compares billing anatomy, not model capability.

A small illustrative OpenAI bill

Suppose a support agent completes 1,000 workflows in a month. Each workflow uses GPT-5.6 Terra for six model calls.

Assumptions per workflow:

Using the listed prices, the approximate provider bill is:

OpenAI line Monthly usage Approximate cost
Cache writes 20M tokens $50
Cache reads 100M tokens $20
New input 12M tokens $24
Output 6M tokens $72
Web search 1,000 calls $10, before search-content tokens
Containers 200 eligible 1 GB sessions $6
Provider subtotal $182

This is an illustration, not a quote or forecast. It uses current public list prices and simplified traffic.

Now add costs that will not appear on the OpenAI invoice:

Internal or other-vendor line Illustrative monthly cost
Application runtime, queue, and database $120
Trace storage and observability $80
Human review: 50 cases × 6 minutes × $30/hour $150
Amortised engineering, eval, and domain-review work $1,500
Fully loaded subtotal before failure loss $2,032

If 900 of the 1,000 workflows pass the business completion rule, the provider subtotal is about $0.20 per accepted result. The fully loaded subtotal is about $2.26 per accepted result.

Both numbers are correct. They answer different questions.

Do not present the provider subtotal as the product’s unit cost. It excludes the application, runtime, evidence, review, and build work that makes the result usable.

The provider number helps engineers optimise API use. The fully loaded number helps the business decide whether the workflow is economically worthwhile.

What caching changes

Without caching, the same 20,000-token prefix would be processed as normal input on all six calls. Under the same assumptions, model-token cost would rise from roughly $166 to $336 per 1,000 workflows before search and containers.

Caching does not change the model. It changes the bill because the harness keeps a stable prefix reusable.

It can also fail silently as an economic strategy. Changing a tool definition, system prompt, timestamped block, or cache breakpoint may invalidate the prefix. Anthropic’s documentation explicitly recommends monitoring cache-creation, cache-read, and uncached-input fields rather than assuming a cache hit occurred.[10]

Why a $200 plan is not $200 of compute

Flat-rate subscriptions create a different kind of bill. Three prominent power-user plans share the same headline price but package usage differently.

Plan Public price How usage is described Important limit or economic detail
ChatGPT Pro $200 / month OpenAI positions it for AI power users Access and availability still depend on model, tool, capacity, and plan terms
Claude Max 20x $200 / month Twenty times the session usage of Claude Pro Five-hour reset plus an additional weekly limit shared across Claude and Claude Code
Cursor Ultra $200 / month Twenty times the usage of Cursor Pro Cursor says long-term model-provider partnerships help support the predictable price

Sources: OpenAI, Anthropic, and Cursor plan documentation.[12][13][14]

These plans are access contracts, not prepaid wallets containing exactly $200 of API credit.

The customer sees one fixed price. The vendor sees a distribution:

The plan works at the cohort level. Revenue from light and typical users can offset heavy users. Limits, routing, caching, negotiated model rates, and capacity management protect the vendor.

If a subscriber consumes $1,000 of usage valued at public API prices, it does not follow that the vendor lost $800. Public API price is not the vendor’s marginal cost. The provider may run its own model more cheaply, negotiate lower rates, route work, reuse cached context, or impose limits. Actual loss requires the vendor’s real cost to serve the account, which public plan pages do not disclose.

There is still a real warning. In January 2025, TechCrunch reported that OpenAI CEO Sam Altman said the then-$200 ChatGPT Pro plan was losing money because subscribers used it more than expected.[15] Treat this as a historical signal about forecasting risk, not evidence of OpenAI’s current per-user economics.

A flat price transfers usage risk from the customer to the vendor.

What this means for an AI startup

Suppose a product charges $100 per customer per month. The average direct AI and runtime cost looks manageable until customers are separated by usage.

Customer group Share Revenue each Direct AI and runtime cost each Contribution each
Light 60% $100 $8 $92
Typical 30% $100 $35 $65
Power 10% $100 $260 −$160

Across 100 customers, revenue is $10,000. Direct AI and runtime cost is $4,130. Contribution before support and fixed cost is $5,870, or 58.7%.

The plan is profitable as a cohort even though every power user loses money on direct service cost. If growth concentrates in that segment, margin falls quickly. If those users also receive the most value, crude throttling may destroy the product’s best use case.

The AI PM must decide:

  1. Which expensive behaviour creates customer value?
  2. Which behaviour creates cost without increasing accepted results?
  3. Should heavy use be included, metered, queued, rate-limited, or moved to a higher tier?
  4. Can routing, caching, smaller context, or fewer failed retries preserve the experience?
  5. Are expensive users strategically valuable because they retain, expand, or teach the product more?

Do not price from the average alone. Track median, 90th-percentile, and 99th-percentile cost to serve, then inspect the workflows behind the tail.

Use the margin terms correctly

Measure Plain meaning Why the AI PM needs it
Gross margin Revenue left after direct cost of serving customers Shows whether inference, tools, runtime, and direct review fit the business model
Contribution margin Revenue left from a customer, workflow, or cohort after attributable variable costs Reveals power users or workflows hidden by the company-wide average
Net margin Revenue left after wider research, product, sales, and administrative expenses Describes the company’s overall profitability

“This user has negative contribution margin” and “this company loses money” are not the same claim.

For enterprise contracts, calculate contribution margin by customer and workflow. A large annual commitment can hide one automation whose marginal cost grows faster than its usage revenue.

05Why cheaper calls can produce a larger bill

An agent workflow often performs many more operations than a chat answer:

  1. Interpret the task.
  2. Retrieve current evidence.
  3. Select a tool.
  4. Process the tool result.
  5. Repair invalid output.
  6. Verify the effect.
  7. Continue after partial failure.
  8. Ask another model or person to judge ambiguity.
  9. Compact or reconstruct context.
  10. Repeat until completion or a stop condition.

The price of each model call can fall while total cost rises because the product handles more traffic, takes more steps, carries more context, invokes paid tools, runs longer, and verifies more often.

That increase is not automatically bad. A more expensive run may produce a result the business can actually accept.

Figure 01 · Concept
One cheap call, one expensive workflow
UNIT PRICE FALLS. CONSUMPTION RISES. CHAT 1–2 calls one answer · usually unverified AGENT WORKFLOW interpret · retrieve · act repair · verify · judge THREE MULTIPLIERS Context What travels into each call Iteration Retry, verify, and continue Parallelism Workers, candidate paths, judges The harness sets every multiplier: model per step, context included, retries allowed, and when the workflow stops. COST CONTROL IS A CONTROL-PLANE RESPONSIBILITY
Read it as Procurement negotiates the price per unit. Product and engineering design how many units one accepted workflow consumes.

The main levers are model routing, context size, cache hit rate, retries, parallel workers, judge calls, paid tools, and stopping rules.

06Five costs the token bill hides

The four buckets tell you where money goes. The five cost centres below identify work teams most often omit from the budget.

Hidden cost Bucket What the organisation is buying
Ground truth and evals Build A testable definition of acceptable behaviour
Senior attention Build and Review Translation of business judgement into system decisions
Observability and replay Run and Build Evidence that makes failure diagnosable and repairs testable
Human review Review Qualified judgement for ambiguity and consequence
Versioning and migration Build The ability to change models, tools, policies, and prompts safely

Runtime remains a separate run-cost line in Section 12. This preserves the season’s count of five hidden cost centres without pretending that runtime is free.

07Ground truth and evals

Before a team can measure an agent, it must decide what acceptable work means.

Consider a ₹10,00,000 invoice with a payment of ₹8,00,000. Should the agent mark it paid, apply a partial match, leave it pending, or create an exception?

The model can propose an answer. Ground truth records the organisation’s approved answer and the conditions behind it. An eval reruns the case after a change and checks whether the system still behaves correctly.

The expensive part is not running the test. It is making business judgement explicit.

For reconciliation, the eval set may cover exact matches, partial payments, duplicate payments, tenant isolation, exception states, final ledger counts, and escalation when evidence is insufficient.

Someone must label cases, resolve disagreement, write checks, define rubrics, create holdouts, and update the set when policy changes.

Features produce visible output. Ground truth may look like a spreadsheet until a silent regression reaches production.

Budget eval work beside feature work, not after it.

08Senior attention

Harness work consumes people with scarce context.

A staff engineer understands architecture and failure propagation. A product manager translates among the user outcome, business rules, security, finance, and engineering. Domain experts resolve cases that a generic annotation team cannot judge.

Their time is real even when it never appears on an API invoice.

Name the capacity:

The estimate may be wrong. An unnamed estimate is guaranteed to be missing.

Do not treat all senior attention as one-time implementation. Models, policies, tools, and regulations continue to change. Some expert capacity is operating cost.

09Observability and replay

Three words are often mixed together:

If an invoice was marked paid incorrectly, the trace shows what evidence reached the model and which tool call followed. The eval says the result violated the partial-payment rule. Replay lets the team test the corrected system against the same case.

A trace without judgement is a detailed mystery. An eval without a trace can say “wrong” without explaining why. An incident that cannot be replayed is expensive to repair with confidence.

Useful evidence may include tenant identity, model and harness versions, context sources, tool arguments, approval decisions, retries, state changes, cost, and completion proof.

That evidence creates its own cost: storage, search, access control, redaction, retention, and deletion.

Measure the return in plain operating terms:

  1. How long does it take to find the first divergence?
  2. What percentage of incidents can be reproduced?
  3. How long does it take to turn a failure into a regression test?

10Human review

“Human in the loop” sounds like a safety principle. At production volume, it is a queue.

Suppose 10,000 sessions run each day. If 3% need review, 300 cases enter the queue. At five minutes each, the team needs 1,500 minutes, or 25 person-hours, every day.

The capacity model is:

daily review hours = (sessions × review rate × minutes per review) ÷ 60

A better model may lower the review rate. Wider deployment raises the number of sessions. Harder cases raise review time. Price all three.

The review system also needs an interface, routing by expertise, service levels, timeout behaviour, calibration, disagreement resolution, and coverage outside business hours.

McKinsey and QuantumBlack made this cost visible in August 2026. For one banking customer-service example, they estimated token costs at 20% to 25% of variable run cost and human oversight at 70% to 75%. For customer onboarding, they estimated expert review for 10% to 20% of agentic runs.[5]

These are estimates for specified banking examples, not universal ratios. Their lesson is still direct: a budget with a token line and no review line may be missing the larger variable cost.

11Versioning and migration

Production agents change on several clocks:

Without version discipline, the team cannot reconstruct which model, prompt, policy, and tool contract produced an action.

Every meaningful artifact needs a version, owner, reason for change, release evidence, rollout plan, rollback path, and retirement condition if temporary.

Migration cost is not an unusual event outside the operating model. It is part of the operating model.

12Runtime is a separate line

Even an open-source harness needs somewhere to run.

Runtime includes durable sessions, queues and workers, checkpoints, sandboxes, multi-tenancy, deployment, scaling, secrets, trace transport, and recovery.

Runtime line Question
Durable state What survives worker failure?
Execution Where do tools and code run?
Scaling What concurrency and queue guarantees are required?
Observability What is retained, searchable, and exportable?
Tenancy How are customers isolated?
Recovery How does work resume without duplicate effects?
Operations Who is on call, and what does the vendor own?

A managed platform may bundle these costs into a seat, session, container, or usage allowance. Bundling simplifies purchasing. It does not remove the cost.

Anthropic’s Managed Agents exposes a hard session-spend cap. When the cap is reached, the system can pause the run with its container preserved and return budget_reached, allowing the same work to resume if the limit changes.[7] Cost becomes an explicit state in the harness.

SpaceXAI’s Grok Bot packages always-on agents with a cloud computer and bundled usage.[8] That creates a different finance problem: the organisation must allocate a seat-level price across accepted workflows if it wants to compare the product with API-based or internally hosted alternatives.

13The value of reliability

Reliability creates value in two places.

First, it reduces the cost of wrong work:

Second, it can increase delegated volume, meaning the amount of eligible work the organisation permits the agent to handle.

Suppose 100,000 support cases are eligible. The company initially delegates only 10,000 because mistakes are hard to detect and reverse. If verification, limits, and rollback reduce exposure, the company may delegate 40,000 without changing the model.

That increase is the reliability dividend: more useful work becomes available because the downside is bounded and visible.

In plain language:

Net value equals value from accepted work, minus run and review cost, minus expected failure loss.

A simple model is:

V = (N × ps × vs) − Crun − Creview − (N × pf × Lf)

Where:

Build cost can be subtracted separately or amortised into C_run for the reporting period. State which treatment you use.

The harness can improve value by increasing accepted completion, reducing failure frequency, reducing loss when failure occurs, and increasing delegated volume.

14Autonomy is a loss decision

An autonomy boundary is a loss decision expressed as product scope.

A team may let the agent draft every refund recommendation but issue only low-value, reversible refunds that pass deterministic policy. High-value or ambiguous cases go to review.

As evidence, rollback, and escalation improve, the issuing boundary can move.

The useful question is not “How autonomous is the agent?”

Ask:

Which actions may it perform, under what limits, with what evidence and recovery?

15A worked business case

The following numbers are illustrative assumptions. They explain the model; they are not external evidence.

A support workflow has 100,000 eligible monthly cases.

System A: cheaper attempt, weak controls

The organisation delegates only 10,000 cases because failure is difficult to detect.

System A line Calculation Amount
Value from accepted work 10,000 × 90% × $3 $27,000
Run cost 10,000 × $0.20 −$2,000
Review cost 10,000 × 10% × $4 −$4,000
Expected failure loss 10,000 × 2% × $30 −$6,000
Net value $15,000

System A produces 9,000 accepted results.

Its run, review, and failure cost is $12,000. That is about $1.33 for each accepted result.

System B: more expensive attempt, stronger controls

Verification, routing, and rollback improve. The organisation delegates 40,000 cases.

System B line Calculation Amount
Value from accepted work 40,000 × 94% × $3 $112,800
Run cost 40,000 × $0.28 −$11,200
Review cost 40,000 × 5% × $4 −$8,000
Expected failure loss 40,000 × 0.5% × $30 −$6,000
Net value $87,600

System B produces 37,600 accepted results.

Its run, review, and failure cost is $25,200. That is about $0.67 for each accepted result.

The attempted case became more expensive: $0.20 rose to $0.28.

The accepted result became cheaper: approximately $1.33 fell to $0.67 when review and expected failure loss are included.

Figure 02 · Practice
The expensive attempt is the better investment
SAME MODEL. DIFFERENT DELEGATED VOLUME. SYSTEM A · WEAK CONTROLS 10,000 delegated · 9,000 accepted $0.20 per attempt NET VALUE $15,000 COST PER ACCEPTED RESULT $1.33 SYSTEM B · STRONGER CONTROLS 40,000 delegated · 37,600 accepted $0.28 per attempt NET VALUE $87,600 COST PER ACCEPTED RESULT $0.67 THE CHEAPER ATTEMPT PRODUCED THE MORE EXPENSIVE ACCEPTED RESULT
Read it as The system with more verification costs more each time it tries. Under these assumptions, it costs less for each result the business can accept.

Do not copy the conclusion without replacing the assumptions. System B wins only because the example assumes controls improve completion, failure, review, and delegated volume enough to cover their cost.

FalsifierIf stronger controls do not improve accepted completion, reduce review or failure cost, or unlock enough additional volume, they are not an economic improvement.

16Named evidence, used carefully

Public examples show what happened inside named organisations. They rarely isolate one causal variable.

Azure SRE Agent

Microsoft reported in April 2026 that Azure SRE Agent had handled more than 35,000 incidents autonomously over nine months and saved more than 50,000 developer hours. It also reported that Azure App Service reduced time to mitigation from a 40.5-hour human-only average to three minutes for live-site incidents.[3][4]

These are Microsoft-reported internal outcomes, not an independent controlled trial. The system included access to logs, metrics, incident records, code changes, specialised tools, role-based access, approval boundaries, and ongoing evaluation. The model call was one component.

McKinsey’s completed-work unit

McKinsey recommends measuring the fully loaded cost to complete a job across people, agents, and deterministic systems. In its illustrative banking-onboarding example, it estimates that cost may fall from roughly $50–$150 per customer to $10–$30, despite a workflow involving several agents, deterministic systems, and oversight teams.[5]

Treat those figures as consultancy estimates, not a forecast. The reusable idea is to price the completed job.

Provider pricing makes harness design visible

OpenAI, Anthropic, and Google all price cached context differently from new context. OpenAI also exposes search, file search, storage, and containers as separate lines. Google adds cache-storage time and grounding requests. Anthropic distinguishes cache writes from cache reads and requires exact reusable prefixes.[9][10][11]

The economic question is no longer only “Which model is cheaper?”

It is also:

17Cost per accepted result

Use this operating metric:

cost per accepted result = (attempts + tools + runtime + review + investigation + rework) ÷ accepted results

If you want a fully loaded business metric, add amortised build and migration cost to the numerator. If you want net value, also subtract expected external failure loss from the value produced.

Do not switch definitions between months. Write the numerator under the metric.

Figure 03 · Framework
Cost per accepted result
THE UNIT THAT CONNECTS ENGINEERING TO FINANCE ALL COST attempts · retries · tools · runtime · review · investigation · rework DIVIDED BY ACCEPTED RESULTS outcomes that passed the business completion rule UNRELIABILITY IS PAID FOUR TIMES Retries Repeated attempts, longer runs Investigation Human time finding the cause Rework Corrections, reversals downstream Trust Work withdrawn from the system Every failed attempt lands above the line. Only accepted outcomes land below it.
Read it as Every failed attempt lands above the line. Only accepted outcomes land below it.

A harness can improve the ratio through smaller current context, higher cache reuse, right-sized models, fewer unproductive retries, early stopping, durable resume, and cheap deterministic checks before expensive judgement.

Beyond some point, another verifier costs more than the failure it prevents. Track both sides.

18Quality is part of cost

LangChain’s 2026 survey received 1,340 responses, mostly from technology organisations. It reported quality as the leading production barrier at 32%. It also reported 89% observability adoption, 52.4% offline-eval adoption, and 59.8% use of human review.[1]

The survey is self-selected and technology-heavy. It is not a census of every enterprise.

Its economic lesson is still useful: poor quality creates retries, review, investigation, rework, incidents, and resistance to delegation.

Finance should ask two questions together:

  1. What does one accepted workflow cost?
  2. What prevents the organisation from delegating more of them?

A program that optimises only the first can become cheap and irrelevant.

19What to rent and own

A restaurant may rent its building, payment terminal, and delivery software. It should not outsource the definition of what its food should taste like.

An agent team can rent models, sessions, execution, search, and trace storage. It should retain control of the definition of good, workflow rules, approval policy, escalation, and failure taxonomy.

Usually sensible to rent or reuse

Usually important to control

The hybrid is often correct: rent operations, own meaning.

20The lock-in test

Harness area Portability question
Identity Can we export and version the operating contract?
Memory Can we inspect context assembly, caching, and compaction?
Orchestration Can we change models or runtimes without rewriting the workflow?
Controls Can we control approvals and policy checkpoints?
Evidence Can we export traces, datasets, rubrics, and results?

Lock-in is not automatically bad. It may be acceptable for a commodity workflow.

The cost appears when the workflow becomes differentiating, regulated, or difficult to migrate after years of accumulated data and rules.

Price the exit before you need it.

21When to stop investing

Harness work can overshoot.

The twelfth verifier may cost as much as the second and improve little. A model-specific workaround may be close to becoming unnecessary. Another month on a mature low-risk workflow may produce less value than applying proven controls to a second workflow.

Ask:

  1. Is the eval curve still moving?
  2. Is this component temporary compensation or a permanent business obligation?
  3. Would the next rupee buy more value through deeper reliability here or wider coverage elsewhere?

Stopping is capital allocation. A harness that only grows is accumulating, not compounding.

22When expansion is earned

Expansion is supported when:

Expansion is not a reward for effort. It is a decision supported by outcomes.

23The monthly value review

Do not start with a dashboard. Start with the decision the meeting must make.

Read the workflow from opportunity to outcome:

  1. How much work was eligible?
  2. How much did the organisation delegate?
  3. How much finished acceptably?
  4. What failed, and what did failure cost?
  5. How much human attention was required?
  6. What did each accepted result cost?
  7. Has the workflow earned wider scope?
Line Plain definition Decision supported
Eligible volume All work that could enter the workflow Size of the opportunity
Delegated volume Work the organisation allowed the agent to attempt Whether trust is expanding
Accepted results Outcomes that passed the business completion rule Whether useful work was produced
Expected failure loss Consequential failures × average loss Whether exposure is acceptable
Review hours Reviewed cases × review time Whether attention is becoming selective
Cost per accepted result Defined full cost ÷ accepted results Whether unit economics are improving
Expansion requirement Evidence needed for wider scope Expand, hold, or narrow
Cost by customer cohort Contribution margin for light, typical, and power users Reprice, meter, route, or redesign

The meeting ends with one of three decisions:

If the review cannot change scope or investment, remove it.

24Try this before Friday

Choose one workflow. Use measured numbers where available. Label estimates and unknowns.

Field Record
Eligible monthly volume
Delegated monthly volume
Accepted completion rate
Model and tool cost
Runtime and observability cost
Human-review rate, time, and loaded cost
Consequential failure rate
Average failure loss or range
Amortised build and migration cost
Cost by light, typical, and power-user cohort
Gross and contribution margin
Cost per accepted result
Evidence required to expand
Evidence qualityMeasured, estimated, or unknown
Owner

Then complete one sentence:

The largest uncertainty is [field]. We will measure it through [method], owned by [person or team], before deciding to [expand / hold / narrow].

If the answer is “a better model,” name the failure that model must remove. If the answer is “more trust,” translate trust into completion evidence, limits, rollback, and observed performance.

25What to remember

Remember four buckets:

  1. Build the system.
  2. Run the workflow.
  3. Review uncertain or consequential cases.
  4. Pay for Failure when wrong work escapes or must be repaired.

Remember four economic measures:

  1. Cost per attempted run.
  2. Cost per accepted result.
  3. Contribution margin by customer and workflow cohort.
  4. Net value after review and expected failure loss.

Remember two sources of reliability value:

  1. Failures become fewer, cheaper, or easier to recover from.
  2. The organisation delegates more work because the downside is bounded and visible.

And remember one rule:

Price the completed business outcome, not the model call.

Episode 06 replaces the assumptions in this model with evidence from one real workflow.

You now hold
Economics 05 Four cost buckets, five commonly hidden cost centres, a separate runtime line, a worked provider bill, power-user cohort economics, the value equation, and a monthly review that ends in expand, hold, or narrow.
The next question
The model is priced. What do you do on Monday to replace the remaining assumptions with evidence?
Continue
Harness 06 Your Monday Morning Harness Kit → — nine days to diagnose, then twelve weeks to turn the diagnosis into owned changes.
Read alongside
Environment 06 Decide What Runs Without You → — the autonomy ladder behind delegated volume.
Sources
  1. LangChain, “State of Agent Engineering,” 12 June 2026; survey of 1,340 respondents.
    langchain.com/state-of-agent-engineering
  2. Rakuten, “Rakuten accelerates development with Claude Code,” company-reported workflow outcomes.
    rakuten.today/blog/rakuten-accelerates-development-with-claude-code
  3. Shamir AbdulAziz, Microsoft, “How we build and use Azure SRE Agent with agentic workflows,” 5 April 2026, updated 9 April 2026.
    techcommunity.microsoft.com/blog/appsonazureblog/azure-sre-agent
  4. Microsoft Learn, “Azure SRE Agent overview.”
    learn.microsoft.com/en-us/azure/sre-agent/overview
  5. Chandana Asif, Dieter Kiewell, Lari Hämäläinen, Tom Kolaja, and Tunde Olanrewaju, McKinsey/QuantumBlack, “Where AI agents pay off,” 24 August 2026.
    mckinsey.com/.../where-ai-agents-pay-off
  6. Anthropic, “Claude Fable,” pricing and retention terms, 1 September 2026.
    anthropic.com/claude/fable
  7. Anthropic, Claude Platform release notes, 7 August 2026; Managed Agents session-spend budgets.
    docs.anthropic.com/en/release-notes/api
  8. SpaceXAI, “Introducing Grok Bot,” early beta announcement, 11 August 2026.
    x.ai/news/introducing-grok-bot
  9. OpenAI, API pricing, accessed 2 September 2026.
    developers.openai.com/api/docs/pricing
  10. Anthropic, “Prompt caching,” Claude Platform documentation, accessed 2 September 2026.
    platform.claude.com/docs/.../prompt-caching
  11. Google, Gemini Developer API pricing, accessed 2 September 2026.
    ai.google.dev/gemini-api/docs/pricing
  12. OpenAI, “Introducing ChatGPT Go, now available worldwide,” 16 January 2026; lists ChatGPT Pro at $200 per month.
    openai.com/index/introducing-chatgpt-go
  13. Anthropic Help Center, “What is the Max plan?” accessed 2 September 2026.
    support.claude.com/.../what-is-the-max-plan
  14. Cursor, “Updates to Ultra and Pro,” 16 June 2025, updated 30 June 2025.
    cursor.com/blog/new-tier
  15. Kyle Wiggers, TechCrunch, “OpenAI is losing money on its pricey ChatGPT Pro plan, CEO Sam Altman says,” 5 January 2025.
    techcrunch.com/.../openai-is-losing-money