- BillWhy the provider invoice is not the total cost of an agent product.
- BucketsHow to separate build, run, review, and failure costs.
- UnitWhy cost per accepted result is more useful than cost per call.
- ValueHow reliability can reduce losses and increase delegated volume.
- DecisionHow to end a monthly review with expand, hold, or narrow.
01Where Episode 04 left us
At 10 AM on Monday, the CFO opens the quarterly spend report.
Model prices fell. The AI line rose.
The dashboard explains tokens in detail. It does not show the finance experts who labelled difficult invoice cases, the engineer who made failures replayable, the reviewers clearing exceptions, or the runtime preserving state between attempts.
The number is accurate. The picture is incomplete.
Episodes 01 to 04 built one chain of reasoning:
- Episode 01 found the first broken contract.
- Episode 02 located the responsible component and control.
- Episode 03 replaced vague choices with specific signals.
- Episode 04 asked what each control gained and what it cost.
This episode puts money around that system.
The economics are not separate from reliability. They are reliability viewed from finance.
02The budget meeting
The CFO asks:
If intelligence became cheaper, why did the agent become more expensive?
The answer begins with a distinction.
A model produces an answer or proposes an action. An agent product must also retrieve evidence, call tools, preserve state, enforce authority, verify the result, recover from failure, and send uncertain cases to a person.
Those activities create different bills.
For the invoice agent, a completed workflow means more than “the model responded.” It means:
- CoverageAll 2,347 eligible invoices were accounted for.
- AccuracyCorrect matches were recorded.
- ExceptionsEvery exception received a valid state.
- AuthorityNo prohibited action occurred.
- ReconciliationThe final count matched the source ledger.
The business is not buying tokens. It is buying that completed outcome.
This gives us the central rule for the episode:
Price the work from request to accepted business result.
The reader map
| Remember | Meaning |
|---|---|
| Four cost buckets | Build, Run, Review, Failure |
| Three unit measures | Cost per attempt, cost per accepted result, net workflow value |
| Two reliability returns | Lower failure cost and more delegated volume |
| One PM rule | Price the completed outcome, not the model call |
The rest of the episode explains each line slowly, then puts them together in one business case.
03Four cost buckets
Use four ordinary words whenever the bill becomes confusing.
What did it cost to create and change the system?
This includes product discovery, engineering, domain rules, eval cases, tool integration, security review, migration, and release work.
Some build cost happens once. Much of it returns whenever the model, tool, policy, or workflow changes.
What did it cost to attempt the work?
This includes input and output tokens, cache writes and reads, web search, file search, code execution, API calls, queues, databases, containers, storage, and observability.
Run cost grows with traffic and with the number of steps inside each workflow.
What did it cost people to inspect or decide cases?
This includes routine approval, expert review, escalation, reviewer tooling, calibration, and service-level coverage.
Review cost usually grows with volume, ambiguity, and consequence.
What did wrong or unfinished work cost?
This includes retries, investigation, rework, reversals, customer support, incidents, regulatory exposure, and work that the organisation refuses to delegate after trust falls.
These four buckets are deliberately plain:
| Bucket | Question | Invoice example |
|---|---|---|
| Build | What did we create or change? | Define partial-payment rules and add regression tests |
| Run | What did the system consume? | Model calls, invoice API, state store, traces |
| Review | What did people inspect or decide? | Finance specialist reviews an ambiguous match |
| Failure | What did wrong work cost? | Reverse a bad update and investigate the incident |
Model tokens sit inside Run. They are not the whole run cost, and run cost is not the whole business cost.
04What a real API bill contains
Provider bills differ because providers meter different resources.
Provider pricing at a glance
The providers meter different resources. This table uses public list prices available on 2 September 2026. Use it to understand the bill’s structure, not to select a vendor from this article alone.
| Provider example | New input | Cached input | Output | Other separately billed items |
|---|---|---|---|---|
| OpenAI GPT-5.6 Terra, short context | $2.00 / 1M tokens | $0.20 / 1M cache reads; $2.50 / 1M cache writes | $12.00 / 1M tokens | Web search, file search, file storage, hosted containers |
| Anthropic Claude Fable 5.1 | $10.00 / 1M tokens | $0.25 / 1M cache reads; $12.50 / 1M five-minute cache writes | $50.00 / 1M tokens | Longer cache duration, web search, code execution, managed runtime |
| Google Gemini 3.7 Flash, promotional standard pricing | $0.75 / 1M tokens | $0.075 / 1M cached tokens plus $0.50 / 1M tokens per cache-storage hour | $3.75 / 1M tokens | Search grounding, maps grounding, file search, live APIs, higher service tiers |
Sources: OpenAI API pricing, Anthropic prompt-caching documentation, and Google Gemini API pricing.[9][10][11]
Three details matter more than the headline rates:
- The same text can have three prices. New input, a cache write, and a cache read may be billed differently.
- Tools sit outside token pricing. Search, retrieval, code execution, containers, and storage can appear as separate line items.
- Storage duration can matter. Google charges for cached-context storage time. Anthropic offers different cache lifetimes. A cache is an economic choice, not merely a latency feature.
The rows are not quality-equivalent. This table compares billing anatomy, not model capability.
A small illustrative OpenAI bill
Suppose a support agent completes 1,000 workflows in a month. Each workflow uses GPT-5.6 Terra for six model calls.
Assumptions per workflow:
- PrefixA stable 20,000-token prefix containing instructions, examples, and tool definitions.
- Input2,000 new input tokens per call.
- Output1,000 output tokens per call.
- SearchOne web-search call.
- ContainersA 1 GB container for 20% of workflows.
- CacheThe first call writes the stable prefix to cache; the next five reuse it.
Using the listed prices, the approximate provider bill is:
| OpenAI line | Monthly usage | Approximate cost |
|---|---|---|
| Cache writes | 20M tokens | $50 |
| Cache reads | 100M tokens | $20 |
| New input | 12M tokens | $24 |
| Output | 6M tokens | $72 |
| Web search | 1,000 calls | $10, before search-content tokens |
| Containers | 200 eligible 1 GB sessions | $6 |
| Provider subtotal | $182 |
This is an illustration, not a quote or forecast. It uses current public list prices and simplified traffic.
Now add costs that will not appear on the OpenAI invoice:
| Internal or other-vendor line | Illustrative monthly cost |
|---|---|
| Application runtime, queue, and database | $120 |
| Trace storage and observability | $80 |
| Human review: 50 cases × 6 minutes × $30/hour | $150 |
| Amortised engineering, eval, and domain-review work | $1,500 |
| Fully loaded subtotal before failure loss | $2,032 |
If 900 of the 1,000 workflows pass the business completion rule, the provider subtotal is about $0.20 per accepted result. The fully loaded subtotal is about $2.26 per accepted result.
Both numbers are correct. They answer different questions.
Do not present the provider subtotal as the product’s unit cost. It excludes the application, runtime, evidence, review, and build work that makes the result usable.
The provider number helps engineers optimise API use. The fully loaded number helps the business decide whether the workflow is economically worthwhile.
What caching changes
Without caching, the same 20,000-token prefix would be processed as normal input on all six calls. Under the same assumptions, model-token cost would rise from roughly $166 to $336 per 1,000 workflows before search and containers.
Caching does not change the model. It changes the bill because the harness keeps a stable prefix reusable.
It can also fail silently as an economic strategy. Changing a tool definition, system prompt, timestamped block, or cache breakpoint may invalidate the prefix. Anthropic’s documentation explicitly recommends monitoring cache-creation, cache-read, and uncached-input fields rather than assuming a cache hit occurred.[10]
Why a $200 plan is not $200 of compute
Flat-rate subscriptions create a different kind of bill. Three prominent power-user plans share the same headline price but package usage differently.
| Plan | Public price | How usage is described | Important limit or economic detail |
|---|---|---|---|
| ChatGPT Pro | $200 / month | OpenAI positions it for AI power users | Access and availability still depend on model, tool, capacity, and plan terms |
| Claude Max 20x | $200 / month | Twenty times the session usage of Claude Pro | Five-hour reset plus an additional weekly limit shared across Claude and Claude Code |
| Cursor Ultra | $200 / month | Twenty times the usage of Cursor Pro | Cursor says long-term model-provider partnerships help support the predictable price |
Sources: OpenAI, Anthropic, and Cursor plan documentation.[12][13][14]
These plans are access contracts, not prepaid wallets containing exactly $200 of API credit.
The customer sees one fixed price. The vendor sees a distribution:
- LightA light user may consume far less inference than the subscription price.
- TypicalA typical user may support a healthy margin.
- PowerA power user running long coding sessions, large repositories, or parallel agents may consume far more.
The plan works at the cohort level. Revenue from light and typical users can offset heavy users. Limits, routing, caching, negotiated model rates, and capacity management protect the vendor.
If a subscriber consumes $1,000 of usage valued at public API prices, it does not follow that the vendor lost $800. Public API price is not the vendor’s marginal cost. The provider may run its own model more cheaply, negotiate lower rates, route work, reuse cached context, or impose limits. Actual loss requires the vendor’s real cost to serve the account, which public plan pages do not disclose.
There is still a real warning. In January 2025, TechCrunch reported that OpenAI CEO Sam Altman said the then-$200 ChatGPT Pro plan was losing money because subscribers used it more than expected.[15] Treat this as a historical signal about forecasting risk, not evidence of OpenAI’s current per-user economics.
A flat price transfers usage risk from the customer to the vendor.
What this means for an AI startup
Suppose a product charges $100 per customer per month. The average direct AI and runtime cost looks manageable until customers are separated by usage.
| Customer group | Share | Revenue each | Direct AI and runtime cost each | Contribution each |
|---|---|---|---|---|
| Light | 60% | $100 | $8 | $92 |
| Typical | 30% | $100 | $35 | $65 |
| Power | 10% | $100 | $260 | −$160 |
Across 100 customers, revenue is $10,000. Direct AI and runtime cost is $4,130. Contribution before support and fixed cost is $5,870, or 58.7%.
The plan is profitable as a cohort even though every power user loses money on direct service cost. If growth concentrates in that segment, margin falls quickly. If those users also receive the most value, crude throttling may destroy the product’s best use case.
The AI PM must decide:
- Which expensive behaviour creates customer value?
- Which behaviour creates cost without increasing accepted results?
- Should heavy use be included, metered, queued, rate-limited, or moved to a higher tier?
- Can routing, caching, smaller context, or fewer failed retries preserve the experience?
- Are expensive users strategically valuable because they retain, expand, or teach the product more?
Do not price from the average alone. Track median, 90th-percentile, and 99th-percentile cost to serve, then inspect the workflows behind the tail.
Use the margin terms correctly
| Measure | Plain meaning | Why the AI PM needs it |
|---|---|---|
| Gross margin | Revenue left after direct cost of serving customers | Shows whether inference, tools, runtime, and direct review fit the business model |
| Contribution margin | Revenue left from a customer, workflow, or cohort after attributable variable costs | Reveals power users or workflows hidden by the company-wide average |
| Net margin | Revenue left after wider research, product, sales, and administrative expenses | Describes the company’s overall profitability |
“This user has negative contribution margin” and “this company loses money” are not the same claim.
For enterprise contracts, calculate contribution margin by customer and workflow. A large annual commitment can hide one automation whose marginal cost grows faster than its usage revenue.
05Why cheaper calls can produce a larger bill
An agent workflow often performs many more operations than a chat answer:
- Interpret the task.
- Retrieve current evidence.
- Select a tool.
- Process the tool result.
- Repair invalid output.
- Verify the effect.
- Continue after partial failure.
- Ask another model or person to judge ambiguity.
- Compact or reconstruct context.
- Repeat until completion or a stop condition.
The price of each model call can fall while total cost rises because the product handles more traffic, takes more steps, carries more context, invokes paid tools, runs longer, and verifies more often.
That increase is not automatically bad. A more expensive run may produce a result the business can actually accept.
The main levers are model routing, context size, cache hit rate, retries, parallel workers, judge calls, paid tools, and stopping rules.
06Five costs the token bill hides
The four buckets tell you where money goes. The five cost centres below identify work teams most often omit from the budget.
| Hidden cost | Bucket | What the organisation is buying |
|---|---|---|
| Ground truth and evals | Build | A testable definition of acceptable behaviour |
| Senior attention | Build and Review | Translation of business judgement into system decisions |
| Observability and replay | Run and Build | Evidence that makes failure diagnosable and repairs testable |
| Human review | Review | Qualified judgement for ambiguity and consequence |
| Versioning and migration | Build | The ability to change models, tools, policies, and prompts safely |
Runtime remains a separate run-cost line in Section 12. This preserves the season’s count of five hidden cost centres without pretending that runtime is free.
07Ground truth and evals
Before a team can measure an agent, it must decide what acceptable work means.
Consider a ₹10,00,000 invoice with a payment of ₹8,00,000. Should the agent mark it paid, apply a partial match, leave it pending, or create an exception?
The model can propose an answer. Ground truth records the organisation’s approved answer and the conditions behind it. An eval reruns the case after a change and checks whether the system still behaves correctly.
The expensive part is not running the test. It is making business judgement explicit.
For reconciliation, the eval set may cover exact matches, partial payments, duplicate payments, tenant isolation, exception states, final ledger counts, and escalation when evidence is insufficient.
Someone must label cases, resolve disagreement, write checks, define rubrics, create holdouts, and update the set when policy changes.
Features produce visible output. Ground truth may look like a spreadsheet until a silent regression reaches production.
Budget eval work beside feature work, not after it.
08Senior attention
Harness work consumes people with scarce context.
A staff engineer understands architecture and failure propagation. A product manager translates among the user outcome, business rules, security, finance, and engineering. Domain experts resolve cases that a generic annotation team cannot judge.
Their time is real even when it never appears on an API invoice.
Name the capacity:
- Engineering0.5 staff engineer for 16 weeks.
- Product0.5 product manager for 12 weeks.
- FinanceTwo finance reviewers for four hours each week.
- GatesSecurity and legal review at two release gates.
The estimate may be wrong. An unnamed estimate is guaranteed to be missing.
Do not treat all senior attention as one-time implementation. Models, policies, tools, and regulations continue to change. Some expert capacity is operating cost.
09Observability and replay
Three words are often mixed together:
- TraceRecords what happened.
- EvalJudges whether it was acceptable.
- ReplayRecreates enough of the run to test a repair.
If an invoice was marked paid incorrectly, the trace shows what evidence reached the model and which tool call followed. The eval says the result violated the partial-payment rule. Replay lets the team test the corrected system against the same case.
A trace without judgement is a detailed mystery. An eval without a trace can say “wrong” without explaining why. An incident that cannot be replayed is expensive to repair with confidence.
Useful evidence may include tenant identity, model and harness versions, context sources, tool arguments, approval decisions, retries, state changes, cost, and completion proof.
That evidence creates its own cost: storage, search, access control, redaction, retention, and deletion.
Measure the return in plain operating terms:
- How long does it take to find the first divergence?
- What percentage of incidents can be reproduced?
- How long does it take to turn a failure into a regression test?
10Human review
“Human in the loop” sounds like a safety principle. At production volume, it is a queue.
Suppose 10,000 sessions run each day. If 3% need review, 300 cases enter the queue. At five minutes each, the team needs 1,500 minutes, or 25 person-hours, every day.
The capacity model is:
daily review hours = (sessions × review rate × minutes per review) ÷ 60
A better model may lower the review rate. Wider deployment raises the number of sessions. Harder cases raise review time. Price all three.
The review system also needs an interface, routing by expertise, service levels, timeout behaviour, calibration, disagreement resolution, and coverage outside business hours.
McKinsey and QuantumBlack made this cost visible in August 2026. For one banking customer-service example, they estimated token costs at 20% to 25% of variable run cost and human oversight at 70% to 75%. For customer onboarding, they estimated expert review for 10% to 20% of agentic runs.[5]
These are estimates for specified banking examples, not universal ratios. Their lesson is still direct: a budget with a token line and no review line may be missing the larger variable cost.
11Versioning and migration
Production agents change on several clocks:
- ModelsLaunch, change behaviour, and retire.
- PromptsPrompts, skills, and context policies evolve.
- ToolsTool schemas and permissions change.
- PolicyBusiness policy changes.
- RuntimeRuntime services upgrade.
- EvalsEval thresholds move.
- RegulationCustomer and regulatory obligations change.
Without version discipline, the team cannot reconstruct which model, prompt, policy, and tool contract produced an action.
Every meaningful artifact needs a version, owner, reason for change, release evidence, rollout plan, rollback path, and retirement condition if temporary.
Migration cost is not an unusual event outside the operating model. It is part of the operating model.
12Runtime is a separate line
Even an open-source harness needs somewhere to run.
Runtime includes durable sessions, queues and workers, checkpoints, sandboxes, multi-tenancy, deployment, scaling, secrets, trace transport, and recovery.
| Runtime line | Question |
|---|---|
| Durable state | What survives worker failure? |
| Execution | Where do tools and code run? |
| Scaling | What concurrency and queue guarantees are required? |
| Observability | What is retained, searchable, and exportable? |
| Tenancy | How are customers isolated? |
| Recovery | How does work resume without duplicate effects? |
| Operations | Who is on call, and what does the vendor own? |
A managed platform may bundle these costs into a seat, session, container, or usage allowance. Bundling simplifies purchasing. It does not remove the cost.
Anthropic’s Managed Agents exposes a hard session-spend cap. When the cap is
reached, the system can pause the run with its container preserved and return
budget_reached, allowing the same work to resume if the limit
changes.[7] Cost becomes an explicit state in the harness.
SpaceXAI’s Grok Bot packages always-on agents with a cloud computer and bundled usage.[8] That creates a different finance problem: the organisation must allocate a seat-level price across accepted workflows if it wants to compare the product with API-based or internally hosted alternatives.
13The value of reliability
Reliability creates value in two places.
First, it reduces the cost of wrong work:
- EscapeFewer failures escape.
- DamageFailures cause less damage.
- RecoveryRecovery becomes faster.
- ReviewFewer cases need review or rework.
Second, it can increase delegated volume, meaning the amount of eligible work the organisation permits the agent to handle.
Suppose 100,000 support cases are eligible. The company initially delegates only 10,000 because mistakes are hard to detect and reverse. If verification, limits, and rollback reduce exposure, the company may delegate 40,000 without changing the model.
That increase is the reliability dividend: more useful work becomes available because the downside is bounded and visible.
In plain language:
Net value equals value from accepted work, minus run and review cost, minus expected failure loss.
A simple model is:
V = (N × ps × vs) − Crun − Creview − (N × pf × Lf)
Where:
- NIs delegated volume.
- psIs accepted completion rate.
- vsIs value per accepted completion.
- CrunIs model, tool, runtime, and observability cost.
- CreviewIs human-review cost.
- pfIs consequential failure rate.
- LfIs average loss per consequential failure.
Build cost can be subtracted separately or amortised into C_run for the
reporting period. State which treatment you use.
The harness can improve value by increasing accepted completion, reducing failure frequency, reducing loss when failure occurs, and increasing delegated volume.
14Autonomy is a loss decision
An autonomy boundary is a loss decision expressed as product scope.
A team may let the agent draft every refund recommendation but issue only low-value, reversible refunds that pass deterministic policy. High-value or ambiguous cases go to review.
As evidence, rollback, and escalation improve, the issuing boundary can move.
The useful question is not “How autonomous is the agent?”
Ask:
Which actions may it perform, under what limits, with what evidence and recovery?
15A worked business case
The following numbers are illustrative assumptions. They explain the model; they are not external evidence.
A support workflow has 100,000 eligible monthly cases.
- ValueEach accepted automated case creates $3 of value.
- Review costA human review costs $4.
- Failure lossA consequential automated failure causes an average $30 loss.
System A: cheaper attempt, weak controls
The organisation delegates only 10,000 cases because failure is difficult to detect.
- CompletionAccepted completion rate: 90%.
- FailureConsequential failure rate: 2%.
- ReviewReview rate: 10%.
- Run cost$0.20 per attempted case.
| System A line | Calculation | Amount |
|---|---|---|
| Value from accepted work | 10,000 × 90% × $3 | $27,000 |
| Run cost | 10,000 × $0.20 | −$2,000 |
| Review cost | 10,000 × 10% × $4 | −$4,000 |
| Expected failure loss | 10,000 × 2% × $30 | −$6,000 |
| Net value | $15,000 |
System A produces 9,000 accepted results.
Its run, review, and failure cost is $12,000. That is about $1.33 for each accepted result.
System B: more expensive attempt, stronger controls
Verification, routing, and rollback improve. The organisation delegates 40,000 cases.
- CompletionAccepted completion rate: 94%.
- FailureConsequential failure rate: 0.5%.
- ReviewReview rate: 5%.
- Run cost$0.28 per attempted case because the system verifies more.
| System B line | Calculation | Amount |
|---|---|---|
| Value from accepted work | 40,000 × 94% × $3 | $112,800 |
| Run cost | 40,000 × $0.28 | −$11,200 |
| Review cost | 40,000 × 5% × $4 | −$8,000 |
| Expected failure loss | 40,000 × 0.5% × $30 | −$6,000 |
| Net value | $87,600 |
System B produces 37,600 accepted results.
Its run, review, and failure cost is $25,200. That is about $0.67 for each accepted result.
The attempted case became more expensive: $0.20 rose to $0.28.
The accepted result became cheaper: approximately $1.33 fell to $0.67 when review and expected failure loss are included.
Do not copy the conclusion without replacing the assumptions. System B wins only because the example assumes controls improve completion, failure, review, and delegated volume enough to cover their cost.
16Named evidence, used carefully
Public examples show what happened inside named organisations. They rarely isolate one causal variable.
Azure SRE Agent
Microsoft reported in April 2026 that Azure SRE Agent had handled more than 35,000 incidents autonomously over nine months and saved more than 50,000 developer hours. It also reported that Azure App Service reduced time to mitigation from a 40.5-hour human-only average to three minutes for live-site incidents.[3][4]
These are Microsoft-reported internal outcomes, not an independent controlled trial. The system included access to logs, metrics, incident records, code changes, specialised tools, role-based access, approval boundaries, and ongoing evaluation. The model call was one component.
McKinsey’s completed-work unit
McKinsey recommends measuring the fully loaded cost to complete a job across people, agents, and deterministic systems. In its illustrative banking-onboarding example, it estimates that cost may fall from roughly $50–$150 per customer to $10–$30, despite a workflow involving several agents, deterministic systems, and oversight teams.[5]
Treat those figures as consultancy estimates, not a forecast. The reusable idea is to price the completed job.
Provider pricing makes harness design visible
OpenAI, Anthropic, and Google all price cached context differently from new context. OpenAI also exposes search, file search, storage, and containers as separate lines. Google adds cache-storage time and grounding requests. Anthropic distinguishes cache writes from cache reads and requires exact reusable prefixes.[9][10][11]
The economic question is no longer only “Which model is cheaper?”
It is also:
- ResendHow much context does the harness resend?
- CacheHow often does the cache hit?
- ToolsHow many paid tools does a workflow call?
- StorageHow long does state or cached context remain stored?
- AttemptsHow many attempts fail before one result is accepted?
17Cost per accepted result
Use this operating metric:
cost per accepted result = (attempts + tools + runtime + review + investigation + rework) ÷ accepted results
If you want a fully loaded business metric, add amortised build and migration cost to the numerator. If you want net value, also subtract expected external failure loss from the value produced.
Do not switch definitions between months. Write the numerator under the metric.
A harness can improve the ratio through smaller current context, higher cache reuse, right-sized models, fewer unproductive retries, early stopping, durable resume, and cheap deterministic checks before expensive judgement.
Beyond some point, another verifier costs more than the failure it prevents. Track both sides.
18Quality is part of cost
LangChain’s 2026 survey received 1,340 responses, mostly from technology organisations. It reported quality as the leading production barrier at 32%. It also reported 89% observability adoption, 52.4% offline-eval adoption, and 59.8% use of human review.[1]
The survey is self-selected and technology-heavy. It is not a census of every enterprise.
Its economic lesson is still useful: poor quality creates retries, review, investigation, rework, incidents, and resistance to delegation.
Finance should ask two questions together:
- What does one accepted workflow cost?
- What prevents the organisation from delegating more of them?
A program that optimises only the first can become cheap and irrelevant.
19What to rent and own
A restaurant may rent its building, payment terminal, and delivery software. It should not outsource the definition of what its food should taste like.
An agent team can rent models, sessions, execution, search, and trace storage. It should retain control of the definition of good, workflow rules, approval policy, escalation, and failure taxonomy.
Usually sensible to rent or reuse
- ModelsModel APIs.
- RuntimeSession and runtime infrastructure.
- SandboxSandbox provisioning.
- DeploymentGeneric deployment and scaling.
- TracesTrace transport and storage.
- LibrariesCommon protocol and tool libraries.
Usually important to control
- WorkflowWorkflow definition.
- Ground truthDomain ground truth.
- EvalsEval cases and thresholds.
- AuthorityAuthority and approval policy.
- EscalationEscalation paths.
- FailureFailure taxonomy.
- ContextTenant and user context rules.
The hybrid is often correct: rent operations, own meaning.
20The lock-in test
| Harness area | Portability question |
|---|---|
| Identity | Can we export and version the operating contract? |
| Memory | Can we inspect context assembly, caching, and compaction? |
| Orchestration | Can we change models or runtimes without rewriting the workflow? |
| Controls | Can we control approvals and policy checkpoints? |
| Evidence | Can we export traces, datasets, rubrics, and results? |
Lock-in is not automatically bad. It may be acceptable for a commodity workflow.
The cost appears when the workflow becomes differentiating, regulated, or difficult to migrate after years of accumulated data and rules.
Price the exit before you need it.
21When to stop investing
Harness work can overshoot.
The twelfth verifier may cost as much as the second and improve little. A model-specific workaround may be close to becoming unnecessary. Another month on a mature low-risk workflow may produce less value than applying proven controls to a second workflow.
Ask:
- Is the eval curve still moving?
- Is this component temporary compensation or a permanent business obligation?
- Would the next rupee buy more value through deeper reliability here or wider coverage elsewhere?
Stopping is capital allocation. A harness that only grows is accumulating, not compounding.
22When expansion is earned
Expansion is supported when:
- CompletionAccepted completion improves.
- LossExpected failure loss remains within tolerance.
- ReviewReview becomes more selective.
- RecoveryIncident diagnosis and recovery become faster.
- Unit costCost per accepted result falls or remains justified by higher value.
- AutonomyA wider autonomy boundary is supported by evidence and rollback.
Expansion is not a reward for effort. It is a decision supported by outcomes.
23The monthly value review
Do not start with a dashboard. Start with the decision the meeting must make.
Read the workflow from opportunity to outcome:
- How much work was eligible?
- How much did the organisation delegate?
- How much finished acceptably?
- What failed, and what did failure cost?
- How much human attention was required?
- What did each accepted result cost?
- Has the workflow earned wider scope?
| Line | Plain definition | Decision supported |
|---|---|---|
| Eligible volume | All work that could enter the workflow | Size of the opportunity |
| Delegated volume | Work the organisation allowed the agent to attempt | Whether trust is expanding |
| Accepted results | Outcomes that passed the business completion rule | Whether useful work was produced |
| Expected failure loss | Consequential failures × average loss | Whether exposure is acceptable |
| Review hours | Reviewed cases × review time | Whether attention is becoming selective |
| Cost per accepted result | Defined full cost ÷ accepted results | Whether unit economics are improving |
| Expansion requirement | Evidence needed for wider scope | Expand, hold, or narrow |
| Cost by customer cohort | Contribution margin for light, typical, and power users | Reprice, meter, route, or redesign |
The meeting ends with one of three decisions:
- ExpandVolume, autonomy, or workflow coverage.
- HoldScope while a named gap is repaired.
- NarrowScope because loss, review, or cost exceeds value.
If the review cannot change scope or investment, remove it.
24Try this before Friday
Choose one workflow. Use measured numbers where available. Label estimates and unknowns.
| Field | Record |
|---|---|
| Eligible monthly volume | |
| Delegated monthly volume | |
| Accepted completion rate | |
| Model and tool cost | |
| Runtime and observability cost | |
| Human-review rate, time, and loaded cost | |
| Consequential failure rate | |
| Average failure loss or range | |
| Amortised build and migration cost | |
| Cost by light, typical, and power-user cohort | |
| Gross and contribution margin | |
| Cost per accepted result | |
| Evidence required to expand | |
| Evidence quality | Measured, estimated, or unknown |
| Owner |
Then complete one sentence:
The largest uncertainty is [field]. We will measure it through [method], owned by [person or team], before deciding to [expand / hold / narrow].
If the answer is “a better model,” name the failure that model must remove. If the answer is “more trust,” translate trust into completion evidence, limits, rollback, and observed performance.
25What to remember
Remember four buckets:
- Build the system.
- Run the workflow.
- Review uncertain or consequential cases.
- Pay for Failure when wrong work escapes or must be repaired.
Remember four economic measures:
- Cost per attempted run.
- Cost per accepted result.
- Contribution margin by customer and workflow cohort.
- Net value after review and expected failure loss.
Remember two sources of reliability value:
- Failures become fewer, cheaper, or easier to recover from.
- The organisation delegates more work because the downside is bounded and visible.
And remember one rule:
Price the completed business outcome, not the model call.
Episode 06 replaces the assumptions in this model with evidence from one real workflow.
-
LangChain, “State of Agent Engineering,” 12 June 2026; survey of 1,340
respondents.
langchain.com/state-of-agent-engineering -
Rakuten, “Rakuten accelerates development with Claude Code,”
company-reported workflow outcomes.
rakuten.today/blog/rakuten-accelerates-development-with-claude-code -
Shamir AbdulAziz, Microsoft, “How we build and use Azure SRE Agent with
agentic workflows,” 5 April 2026, updated 9 April 2026.
techcommunity.microsoft.com/blog/appsonazureblog/azure-sre-agent -
Microsoft Learn, “Azure SRE Agent overview.”
learn.microsoft.com/en-us/azure/sre-agent/overview -
Chandana Asif, Dieter Kiewell, Lari Hämäläinen, Tom Kolaja, and Tunde
Olanrewaju, McKinsey/QuantumBlack, “Where AI agents pay off,” 24 August
2026.
mckinsey.com/.../where-ai-agents-pay-off -
Anthropic, “Claude Fable,” pricing and retention terms, 1 September
2026.
anthropic.com/claude/fable -
Anthropic, Claude Platform release notes, 7 August 2026; Managed Agents
session-spend budgets.
docs.anthropic.com/en/release-notes/api -
SpaceXAI, “Introducing Grok Bot,” early beta announcement, 11 August
2026.
x.ai/news/introducing-grok-bot -
OpenAI, API pricing, accessed 2 September 2026.
developers.openai.com/api/docs/pricing -
Anthropic, “Prompt caching,” Claude Platform documentation, accessed 2
September 2026.
platform.claude.com/docs/.../prompt-caching -
Google, Gemini Developer API pricing, accessed 2 September 2026.
ai.google.dev/gemini-api/docs/pricing -
OpenAI, “Introducing ChatGPT Go, now available worldwide,” 16 January
2026; lists ChatGPT Pro at $200 per month.
openai.com/index/introducing-chatgpt-go -
Anthropic Help Center, “What is the Max plan?” accessed 2 September
2026.
support.claude.com/.../what-is-the-max-plan -
Cursor, “Updates to Ultra and Pro,” 16 June 2025, updated 30 June
2025.
cursor.com/blog/new-tier -
Kyle Wiggers, TechCrunch, “OpenAI is losing money on its pricey ChatGPT Pro
plan, CEO Sam Altman says,” 5 January 2025.
techcrunch.com/.../openai-is-losing-money