Executive Summary
Outcome as a Service is sold on the claim that it moves 100 per cent of delivery and performance risk to the vendor. Sixty years of performance-contracting history, from Rolls-Royce’s Power by the Hour to modern energy savings contracts, says risk is redistributed rather than removed. Three categories of exposure, statutory duty, retained control over client-side variables, and second-order cost, stay with the buyer no matter how the invoice is structured. Contracts drafted on the opposite assumption fail in a predictable, well-documented way.
£467M
additional cost to the UK taxpayer after Transforming Rehabilitation’s volume-risk assumptions failed, National Audit Office
$234B
of enterprise application software spend Gartner estimates is exposed to agentic arbitrage by 2030
40%+
of agentic AI projects Gartner forecasts will be cancelled by the end of 2027 on cost, value or risk grounds
$0.99 vs $1.21
Fin’s launch price per resolution against its reported cost to serve, before subsequent efficiency work
Core conclusions
- Risk is redistributed, not extinguished. Non-transferable statutory duty, client-retained-control variables, and second-order costs such as switching and capability erosion stay with the buyer regardless of the payment mechanism chosen.
- The pattern across six decades of performance contracting is consistent: outcome pricing works only where the outcome is observable within a useful period, the vendor controls the dominant causal variable, measurement is standardised and verifiable, and the caseload is homogeneous or priced by segment.
- The commercial model matters less than the measurement and verification architecture underneath it. An outcome contract without an independently verifiable measurement basis is a fixed-price contract with a dispute attached.
The premise, the evidence, and the contract checklist, ten slides
Save it, share it, or send it to whoever is about to sign an outcome contract on the assumption it removes their risk.










1. Scope and definition
Outcome as a Service (OaaS) describes a commercial arrangement in which the buyer pays for a verified result rather than for licences, seats, hours, or consumption. The venture firm Foundamental claims authorship of the term for the architecture, engineering, construction and supply chain sectors. Sequoia Capital popularised an adjacent framing, “service as a software”, on the basis that agentic systems convert labour into software and therefore address the services market rather than the software market.
The distinction from prior models is narrower than the terminology suggests. Outcome-linked payment is not new. Rolls-Royce introduced Power by the Hour in 1962 for the Viper engine on the de Havilland (later Hawker Siddeley) 125, charging a fixed rate per flying hour and absorbing maintenance cost variance. Energy savings performance contracting has operated on guaranteed savings since the 1980s. United States defence acquisition has used performance-based logistics for two decades. United Kingdom central government has spent at least £15 billion through payment-by-results schemes, according to the National Audit Office.
What is new is the class of work now addressable. Agentic systems can execute discrete, high-frequency transactions end to end, which makes per-unit outcome metering feasible at volumes that would previously have required a business process outsourcing contract with a large human cost base. Gartner projects that 40 per cent of enterprise applications will embed task-specific agents by the end of 2026, against under 5 per cent in 2025, and estimates that US$234 billion of enterprise application software spend, roughly 20 per cent of global SaaS spend, is exposed by 2030 to what it calls agentic arbitrage.
2. Correcting the premise
A common characterisation of OaaS is that it removes risk from the client and places 100 per cent of delivery and performance risk on the vendor. That characterisation does not survive examination and is worth addressing directly, because contracts drafted on that assumption fail predictably.
Risk in an outcome contract is redistributed, not extinguished. Three categories of client exposure survive any payment mechanism.
Non-transferable risk. Statutory duties, regulatory obligations, safety accountability and duty of care generally cannot be delegated. A water utility remains accountable to its economic and environmental regulators for compliance breaches produced by a vendor’s agent. A transport operator remains accountable for a safety event. The payment mechanism determines who bears the cost; it does not determine who bears the duty. This asymmetry is the single most important design constraint for public sector and regulated infrastructure buyers.
Retained-control risk. The party controlling a variable must bear the risk attached to it, or the contract becomes unpriceable. The US Department of Energy’s Federal Energy Management Program states the position plainly in its guidance on energy savings performance contracts: if no values are fixed and savings are verified entirely on measurement, all risk sits with the energy services company; where parameters are stipulated at agreed fixed values, the customer assumes the risk that those values misstate reality. Usage variables such as operating hours, occupancy and weather are normally stipulated precisely because the contractor cannot control them. The same logic applies to an AI outcome contract. Data quality, upstream process design, system availability, policy changes and demand volume are typically client-controlled, and a vendor asked to carry them will either price the exposure heavily or fail.
Second-order risk. Opportunity cost of a failed programme, internal capability erosion, integration and change costs, switching cost at exit, and counterparty solvency all remain with the buyer. These are not covered by a no-cure-no-pay clause.
The accurate framing is that OaaS transfers a defined slice of performance risk in exchange for a price premium, and in doing so creates a new set of measurement, attribution and governance risks that did not previously exist. The Transforming Rehabilitation programme illustrates the consequence of assuming otherwise. The Ministry of Justice sought to transfer volume risk to Community Rehabilitation Companies but tested only a 2 per cent reduction in volumes. Actual volumes came in between 16 and 48 per cent below assumption. The Ministry subsequently renegotiated, revising the contractual fixed-cost assumption from 20 per cent to 77 per cent at a cost the NAO estimated at £342 million, and terminated the contracts 14 months early. Total additional cost to the taxpayer was assessed at £467 million. Risk that a supplier cannot absorb returns to the buyer, usually at a worse price than if it had never been transferred.
3. Market evidence as at mid-2026
The clearest live implementations are in customer service, where the unit of work is discrete, high-frequency and instrumented.
| Vendor | Unit | Published rate | Notes |
|---|---|---|---|
| Fin (formerly Intercom) | Outcome | US$0.99 (US$9.99 for a sales qualification) | US$49/month base including 50 outcomes; charged at most once per conversation; Salesforce agreed acquisition for approximately US$3.6bn in June 2026 |
| Zendesk | Automated resolution | US$1.50 committed, US$2.00 pay-as-you-go | Resolution defined with reference to a 72-hour quiet period |
| HubSpot | Resolved conversation | US$0.50 | Reduced from US$1.00 in April 2026 |
| Sierra | Resolution and other defined outcomes | Not published | Third-party estimates cluster around US$1.50 per resolution; year-one budgets of US$200,000 to US$350,000 including implementation are commonly reported |
| Salesforce Agentforce | Migrating | Not directly comparable | Moved from per-conversation to Flex Credits, then introduced pay-per-resolution for Help Agent; reporting an “agentic work unit” measure |
Rates as published or reported as at mid-2026. See sources below.
Two observations follow from the table. First, the headline rates are converging on a narrow band and falling, which indicates competitive pressure rather than value-based pricing. Second, the pure model is already hybridising. Sierra states publicly that outcome-based pricing is not always the right fit and that low-value interactions such as routing or greeting may be better priced on consumption. That is not a retreat; it is the correct response to the fact that outcome pricing only works where the outcome is discrete, attributable and verifiable.
The economic case rests on labour substitution. Sierra’s Bret Taylor has cited a human-handled customer service contact cost of US$10 to US$20, predominantly labour, against a per-resolution charge an order of magnitude lower. Vendor-published case studies report AI resolution rates in the range of 42 to 50 per cent, which is the figure buyers should model against rather than the theoretical maximum.
Free tool
AI Vendor Evaluation Scorecard
Rate an outcome-pricing vendor across the same due-diligence dimensions this article argues for, before the retained-control and second-order risks show up in your own budget.
4. What the historical record shows
Sixty years of performance contracting produces a reasonably consistent pattern. Outcome-based models perform where four conditions hold together, and degrade where any one is absent.
Condition one: the outcome is observable within a commercially useful period. Engine flying hours are observable continuously. Reoffending data takes two years to mature. The NAO concluded that this lag, combined with the impossibility of attributing changes in reoffending to Community Rehabilitation Company interventions, made payment by results inappropriate for probation services.
Condition two: the supplier controls the dominant causal variable. Rolls-Royce designs, manufactures, monitors and maintains the engine, and operates health monitoring across the installed fleet. Approximately 90 per cent of Trent engines are enrolled on TotalCare. A welfare-to-work provider does not control the labour market.
Condition three: measurement is standardised and independently verifiable. Energy performance contracting has the International Performance Measurement and Verification Protocol, with four defined options ranging from isolated key-parameter measurement to calibrated simulation, and an established practice of adjusting the baseline to reporting-period conditions rather than comparing raw consumption. AI outcome contracts currently have no equivalent protocol.
Condition four: the cohort is homogeneous, or pricing is calibrated to heterogeneity. The Work Programme’s nine payment groups, defined on prior benefit type, proved too crude. Academic analysis found that most variation in outcome probability was within payment groups rather than between them, which meant the payment structure designed in rather than designed out the practice of concentrating effort on the most tractable cases and neglecting the rest. Early performance was 3.6 per cent against a target of 11.9 per cent.
The failure modes are therefore not accidents of implementation. They are structural properties of the incentive, and they will reappear in AI outcome contracts wherever the four conditions are not met.
5. Client-side risks
Definitional capture. The vendor typically owns the telemetry that determines whether an outcome occurred. Reported practice includes counting a conversation as resolved when the customer went silent, counting any conversation not escalated to a human as resolved, and counting a click on a served article as a resolution. Where the measurement instrument sits inside the vendor’s stack and the vendor’s revenue is a function of the count, the buyer is auditing an interested party’s arithmetic.
Goodhart effects. An agent optimised against the payment metric will drift towards behaviours that raise the metric without raising value. Resolution-rate optimisation absent a counter-metric produces premature closure, discouraged escalation, and suppressed reopening. The countermeasure is a paired-metric structure in which the payment metric is gated by a quality metric.
Adverse selection within the caseload. The AI equivalent of creaming and parking is the routing of complex, ambiguous or low-confidence cases to human escalation while claiming the tractable volume. Because escalation is typically unbilled, this looks efficient on the invoice while transferring the expensive residual to the client’s own cost base. Net cost per contact across the whole population, not per AI resolution, is the correct measure.
Cost non-linearity. Outcome pricing means successful deployment increases spend. Reported migrations to per-resolution billing include increases from US$4,000 to US$9,000 per month, and from US$119 to US$854 per month. Budget holders accustomed to a fixed licence line should expect a variable line that scales with adoption and with seasonal demand peaks.
Counterparty concentration and solvency. Outcome contracts require the vendor to fund delivery ahead of payment, which means the buyer is underwriting the vendor’s balance sheet. Gartner estimates that only around 130 of the thousands of self-described agentic AI vendors are genuinely agentic, and forecasts that more than 40 per cent of agentic AI projects will be cancelled by the end of 2027 on grounds of cost escalation, unclear value or inadequate risk controls.
Lock-in. An outcome contract concentrates configuration, tuning, evaluation data and accumulated operational knowledge inside the vendor. Where changes to workflows or agent behaviour require vendor coordination rather than self-service, switching cost rises over the contract term.
Capability erosion. Outsourcing the outcome outsources the learning. Organisations that lose the internal ability to evaluate agent performance also lose the ability to contest the vendor’s measurement.
6. Vendor-side risks
Cost-to-serve volatility. Inference is a variable cost in cost of goods sold, not a fixed cost. ICONIQ’s January 2026 survey of approximately 300 software executives put inference at around 23 per cent of revenue at scaling-stage AI companies, with AI product gross margins projected at approximately 45 to 53 per cent for 2026 depending on architecture, against 75 to 85 per cent for conventional software. Bessemer has reported that the fastest-scaling AI companies run materially lower, some at negative gross margin. Under outcome pricing, revenue and cost move together but not proportionally: a hard case consumes many times the tokens of an easy one and yields the same fee. Intercom is reported to have launched Fin at US$0.99 against a cost per resolution of US$1.21, reaching positive unit economics only through subsequent efficiency work.
Revenue recognition. Outcome fees are variable consideration under IFRS 15 and ASC 606. They may be included in the transaction price only to the extent it is highly probable that a significant reversal of cumulative revenue will not occur. New offerings launch without usage history, which is precisely the condition that triggers the constraint. The IASB’s post-implementation review of IFRS 15 concluded in September 2024 with a decision to take no further action on variable consideration, and the FASB technical agenda as at mid-2026 contains no project addressing outcome-based AI revenue. Vendors are applying existing principles to a fact pattern the standards were not written for, and the “series” determination, which governs whether the variable-consideration allocation exception is available, can turn on minor contract details.
Working capital and duration. Delivery cost is incurred before outcome revenue is earned, and enterprise onboarding of four to ten weeks is common with larger programmes running three to seven months. The vendor funds the gap.
Exogenous dependency. The vendor’s revenue is a function of variables it does not control: the client’s demand volume, data quality, upstream process design, system availability and product changes. A client that improves its knowledge base reduces contact volume and reduces vendor revenue, which is an incentive misalignment in the opposite direction to the one usually advertised.
Attribution. Where the outcome has multiple contributing causes, credit allocation has no rigorous general solution. This is manageable for a resolved ticket and unmanageable for a revenue uplift, a retention improvement or an availability gain in a multi-vendor estate.
Contract incompleteness. Outcomes cannot be perfectly specified in advance. Edge cases return the parties to negotiation, and the cost of that negotiation is a recurring operating expense for both sides.
7. Contract architecture
The following controls address the failure modes above. They are drawn from established practice in energy performance contracting and performance-based logistics, adapted to agentic systems.
Outcome taxonomy. Classify each candidate outcome as hard (binary, provable from system logs, for example a ticket closed without escalation, a filing accepted by a regulator, a record validated against a reference database) or soft (requiring human judgement, for example accuracy, satisfaction or quality). Price hard outcomes per unit. Do not price soft outcomes per unit; use them as gates or as a fixed-fee scope.
Baseline and normalisation. Establish a pre-deployment baseline over a period long enough to capture seasonality, and specify in the contract how the baseline is adjusted for changes in volume, mix, product, policy and channel. The IPMVP convention of comparing reporting-period performance against an adjusted baseline rather than raw historical data is directly applicable.
Stipulation register. Enumerate every parameter held at a fixed agreed value, identify which party bears the risk of each stipulation being wrong, and set a review cadence. An unstated stipulation is an unpriced risk transfer.
Measurement and verification architecture. Specify the instrumentation, the data source of record, retention periods, and the buyer’s audit rights over raw event data rather than over vendor-produced summaries. Where contract value justifies it, appoint an independent verifier. Set a target measurement uncertainty and a confidence level, as energy performance contracts routinely do.
Paired metrics. Every payment metric requires a counter-metric that detects the gaming behaviour it invites. Resolution rate pairs with reopen rate within a defined window, escalation-after-closure rate, and downstream complaint volume. Payment is earned only where the counter-metric threshold is also met.
Cohort segmentation. Segment the caseload by tractability using observable characteristics, and price by segment. Set minimum service levels for the least tractable segment, and monitor the distribution of AI-handled versus escalated cases by segment to detect selective abandonment. Progressive pricing, in which the unit rate rises as the provider works deeper into the caseload, is the established countermeasure.
Payment mechanism. A pure outcome mechanism is rarely optimal. A three-layer structure is generally more robust: a fixed platform and availability fee sized to the vendor’s committed fixed cost; a variable outcome fee within a collar, with a floor that protects the vendor against client-caused volume collapse and a cap that protects the buyer against uncontrolled spend; and a retention or clawback element linked to sustained performance measured after a defined stabilisation period.
Financial security. Where the vendor is early-stage, require a parent company guarantee, a performance bond, or an escrowed retention proportionate to the switching cost. Insurance-backed savings guarantees are established practice in the ESCO market and offer a precedent.
Exit and portability. Contract for portability of configuration, prompts, tool definitions, evaluation datasets, conversation logs and outcome telemetry in a documented format. Include escrow, transition assistance obligations, and a defined step-in right on insolvency or sustained failure. The absence of these provisions converts an outcome contract into an indefinite dependency.
Dispute resolution. Establish a tiered process with a defined escalation ladder and expert determination for measurement disputes, with the expert’s remit and the evidential standard specified in advance. Measurement disputes are a normal operating feature of these contracts and should not require litigation.
Term. The term must be long enough for the vendor to recover deployment investment and short enough to preserve contestability. Break points tied to performance thresholds achieve both.
Free tool
Outcome as a Service Qualification Checklist
The contract architecture above, operationalised: 8 threshold gates plus 55 scored criteria across 9 domains, with a live recommended commercial structure.
8. Considerations specific to government and infrastructure operators
Accountability is not transferable. Where a statutory duty, licence condition or safety case attaches to the operator, the contract can transfer cost but not responsibility. Outcome payment mechanisms should not be applied to safety-critical decisions, and should not be structured so that a vendor bears financial consequence for an event whose prevention requires operator action.
Attribution lag defeats social outcomes. Where the outcome of policy interest matures over years and is subject to confounding, outcome payment is inappropriate regardless of how attractive the risk transfer appears. The NAO’s finding on probation is the reference case. Intermediate, controllable outcomes are the alternative.
Budgetary mechanics. Outcome payments are variable; appropriations are generally fixed and annual. A successful deployment can overrun its budget line. Collars, in-year monitoring and pre-agreed variation procedures are required, and the finance function should be involved in the payment mechanism design rather than presented with it.
Market contestability. Concentrating a critical operational process in a single outcome vendor reduces the number of credible future bidders. Where the vendor accumulates proprietary operational knowledge, the incumbent advantage at re-tender may be decisive. Data and configuration portability provisions are the principal remedy, and they need to be enforced during the term rather than invoked at exit.
Central capability. The NAO observed that payment by results is technically challenging, is not suited to all public services, and that neither the Cabinet Office nor the Treasury monitored its use across government, with the result that commissioners repeatedly rebuilt the same design from scratch. Organisations intending to use OaaS at scale should establish a single commercial standard, a reusable measurement and verification specification, and a central record of what has and has not worked.
9. Suitability assessment
The following test determines whether an outcome contract is appropriate. A negative answer to any item indicates that a hybrid or conventional model is the better structure.
- Is the outcome binary and provable from system-generated evidence?
- Does the evidence mature within the billing period?
- Does the vendor control the dominant causal variable?
- Can the buyer independently verify the count from raw data?
- Is a counter-metric available that detects the gaming behaviour the payment metric invites?
- Is the caseload homogeneous, or can it be segmented and priced by tractability?
- Can the vendor’s cost-to-serve be estimated with sufficient confidence to price a sustainable rate?
- Is the vendor financially capable of absorbing the transferred risk for the full term?
- Is the residual accountability that cannot be transferred identified, owned and resourced internally?
- Are portability and step-in arrangements enforceable during the term?
10. Assessment
Outcome as a Service is a legitimate and, in bounded applications, an efficient contracting model. It is not a novel one, and the evidence base for its behaviour is sixty years deep. The applications where it has worked share a common profile: a discrete, observable outcome; a vendor with control over the dominant variable; a standardised measurement protocol; and a payment mechanism that allocates each risk to the party best placed to manage it.
The proposition that the model eliminates client risk is inaccurate and, if relied upon in contract design, is the mechanism by which these arrangements fail. What the model does is exchange a known input cost for a set of measurement, attribution, incentive and counterparty risks. Whether that exchange represents value depends almost entirely on the quality of the contract, and specifically on the measurement and verification architecture, which is the element most often specified last and least.
An outcome contract without an independently verifiable measurement basis is a fixed-price contract with a dispute attached.
For buyers in government and regulated infrastructure, the practical conclusion is that the commercial model deserves less attention than the instrumentation beneath it.
Free tool
Board AI Oversight Checklist
Non-transferable risk in practice: whether your board can name who is accountable for a vendor’s outcome contract when the measurement is disputed.
Sources
- National Audit Office. (2019, February). Transforming rehabilitation: Progress review.
- National Audit Office. (2015, June). Outcome-based payment schemes: Government’s use of payment by results.
- Carter & Whitworth. (2015). Creaming and parking in quasi-marketised welfare-to-work schemes. Journal of Social Policy, 44(2).
- US Department of Energy Federal Energy Management Program. Using measurement and verification to manage risk in federal energy and water saving projects.
- US Department of Energy. M&V guidelines: Measurement and verification for performance-based contracts.
- Rolls-Royce. TotalCare.
- Gartner. (2026, July 1). $234 billion in enterprise application software spend is at risk from agentic AI.
- Gartner. (2025, June 25). Over 40% of agentic AI projects will be canceled by end of 2027.
- Sierra. Outcome-based pricing for AI agents.
- Fin (Intercom). Fin pricing: Outcomes.
- Zendesk. Understanding outcome-based pricing.
- Deloitte. (2026, June). Technology spotlight: Accounting for outcome-based pricing in an agentic AI software product.
- IFRS Foundation. IFRS 15 revenue from contracts with customers.
- BCG. (2025, August). Rethinking B2B software pricing in the era of AI.
- Siena AI. Why outcome based pricing in AI hurts customer service.
- Parloa. (2026, January). Outcome-based pricing: The most expensive myth in enterprise AI. Forbes.
- ICONIQ Capital. (2026, January). State of AI: Bi-annual snapshot, as reported in secondary analysis.
- Foundamental. Outcome-as-a-service.
The AI Governance & ROI Executive Programme covers exactly this, structuring the measurement and verification basis before the contract is signed rather than litigating it after. Details are on the workshops page.
