When AI pricing depends on the outcome, who verifies that outcome?
Outcome-based pricing is arriving in customer service software. It asks boards a question that license-based software never did: who decides whether the outcome happened?
AI is changing how enterprises pay for customer service software, not only how they deliver it.
Traditional software pricing was about access: seats, licenses, capacity. Agentic AI is pushing the market toward paying for what the software does. Salesforce’s new Agentforce Help Agent charges $2 per autonomous resolution. Intercom charges $0.99 per Fin resolution. The appeal is straightforward. Pay for the result rather than the tool, and the vendor’s incentives start to look more like yours.
Gartner believes up to $234 billion of enterprise application spending is exposed to what it calls agentic arbitrage between now and 2030, roughly 20% of enterprise application SaaS spend, as AI agents complete tasks across multiple systems and reduce the need for anyone to sit in front of a traditional interface. That’s a related trend to outcome pricing, but not the same one. Gartner’s number is about software becoming invisible, not about how vendors bill for it. Both point the same way: less patience for paying for access when an agent can deliver the result.
Which brings us to the question a license never had to answer. Who decides whether the outcome happened?
What counts as a resolution
The price isn’t the only thing to watch with outcome pricing. The devil is in the details, and the details are the definitions.
Salesforce’s Help Agent bills when the agent resolves an issue autonomously from start to finish, and the model has real buyer protections built in. If the customer gives negative feedback, or asks for a human, there’s no charge.
Intercom counts a resolution when, following Fin’s last answer, the customer either confirms the answer was satisfactory, which Intercom calls a confirmed resolution, or leaves the conversation without asking for further help, which it calls an assumed resolution. Human escalations don’t count. Some other outcomes, including procedure handoffs and disqualifications, are billed separately.
Both are reasonable engineering answers to a genuinely hard measurement problem, and one of them goes out of its way to protect the buyer. Neither is quite the same thing as proving the customer’s problem got solved.
An interaction is not always an outcome
A customer contacts an AI agent to change an account setting. The agent answers. The customer asks a follow-up. The agent answers again. The customer closes the conversation and doesn’t ask for a person.
Under an assumed-resolution definition, that’s a resolution. The meter ticks and it lands on next month’s invoice.
To the customer, the answer might have been confusing, wrong, or impossible to act on. They didn’t ask for a human because they ran out of time, lost patience, or gave up and went somewhere else. They abandoned the journey and you got billed for a successful one.
That’s not a resolution. It’s pay-per-abandonment.
Grading your own homework
None of this is anyone behaving badly. Intercom and Salesforce are solving a real problem: how to measure autonomous interactions at scale, programmatically, without a person reading every conversation. Silence is one of the few signals available, and it isn’t a foolish one to use.
When the party that builds the AI also defines the metric, reads the telemetry and issues the invoice, the buyer is relying heavily on the vendor’s measurement of its own service. Even a buyer-friendly definition rests on signals the vendor collects and interprets.
HFS Research found the gap this creates. In its 2026 customer experience research, 76% of enterprise buyers say they are open to outcome-based or gainshare pricing, and only 5% are in a contract built that way today. That 71-point gap between appetite and adoption is too large to ignore. Buyers are happy to pay for results, but they may not have a way to check the definition of a result.
Info-Tech Research Group names the mechanism plainly in its blueprint on negotiating AI contracts, which is titled, with no ambiguity at all, Negotiate Safe AI Contracts to Prevent Bill Shock. In this pricing model, vendors “define the billable events, control the meters, and reserve the right to change the pricing logic mid-cycle.”
When the billing definition and the customer’s experience diverge, the pricing model stops driving alignment and starts driving cost you find out about on the invoice instead of at signing. That is the ROI of customer experience question at its sharpest: proof of what you actually got for the money, and whether the invoice matches it.
Testing the path, not observing the call
The answer isn’t to reject outcome pricing. It’s to separate observability from assurance.
Observability works from passive telemetry: logs, traces and metrics gathered after a call or a chat has already happened. That tells you every component reported that it worked. It does not tell you whether a path will fail for the next customer to take it, or whether a dropped session was counted as a successful assumed resolution.
Assurance works the other way around. Synthetic journeys are run from the outside in, dialing carrier lines, walking IVR flows, talking to voicebots and following agentic paths the way a customer would, on a schedule, so you find out whether the journey holds up before a paying customer does.
That is the difference between a description of the system and evidence of what it did. It is also a check that does not come from the party sending the invoice.
Before you sign the next contract
Outcome pricing has real potential to align cost with value. It needs an independent, outside-in layer to check that an assumed resolution matches a resolved customer. Five things worth doing before signing or expanding an outcome-based AI contract.
Push the definition past the session. Where the AI is supposed to cause a business action, tie the definition of success to a verifiable result, a backend state change, confirmed record update, completed transaction or other agreed business outcome, rather than conversation length or a customer going quiet.
Test continuously, not once. Run assurance across the whole journey, end to end, across carriers, IVRs, voicebots, digital agents and desktop integrations, so you know it’s working now, not just that it launched working.
Get notice when the model changes, and re-verify when it does. A vendor can change the model, prompt, knowledge base or workflow behind the agent without changing the billing definition on paper. Build in the right to be told when that happens, and to re-test before you take the next invoice on faith.
Get audit rights in writing. Build in the right to reconcile vendor billing logs against your own independent test records, so a gap between the meter and reality has somewhere to go other than your P&L.
Ask how test traffic is excluded from billing, and get the exclusion in the contract. If you’re going to verify the journey, you need to know your own testing won’t appear on the invoice as customer resolutions. Every vendor handles this differently. Settle it before signing, not in a dispute afterwards.
Boards are being asked to move quickly on AI, and most of the people I talk to are doing their best with the information in front of them. The question I would want answered before the next renewal is a simple one.
If your AI vendor billed you for 10,000 assumed resolutions last month, how many would your customers call solved? And how would you know?
