Worked examples

How Business Leaders Can De-Risk AI Vendor Promises

For most companies, the biggest AI risk is the gap between what a vendor shows in a demo and what the tool does in production. Closing it means buying AI the way you would hire: each step up in commitment waits for evidence from your own data, in your own workflow.

For most companies, the biggest AI risk is usually the gap between what a vendor shows in a demo and what the business gets once the tool is in production.

Most business unit leaders have probably sat through some version of this meeting by now. A vendor opens a laptop, uploads a clean sample file, and an AI agent drafts the proposal, reconciles the invoices, or forecasts demand in a few seconds. The room is impressed (and to be fair, it often is impressive). Six months later, the pilot is still a pilot, accuracy is lower than promised, the team is quietly checking most outputs by hand, and the invoice is higher than the business case assumed.

I don’t think this means AI doesn’t work. I think it usually means AI is being bought the way enterprise software used to be bought: feature checklists, reference calls, annual licenses, etc. That approach tends to break down for AI because AI products behave differently from traditional software, and leaders who buy with that difference in mind generally capture value faster and waste less money.

Why AI promises tend to break differently

Traditional software is deterministic. If a feature works in the demo, it will usually work the same way in your environment. AI is probabilistic, so its performance depends heavily on your data, your edge cases, your workflows, and the people using it. Generally speaking, five gaps explain most of the distance between the pitch and the result.

The demo gap. Demos usually run on curated inputs. Most businesses run on scanned PDFs, inconsistent product codes, regional language variations, and a long tail of exceptions. A tool that handles the easy 80 percent of cases well can still struggle with the 20 percent that take up most of your team’s time.

The data gap. The vendor’s model has most likely never seen your data. Accuracy claims are often measured on benchmark datasets or on other customers’ data, which might look very different from yours. Until the tool is tested on your own records, the claimed accuracy is closer to a hypothesis than a fact.

The workflow gap. AI output is only valuable if someone acts on it. If a recommendation still needs three approvals, manual re-entry into the ERP, and a line-by-line human review, most of the time saving disappears. Many vendors sell the model, and fewer take responsibility for the workflow around it.

The economics gap. The license fee is rarely the full cost. Integration, data preparation, usage-based charges, human review, retraining, and change management can add up to more than the subscription itself, so business cases built on license cost alone often overstate the return.

The durability gap. AI products change over time, sometimes without much notice. Vendors swap the underlying model, retrain on new data, adjust pricing, get acquired, etc. Performance that looked fine in March might drift noticeably by September.

Across all five, the vendor usually controls the conditions under which its promise looks true, while the buyer carries the risk when those conditions change. Most of de-risking is about rebalancing that.

A five-step approach: proof before payment

The principle I’d suggest (super scientifically named “proof before payment”) is fairly simple: each increase in commitment should be matched by evidence gathered on your own data, in your own workflow. It’s a bit like hiring. You wouldn’t hire someone permanently based on a polished interview answer to a question they prepared in advance; you’d give them real work and a probation period first. AI vendors can be treated the same way.

1. Turn claims into a measurable outcome agreement

Vendor language is often hard to disprove, and therefore hard to enforce. “Up to 40 percent productivity gains” or “enterprise-grade accuracy” can mean almost anything.

Before any pilot, it usually helps to translate each claim into a baseline, a metric, and a threshold. For example: “Our team currently processes 120 supplier quotes a week with a 6 percent error rate. Success means 200 quotes a week at or below the same error rate, measured over four consecutive weeks.” If a vendor won’t agree to a specific definition of success, that hesitation is probably one of the more useful data points in the whole evaluation.

The baseline matters more than it might seem. Without a documented starting point, most pilots look like successes, and it becomes hard to tell whether the gains came from the tool, from the extra attention the process got during the pilot, or from a seasonal dip in volume.

2. Test on your own data, including the messy parts

Before talking to vendors seriously, build a “golden dataset” - a representative sample of real cases, deliberately including the hard ones (incomplete records, unusual formats, ambiguous requests, cases where your best people disagreed, etc.). Label the correct answers in advance.

Then run each vendor against it under controlled conditions and, where possible, blind: the vendor doesn’t see the answer key, and your reviewers don’t know which output came from which tool or from a person. Compare vendors against each other and against your current process, rather than against an ideal.

A few rules help keep these tests honest. The vendor shouldn’t be allowed to tune the system on the test set. It’s worth looking at where the errors happen, and not just the average; a tool that is 95 percent accurate but confidently wrong on high-value cases might be worse than one that is 90 percent accurate but flags when it’s unsure. And track how much human review the output needs, since that is often where a lot of the hidden cost sits.

In the industrial sectors I work in, the typical failure modes are fairly concrete: the tool misreads scanned supplier quotes, skips handwritten notes on drawings, or invents part numbers that look plausible. A golden dataset built from your own documents tends to surface these within the first week.

3. Price the whole system, not just the license

Ask vendors to help build a three-year total cost of ownership, and then pressure-test it internally. The categories most often missed include integration with core systems, ongoing data cleanup, usage charges that grow with adoption, human review effort, training and change management, and the cost of switching if things don’t work out.

Usage-based pricing deserves a closer look. A tool priced per query can look inexpensive in a pilot with 20 users and turn into a significant line item at 2,000. It’s worth asking for pricing at your projected scale, and what happens to your price if the vendor’s own model costs change.

The aim here isn’t to minimize cost; it’s to make sure the business case reflects what the business will pay, so the decision is made on realistic numbers.

4. Stage commitments to evidence

Probably the strongest lever a buyer has is the order in which money gets committed. Instead of signing a multi-year agreement after a good demo, the relationship can be structured as a series of gates, each unlocked by evidence.

A typical version has three stages. First, a short, tightly scoped paid proof of value that tests the tool on your golden dataset against the agreed outcome. Second, a limited production rollout in one team or region, measuring adoption, accuracy, and time savings over a full business cycle. Third, a scaled rollout once the numbers hold, with pricing tied at least partly to delivered results.

Where the vendor is confident, it’s reasonable to ask for outcome-linked terms, e.g. fee holdbacks released when thresholds are met, credits if performance drops below agreed levels, or pricing based on documents processed correctly rather than seats. Vendors who believe in their product will often accept some of this risk. If a vendor refuses every version of it, that usually tells you something about how confident they are in the pitch.

Paying for the proof of value is usually worth it. Free pilots often get junior vendor attention and weaker commitment on both sides, while a modest paid pilot tends to get the vendor’s stronger team and a more serious test.

5. Plan for change and exit from the start

Since AI products keep changing, the contract should probably address change directly. The provisions that matter most are generally advance notice before material changes to the underlying model (with the right to re-test on your golden dataset), clear ownership and portability of your data, including prompts, labeled examples, and any fine-tuned models your team helped create, audit rights over how your data is used, including whether it trains models that also serve your competitors, and a defined exit path with usable data exports and a reasonable transition period.

On the operational side, keep the golden dataset alive and rerun it every quarter, and whenever the vendor announces an update. That way, performance drift shows up as a number you’re tracking rather than as a surprise from a frustrated team.

Five questions to ask every AI vendor

Before the next vendor meeting, these questions can help separate substance from sales polish:

  • What accuracy have you achieved on data similar to ours, and can we test that on our own sample before signing?
  • Which cases does your system handle poorly, and how does it signal when it’s unsure?
  • What does the full cost look like at our projected scale in year three, including usage charges?
  • What share of your fees are you willing to tie to measured outcomes?
  • What happens to our data, configurations, and performance if you change models or if we leave?

Downloads

Views are my own and do not represent my employer.