A demo shows a product working; before buying, find out how it fails, what its accuracy figures rest on, and what it costs to run and to leave.
- Ask what it does when it is wrong. What happens on the failures, who notices and how fast tells you more than an accuracy figure.
- Question any accuracy figure. Ask what it was measured on, when, what counted as correct, and for the false positives rather than a single percentage.
- Test it on your own data. Ask to see a failure and to run a trial on your data; the demo dataset was chosen, yours was not.
- Pin down the model, the data terms, the cost and the exit. Ask which model is underneath and what happens when it changes, get data terms in the contract, count the all-in cost at your volume, and ask how you leave before signing.
- A red flag means unverified, not bad. An undated accuracy figure, a claim that it does not hallucinate or a refused trial leaves the claim unchecked.
The prices here come from a third-party survey, not from the vendors’ own pages, and were correct when this site read it on 17 Sep 2026. Prices change often and this site no longer updates them, so check each vendor’s own site before you rely on one. Survey read: CloudTalk, Retell AI pricing.
Every vendor demo works. That is what a demo is, and a benchmark in a launch post is a claim rather than a measurement. These are the questions that separate a product from a demonstration, and most of them can be asked in a first call.
"What does it do when it is wrong?" Not "how accurate is it" — every vendor has a number for that. What happens on the failures, who notices, and how fast.
A vendor who has not thought about this has not run it in production.
On accuracy claims
- Measured on what? A benchmark, their own test set, or customer data. Their own test set is not evidence, because they chose it.
- Measured when? A figure from a model version that no longer exists describes nothing you can buy.
- What counts as correct? Binary scoring hides the failure mode that matters — a confidently wrong answer scores the same as a blank one on some rubrics and better on others.
- What is the false-positive rate? Vendors quote accuracy; the cost usually lives in false positives. Independent testing of AI-text detectors found 61.3% false-positive rates on non-native English writers.Liang, Yuksekgonul, Mao, Wu & Zou, Patterns 4(7):100779, July 2023, read at source 23 Sep 2026: “they incorrectly labeled more than half of the TOEFL essays as ‘AI-generated’ (average false-positive rate: 61.3%)”.
Ask for the confusion matrix, not the accuracy. A single percentage is a summary of four numbers, and the vendor chose which one to show you.
Five questions worth asking
What is underneath, and what happens when it changes?
If the product is built on another company’s model, that is fine — but ask which one, and what happens when it is deprecated. OpenAI discontinued the Sora web and app experiences on 26 April 2026, and the Sora API shuts down on 24 September 2026; anything built on it inherits that date.OpenAI Help Center, What to know about the Sora discontinuation, read at source 17 Sep 2026: “The Sora web and app experiences were discontinued on April 26, 2026.” “The Sora API will be discontinued on September 24, 2026.”
Where does our data go, and is it trained on?
Ask for it in the contract, not the sales call. "We don't train on customer data" and "we don't train on customer data by default" are different sentences.
Show me a failure
Ask them to demonstrate the product getting something wrong. A vendor who cannot produce a failure on request either has not looked or will not say.
What is the all-in cost at our volume?
Per-seat pricing usually excludes the API meter. For voice, the advertised per-minute rate is a floor — the all-in runs $0.13–$0.31 per minute once speech-to-text, text-to-speech, the LLM and telephony are counted.CloudTalk, Retell AI pricing, read at source 17 Sep 2026: “the advertised $0.07/min covers the voice infrastructure layer only”, with real production costs “landing most teams at $0.13–$0.31/min once a working agent is configured.” Autocalls, Aircall, Trillet, Kommunicate and WhiteLabelAI analyses, Mar–Aug 2026, were also cited when this was written and were not re-read.
How do we leave?
Data export format, notice period, what happens to anything they fine-tuned on your behalf. Ask before signing, because the answer is much worse after.
Five things that should stop the conversation
- An accuracy figure with no date and no test set named.
- "It doesn't hallucinate." That is a claim to test on your own data, not to take on trust; a vendor who makes it and will not let you test it is counting on you not checking.
- No answer on what happens when the underlying model changes.
- Case studies with no numbers, or numbers with no denominator. “Four times the output” without “of what” is a shape, not a result.
- Refusal to run it on your data in a trial. The demo dataset is chosen; yours is not.
None of these means the product is bad. All of them mean the claim is unverified, and the distinction is the whole job — see auditing AI-generated work for the same discipline applied to output.