If you bought an AI assistant in the last year, you have probably not seen it be wrong. That is not reassuring, and it is worth being precise about why.

An assistant that is wrong is not obviously wrong. It does not produce an error. It produces a fluent, confident, incorrect sentence, and the conversation carries on. The patient does not complain, because they do not know. They fast when they did not need to, or turn up at the branch that cannot do their test, or quietly book somewhere else because they were told you do not offer something you do offer.

None of that arrives at your front desk as a fault. It arrives as an appointment that did not happen.

Why a demo cannot show you this

Every assistant demos well. The demo is a conversation about services as they were configured on the day it was built, asked by someone who already knows the answer.

The failures that matter are the ones that appear afterwards, when your clinic moves and the assistant does not. A package changes what is inside it. A price moves. A test becomes available at one site and not another. A doctor leaves. Someone asks a question nobody ever wrote down.

A twenty-minute demonstration cannot show you whether the thing will state a wrong preparation rule in month six. Nothing can, except reading what it actually said.

The three shapes this takes

It has the answer and says it does not. A patient asks for something under a slightly different name than your catalogue uses, and the assistant tells them it is not offered. The search found it. The assistant rejected it. There is no gap, no escalation, and no trace — just a patient who went elsewhere.

It is cautious in the wrong direction. A patient asks something the assistant does not have a fact for, so it defers to your team. Sensible, until you notice the patient was left waiting overnight on a question you could have answered in one line, and was steered toward cancelling something they had no reason to cancel. Nothing false was said. The service was still bad.

It has learned something that was never policy. Somebody answered one patient, once, with a hedge — "the lab will typically run this" — and that answer became a standing rule stated confidently to everyone who asks. This is the failure mode a busy correction process walks straight into, and it is invisible from outside.

What I would ask your vendor

Not "is it accurate." Everyone says yes. Ask instead:

"How many wrong answers have you found in mine, and can I see them?"

An assistant that has generated no corrections is not a better assistant. It is an unread one. On the deployment I run, that loop turns up somewhere between twenty and seventy corrections a month, and a good share of them are facts the assistant already held and failed to use. Those are the ones nobody would ever have noticed, because nothing broke.

I am not claiming my assistant is right and other assistants are wrong. Mine produces these too. The only thing that varies between suppliers is whether somebody is reading the transcripts and doing something about what they find.

How to find out, without changing supplier

You do not need to replace anything to answer this question. It can be checked from outside, with a phone, a browser and your own booking screens — no access to your vendor's code and no cooperation required from them.

What it establishes: whether what the assistant tells patients about your services matches what you actually deliver and charge; whether "booked", "moved" and "cancelled" are true when it says them; whether it refuses what it must refuse; and whether asking for a person actually reaches one.

What you get is a findings list ranked by severity, each with the evidence attached — the exact exchange, the timestamp, and what your own records showed at that moment — plus a short list of questions to put to your vendor about the things that cannot be established from outside.

I am direct about the boundary. Fixing anything inside another supplier's product is that supplier's work, not mine. What I supply is the evidence that makes the conversation specific, instead of "it sometimes gets things wrong." Where the fix is corrected source information rather than code — a preparation and eligibility rule set, say — I can supply that in a form any system can use, and it stays yours.

It certifies nothing. It is not a security certification, not a penetration test, and not regulatory advice.

If you have an assistant answering your patients and you have never read a week of what it said, that is the thing worth doing next. Get in touch and I will tell you what it would involve for your setup — including if the answer is that you do not need me to do it.

Need an assistant that acts, not just answers?

Tell me what you're scoping. I'll tell you honestly whether it's worth building, and what it takes to let it act on your live systems without breaking them.

Get in touch

More notes from production →