Insights
How many callers notice they're talking to an AI? We counted all 102,154.
Callable ·
The objection that kills most AI voice deals isn't price or capability. It's will my customers know? — and behind that, will they mind?
We had never actually measured it. So we ran a read-only query across 30 days of production call logs and counted every caller who asked.
The answer
253 of 102,154 connected human conversations contained an explicit challenge to the agent's humanity. That's 0.25% — about one in 404.
Measured 19 July to 17 August 2026, across 11 business tenants. Wilson 95% interval 0.22%–0.28%, and stable week to week between 0.21% and 0.32%.
Three quarters of those aren't what you'd assume
This is the part that changed how we read the result.
| How the question got asked | Conversations | Share |
|---|---|---|
| Human / real-person challenge | 187 | 73.9% |
| Called a machine or computer | 32 | 12.6% |
| Asked if it was a recording | 22 | 8.7% |
| Named explicitly as AI | 7 | 2.8% |
| Called a robot or bot | 6 | 2.4% |
Three quarters of detections are the generic "are you a real person?" — a question people also ask human offshore call centres, and have for twenty years. It signals scepticism about the call, not a conclusion about the caller.
Only 7 conversations in 102,154 — 0.007% — named the agent as AI unprompted.
That matters for how the headline should be read. Because our rule matches phrasing rather than intent, it counts ordinary call-centre scepticism identically to genuine detection. 0.25% is therefore a conservative ceiling on confident detection, not a floor.
A few verbatim, to show the texture:
- "Are you a real person?" — 14 conversations
- "Are you real?" — 6 conversations
- "Is this a recording?" — 4 conversations
- "Alex, is this a prerecorded message, or are you a real person?"
- "Am I talking to a real person or AI?"
Note the fourth. They used the AI Voice Agent's name, then asked what it was — that is someone comfortable enough in the conversation to be curious about it.
The longer they talk, the more likely they ask
| Conversation length | Detections | Rate |
|---|---|---|
| 25–59 seconds | 128 of 75,262 | 0.17% |
| 60–119 seconds | 79 of 21,611 | 0.37% |
| 2–5 minutes | 43 of 5,207 | 0.83% |
| Over 5 minutes | 3 of 76 | 3.95% |
A clean gradient, with one confound worth naming: the small population of very long conversations is enriched with callers who stayed on precisely because they were probing the agent. Read the 3.95% cell — three conversations out of 76 — as directional, not as a stable rate.
How we counted
Every row in the production call log for the 30-day window, filtered down to conversations that actually happened:
| Filter | Why |
|---|---|
| Has a transcript | Nothing to scan otherwise |
| Duration ≥ 25 seconds | Below that, nobody has formed an impression |
| At least 3 caller turns | A real exchange, not a hang-up |
We then excluded the AI Voice Agent's own speech and scanned only the caller's turns, for AI, robot, machine and "are you real" language.
The exclusion that changes the number
The raw scan flagged far more than expected — until we looked at what was flagging.
132,276 conversations met the inclusion criteria. 28,787 of them — 21.8% — weren't humans at all. They were voicemail greetings, carrier and Google Call-Assist screening bots, and "this call is being recorded" announcements. Those announcements trip the same keywords. But that is one machine talking to another, and it tells you nothing about whether a person noticed anything.
Removing them left 103,489 genuine human conversations, of which 102,154 had machine-parseable caller turns. Against that population: 253 detections, or 0.25%.
Agent speech was discarded before matching, so the AI Voice Agent's own wording can never trigger a detection — and a conversation counts once, however many times the caller raises it.
Why we didn't use a sample
We ran a 500-call random sample first. It returned 1 detection in 493 usable conversations — 0.2%.
That looked close to the eventual answer, but it was luck. At a base rate this low, a single call moves a 500-call sample by 0.2 percentage points. The sample had no power to distinguish 0.1% from 0.4%, which is the entire range we cared about.
So we discarded it and ran the identical detector over the full population. 0.25% is a census, not an estimate. There is no confidence interval on it because we didn't sample — we counted.
The single detection it did find is a nice artefact in its own right: that caller said "Are you real?" and then, moments later, "You sound like a real person, mate."
What this number does not say
This is the part most vendors would leave out.
It measures what callers said, not what they thought. Someone can suspect they're talking to an AI and never mention it — most people don't interrogate a caller, they just get on with the conversation. The honest claim is that about one caller in 404 asked. It is emphatically not that the other 403 couldn't tell, and we'd ask people not to quote it that way.
Keyword matching isn't comprehension. The rule matches phrasing, not intent — which is why 0.25% is a ceiling on confident detection rather than a floor.
Transcription error cuts both ways. Detections rest on speech recognition. Spot-checking found occasional garbles that inflate counts (one transcript reads "you and I are bot" in an unrelated context) and near-misses that deflate them. Small relative to the effect, and running in both directions.
Single vendor, single market. Every conversation was placed by a CallableAI AI Voice Agent to an Australian number, largely in solar, finance and services outbound. Don't read it as a property of conversational AI in general.
And the most important one: our AI Voice Agents introduce themselves as AI. This isn't a measure of whether a synthetic voice passes for human. It's a measure of what happens after disclosure — which is the number we actually wanted, because disclosure is our position either way.
Why this is the more useful finding
If the claim were "callers can't tell", the interesting fact would be about voice quality. It isn't, and we're not making that claim.
The finding is that you can tell people they're speaking to an AI and the conversation still works. Across 102,154 conversations — through qualification, questions, objections and bookings — the question came up 253 times, and only seven of those named it as AI.
That is a much better position to build on than hoping nobody asks. It's also the only one that survives contact with the Australian Consumer Law, where actively denying you're an AI is a materially different risk from simply not being asked.
The method is the point
We've published this with the population, the filters, the 28,787 exclusions and the discarded sample because a number without a method is a number nobody can check. The full write-up, including the reproducibility section, is Research Note CA-2026-01.
Every figure CallableAI publishes lives on our research page with the same treatment — including the ones that are still product targets rather than measurements, which are labelled as such.
If you're evaluating AI voice vendors, ask each of them how they measured whatever they're claiming, and how many calls it was measured over. The answers are informative.
Deployed white glove. Backed by an SLA.
CallableAI works with organisations that need quality, volume and compliance to hold at once. Bring us your requirements — we'll scope the deployment, the rollout and the SLA with you.