Speech editions

The Speech Edition (v2) of Bot or Not? (est. 2024) examines a simple question with increasingly important consequences:

How well can listeners distinguish between authentic and AI-generated speech, and how confident are they in doing so?

Led by Professor Claire Hardaker and Dr Georgina Brown, with major research assistance from Hope McVean, this strand of the wider Bot or Not? project focuses on short speech recordings.

Why synthetic speech?

QR code linking to the Bot or Not Speech Edition v2 quiz
Want to share the quiz?
Feel free to use this QR code.

Voices help us decide who we’re listening to, whether to trust a caller or recording, and sometimes whether to disclose information, transfer money, or grant access. Synthetic speech has legitimate applications in areas such as accessibility, entertainment and translation, but it also creates opportunities for fraud, impersonation, defamation, disinformation, harassment, and other even more sinister crimes.

AI-generated speech can now reproduce highly naturalistic prosody, intonation and timbre. Short samples may reach us as voice notes, audio clips, parts of longer calls, sections from CCTV and so forth, and sometimes in circumstances where a decision has to be made quickly. The issue is therefore not simply whether a voice sounds synthetic. It is whether listeners can identify a synthetic voice reliably enough in the first place to act on that judgement.

The central research questions for the Speech Edition are:

  • What is the overall detection accuracy for AI-generated versus human speech?
  • How well calibrated are participants’ confidence judgements?
  • What acoustic or perceptual cues do listeners report relying on?
  • Are those cues genuinely diagnostic, or merely persuasive?
  • How do voice quality and line quality affect how detectable AI-generated speech is?

Current performance snapshot (v2)

As at 20 July 2026, the Speech Edition had recorded 1,681 scored responses. The mean score was 10.48 out of 15, equivalent to approximately 69.8% – a slight increase from a mean of 10.46 out of 15, or about 69.7%, in February 2026.

Scored responses Mean SD Variance
1,681 10.48 (69.8%) 2.11 4.44

The most common result was 11 out of 15, recorded in 324 responses. The mean is better than the 7.5 expected from blind 50:50 guessing, but it still amounts to roughly three incorrect judgements in every ten. In practical terms, people appear able to detect some signals, but not with anything approaching consistent reliability.

A note on interpretation

These are live, descriptive results from an open, self-selecting public quiz. Because of this, they can’t estimate how the general population would behave. The quiz is also a research and public-engagement activity, not a diagnostic tool: neither an individual score nor the cues discussed here should be used to decide whether a particular recording contains human or AI-generated speech. Indeed, if these results show us anything, it’s that detecting AI-generated speech from short samples is fraught with difficulty.

Another issue is that, as the quiz is taken by various groups, differing participant pools can approach it with more or less prior knowledge. Some groups have arrived as part of their training. Others are taking this for fun. The size of the dataset still allows for extrapolations, but these must be made with care.

The Guardian and a new wave of participation

On 4 July 2026, David Shariatmadari’s Guardian feature, How AI is changing language, opened with a miniature version of the Text Edition and linked directly to that quiz. The article also discussed the wider Bot or Not? suite, including its expansion into speech and music.

The feature prompted a major new wave of participation in the Text Edition and a smaller but still clear increase here. In February 2026, the Speech Edition had 1,161 scored responses and a cumulative mean of 10.46 out of 15 (69.7%). By 20 July, the response count had risen by 520 – an increase of 44.8% – to 1,681, while the cumulative mean had barely changed at 10.48 (69.8%).

The new figures expand the dataset considerably, but the public response matters beyond the raw total. It shows how strongly questions of voice authenticity, detection and misplaced certainty now resonate. The next question is what people do with that uncertainty: whether it affects how they evaluate calls, voice notes and recordings, make decisions, teach others or approach claims that a voice “sounds AI-generated”.

Design overview (v2)

Each participant hears 15 speech samples of around three to five seconds, randomly selected from a curated bank of 401 recordings:

  • 200 authentic human speech recordings; and
  • 201 AI-generated speech recordings.

All materials are drawn from ASVspoof 2019, an international benchmark dataset developed for automatic speaker verification and anti-spoofing research. This allows the human-perception findings to sit alongside computational work on synthetic-speech detection.

The 201 AI-generated recordings are distributed across two dimensions: perceived voice quality, ranging from poor or obviously synthetic to highly convincing; and line quality, ranging from phone quality to studio quality.

Line quality ➡️
⬇️ Voice quality
Phone
(poor)
Internet
(medium)
Studio
(excellent)
Total
Poor 22 22 23 67
Medium 22 23 22 67
Excellent 23 22 22 67
Total 67 67 67 201

The extra AI-generated recording is deliberate: 201 allows every voice-quality band and every line-quality band to contain exactly 67 recordings. This design lets us investigate whether, for instance, a poor-quality line masks artefacts that might otherwise reveal synthetic speech, or whether studio-quality audio makes those artefacts easier to hear.

A participant’s set can be entirely human, entirely AI-generated or, more usually, a mixture.

The task proceeds in four stages:

  1. Before hearing the recordings, participants rate their overall confidence in their ability to distinguish human from AI-generated speech.
  2. For each of the 15 samples, they make a binary human or AI-generated judgement.
  3. After completing those judgements, they explain the cues or reasoning they relied on and rate their overall confidence again.
  4. Only after those responses have been submitted do they receive their score.

A key methodological refinement in v2 is the separation of accuracy and confidence. Scores are based solely on correct binary classifications, while confidence is measured independently. This enables us to examine:

  • calibration between confidence and accuracy;
  • patterns of overconfidence and underconfidence;
  • changes in perceived ability after exposure to the task;
  • the effects of voice quality and line quality on detection; and
  • the relationship between the cues people report using and the features that are genuinely predictive.

This design also brings the Speech Edition into alignment with the Text and Music Editions, enabling cautious comparison across spoken language, written language and singing.

What can the dataset tell us?

The Speech Edition supports research into:

  • human judgement and metacognition under uncertainty;
  • the acoustic and perceptual cues associated with perceived AI generation;
  • the difference between perceived and statistically predictive cues;
  • the effects of voice quality and line quality on detection;
  • forensic phonetics and speaker comparison;
  • fraud, impersonation and social-engineering risk; and
  • AI literacy and public resilience.

Of particular interest is the gap between perceived cues and reliable discriminators. Features such as accent, flat intonation, breathing and unusual pacing may feel revealing, but human speech can contain all of them, while synthetic speech can reproduce or avoid them. A plausible explanation is not necessarily an accurate one.

Has Bot or Not? been useful?

If you have used Bot or Not? in teaching, training, public engagement, research, policy or professional practice – or if taking part affected how you assess online material – we would like to hear what changed, if anything. Examples are extremely welcome. Please email the FACTOR team at factor@lancaster.ac.uk.

What about version 1?

The first iteration of the Speech Edition was created at short notice for an event as a modest online form comprising 12 speech samples. You can still play it here.

Participants received a score at the end, but v1 did not collect:

  • confidence measures;
  • qualitative explanations; or
  • item-level responses or scores for later analysis.

Although several hundred people completed the task, the absence of retained response data limited its research utility. Being present while people undertook the quiz still taught us a great deal, and the simplicity of the instrument worked well for that particular event, but we quickly missed the methodological depth.

For these reasons, the project was redeveloped into the current Qualtrics-based v2. This expanded the stimulus bank to 401 recordings, captured richer response data and aligned the Speech Edition methodologically with the Text and Music Editions. There is no like-for-like v1 benchmark against which the current results can be compared.

Across both versions, the Speech Edition continues to ask a simple but increasingly consequential question: when we hear a voice in a short recording, what convinces us that a human is behind it? And how often are we mistaken?