Original editions

The Original Editions are three open public perception tests asking the same basic questions across written language, spoken language and singing: can people distinguish human-created content from AI-generated content, how confident are they, and what cues do they think give the game away?

As at 20 July 2026, the current editions had recorded 9,630 scored response sets, comprising 139,152 individual human-or-AI judgements.

The headline is uncomfortable but useful: mean accuracy ranges from 54.6% to 69.8%. People perform better than blind 50:50 guessing, but not well enough for “it looks AI-generated” or “it sounds AI-generated” to count as proof.

Led by Professor Claire Hardaker and Dr Georgina Brown, the three editions were developed with major research assistance from Amy Dixon, Hope McVean and Lydia Cooper.

Choose an edition

Text Edition (v2): hotel reviews

Participants classify 15 positive hotel reviews. Across 7,066 scored responses, mean accuracy is 64.8%: better than chance, but roughly one judgement in three is still wrong. This is the largest of the three datasets and sits directly within current questions about fake reviews, consumer protection and platform responsibility.

Take the Text Edition  |  Read the full Text overview and results

Speech Edition (v2): short recordings

Participants classify 15 speech samples of around three to five seconds. Across 1,681 scored responses, mean accuracy is 69.8%, the highest of the current editions – but roughly three judgements in ten are still wrong. In calls, voice notes and other situations where a voice is being used as evidence of identity, that is a sizeable margin for error.

Take the Speech Edition  |  Read the full Speech overview and results

Music Edition (v1): singing voices

Participants classify nine singing samples. Across 883 scored responses, mean accuracy is 54.6%: only slightly above chance, with almost one judgement in two wrong. Music also takes the project beyond detection into questions of creativity, identity, consent, copyright, remuneration and whether audiences respond differently when the performer behind a compelling voice may not exist.

Take the Music Edition  |  Read the full Music overview and results

Results in one view

The raw mean scores can’t be placed directly beside one another because the Text and Speech Editions are scored out of 15, while the Music Edition is scored out of nine. Converting each mean to a percentage provides a common descriptive scale. It does not turn the three quizzes into a controlled head-to-head experiment.

Current public quiz results as at 20 July 2026
Edition Scored responses Individual judgements Mean correct Mean incorrect Above chance
Text (v2) 7,066 105,990 64.8% 35.2% +14.8 points
Speech (v2) 1,681 25,215 69.8% 30.2% +19.8 points
Music (v1) 883 7,947 54.6% 45.4% +4.6 points
Total 9,630 139,152
Mean classification accuracy
All three editions shown on the same 0-100% scale. The middle of the scale, 50%, marks the result that would be expected from randomly guessing at every question.
Text64.8%

 

Speech69.8%

 

Music54.6%

 

How the totals were calculated: each scored Text and Speech response contains 15 classifications, and each scored Music response contains nine. Response sets are quiz submissions, not a count of unique people; someone may complete more than one edition or take a quiz more than once.

Important caveats

These are live, descriptive results from open, self-selecting public quizzes, not estimates of how the general population would perform. The editions also differ in their stimulus types, item counts, versions, source banks, and participant pools. Speech currently has the highest mean and Music the lowest, but given their underlying differences, that shouldn’t be treated as controlled proof that one medium is inherently easier to classify than another.

The quizzes are research and public-engagement activities, not diagnostic tools. Neither a score nor a cue reported by participants should be used as proof that a particular review, recording, or singing voice is AI-generated.

The common thread

Across all three editions, one distinction keeps returning: a cue that feels persuasive is not necessarily a cue that reliably predicts provenance. Participants identify issues such as fluency, blandness, and a lack of specific details that make a review feel synthetic. Flat intonation, unusual pacing, and a lack of breaths may do the same for speech. Pitch, vibrato, lyrics, and emotional expression may influence judgements about singing. However, human content can also contain all of these features, whilst AI-generated content can reproduce or avoid them.

The current designs therefore separate binary accuracy from pre- and post-task confidence, and collect participants’ explanations before revealing their scores. This lets us examine not only whether people are right, but whether they know when they are likely to be wrong, and whether the cues they describe are genuinely useful.

The consequences differ by domain. For text, the immediate issues include consumer decisions, fake-review regulation, and platform responsibility. For speech, they include fraud, impersonation, identity, and the evidential status of a voice. For music, they extend into authorship, consent, creative labour, copyright, and disclosure. Generative AI can support accessibility, experimentation and new forms of participation; it can also be used for manipulation, extraction and mass production. We don’t start from the premise that the technology is simply good or bad. We simply want to understand what people can actually perceive, how certain they should be, and by extension, what happens when those judgements are used in the real world.

Why now?

On 4 July 2026, David Shariatmadari’s Guardian feature, How AI is changing language, opened with a miniature version of the Text Edition, linked directly to the full quiz and discussed the wider suite. It prompted a major new wave of participation, especially in the Text Edition, with smaller but still clear increases in Speech and Music. The response matters beyond the additional data: questions of AI authorship, authenticity, and misplaced certainty are issues for everyone, from consumers and school teachers to banks and politicians.

Want more?

Discussions around AI are happening at local, national, and global levels. If you’d like to read more from within the UK policy landscape, however, then useful starting points include the Competition and Markets Authority’s fake-reviews guidance, the Home Office’s National Assessment Centre Fraud Assessment 2025, the Department for Science, Innovation and Technology’s report on deepfake detection technology, the UK Government’s report on copyright and artificial intelligence, and its AI Adoption Plan for the Creative Industries.

Has Bot or Not? been useful?

If you have used Bot or Not? in teaching, training, public engagement, research, policy or professional practice – or if taking part affected how you assess online material – we would like to hear what changed, if anything. Examples are extremely welcome. Please email the FACTOR team at factor@lancaster.ac.uk.