Text editions

The Text Edition (v2) of Bot or Not? (est. 2023) examines a simple question with increasingly important consequences:

How well can readers distinguish between authentic and AI-generated online reviews, and how confident are they in doing so?

Led by Professor Claire Hardaker and Dr Georgina Brown, with major research assistance from Amy Dixon, this strand of the wider Bot or Not? project focuses on hotel reviews.

Why hotel reviews?

QR code linking to the Bot or Not Text Edition v2 quiz
Want to share the quiz?
Feel free to use this QR code.

Online reviews help people decide where to stay, what to buy, which services to trust, and especially, when to click X and keep their money safe. As this suggests, fake and misleading reviews can distort those decisions. They can also disadvantage legitimate businesses, and inevitably, they erode our trust. Generative AI intensifies the problem. It’s now quick and easy to flood the internet with fluent, tailored reviews, whether for genuine products and services or entirely fictitious ones.

In recent years the problem has taken on regulatory significance. Under the UK’s Digital Markets, Competition and Consumers Act 2024, practices involving fake reviews are prohibited, whilst businesses publishing consumer reviews are obliged to take reasonable and proportionate steps to prevent and remove fake reviews, reviews with concealed incentives, and false or misleading review information. (If you want to read more, and who doesn’t, the Competition and Markets Authority’s fake-reviews guidance explains these duties in more detail.)

The central research questions for the Text Edition are:

  • What is the overall detection accuracy for AI-generated versus human-written reviews?
  • How well calibrated are participants’ confidence judgements?
  • What linguistic or stylistic cues do readers report relying on?
  • Are those cues genuinely diagnostic, or merely persuasive?
  • Does the amount of prompt engineering affect how detectable an AI-generated review is?

Current performance snapshot (v2)

As at 20 July 2026, the Text Edition had recorded 7,066 scored responses. The mean score was 9.71 out of 15, equivalent to approximately 64.8% – a slight increase from a mean of 9.10 out of 15, or about 60.7%, in February 2026.

Scored responses Mean SD Variance
7,066 9.71 (64.8%) 2.24 5.03

The most common result was 10 out of 15, recorded in 1,285 responses. The mean is better than the 7.5 expected from blind 50:50 guessing, but it still amounts to roughly one incorrect judgement in every three. In practical terms, people appear able to detect some signals , but not with anything approaching consistent reliability.

A note on interpretation

These are live, descriptive results from an open, self-selecting public quiz. Because of this they can’t estimate how the general population would behave. The quiz is also a research and public-engagement activity, not a diagnostic tool: neither an individual score nor the cues discussed here should be used to decide that whether or not someone has used AI. Indeed, if these results show us anything, it’s that detecting AI is fraught with difficulty.

Another issue is that as the quiz is taken by various groups, the differing participant pools can approach the quiz with more or less prior knowledge. Some groups have arrived as part of their training. Others are taking this for fun. The size of the dataset still allows for extrapolations, but these must be made with care.

The Guardian and a new wave of participation

On 4 July 2026, David Shariatmadari’s Guardian feature, How AI is changing language, opened with a miniature version of the Text Edition, linked directly to the full quiz and discussed our then-current finding that people were identifying the source correctly only around 60% of the time.

The feature prompted a major new wave of participation. In February 2026, the Text Edition had 589 scored responses and a cumulative mean of 9.10 out of 15 (60.7%). By 20 July, the response count had risen almost twelvefold to 7,066, while the cumulative mean had increased to 9.71 (64.8%). The Text Edition received by far the largest increase, although the coverage also brought a smaller rise in participation in the Speech Edition and Music Edition.

The new figures expand the dataset considerably, but the public response matters beyond the raw total. It shows how strongly questions of AI authorship, detection and misplaced certainty now resonate. The next question is what people do with that uncertainty: whether it affects how they evaluate online material, make decisions, teach others or approach claims that a text “looks AI-generated”.

Design overview (v2)

Each participant sees 15 positive hotel reviews, randomly selected from a bank of 400:

A participant’s set can therefore be entirely human-authored, entirely AI-generated or, more usually, a mixture. The AI-generated reviews were produced at three broad levels of prompt engineering:

  • Minimal: for example, “Write a positive review of Hotel XYZ.”
  • Moderate: for example, “Write a positive review of Hotel XYZ and include specific details.”
  • Extensive: for example, “Write a positive review of Hotel XYZ, include an array of specific details, and include some typographical errors.”

All hotel names are replaced with Hotel XYZ to reduce brand-familiarity effects.

The task proceeds in four stages:

  1. Before seeing the reviews, participants rate their overall confidence in their ability to distinguish human from AI-generated writing.
  2. For each of the 15 reviews, they make a binary human or AI-generated judgement.
  3. After completing those judgements, they explain the cues or reasoning they relied on and rate their overall confidence again.
  4. Only after those responses have been submitted do they receive their score.

A key methodological refinement in v2 is the separation of accuracy and confidence. Scores are based solely on correct binary classifications, while confidence is measured independently. This enables us to examine:

  • calibration between confidence and accuracy;
  • patterns of overconfidence and underconfidence;
  • changes in perceived ability after exposure to the task; and
  • the relationship between the cues people report using and the features that are genuinely predictive.

This design also brings the Text Edition into alignment with the Speech and Music Editions, enabling cautious comparison across written language, spoken language and singing.

What can the dataset tell us?

The Text Edition supports research into:

  • human judgement and metacognition under uncertainty;
  • the linguistic cues associated with perceived AI authorship;
  • the difference between perceived and statistically predictive cues;
  • prompt engineering and adversarial text generation;
  • platform governance and consumer protection; and
  • AI literacy and public resilience.

Of particular interest is the gap between perceived cues and reliable discriminators. Features such as polish, repetition, detail, punctuation, and formulaic structure may feel revealing, but human writing can contain all of them, while AI-generated writing can be prompted to avoid them. A plausible explanation is not necessarily an accurate one.

Has Bot or Not? been useful?

If you have used Bot or Not? in teaching, training, public engagement, research, policy or professional practice – or if taking part affected how you assess online material – we would like to hear what changed, if anything. Examples are extremely welcome. Please email the FACTOR team at factor@lancaster.ac.uk.

What about version 1?

The first iteration of the Text Edition – and the first quiz in the entire Bot or Not? suite – closed on 13 May 2025 with 957 completed responses. In v1, confidence was embedded directly within each answer using a five-point categorical scale:

  • Definitely human
  • Maybe human
  • Not sure
  • Maybe bot
  • Definitely bot

Scoring was weighted:

  • 1 point for a correct “definitely” judgement; and
  • 0.5 points for a correct “maybe” judgement.

This provided useful gradience data, but it also introduced a possible behavioural distortion: stronger answers attracted more points, potentially encouraging participants to overstate their certainty. In other words, the measurement instrument risked shaping the phenomenon it was intended to observe.

Version 2 therefore separates binary classification from task-level confidence. Because the two versions use different response and scoring systems, their mean scores should not be treated as directly comparable.

Final scores snapshot (v1)

Completed responses Mean SD Variance
957 7.19 2.34 5.48

Final Bot or Not Text Edition v1 score distribution across 957 completed responses

Across both versions, the Text Edition continues to ask a simple but increasingly consequential question: when we read persuasive online prose, what convinces us that a human is behind it? And how often are we mistaken?