Music editions

The Music Edition (v1) of Bot or Not? (est. 2025) examines a simple question with increasingly important consequences:

How well can listeners distinguish between authentic human singing and AI-generated singing voices, and how confident are they in doing so?

Led by Professor Claire Hardaker and Dr Georgina Brown, with major research assistance from Lydia Cooper and Hope McVean, this ESRC-supported strand of the wider Bot or Not? project focuses on singing voices.

Why singing?

QR code linking to the Bot or Not Music Edition v1 quiz
Want to share the quiz?
Feel free to use this QR code.

Music is not only an acoustic signal. It’s a significant component of our identity, memory, emotion, embodiment, community, and beliefs about creativity. Singing presents a particularly complex test. A lot of modern music involves synths, auto-tune, voice effects, and more besides. But even if it didn’t, singing is as different to speaking as sprinting is to strolling. Singing combines pitch, vibrato, breath, phrasing, timbre, style, and emotional projection, and any of these could feel like a signal of “humanness”. However, whether they are genuinely reliable is another matter.

This may help to explain why AI-generated music provokes such passionate and sometimes visceral responses. For some listeners, the possibility of a compelling performance without a human singer touches on what makes people sentient. For others, the more immediate question is whether the music does something worthwhile, regardless of how it was made.

The industrial questions are no simpler. Longstanding arguments about gatekeeping, ownership and how performers are paid did not begin with generative AI. These systems could lower barriers created by cost, specialist training, disability, geography, and access to professional networks. They could allow people to make highly personal or niche music that commercial gatekeepers have rarely supplied.

At the same time, these systems could also introduce new forms of mass exploitation. Musicians and other creators have raised concerns about consent, control and payment when creative work is used to train generative models, as well as synthetic imitation, automated content at enormous scale, and the concentration of rewards among model and platform owners. The UK’s continuing copyright and AI debate reflects how unsettled these questions remain.

Both sets of possibilities can be true. Generative AI may widen participation and enable experimentation while also making extraction and mass production easier. It may support intensely personal creativity and produce oceans of generic white noise. The useful questions are therefore not simply whether AI-generated music is good or bad, but under what conditions it is made, whose work it relies on, who benefits, who is paid, and what audiences are told. Controversies such as The Velvet Sundown illustrate how quickly authorship and its link with authenticity become socially and legally charged.

The central research questions for the Music Edition are:

  • What is the overall detection accuracy for AI-generated versus human singing voices?
  • How well calibrated are participants’ confidence judgements?
  • What acoustic or perceptual cues do listeners report relying on?
  • Are those cues genuinely diagnostic, or merely persuasive?
  • How do voice type, genre and production style affect how detectable AI-generated singing is?

Current performance snapshot (v1)

As at 20 July 2026, the Music Edition had recorded 883 scored responses. The mean score was 4.92 out of 9, equivalent to approximately 54.6% – essentially unchanged from a mean of 4.90 out of 9, or about 54.4%, in February 2026.

Scored responses Mean SD Variance
883 4.92 (54.6%) 1.65 2.71

July 2026 Bot or Not Music Edition v1 score distribution across 883 responses

The most common result was 5 out of 9, recorded in 241 responses. The mean is only slightly better than the 4.5 expected from blind 50:50 guessing, and it still amounts to almost one incorrect judgement in every two. In practical terms, people do not appear able to distinguish human from AI-generated singing with anything approaching consistent reliability.

A note on interpretation

These are live, descriptive results from an open, self-selecting public quiz. Because of this, they can’t estimate how the general population would behave. The quiz is also a research and public-engagement activity, not a diagnostic tool: neither an individual score nor the cues discussed here should be used to decide whether a particular singing voice is human or AI-generated. Indeed, if these results show us anything, it’s that detecting AI-generated singing from short samples is fraught with difficulty.

Another issue is that, as the quiz is taken by various groups, differing participant pools can approach it with more or less prior knowledge. Some groups have arrived as part of their training. Others are taking this for fun. The size of the dataset still allows for extrapolations, but these must be made with care.

The Guardian and a new wave of participation

On 4 July 2026, David Shariatmadari’s Guardian feature, How AI is changing language, opened with a miniature version of the Text Edition and linked directly to that quiz. The article also discussed the wider Bot or Not? suite, including its expansion into music and the strength of people’s reactions to AI-generated songs.

The feature prompted a major new wave of participation in the Text Edition and a smaller but still substantial increase here. In February 2026, the Music Edition had 374 scored responses and a cumulative mean of 4.90 out of 9 (54.4%). By 20 July, the response count had risen by 509 – an increase of 136.1% – to 883, while the cumulative mean had barely changed at 4.92 (54.6%).

The new figures expand the dataset considerably, but the public response matters beyond the raw total. It shows how strongly questions of musical authenticity, creativity, and misplaced certainty now resonate. The next question is what people do with that uncertainty: whether it affects how they listen, credit artists, evaluate provenance, teach others, or respond to a track whose singer may not exist.

Design overview (v1)

Each participant hears nine singing samples, randomly selected from a curated and expanding bank of human and AI-generated singing voices. A participant’s set can therefore be entirely human, entirely AI-generated or, more usually, a mixture.

Participants are instructed to focus on the singing voice only. Instrumentation may be synthetic in both human- and AI-voiced tracks, while human performances may legitimately include processing such as pitch correction. We’re not asking whether the production sounds modern or manipulated. What we want to know is whether the origin of the singing voice is perceptually identifiable.

The task proceeds in four stages:

  1. Before hearing the samples, participants rate their overall confidence in their ability to distinguish human from AI-generated singing.
  2. For each of the nine samples, they make a binary human or AI-generated judgement.
  3. After completing those judgements, they explain the cues or reasoning they relied on and rate their overall confidence again.
  4. Only after those responses have been submitted do they receive their score.

As in the current Text and Speech Editions, accuracy and confidence are separated. Scores are based solely on correct binary classifications, while confidence is measured independently. This enables us to examine:

  • calibration between confidence and accuracy;
  • patterns of overconfidence and underconfidence;
  • changes in perceived ability after exposure to the task;
  • error patterns across voice types, genres and production styles; and
  • the relationship between the cues people report using and the features that are genuinely predictive.

This design brings the Music Edition into alignment with the Text and Speech Editions, enabling cautious comparison across singing, written language and spoken language.

What can the dataset tell us?

The Music Edition supports research into:

  • human judgement and metacognition under uncertainty;
  • the acoustic and perceptual cues associated with perceived AI generation;
  • the difference between perceived and statistically predictive cues;
  • the effects of voice type, genre and production style on detection;
  • audience assumptions about authenticity, authorship and creativity;
  • performer identity, attribution, consent, remuneration and disclosure;
  • creative-industry practice and policy; and
  • AI literacy and public understanding.

Of particular interest is the gap between perceived cues and reliable discriminators. Features such as breathing, pitch, vibrato, timbre, phrasing, and emotional expression may feel revealing, but processed human vocals can sound synthetic, while AI-generated singing can reproduce or avoid many of the same features. A performance can feel deeply human without being correctly identified, while an unusual human performance may be mistaken for AI. A plausible explanation is not necessarily an accurate one.

Detection is also only part of the issue. A listener may value the same track differently once its provenance is disclosed, even though nothing in the sound has changed. That response connects perceptual judgement with authorship, consent, labour, and beliefs about creativity. It does not follow that AI-generated music is inherently good or bad. It does mean that questions about what it enables, whose work it relies on, and how its consequences are distributed cannot be answered by an accuracy score alone.

Has Bot or Not? been useful?

If you have used Bot or Not? in teaching, training, public engagement, research, policy or professional practice – or if taking part affected how you assess online material – we would like to hear what changed, if anything. Examples are extremely welcome. Please email the FACTOR team at factor@lancaster.ac.uk.

At its core, the Music Edition continues to ask a simple but increasingly consequential question: when we hear singing, what convinces us that a human is behind it? How often are we mistaken? And what changes when the answer matters not only to perception, but to identity, livelihoods, rights, and our sense of what creativity is?