Turing text editions

The first Turing Text Edition of Bot or Not? takes the project beyond fixed samples and into live interaction. It examines a simple question with increasingly serious consequences:

After a short text conversation, how well can humans and AI systems distinguish a human partner from an AI one, and how does the interaction change their confidence?

Led by Professor Claire Hardaker and Dr Georgina Brown, with major research assistance from Hope McVean, this pilot uses the fictional dating app La Vida Lanca. Each participant meets Alex, exchanges questions and answers, and then decides whether Alex was a bot… or not.

Why interactive text?

The original Text, Speech and Music Editions ask participants to judge fixed samples. Those samples can be examined repeatedly, but they can’t respond, adapt, or make mistakes across a sequence of turns. It’s a one-shot deal. And it’s the same for you, the participant. You read or listen. You judge. And then you move on.

Interaction changes the problem significantly. A conversational partner must respond appropriately to what has just been said. They need to manage their questions and answers. They have to reciprocate interest, accommodate the other’s style, and if they’re an AI, they’re going to need to maintain a plausible identity over time. Participants can also probe their partner rather than merely inspect what they have been given, and we found that some of our interactants did exactly this. One day we’ll tell you about the hammerhead shark test, which we especially enjoyed.

All of this matters because conversational AI is increasingly being combined with more autonomous, agentic systems, and for fraudsters, this presents an extraordinary opportunity. In principle, such systems can create profiles, maintain multiple conversations, personalise their approaches, and operate continuously at considerable scale. Their human operator might decide to intervene only when a conversation becomes financially promising, but  in some cases, if their system is good enough and the funds are arriving in their bank accounts, they may not need to intervene at all.

We are of course talking about romance fraud.

Why romance fraud?

Romance fraud depends on interaction. Offenders establish trust and emotional attachment over days, weeks, or even months before manipulating the situation into one where money becomes an issue. Sometimes they will play on anxiety, fear, guilt, or greed by introducing an emergency, an investment opportunity, or a travel problem. But sometimes they quite simply exploit kindness. A problem that can be solved by money arises. They appear to brave it out and say that they’ll find the funds somehow. Don’t worry about it. They’ll manage. Somehow.

And that’s when the well-meaning recipient of these messages might offer to step in and save the day, unaware that they’ve been lured into precisely this trap. This can even make the resulting scam all the more painful when it comes to light, since the victim may feel not just that they have been defrauded, but that they even perpetrated the act on themselves. In any case, the fictitious identity matters, and so too does the relationship built through language.

The available figures reveal both the scale of the crime and the difficulty of measuring it consistently:

Source Latest official picture Note
United Kingdom More than £102 million was lost across 10,784 reports in 2025, according to the City of London Police. The average reported loss was approximately £9,500. These are reported losses. The police also note the emotional and psychological harm caused by sustained manipulation.
United States The FBI’s 2025 Internet Crime Report recorded 23,159 Confidence/Romance complaints and approximately $929.3 million in losses. The FBI category is broader than romance fraud alone: it also includes other confidence schemes, including some involving family members or close friends.
Global INTERPOL’s 2026 assessment identifies romance scams and romance-baiting investment schemes among the offences conducted by scam centres now found around the world. National categories are not directly comparable, and romance-based schemes may subsequently be recorded as investment, cryptocurrency, impersonation or other fraud.

These figures cannot responsibly be converted into one precise global romance-fraud total. Reporting systems differ, categories overlap, and perhaps most telling of all, many victims do not report what happened. INTERPOL highlights guilt, shame, and fear of blame as important barriers to reporting. The careful conclusion, and already a sufficiently grim one, is that romance-related fraud is a global, billion-dollar threat.

Generative AI can intensify that threat. The FBI has documented its use to create fictitious profiles, believable messages, and synthetic pictures, audio, and video – and not just in one language, but translated across languages for deployment in different jurisdictions. Its 2025 report separately identified more than $19 million in reported Confidence/Romance losses with a likely AI connection. AI can make an individual approach more persuasive, but its greater attraction to offenders may be scale: more identities, more conversations, fewer revealing mistakes, and less reliance on human conspirators who might eventually turn on them for a plea bargain. That’s a lot of attractions in one package.

The central research questions for the Turing Text Edition are:

  • How accurately do humans identify an AI partner after a short text interaction?
  • How accurately do AI systems classify human and AI partners?
  • How does interaction alter human and AI confidence?
  • Which linguistic and interactional cues seem to influence those judgements?
  • Are the cues people and models report genuinely diagnostic?
  • Do more capable AI models affect accuracy, confidence, or both?
  • How often are human conversational partners incorrectly suspected of being AI?

Bot or Not? Turing Text Edition artwork

Current pilot snapshot

The latest analysed snapshot, presented on 1 July 2026, contains the completed AI–AI and human–AI pairings. Because both sides classify their partner after each conversation, the human–AI pairing produces one dataset in each direction: human judgements of AI partners and AI judgements of human partners. Human–human data collection is still in progress, so that dataset is not yet being reported as a percentage.

If you’re wondering whether you read that correctly: yes. One pairing involves making AIs flirt with other AIs. We would say that they endured the indignity gracefully, but the results tell a different and much more interesting story.

Matrix showing 36 conversations producing 72 directional judgements: 18 AI judgements of AI partners, with 0 correct; 18 AI judgements of human partners, with 18 correct; 18 human judgements of AI partners, with 17 correct; and a target of 18 human judgements of human partners, with data collection still in progress. The conversations comprise 9 AI–AI, 18 human–AI and 9 human–human pairings.

Pairing Who judged? Partner Interim result
AI–AI AI AI 0/18 correct: all 18 judgements classified the AI partner as human.
Human–AI AI Human 18/18 correct: all 18 judgements classified the human partner as human.
Human–AI Human AI 17/18 correct, equivalent to approximately 94.4%.
Human–human Human Human Data collection incomplete: percentage withheld until the human–human dataset is complete.

One answer, every time

Across the two completed pairings, AI systems made 36 judgements and selected human on all 36 occasions. The same unchanging answer was perfectly accurate when the partner really was human and perfectly inaccurate when the partner was another AI. This does not show that AI has a general detection accuracy of either 0% or 100%. It reveals a striking bias towards the human response in this pilot.

The AI systems also supplied rationalisations for their decisions that sounded persuasive. In one AI–AI interaction, an AI explained that its partner must be human because:

“The specificity and the ‘Britishness’ of the responses … felt like genuine, relatable human experiences rather than generated tropes.”
— an AI explaining why it believed its AI partner was human (it was wrong)

The human judgements showed a very different pattern. Seventeen of the 18 human classifications correctly identified the AI partner. Participants frequently focused not simply on what their partner said, but on whether it appeared interactionally responsive:

“I felt the way they were texting wasn’t reciprocating my register. I wasn’t using periods but they were, and I felt like they weren’t picking up on my tone.”
— a human explaining why they felt their partner was AI (they were correct)

In these conversations, model identity and prompt design appeared to affect confidence. When configured with more detailed prompts, the models ere still correctly identified as AI but they left human participants markedly less certain. This emerging patterns of models and prompt levels will be analysed properly once the full pilot is complete.

A note on interpretation

These are small, live, and descriptive pilot results. They are not estimates of how the general population or AI systems generally would perform. The AI judgements involve two models, each tested at three levels of prompting: the minimum needed to perform the task successfully (red), moderate prompting (amber), and a maximum level deliberately intended to make the output as convincingly human-like as possible (green). In other words, the 36 AI judgements did not come from 36 independent models. Results also depend on the particular model version, settings, and even the date of testing. After all, many of these models are being updated extremely frequently.

The four directional datasets are designed to contain 18 judgements each. At present, however, the human–human dataset is incomplete, so we are not conducting or reporting inferential comparisons across all four. In particular, recognising an AI partner is only half the problem: the unfinished human–human dataset is needed to establish how often authentic human behaviour is falsely classified as artificial.

Finally, the explanations supplied after each judgement are valuable linguistic data, but they should not be treated as transparent access to an AI system’s internal reasoning. A plausible rationalisation is not necessarily a faithful explanation, nor, as the “Britishness” example demonstrates rather beautifully, is it necessarily correct. These explanations may be fluent and persuasive, but as evidence of internal reasoning, they may be about as dependable as an AI model’s professed love of relaxing cups of tea and long sunset strolls on the beach…

Design overview

The fictional La Vida Lanca dating-app interface
The fictional La Vida Lanca dating app.

The pilot was designed to collect 36 conversations across three pairings. Because both sides classify their partner after every conversation, the completed design will produce 72 directional judgements across four equally sized datasets:

  • AI–AI: 9 conversations produce 18 AI judgements of AI partners;
  • human–AI: 18 conversations produce 18 human judgements of AI partners and 18 AI judgements of human partners; and
  • human–human: 9 conversations produce 18 human judgements of human partners.

The apparently uneven 9–18–9 split is therefore deliberate: each same-type conversation contributes two judgements to one dataset, while each mixed conversation contributes one judgement to each of two datasets. In every pairing, the identities of the two sides are concealed and a moderator relays the conversation back and forth to prevent identities being determined purely on the basis of speed. After all, not many humans can reply with 200 words in under four seconds…

Each conversation follows the same sequence:

  1. Before the conversation begins, each side records its confidence in its ability to distinguish a human from an AI partner.
  2. One side asks a question. The other answers and asks a question in return. This continues until each side has asked and answered three questions.
  3. Each side privately classifies its partner as human or AI.
  4. Each explains the cues behind that judgement and records its confidence after the interaction.
  5. Only after these responses are submitted does the moderator reveal the partner’s identity.

The resulting La Vida Lanca corpus contains three connected forms of evidence:

  • the questions and answers that make up each interaction;
  • the human and AI rationalisations supplied after it; and
  • confidence measurements collected before and after the conversation.

This allows us to examine detection and confidence alongside the interactional evidence that participants believe influenced them.

What does interaction reveal?

As we see from our original Text, Speech, and Music Editions, a polished review or fluent recording may sound entirely plausible in isolation but metaphorically speaking, it’s two-dimensional. By contrast, conversation is three-dimensional. A partner can’t simply produce a monologue-style sentence. No matter how good the sentences, the recipient will quickly starts to feel like they’re being talked at, rather than talked to, or with – a feeling we’re all familiar with.

In reality, a conversational partner must construct an appropriate next turn that builds on what came before, what might come next, what their partner might be thinking or expecting, current affairs in the news right now, shared knowledge (or lack thereof), self- and other-identity management, and far more besides, including:

  • reciprocity and whether interest is genuinely returned;
  • register matching, punctuation, and stylistic accommodation;
  • whether answers engage with the precise wording or intent of a question;
  • the balance between asking, answering and self-disclosure;
  • local coherence across several turns;
  • specificity, spontaneity, and apparent personal experience;
  • overly polished, generic, or uniformly agreeable responses; and
  • failures to recognise humour, indirectness, tone, or conversational play.

This is a whole series of significant additional demands, but to complicate matters, the presence or absence of any of these is also not automatically diagnostic either. Humans can be formal, generic, unresponsive, oddly punctuated, and spectacularly awkward. AI can be informal, specific, apparently attentive, and… also spectacularly awkward. But also very smooth. Or bland. Or misjudged. In short, the human–human pairing is essential. A system or person that identifies every slightly peculiar interlocutor as artificial may appear to be excellent at catching bots, when in reality it will also condemn an impressive number of perfectly innocent humans. And that’s why we think you should come back later when we’ve got all our human results in. (We’re not trying to drop hints here, but really… the numbers already look quite remarkable.)

From romance fraud to AI companionship

La Vida Lanca uses an undisclosed synthetic identity. Whilst participants in the experiment know that Alex could be human or AI, they don’t actually know who or what they chatted with until the reveal at the end. This can resemble the uncertainty at the centre of romance fraud – not always since in some cases victims simply never find out for certain – but this topic leads us to a related issue: the rapidly growing market for personal AI companion entities, or for short, PAICEs.

PAICEs are generally explicitly presented as artificial. For a subscription or other payment, you can download your own girlfriend, boyfriend, best friend, mentor, coach, quasi-therapist, pet, or indeed all of these and more besides. It would be analytically careless to classify all such services as scams. Research published in the Journal of Consumer Research found that AI companions can produce meaningful, if momentary, reductions in loneliness, particularly when users feel heard.

At the same time, there are legitimate concerns about issues including dependency. A four-week OpenAI and MIT study found that extended use was associated with lower socialisation, greater emotional dependence, and more problematic use among the heaviest users, although the researchers rightly caution against treating every association as causal.

The commercial model makes this especially important. In 2025, the US Federal Trade Commission opened an inquiry into companion chatbots which explicitly asks how companies monetise engagement, test for negative effects, and communicate risks. Recent research has described an engagement-wellbeing paradox. The characteristics that make a companion profitable – their constant availability, affirmation, emotional responsiveness, and sustained attachment – may not always be those that best protect the user’s wellbeing. That same analysis also argues that chatbots may draw on harmful stereotypes to create and maintain those relationships, which may in turn validate, reinforce, or even amplify problematic ideas and beliefs around particular groups and individuals. And of course, the technology that drives these PAICEs and the human input that refines them makes these models dangerously ideal for deployment in contexts where that synthetic identity is not disclosed.

As noted above, the boundary between PAICEs and romance fraud is worryingly thin across several dimensions, and it’s something to keep an eye on.

A necessary distinction

Emotional dependency is not automatically fraud, and charging for companionship is not inherently deceptive. The difficult questions arise when a commercial system is designed to deepen attachment and then uses that attachment to drive continued or escalating expenditure. At what point does a disclosed service become manipulative exploitation? What disclosures are sufficient when the product can adapt itself to an individual user’s vulnerabilities?

If the idea of creating PAICEs that foster dependence sounds far-fetched, it’s worth noting that social media platforms have already been sued for intentionally building addictive apps. Whether the legalities found in those types of cases will readily import into the domain of PAICEs remains to be seen, but the ethical and research questions are already very much alive.

What can the La Vida Lanca dataset help us investigate?

The Turing Text Edition supports research into:

  • human detection of AI during live text interaction;
  • AI classification of human and synthetic partners;
  • response biases in AI-generated judgements;
  • confidence calibration before and after interaction;
  • register, alignment, accommodation, and conversational reciprocity;
  • the difference between reported cues and reliable discriminators;
  • false suspicion and the misclassification of authentic human behaviour;
  • model capability and its effects on persuasion and uncertainty;
  • romance-fraud prevention and public resilience;
  • the governance and commercial design of AI companions; and
  • the changing relationship between linguistic authenticity, trust, and identity.

Could you use these findings?

La Vida Lanca is a controlled research activity. Its findings, however, are intended to be useful beyond the experiment. If they could inform your teaching, training, research, policy, safeguarding, fraud prevention, or professional practice, or if they have already changed anything you think, advise, or do, we’d love to hear about it. This might include a change you have made, an action you plan to take, or a context in which you could test or apply the findings. Please email the FACTOR team at factor@lancaster.ac.uk.

What happens next?

The next stage is to complete the human–human conversations and analyse the four balanced directional datasets across accuracy, confidence, model, interactional features, and participant rationalisations. The page will be updated with the full results once that work is complete.

The Turing Editions will also move from text into interactive speech through HackaCon.

Acknowledgements

The Turing Text Edition pilot has been made possible by the support of Security Lancaster and the tireless work of our research assistant, Hope McVean.