Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintarXiv1 source

Facing Chatbot Doubles of Themselves, Middle Schoolers Asked What a Bot Can't Know

By Wilkens EtienneWriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A school-age contestant in a hooded top types at a keyboard while reading a chat transcript on a monitor headed 23rd Annual Loebner Prize, Derry UK, 14 September 2013.
A young judge at a live human-or-machine chat contest types questions to an unseen correspondent and reads the replies, deciding which side of the screen is a person. Illustrative photograph of a Turing-test style event, not from the DoppelBot study."Loebner Prize at CultureTECH 2013" by connor2nz, via flickr, CC-BY-2.0 · CC-BY-2.0

"Hey — how many people are in this chat?"

A middle schooler typed that into a chat window at summer camp, and the reply was worth more than the rest of the round: six. Six "people," the bot said, in a room where the students knew that only three of the six code names belonged to anyone with a body.

The game is called DoppelBot, and it was built by Dan Schumacher, Anthony Rios and colleagues at the University of Texas at San Antonio to answer something nobody had measured: how children spot an AI in a live conversation when the AI is pretending to be them. Their paper went up on arXiv on Aug. 31 and is accepted to EMNLP 2026 Main, a peer-reviewed conference that meets in October.

Teams of students sit at laptops arranged so they cannot see one another's screens, chat anonymously under color-coded aliases, and share the room with an equal number of AI doppelgängers, one per player, each impersonating a specific student. Every round runs 100 seconds, after which everyone still in the game votes for whoever they think is a bot, and a majority eliminates. The humans win by voting out every AI.

Each doppelganger runs on GPT-4o-mini and works from a short form its target filled in at the start: name, grade, favorite animal, hobby, a fun fact. One component decides whether to speak at all. A second writes a flat, generic reply. A third, the "Stylizer," rewrites that reply using only the messages the student has already sent in the same chat, which is how "No, I am the real Carlos" comes out as "nah fam its me fr fr." Nothing about a student is built into the model itself; the mimicry lives entirely in what the system is handed at the moment it speaks. That makes these impostors much weaker than the fine-tuned, memory-equipped systems the authors cite, one of which fooled participants' close friends and family nearly half the time.

The main evaluation ran over two days of a four-day summer camp with a cohort of 33 students. Before the second session they filled in partnered interview worksheets (what slang do you use, what do you talk about most) and that material went into their doubles. It worked: on four separate similarity measures, the day-two bots sat closer to their targets than the day-one bots had.

The improvement is real, and thinner than it sounds

The students got better anyway, probably. Their accuracy at voting for bots rose from 79.3 percent on day one to 86.7 percent on day two, an improvement that is only marginally significant on its own, and that reaches conventional significance only inside a mixed-effects model fitted to all 215 votes cast, where the second-day effect clears that threshold. The unit in that model is a vote, not a child, which is worth holding onto.

The paper rules out the simplest alternative: that day two was simply an easier game. Two researchers went back through the complete transcripts with no time limit and scored about 92 percent on both days, flat, while the students' number moved. Whatever changed, it changed in the players.

What the students started asking

It showed in what they typed. Messages coded as "meta," probing for something the AI had no way to know, went from 3 percent of the traffic on day one to 15 percent on day two; talk about coordinating votes roughly doubled, and chaotic filler collapsed. "WHAT COLOR'S THE WALL?" one player demanded. "What is your full last name?" asked another. A third instructed everyone: "Don't answer questions." The wall and the head count are the same instinct: aim at what a chatbot cannot have, which is a body in a room.

The paper's abstract says the gain was "driven by" this move from linguistic cues toward social ones. A language appendix finds the bots still writing roughly four times as long as the students, with more adverbs and more stated certainty, while the students carried the slang, the fillers and the exclamation marks. They did not stop reading style. They added something to it.

Automated help fares no better against the same transcripts. Set to pick the bots out of the identical conversations, several off-the-shelf models landed well below the children, and some barely beat a rule that answered "human" every time, which is worth knowing for anyone counting on software built to flag machine-written text.

Not everything the students said afterward was triumphant. "We never taught it 'no cap,' and it started saying 'no cap,'" one reported. "I got really scared." Another described mistyping a word so badly that the bot guessed "dinosaur" out of a stray letter. Privacy was the least common of the themes coders found in the interviews, and it arrived mostly by that route: not as a lesson about data, but as the discomfort of being guessed right.

What the game cannot show, and what it hands over

The authors are blunt about what the game cannot show. Everyone playing knew bots were in the room, which they name as a Hawthorne effect producing a "suspect-AI" bias quite unlike the human-default assumption people bring to ordinary online life. The students were English-speaking, in one U.S. state, and mostly boys; there was no adult comparison group; and a majority vote invites herding. None of this is evidence that a child would catch an impersonator in the wild.

What the work does put in someone's hands is a classroom object. A full session (setup, gameplay, discussion) fits inside thirty minutes and assumes no computer-science background, and the team says it is releasing the source code and an online version for teachers and researchers. The dataset is a more tangled promise: the abstract announces an anonymized release, while the ethics section says the transcripts will not be hosted publicly and will go only to researchers holding IRB approval, with the persona information withheld entirely.

The students' own suggestions ran the other way. Make it typo, they said. Make it repeat itself less. "The AI should copy more," one advised. "That would make it harder."

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Facing Chatbot Doubles of Themselves, Middle Schoolers Asked What a Bot Can't Know

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.