Skip to content
See the World Through Science
Source: PreprintarXiv2 sources

What Mattered Wasn't How Often People Asked for Help, but How Long They Thought First

By Wilkens EtienneWriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A hand-drawn deduction grid for Einstein's logic puzzle on paper, its rows labelled with nationalities, pets, cigarette brands and drinks and its cells filled in with pencil ticks and crosses.
A logic puzzle worked out by hand. What predicted later performance in the experiment was not how often people asked for help, but how long they spent on a problem before asking."'Einstein's' Logic Puzzle" by Myrmi, via flickr, CC-BY-SA-2.0 · CC-BY-SA-2.0

Put a price on asking for help and people ask for less of it. That part of a new experiment out of the University of California, Irvine behaved exactly as you would expect. The interesting part came twenty minutes later, when the help was switched off and everyone went back to solving puzzles alone.

Shang Wu, Padhraic Smyth and colleagues recruited adults through Prolific and sat them down with a logic puzzle. Six objects, five constraints, one correct ordering; credit only if all six ended up in the right place. Each person did a short round with no help at all, then twenty minutes in which help could be requested, then a final round with no help again. The paper was posted to arXiv on August 24 and is due at ACM's HCOMP conference in Alexandria, Virginia, at the end of September. It has not been presented yet, and it has not been through journal peer review.

What "help" meant here matters more than anything else in the study, and it is not what most readers will picture. There was no chatbot and no language model. The assistant was simulated and hard-coded to be right every time. Each request revealed the position of one randomly chosen object: no reasoning, no explanation, just one square of the answer. Participants were never told it was perfect. That was deliberate: the only thing varying between groups was the price of a request, 0.1 points for one group and 0.18 for another, taken out of real bonus money, while a third group had no assistant at all.

The pricing worked in the expected direction. The cheap-help group made 6.67 requests on average during the middle round against 3.33 in the expensive-help group, about twice as many. The paper's abstract states this flatly. The result behind it is a one-sided test at p < 0.10, in groups of roughly forty, which is a trend pointing the right way rather than a settled difference.

Then comes the number the abstract leaves out. In the final round, with the assistant gone, accuracy did not differ between the three groups at all: an analysis of variance returns p = 0.91, about as close to nothing as a comparison gets. That is the study's cleanest test (the price was assigned at random, so keener participants could not have sorted themselves into one group), and it found no effect on how many objects anyone placed correctly.

Where the groups did differ was speed. In that final round the cheap-help group took about 112 seconds a problem, against 87 seconds for the expensive-help group and 92 for the group that never had an assistant. The study's summary measure, which it calls reward rate, is accuracy divided by time, so a slower group with identical accuracy scores lower on it. Gains on that measure ran 1.95 for the no-help group, 1.66 for expensive help and 1.32 for cheap help, a one-sided trend at p < 0.10 driven by the clock rather than by correctness. The cheap-help group was not less accurate. It was slower.

The paper also reports that the people who asked for help did worse afterwards than the people who did not. That comparison is descriptive, the paper's own word, and it sorts people by a choice they made rather than by the group they were assigned to.

The most useful thing in the paper is another null. The authors fit a Bayesian model that treats accuracy and response time as noisy readings of an underlying ability, then ask what predicts a change in that ability between the first round and the last. Request frequency predicts nothing at all. The coefficient lands at 0.0004, its credible interval runs from -0.122 to 0.124, and the probability that it is positive at all is a coin flip. The rival candidate (solo share, the fraction of each problem a person worked through before reaching for the assistant) comes out positive, with a credible interval that only just clears zero. It is also a behavior people chose rather than one assigned, which the authors say plainly. On their own arithmetic, one extra minute of independent work out of twenty corresponds to roughly 1.5% higher predicted accuracy at the end.

So the line the study draws is not between people who used the assistant and people who did not. It is between help that arrives on top of your own thinking and help that arrives instead of it.

One more finding sets the scale for all of the above: everyone got much better. Participants who never asked for help went from about two correctly placed objects per minute at the start to nearly four at the end, a 90.2% improvement and the largest effect in the paper. The puzzles carried learnable structure (recurring visual markers tied to particular positions), and people picked it up.

One result carries a sharper practical edge. When the researchers used performance during the assisted round to predict final unassisted performance, the prediction missed in opposite directions for the two groups: it undershot the people who never asked by 0.15 reward-rate units and overshot the assistant users by 0.22. How well someone does with a tool in hand systematically flatters how well they will do without it. No significance test is reported for it, so this is a described pattern rather than a tested effect; it is the one that matters wherever assisted work is used to judge unassisted competence.

The authors keep the scope small and it should stay small: 124 people, one artificial puzzle, a single sitting of under an hour. Throughout, they call it "short-term, task-specific skill development." The same team ran an earlier experiment on the same puzzle, and other groups have reported comparable patterns in mathematics, programming and essay writing. None of it makes this a verdict on what AI assistants do to people over months, in classrooms or at work. It is an hour of puzzles, priced.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

What Mattered Wasn't How Often People Asked for Help, but How Long They Thought First

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.