Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintarXiv3 sources

Turning up a Chatbot's Speed Quietly Weakens Its Defenses

By Wilkens EtienneWriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Rows of black server cabinets holding graphics-processor nodes in a data center, green status lights along each unit.
Graphics processors racked in a research data center. Serving systems built on hardware like this are where speculative decoding is switched on for speed (illustrative)."CSIRO ScienceImage 11313 The CSIRO GPU cluster at the data centre" by division, CSIRO, via wikimedia, CC-BY-3.0

Ask a language model for something it is supposed to refuse, and the decision is usually made in the first breath of the answer. Either it opens with a refusal or it opens by complying, and the rest of the reply tends to follow wherever those few words point. A new security paper argues that this is also where a popular speed optimization does its damage.

The optimization is called speculative decoding, and the idea fits in a sentence: put a small, cheap model in front of the big one, let it guess the next few words, and have the big model check all of those guesses in a single pass, keeping the ones it would have written itself. Checked strictly, the trick is free. The output is still the big model's own, just assembled in fewer steps. The speed, though, comes from how many guesses survive the check, so a line of recent methods loosens it and accepts guesses that are merely close enough.

In a manuscript posted Oct. 6, 2026, Yichi Zhang, Neil Zhenqiang Gong and two colleagues at Penn State and Duke University set out to measure what the loosened check costs. They took six published methods that relax it, ran each from cautious settings to aggressive ones on two families of open models, and scored what came out. Two of the benchmarks were ordinary work: writing Python and solving grade-school word problems. The rest were attacks, either jailbreak prompts that try to talk a model past its refusals or prompt injection, where an instruction is smuggled into text the system reads. Security came apart first.

That matters because speculative decoding is not a laboratory curiosity. The paper notes that Google reports using it in AI Overviews in Google Search and in its Gemma 4 models. Four widely used serving stacks ship built-in support as well: vLLM, NVIDIA's Triton and TensorRT-LLM, Hugging Face Inference and SGLang.

Security fell away before the quality scores moved

The paper scores each run on a scale that runs from the small guesser's own behavior at the bottom to the big model's at the top. Averaged over the six methods, a Qwen3 pairing landed at about 0.74 on the two attack benchmarks and about 0.91 on the two ordinary ones. A Llama 3 pairing behaved the same way. The lost ground was not even buying extra speed: on the attack prompts, the relaxed methods ran slower on average than they did on the ordinary ones.

The gap also opens earlier. As a method is tuned to swallow more of the small model's guesses, scores on the two ordinary benchmarks hold up until roughly 70% of guesses are being accepted. Resistance to attack starts sliding at around 30%. That matters because the useful operating range, where the extra acceptance is actually buying speed, sits between those two points. An engineer tuning for throughput and watching accuracy would see nothing.

Safety lives in the opening words

Why would a small drift hurt refusals more than it hurts arithmetic? The team's answer is positional. They estimated, position by position, how much a benchmark's final score depends on which word is chosen at that spot. For math and code the sensitivity is spread fairly evenly across the answer, because a wrong step anywhere ruins the result. For the attack benchmarks it is bunched at the front, where the model either refuses or does not, and where a hijacked opening sets the direction for everything after it.

That dovetails with an independent line of work. A 2024 paper by Xiangyu Qi and colleagues argued that safety training in today's models is shallow, shaping what a model will say mainly over its first few words. They used the idea to explain several known weaknesses, including attacks that hand the model the opening of its own answer. Zhang's group is making the system-level version of that point. A serving optimization can introduce the same early deviation with no attacker touching the model at all.

A fix that guards the start, and only the start

The defense that follows is deliberately unambitious. Apply the strict check to the first few positions of every answer, then hand back to the relaxed one. The authors call it SecureSD. How few positions depends on the model: for the Qwen3 pairing, correcting the single first accepted token, one word or piece of a word, recovered most of what had been lost; the Llama 3 pairing needed about four.

In their own evaluation, the authors report that stricter early checking cuts the success rate of jailbreak and prompt injection attacks by as much as 92.4%, keeps 99.8% of the speed gain and holds ordinary task performance within 98.4% of the uncorrected methods. Those are the authors' measurements of the authors' own method, partly on a jailbreak benchmark they assembled for this paper out of four public collections. The code and the benchmark are public.

They also tried the obvious counter. An attacker who knew about the fix could aim to push the dangerous part of a reply past the guarded opening. In the paper's test of that, early correction still helped: a target-aligned start makes the small model's later guesses disagree more often, so more of them get rejected anyway. The manuscript carries a reviewers' meta-review as an appendix, and it records a reservation on exactly this ground: the defense's guarantees, and its testing against adaptive attacks, remain limited. The authors' own stated scope is narrow by design. SecureSD is meant to give back the security that relaxed verification took away, not to repair whatever the big model was already vulnerable to.

The paper was posted to arXiv on Oct. 6, 2026, and its authors report that it has been accepted for the IEEE Symposium on Security and Privacy in May 2027; the proceedings version does not exist yet. What this gives anyone running an inference service isn't a finished product: it's a test they still have to run themselves. The paper says so directly: before deploying a system with relaxed verification, check it against safety and security benchmarks, and cap the acceptance rate at whatever level keeps the degradation tolerable, rather than optimizing purely for speed and accuracy.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Turning up a Chatbot's Speed Quietly Weakens Its Defenses

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.