A paper submitted to arXiv at 17:20 UTC on 17 August 2026 by Enric Boix-Adsera and Benedict Tessler describes a phenomenon the authors call model hypnosis: individually weak and seemingly irrelevant cues in a prompt — a typo, a paraphrase, an unrelated aside — can be systematically combined to strongly control what a model says.

The core claim

In the authors' words, the effect "occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models." Transfer is the part that matters most: a prompt tuned to steer one model keeps its directional effect on a different model, built by a different organisation, that the authors never optimised against. The control does not come from a single powerful instruction but from the accumulation of many inconspicuous textual choices, each of which would look like noise on its own.

What the common framing gets wrong

This will be written up as a new jailbreak. It is not one, and the distinction matters. A jailbreak defeats a refusal — it makes a model produce content it was trained to withhold. Model hypnosis is steering: it moves which answer a model gives on open-ended and self-report questions. Nothing here is described as bypassing a safety boundary. The second likely error is treating the effect as uniformly overwhelming. Large movement in a model's internal preference only flips the visible answer when the model was near-undecided to begin with; on questions where it holds a firm position, the same cues shift the margin without changing the output.

Why interpretability researchers should care most

The authors name interpretability as the field with the biggest problem here, and the reasoning is uncomfortable. A great deal of published work asks models about themselves — do you have a goal, are you aware, why did you answer that way — and treats the reply as a measurement of the model. If typo-level choices in how the question is phrased can be stacked to control the answer, then a single-template evaluation is measuring the template as much as the model. Any result built on one phrasing of one question is unfalsifiable until it has been tested against a distribution of paraphrases.

The practical consequence

Benchmark methodology has generally assumed that incidental prompt wording is noise that averages out. This paper argues it is signal that accumulates. That does not invalidate existing evaluations, but it does mean that reported differences between models can be smaller than the variation a determined prompt-writer can manufacture within a single model.