A British lab at King's College London (Stein et al., 2026, preprint available here, not yet peer reviewed) tested whether administering inert medical air, labeled as nitrous oxide (N2O), to healthy non-clinical volunteers would produce self-reported psychedelic-like effects (placebo) or side effects (nocebo). N2O itself can produce dissociation and some psychedelic-like effects on consciousness, which is why the researchers chose it as their cover story. The design was a pre-registered, counterbalanced, within-subject crossover: 78 participants, drawn from an initial screened pool of 167, each completed two 20-minute inhalation sessions, one correctly presented as medical air and one falsely presented as Entonox (50% N2O/50% O2).
Relative to the control condition, the placebo condition produced significantly higher scores on every phenomenological measure used: the 5D-ASC, the CADSS and its Amnesia, Depersonalization, and Derealization subscales, the EDI, and the GASE. Depersonalization showed the largest effect (Hedges' g = 0.60); most other measures fell in the moderate range (g approximately 0.30 to 0.54); dread of ego dissolution was the weakest (g = 0.19). Side-effect scores on the GASE rose by a comparable moderate margin (g = 0.39), essentially the same order of magnitude as the placebo effects. None of this moved the only one objective, non-self-report measure in the study: performance on a temporal reproduction task, a timing task previously shown to be sensitive to actual psychedelic drugs, did not differ between conditions (g = 0.04, not significant). That finding is worth sitting with. The placebo effect here lived entirely in what people said they felt in self-reported measures, not in a behavioral task readout.
Only 45% of participants in the placebo condition (n = 35) believed they had actually received N2O. The other 55% correctly guessed medical air, mostly because they didn't feel the suggested effects, not because they detected a flaw in the cover story. Unblinded participants still showed placebo and nocebo effects, numerically smaller than in the blinded group but not statistically distinguishable from it after correction for multiple comparisons; the authors are explicit that this comparison was underpowered. The important note is that unblinding attenuates the effect, but does not eliminate it, and this study cannot tell us the true size of that attenuation.
Four of twelve outcome measures (EDI, CADSS total, and the CADSS Depersonalization and Derealization subscales) showed a significant Condition x Order interaction. Participants who received the fake-N2O condition first showed large placebo effects (g range 0.60 to 0.77); participants who received the real-air condition first showed effects that were small and non-significant. Once someone has felt what plain air actually feels like, the false label works measurably less well on the next exposure.
Trait predictors were real but modest, and specific rather than general. Suggestibility (REVS, via the Brief Suggestibility Scale) predicted placebo-induced 5D-ASC scores and little else. Trait dissociative absorption was the more consistent predictor across several outcomes and the only trait variable correlated with the nocebo (GASE) effect, though it dropped out of the final regression model for that outcome. Compliance (Gudjonsson Compliance Scale) predicted dissociative responses specifically. None of the regularized regression models explained much variance (in-sample R² of .11 to .15; cross-validated R² near zero or slightly negative for two of the four outcome models). These predictors are real but weak. They should not be oversold as a screening tool for identifying who will placebo-respond in a psychedelic trial.
Two things, and a third that follows from them. First, what is measured here is self-reported placebo response in a single acute exposure among healthy, non-clinical volunteers who believed they were in a genuine N2O study. This is not a psychedelic-trial population, not an oral-dosing paradigm (mode of administration itself moderates placebo magnitude, by the authors' own account), and not paired with therapeutic framing, psychotherapy, or repeated dosing. The compliance finding, that compliance independently predicted dissociative scores, raises the same concern from the authors' own data: some portion of what is captured may be participants narrating an ambiguous sensory experience in the direction they believe is expected, rather than a genuinely felt altered state. Second, none of this speaks to therapeutic effect. A chronic, outcome-linked design would be needed to know how much this contextual variable interferes with, or contributes to, actual clinical response, as distinct from the acute subjective report captured here. Third, and this is the point that should not go unnoticed: the objective measure showed nothing. If we cite this paper anywhere external, that null result belongs in the same sentence as the effect sizes, not in a footnote.
Read against FDA's July 2026 guidance, this study is best treated as experimental support for concerns FDA has already flagged. FDA's guidance (Section III.E.1, Trial Design) names functional unblinding and expectation bias as the central interpretive problem in psychedelic trials and recommends active or low-dose comparators, central raters blinded to treatment allocation and visit number, blinding questionnaires for both participants and investigators or raters, and an expectancy evaluation questionnaire administered pre-randomization and at end of treatment. Stein et al. ran a stripped-down, single-session version of essentially that measurement package (a Guess of Treatment Questionnaire plus a compliance item) and found it detects a real, moderate, reproducible signal with no drug involved at all. That is useful confirmation. It is not confirmation that effects observed in actual psychedelic trials are mostly placebo though.
Implication 1: Prespecify order and repeat-exposure as covariates
The significant Condition x Order interaction has a direct, practical implication for any multi-dose or crossover-adjacent protocol. Familiarity with a drug context measurably changes the size of the contextual response on a second exposure. FDA's guidance already recommends stratified randomization by prior psychedelic exposure; this data point extends that logic to within-trial exposure. Our protocols should prespecify session number and order as covariates, which is consistent with FDA's own language about prespecified adjustment for site, therapist team, and session number.
Implication 2: Pair self-report with an objective endpoint
This is arguably the single most actionable point in the paper for us. Context moved every self-report measure but did not move a timing task. FDA's guidance asks sponsors to provide a quantitative and qualitative assessment of drug effects, including orientation to time and place and impact on driving, and to consider a formal driving study. A program relying solely on self-report primary endpoints is more exposed to exactly the bias FDA is worried about than a program that pairs self-report with at least one objective or performance-based measure.
Implication 3: What this study can't tell us — durability, dose-response, drug vs. psychotherapy
Durability, dose-response, and the drug-versus-psychotherapy contribution question, all which FDA explicitly asks sponsors to address through 12-week double-blind primary endpoints, 12-month blinded follow-up, and factorial designs to separate drug and psychotherapy contributions.
Implication 4: Build contextual attribution into the safety CRF
FDA states that expected psychoactive effects, euphoria, hallucinations, perceptual distortions, and cognitive alterations, must be recorded as adverse events for psychedelic drugs even when the participant does not experience them as adverse, and that onset, duration, severity, and resolution should be documented throughout the session. The nocebo finding here, a GASE effect of g = 0.39, essentially the same order of magnitude as most of the placebo effects, is direct evidence that side-effect reporting in a psychedelic-context trial will carry a real contextual component that has nothing to do with pharmacology. Practically, our CRF and safety-management plan need to separate, at minimum: (1) observed or participant-reported phenomenology, (2) which scale captured it, (3) timing relative to dosing and session, (4) action taken and resolution, and (5) investigator causality assessment. Given the order-effect data, we would add a sixth field (6): whether the report came from a first or repeat exposure within the protocol, since exposure order measurably changed reporting magnitude in this study.
For efficacy, a psychedelic development program should seek defensible control strategy, blinded central raters primary outcome assessment, assess durability, dose-response, and employ statistical analyses showing that treatment effects are not explained by measured expectancy or functional unblinding.
For safety, the placebo/pseudoplacebo-controlled arm remains valuable precisely because, as FDA notes, it better contextualizes safety findings. The implications coming from this study is that both drug-related and context-related factors contribute to efficacy and safety, and that we need to measure them separately enough, to strive for the most precise benefit-risk interpretation and noninflated detailed safety labeling later.
Taken together, this points to three protocol decisions companies should already be planning, and a broader point about how safety signals get read: