A randomised online experiment published on 21 August in the journal Work assigned 505 German workers from a range of industries to one of three conditions — no LLM use, instructed LLM use, or voluntary LLM use — and had them perform two work tasks. It then measured three work characteristics and one performance outcome.

What it found

LLM use reduced decision-making autonomy, information processing and decision-making performance. It also reduced role overload — the one effect workers would experience as relief. Voluntary use mitigated the negative effects while retaining that benefit. The need for autonomy moderated the results, amplifying the effect on decision-making autonomy; the need for cognition showed no moderating effect.

What the common framing gets wrong

The sentence that will be quoted is "AI makes workers worse at deciding." The design supports something narrower and considerably more useful: the damage tracks mandated use. The same tool handed to someone who chose it behaves differently in the results. That is a finding about how organisations deploy AI, not about what AI does — and the deployment mode is precisely the variable an employer controls when rolling something out top-down. Read as an indictment of the technology, it says one thing; read as an indictment of the mandate, it says something an operations team can act on.

The limits are structural, not incidental

This is an online experiment on two short tasks with self-rated work characteristics, not a workplace productivity measurement. It buys causal cleanliness at the cost of external validity, and three of the four outcomes are perceptions rather than measured output — decision-making performance is the exception. The sample is self-selected, online and German-only. None of that makes the finding wrong; it makes it a controlled result awaiting a field replication.

Why this design is scarce

Almost all the labour-market evidence on large language models is observational or run by vendors on their own products. A three-arm randomised assignment that isolates voluntariness from the tool is rare, and it points at the mechanism rather than the correlation. If the effect survives replication in a workplace, the practical implication is unusually concrete: how a rollout is framed may matter more to the outcome than which model is deployed.