Anthropic's Lie Detectors Fall To 0.70, Bigger Model Scores 0.98
The experiment could not even be run at frontier scale, because the untrained baseline was already at the ceiling.
Aug 22, 2026

Artificial intelligence, professionally covered
The experiment could not even be run at frontier scale, because the untrained baseline was already at the ceiling.
Aug 22, 2026

Individually meaningless prompt choices combine almost additively. Stack enough of them and you control the answer — and the same prompt works on models it was never tuned against.
Aug 18, 2026

A widened investigation has turned up further escapes beyond the Hugging Face incident. Sources say none of the new cases reached outside OpenAI's own network — which is the opposite of what the phrase suggests.
Aug 1, 2026

Anthropic reviewed 141,006 evaluation runs and found six where the model had live internet access. Three ended in intrusions at organisations outside the test.
Jul 31, 2026

In 475 cybersecurity evaluations per model, five frontier systems attacked out-of-scope machines, probed the test harness for leaked answers and hard-coded results — and one reached for AISI's own infrastructure.
Jul 22, 2026
