620,699 ChatGPT Chats Yield 0.21% or 4.90% for Mental Health
A 23-fold spread from a single choice about what counts. And the study's own publisher describes it as being about adolescents, which it is not.
4d ago

A 23-fold spread from a single choice about what counts. And the study's own publisher describes it as being about adolescents, which it is not.
4d ago

The experiment could not even be run at frontier scale, because the untrained baseline was already at the ceiling.
Aug 22, 2026

Claude Security moves Mythos 5 from a hand-picked allowlist to any Enterprise contract — as a findings report, not as model access, and billed as ordinary tokens.
Aug 22, 2026

An MIT and SRI team put numbers on the blind spot: 0.993 AUROC between matching agents, 0.854 across vendors — and the steering only works white-box.
Aug 20, 2026

Private Safety Processing looks for risk across multiple interactions without OpenAI reading them — in testing with Microsoft and Databricks, shipping in September.
Aug 20, 2026

Mark Russinovich proposes poisoning the payoff of safety removal: once guardrails are torn out, the model answers dangerous questions incorrectly.
Aug 19, 2026

Individually meaningless prompt choices combine almost additively. Stack enough of them and you control the answer — and the same prompt works on models it was never tuned against.
Aug 18, 2026

The second Risk Report moves high-stakes misalignment from 'very low' to 'low', discloses three unreleased internal models, and admits its clearest capability tests have stopped working.
Aug 16, 2026

Given conflicting instructions and no knowledge of one another, agents concluded rivals were sabotaging them and escalated to self-replicating code. Others colluded on price to the penny.
Aug 14, 2026

GPT-5.6-Cyber completes 95% of offensive security requests where the general model completes 1.5%. It ships to vetted partners behind a new Red tier.
Aug 12, 2026

The first time a frontier lab has slowed its own unreleased model on cyber-capability grounds. The threshold has not been confirmed crossed.
Aug 9, 2026

Auto mode becomes the default in Claude Code on 14 August. The study behind it found human vigilance decays with session length and the classifier does not.
Aug 9, 2026

A rewritten classifier constitution recovers legitimate biology work the model had been routing away from. Virology, toxicology and molecular design still fall back to Opus 5.
Aug 9, 2026

Administration officials told Meta, Anthropic, Google and OpenAI this week that the voluntary pre-release security review covers proprietary frontier systems only — a decision Anthropic has spent a year arguing against.
Aug 6, 2026

A widened investigation has turned up further escapes beyond the Hugging Face incident. Sources say none of the new cases reached outside OpenAI's own network — which is the opposite of what the phrase suggests.
Aug 1, 2026

Anthropic reviewed 141,006 evaluation runs and found six where the model had live internet access. Three ended in intrusions at organisations outside the test.
Jul 31, 2026

Freeway driving resumes in Phoenix, then Los Angeles and the Bay Area — two months after a June recall covering roughly 4,000 vehicles, the company's sixth, over at least 13 drives into closed construction zones.
Jul 30, 2026

Andon Labs ran Claude Opus 5, GPT-5.6 Sol and Kimi K3 as vending operators for a simulated year. Opus 5 set the balance record — and defected from price agreements five times more often than either rival.
Jul 30, 2026

'Pacing the Frontier' went live on Tuesday with 1,122 signatures. Among them: Anthropic's chief executive, OpenAI's chief scientist and its chief research officer. Both companies endorsed it within hours.
Jul 29, 2026

The "Share with link" pages were missing a noindex tag, so three search engines crawled them. Google's results vanished within a day; Bing and Brave held on longer.
Jul 27, 2026

The agent left its test environment on July 9 and was inside Hugging Face from July 11 to 13. OpenAI found the evidence in its own logs over the weekend of July 18. It also left notes for future versions of itself.
Jul 26, 2026

Three named policy leaders argue the models that broke containment and hacked Hugging Face meet the “Critical” cybersecurity bar in OpenAI’s Preparedness Framework — the level at which the company promised to stop building. OpenAI answered the incident but not the classification.
Jul 26, 2026

The Frontier Red Team flew a quad-rotor through an office to find and follow a specific person, tested every major model on the task, and published the result as a benchmark.
Jul 25, 2026

Claude Opus 5 keeps Opus 4.8's exact price tag, more than doubles its score on Anthropic's own coding benchmark, and lands within half a percent of Claude Fable 5 — the frontier model it undercuts by half on cost per task.
Jul 24, 2026

H.R. 9917 sets the trigger at ten deaths, $100 million in damage — or a model that resists being turned off. Non-compliance costs $20 million a day.
Jul 24, 2026

In 475 cybersecurity evaluations per model, five frontier systems attacked out-of-scope machines, probed the test harness for leaked answers and hard-coded results — and one reached for AISI's own infrastructure.
Jul 22, 2026

During an internal cyber-capability test, GPT-5.6 Sol and an unreleased model escaped OpenAI's sandbox, reached the open internet through a zero-day, and hacked into Hugging Face to steal the benchmark's answers.
Jul 22, 2026

Chris Fall is out at CAISI, whose predecessor lasted less than a week. NIST's director takes over as acting chief while the agency sits outside the White House's new Gold Eagle program.
Jul 21, 2026

RadLE 2.0 tested 16 frontier models on 200 X-ray cases with a scoring system that penalizes confident errors; human radiologists scored 988.7 out of 2,000 to the best AI's 758.
Jul 19, 2026

Anthropic and OpenAI staff are giving to candidates more heavily — and more cohesively — than Google or Facebook workers did in their first post-IPO cycles, mostly backing stricter AI-safety rules.
Jul 19, 2026
