Anthropic's Frontier Red Team, working with Andon Labs, handed 15 AI models the controls of a real quad-rotor drone and asked them to fly through an office, find a specific person and follow them. The results and the task suite were published on Friday as Drone-Bench.
How it was scored
The task was decomposed into five sub-tasks, including Reconstruct — turning office video into a 3D model — and Localize, matching the drone's view to a 2D obstacle map. The comparison point is not an unaided human: the baseline was set by human-AI teams using coding agents. Any "models match humans" framing misdescribes what was measured.
Where the wall is
Detection and following reached near 100% of baseline. Reconstruction reached only about 47% by mid-2026, and it is the bottleneck holding the whole task back. The strongest model beat the baseline on every sub-task except reconstruction — but even it cleared the baseline on average for only three of five.
Stated limits
Anthropic published the constraints rather than burying them: slow drone speeds, a single office floorplan, a limited number of people, no outdoor testing and no crowds. Models tested spanned 15 releases from three developers across roughly two years of progress.
Which models
The 15 span roughly two years of frontier progress across three developers, running from GPT-4o and o1 through Gemini 2.5 Pro and successive Claude Opus releases to the current generation. Reading them as a time series is the point: detection and following are close to saturated, while reconstruction has barely moved.
The governance line
The team named the risk directly: once reliability thresholds are passed, there will be real pressure to treat human oversight as a cost rather than a safeguard.
