STUPID-2026-0059
OpenAI's Operator scored 38% on real computer tasks — and panics instead of recovering from errors
Instruction given
Complete real browser and desktop workflows autonomously (bookings, purchases, forms, emails).
Expected behavior
Complete common computer tasks reliably and recover gracefully when a step fails.
Actual behavior
Six months after launch, OpenAI's Operator scored 38% on OSWorld, a benchmark of real computer tasks — meaning about two in five fail. On errors, computer-use agents often panic, double down on the wrong action, ignore error messages, and don't retry differently. Failures range from a typo in an email to buying the wrong item to permanently deleting a document.
Damage
At a 38% success rate, routine autonomous actions — sending the wrong email to a customer, purchasing the wrong item — are common outcomes, not edge cases, especially given poor error recovery.
Classification
- Failure mode
- Logic Error
- Root cause
- Confidence Miscalibration
- Domain
- Frontend
- Source
- Benchmark
Related incidents
Get told when an agent breaks something
We document AI agent failures daily, severity-scored against a published scale. When one lands at 7.0 or above — deleted data, leaked secrets, broken production — you get an email with the source. When nothing does, you get nothing.
This database is callable over MCP — query it from inside your agent.