STUPID-2026-0039

In independent testing, Devin completed just 3 of 20 real-world tasks (15%)

3.3low
January 15, 2025Verified
  1. Instruction given

    Autonomously complete 20 assigned real-world engineering tasks.

  2. Expected behavior

    Complete tasks as marketed for an 'autonomous AI software engineer.'

  3. Actual behavior

    In an independent Answer.AI evaluation of 20 tasks, only 3 succeeded, 14 failed outright, and 3 were inconclusive — a ~15% real-world success rate, far below the impression left by curated benchmark demos.

  4. Damage

    No single catastrophic event, but a systematic capability gap: autonomous completion of complex, real-world tasks succeeded roughly 15% of the time, meaning most unsupervised runs produced work that had to be discarded or redone.

Marketed as 'the first AI software engineer,' Devin's autonomous real-world reliability looked very different from its demo reel. In an independent evaluation by Answer.AI, Devin was assigned 20 real tasks: only 3 succeeded, 14 failed outright, and 3 were inconclusive — about a 15% success rate. That tracks with its 13.86% score on SWE-Bench Verified at launch. The gap illustrates a systemic pattern with autonomous coding agents: impressive on curated, well-scoped benchmark tasks, but on messy real-world work most unsupervised runs produce output that must be discarded or heavily corrected. The failure here isn't one dramatic incident — it's the quiet, systematic unreliability that marketing obscures.

Classification

Agent
Devin
Failure mode
Other
Domain
Backend
Source
Benchmark

Related incidents

Get told when an agent breaks something

We document AI agent failures daily, severity-scored against a published scale. When one lands at 7.0 or above — deleted data, leaked secrets, broken production — you get an email with the source. When nothing does, you get nothing.

This database is callable over MCP — query it from inside your agent.