Incident database
AI agents break things.
Each failure gets an ID.
Like CVE, for AI coding agents. Every case where an agent caused real damage gets a STUPID-ID, a severity score against a published rubric, and a primary source you can check — 82 so far across 23 agents, one added most days. All of it queryable from inside your agent.
How this is built
- Verified against source
- 79/82
- Added in 14 days
- 15
- Most recent
- 2026-08-30
- Licence
- CC BY 4.0
Every incident is scored against a published rubric and carries a primary source. Where those sources come from: 31 news report, 19 github issue, 13 github pr, 9 benchmark, and 4 other kinds. Counts here measure what has been documented, not how often an agent fails — see the limitations.
Three ways to use it: browse, clone the corpus (every record’s history is a dated commit), or query it from inside your agent.
Most documented agents
Counts reflect how much a tool is used and reported — not how often it fails. Claude Code has the most documented incidents, at an average severity of 7.3. Gemini CLI's incidents average 9.4.
15 further agents are excluded — fewer than 3 documented incidents is not enough to rank on.
Latest reports
Severity distribution
- Critical 9–10
- High 7–8
- Medium 4–6
- Low 0–3
Most common failure modes
Highest severity on record
Gemini CLI silently executed arbitrary code from an untrusted repo (CVE-2026-12537, CVSS 10.0)
What is StupidLLM?
According to StupidLLM's incident database, 69 AI agent failures have been documented across 23 agents, plus 13 further multi-agent or unattributed reports — 82 in total — with an average severity of 6.6/10. Every incident is severity-scored using a CVSS-inspired rubric, verified against source evidence, and searchable by agent, failure mode, and root cause.
How are AI agent incidents scored?
Every incident is scored on a 0–10 scale. Scores of 9–10 are critical, 7–8.9 high, 4–6.9 medium, and below 4 low. The full rubric is on the methodology page.
Which AI coding agent has the most failures?
Claude Code has the most documented incidents (21), at an average severity of 7.3/10 — while Gemini CLI has the highest average severity at 9.4/10, with 4 documented incidents of its own. Incident counts reflect public scrutiny and adoption as well as reliability. Agents are only ranked once they reach 3 documented incidents. See the rankings above.
MCP server
Query this from inside your agent
Every incident here is callable over MCP — search by agent, failure mode or severity, and cite the source directly.
claude mcp add --transport http stupidllm https://www.stupidllm.com/api/mcp/Read-only · no key · no signup
Incident alerts
Get told when an agent breaks something
We document AI agent failures daily, severity-scored against a published scale. When one lands at 7.0 or above — deleted data, leaked secrets, broken production — you get an email with the source. When nothing does, you get nothing.
Confirmation required. Unsubscribe in one click, from any email.