Incident database

AI agents break things.
Each failure gets an ID.

Like CVE, for AI coding agents. Every case where an agent caused real damage gets a STUPID-ID, a severity score against a published rubric, and a primary source you can check — 82 so far across 23 agents, one added most days. All of it queryable from inside your agent.

How this is built

Verified against source
79/82
Added in 14 days
15
Most recent
2026-08-30
Licence
CC BY 4.0

Every incident is scored against a published rubric and carries a primary source. Where those sources come from: 31 news report, 19 github issue, 13 github pr, 9 benchmark, and 4 other kinds. Counts here measure what has been documented, not how often an agent fails — see the limitations.

Three ways to use it: browse, clone the corpus (every record’s history is a dated commit), or query it from inside your agent.

82
documented failures
6.6
avg severity /10
23
vendors tracked
8
ranked agents

Most documented agents

Counts reflect how much a tool is used and reported — not how often it fails. Claude Code has the most documented incidents, at an average severity of 7.3. Gemini CLI's incidents average 9.4.

15 further agents are excluded — fewer than 3 documented incidents is not enough to rank on.

Latest reports

Severity distribution

Distribution of 82 incidents by severity score (0–10)
  • Critical 9–10
  • High 7–8
  • Medium 4–6
  • Low 0–3

Most common failure modes

Distribution of failure modes across all documented incidents.

Highest severity on record

STUPID-2026-0027
10.0critical

Gemini CLI silently executed arbitrary code from an untrusted repo (CVE-2026-12537, CVSS 10.0)

What is StupidLLM?

According to StupidLLM's incident database, 69 AI agent failures have been documented across 23 agents, plus 13 further multi-agent or unattributed reports — 82 in total — with an average severity of 6.6/10. Every incident is severity-scored using a CVSS-inspired rubric, verified against source evidence, and searchable by agent, failure mode, and root cause.

How are AI agent incidents scored?

Every incident is scored on a 0–10 scale. Scores of 9–10 are critical, 7–8.9 high, 4–6.9 medium, and below 4 low. The full rubric is on the methodology page.

Which AI coding agent has the most failures?

Claude Code has the most documented incidents (21), at an average severity of 7.3/10 — while Gemini CLI has the highest average severity at 9.4/10, with 4 documented incidents of its own. Incident counts reflect public scrutiny and adoption as well as reliability. Agents are only ranked once they reach 3 documented incidents. See the rankings above.

MCP server

Query this from inside your agent

Every incident here is callable over MCP — search by agent, failure mode or severity, and cite the source directly.

claude mcp add --transport http stupidllm https://www.stupidllm.com/api/mcp/

Read-only · no key · no signup

What you get back, and how to cite it

Incident alerts

Get told when an agent breaks something

We document AI agent failures daily, severity-scored against a published scale. When one lands at 7.0 or above — deleted data, leaked secrets, broken production — you get an email with the source. When nothing does, you get nothing.

Confirmation required. Unsubscribe in one click, from any email.