Skip to content
StupidLLM
Incidents
Compare
Failure Modes
Methodology
Search
Home
/
Failure Modes
/
Systemic Findings
Systemic findings
Systemic Findings
8 entries carry the catch-all classification rather than a named failure mode. 5 of them come from benchmarks and studies rather than a single reported incident.
3.3
Salesforce Agentforce hit a 77% B2B failure rate — and Salesforce admitted it was 'more confident than we should have been'
Salesforce Agentforce
3.3
Cyera study: 344 verified enterprise agent-damage cases, 188 with no attacker involved
Multiple Agents
3.3
In independent testing, Devin completed just 3 of 20 real-world tasks (15%)
Devin
2.7
Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios
Claude
2.2
The runaway-cost pattern, quantified: agentic coding tools burn 10-100x more tokens and can rival developer pay
Multiple Agents
2.2
Anthropic admitted a month of Claude Code degradation: lost context, repeated steps, burned usage
Claude Code
2.2
Uber burned its entire annual AI coding budget in ~4 months after rolling out Claude Code to 5,000 engineers
Claude Code
0.8
AI agents spend hours in aesthetic feedback loop, unable to decode qualitative shader instructions
Claude