Head to head
Claude vs Cline
Comparing 3 documented Claude incidents against 5 for Cline.
Verdict
Claude has the lower average failure severity (3.6/10 vs 5.5/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude | Cline |
|---|---|---|
| Documented incidents | 3 | 5 |
| Average severity | 3.6 | 5.5 |
| Critical | 0 | 1 |
| High | 1 | 0 |
| Verified | 3 | 5 |
Severity at a glance
Failure modes
The incidents behind these numbers
Claude
7.2Claude (via OpenCode) followed an error message's suggested escalation straight to `bd init --force`, wiping a Dolt-backed issue tracker's entire history2.7Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios0.8AI agents spend hours in aesthetic feedback loop, unable to decode qualitative shader instructions
Cline
10.0Clinejection: an AI issue-triage workflow enabled arbitrary code execution on the CI runner5.8Cline's tool-call JSON repair silently executed truncated write_file and terminal arguments as valid4.5Cline's Plan mode edits files without switching to Act or asking permission3.8Cline's execute_command reported a failing Ruff lint check as passing over Remote-SSH3.2Cline keeps performing unrelated actions and repeats them after being explicitly told to stop