Agent report

Gpt 5.6 Sol failures

1 documented incidents, average severity 9.0/10. Most common failure mode: destructive action.
n = 1low confidence

Not ranked

With fewer than 3 documented incidents, this agent is excluded from comparative rankings. These reports are published, but they are not a basis for judging the product.
1
Incidents
9.0
Avg severity
1
Critical
1
Verified

Severity breakdown

Stacked severity breakdown per agent. Longer bar = more incidents.

Failure modes

Distribution of failure modes across all documented incidents.

Worst documented incident

9.0critical

GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model

All 1 incidents