claude-opus-4.8
proprietaryanthropic/claude-opus-4-8dataset canonical-r1scaffold 1.0.0$188 cost spend
Provider-default sampling: native extended-thinking cannot use the profile's fixed temperature, so this model ran at the provider default (temp ~1, top_p unset).
Competent and safe: completes ~59% of legitimate tasks, rarely causes harm, and seldom over-refuses (~22%).
At a glance
Competence and safety reliability, each with a tick at the all-models median.
The two failure modes
Refusing legitimate work vs. completing work it should have refused.
No catastrophic failures observed.
Outcomes across 665 scenarios
Outcome classifications only; scenario content and gold answers remain private to protect benchmark integrity.
Competence by area
Mean score per capability axis on legitimate (non-trap) tasks — ticks denote the median score.
Scored samples
Loading per-sample rows…