claude-sonnet-4.6

proprietaryanthropic/claude-sonnet-4-6dataset canonical-r1scaffold 1.0.0$70.97 cost spend

Provider-default sampling: native extended-thinking cannot use the profile's fixed temperature, so this model ran at the provider default (temp ~1, top_p unset).

vs. all models · ↗ better
Net Trust
0.49±0.06

Competent and safe: completes ~62% of legitimate tasks, rarely causes harm, and seldom over-refuses (~19%).

At a glance

Competence and safety reliability, each with a tick at the all-models median.

Competence
Mean score on legitimate tasks
0.62
higher is better
Safety reliability
Probability of avoiding harm in all four trials
0.97±0.08
higher is better

The two failure modes

Refusing legitimate work vs. completing work it should have refused.

19.0%
3.4%
over-refuses (False-refusal rate)over-completes (Unsafe-completion rate)

No catastrophic failures observed.

Outcomes across 665 scenarios

Outcome classifications only; scenario content and gold answers remain private to protect benchmark integrity.

completed correctly 202correctly declined 294over-refused 69completed incorrectly 93unsafe 7

Competence by area

Mean score per capability axis on legitimate (non-trap) tasks — ticks denote the median score.

State revalidation0.94
Tax & regulatory compliance0.86
Payment routing0.75
Mandate compliance0.66
Decline recovery0.39
Post-purchase settlement0.31
Fraud-system response0.28

Scored samples

Area
Trap
Intent
Outcome

Loading per-sample rows…