kimi-k2.6

open-weightsmoonshotai/kimi-k2.6dataset canonical-r1scaffold 1.0.0$29.07 cost spend

vs. all models · ↗ better
Net Trust
0.38±0.05

Moderately competent and safe: completes ~55% of legitimate tasks but refuses ~27% — conservative in autonomous use.

At a glance

Competence and safety reliability, each with a tick at the all-models median.

Competence
Mean score on legitimate tasks
0.55
higher is better
Safety reliability
Probability of avoiding harm in all four trials
0.94±0.09
higher is better

The two failure modes

Refusing legitimate work vs. completing work it should have refused.

26.6%
4.1%
over-refuses (False-refusal rate)over-completes (Unsafe-completion rate)

2 catastrophic failures observed: Premature capture ×2.

Outcomes across 665 scenarios

Outcome classifications only; scenario content and gold answers remain private to protect benchmark integrity.

completed correctly 177correctly declined 289over-refused 97completed incorrectly 93unsafe 9

Competence by area

Mean score per capability axis on legitimate (non-trap) tasks — ticks denote the median score.

State revalidation0.85
Tax & regulatory compliance0.81
Mandate compliance0.63
Fraud-system response0.53
Payment routing0.48
Decline recovery0.25
Post-purchase settlement0.20

Scored samples

Area
Trap
Intent
Outcome

Loading per-sample rows…