kimi-k2.6
open-weightsmoonshotai/kimi-k2.6dataset canonical-r1scaffold 1.0.0$29.07 cost spend
vs. all models · ↗ better
Net Trust
0.38±0.05
Moderately competent and safe: completes ~55% of legitimate tasks but refuses ~27% — conservative in autonomous use.
At a glance
Competence and safety reliability, each with a tick at the all-models median.
CompetencePerformance on legitimate (non-trap) work.
Mean score on legitimate tasks
0.55
higher is better
Safety reliabilityConsistency of safe behavior on traps across repeated trials.
Probability of avoiding harm in all four trials
0.94±0.09
higher is better
The two failure modes
Refusing legitimate work vs. completing work it should have refused.
26.6%
4.1%
over-refuses (False-refusal rate)Over-caution — legitimate work the model declined.over-completes (Unsafe-completion rate)Payments executed when the correct action was to abort or escalate.
2 catastrophic failures observed: Premature capture ×2.
Outcomes across 665 scenarios
Outcome classifications only; scenario content and gold answers remain private to protect benchmark integrity.
completed correctly 177correctly declined 289over-refused 97completed incorrectly 93unsafe 9
Competence by area
Mean score per capability axis on legitimate (non-trap) tasks — ticks denote the median score.
State revalidation0.85
Tax & regulatory compliance0.81
Mandate compliance0.63
Fraud-system response0.53
Payment routing0.48
Decline recovery0.25
Post-purchase settlement0.20
Scored samples
AreaTrapIntentOutcome
Loading per-sample rows…