mistral-medium-3-5

open-weightsmistralai/mistral-medium-3-5dataset canonical-r1scaffold 1.0.0$24.60 cost spend

vs. all models · ↗ better
Net Trust
0.20±0.04

Moderately competent and safe: completes ~41% of legitimate tasks but refuses ~48% — conservative in autonomous use.

At a glance

Competence and safety reliability, each with a tick at the all-models median.

Competence
Mean score on legitimate tasks
0.41
higher is better
Safety reliability
Probability of avoiding harm in all four trials
0.89±0.10
higher is better

The two failure modes

Refusing legitimate work vs. completing work it should have refused.

47.5%
4.8%
over-refuses (False-refusal rate)over-completes (Unsafe-completion rate)

3 catastrophic failures observed: Premature capture ×3.

Outcomes across 665 scenarios

Outcome classifications only; scenario content and gold answers remain private to protect benchmark integrity.

completed correctly 129correctly declined 285over-refused 173completed incorrectly 65unsafe 13

Competence by area

Mean score per capability axis on legitimate (non-trap) tasks — ticks denote the median score.

State revalidation0.77
Tax & regulatory compliance0.69
Mandate compliance0.55
Payment routing0.33
Decline recovery0.32
Fraud-system response0.27
Post-purchase settlement0.10

Scored samples

Area
Trap
Intent
Outcome

Loading per-sample rows…