llama-4-maverick

open-weightsmeta-llama/llama-4-maverickdataset canonical-r1scaffold 1.0.0$5.28 cost spend

vs. all models · ↗ better
Net Trust
0.00±0.00

Safe but over-cautious: rarely causes harm, but refuses ~96% of legitimate tasks and completes little work autonomously.

At a glance

Competence and safety reliability, each with a tick at the all-models median.

Competence
Mean score on legitimate tasks
0.03
higher is better
Safety reliability
Probability of avoiding harm in all four trials
1.00±0.06
higher is better

The two failure modes

Refusing legitimate work vs. completing work it should have refused.

96.4%
0.0%
over-refuses (False-refusal rate)over-completes (Unsafe-completion rate)

No catastrophic failures observed.

Outcomes across 665 scenarios

Outcome classifications only; scenario content and gold answers remain private to protect benchmark integrity.

completed correctly 10correctly declined 301over-refused 351completed incorrectly 3

Competence by area

Mean score per capability axis on legitimate (non-trap) tasks — ticks denote the median score.

State revalidation0.43
Tax & regulatory compliance0.42
Fraud-system response0.42
Mandate compliance0.29
Decline recovery0.03
Payment routing0.01
Post-purchase settlement0.00

Scored samples

Area
Trap
Intent
Outcome

Loading per-sample rows…