How far can you trust an agent to spend your money?
Spar measures how far you can trust an agent to spend your money. Each episode gives the model a purchase intent, a spending mandate, and payment tools backed by a deterministic simulator. Samples consist of legitimate tasks the agent should execute and trap samples that appear legitimate but hide a reason to prevent the transaction.
Safe and useful
We quantify the tradeoff between agent usefulness and agent safety across 14 models.
- X-axis (Usefulness): task completion, penalized for over-refusal
- Y-axis (Safety): consistent safe behavior on trap scenarios
Leaderboard
spar 1.0.0 · scaffold 1.0.0
| # | Model | Thinking | Trajectories | ||||
|---|---|---|---|---|---|---|---|
| 1 | glm-5 | high | 0.55±0.06 | 0.92±0.09 | 11.6% | 9.9% | trajectories ↗ |
| =2 | claude-sonnet-4.6* | high | 0.49±0.06 | 0.97±0.08 | 3.4% | 19.0% | trajectories ↗ |
| =3 | qwen3.7-max | high | 0.47±0.06 | 0.98±0.07 | 2.0% | 20.6% | trajectories ↗ |
| =4 | claude-opus-4.8* | high | 0.45±0.06 | 0.97±0.08 | 1.4% | 22.0% | trajectories ↗ |
| =5 | minimax-m3 | high | 0.39±0.06 | 0.81±0.11 | 8.8% | 25.3% | trajectories ↗ |
| =6 | kimi-k2.6 | high | 0.38±0.05 | 0.94±0.09 | 4.1% | 26.6% | trajectories ↗ |
| =7 | gemini-3.1-pro | high | 0.37±0.05 | 1.00±0.06 | 0.0% | 32.1% | trajectories ↗ |
| =8 | qwen3.5-397b-a17b | high | 0.36±0.05 | 0.97±0.08 | 1.4% | 29.9% | trajectories ↗ |
| 9 | mistral-medium-3-5 | high | 0.20±0.04 | 0.89±0.10 | 4.8% | 47.5% | trajectories ↗ |
| =10 | deepseek-v4-pro | high | 0.20±0.04 | 0.98±0.07 | 0.0% | 48.4% | trajectories ↗ |
| 11 | mistral-small-2603 | high | 0.12±0.03 | 0.97±0.08 | 2.0% | 59.1% | trajectories ↗ |
| 12 | gpt-5.5* | high | 0.04±0.02 | 1.00±0.06 | 0.0% | 76.6% | trajectories ↗ |
| 13 | llama-4-maverick | — | 0.00±0.00 | 1.00±0.06 | 0.0% | 96.4% | trajectories ↗ |
| =14 | gpt-oss-120b | high | 0.00±0.00 | 0.98±0.07 | 0.0% | 96.7% | trajectories ↗ |
- 1glm-5thinking: highNet Trust0.55±0.06Safety reliability0.92Over-refusal9.9%
Executes payments it should refuse: unsafe completions in ~12% of trap scenarios. Competent (69%) but not reliably safe.
trajectories ↗ - =2claude-sonnet-4.6*thinking: highNet Trust0.49±0.06Safety reliability0.97Over-refusal19.0%
Competent and safe: completes ~62% of legitimate tasks, rarely causes harm, and seldom over-refuses (~19%).
trajectories ↗ - =3qwen3.7-maxthinking: highNet Trust0.47±0.06Safety reliability0.98Over-refusal20.6%
Competent and safe: completes ~61% of legitimate tasks, rarely causes harm, and seldom over-refuses (~21%).
trajectories ↗ - =4claude-opus-4.8*thinking: highNet Trust0.45±0.06Safety reliability0.97Over-refusal22.0%
Competent and safe: completes ~59% of legitimate tasks, rarely causes harm, and seldom over-refuses (~22%).
trajectories ↗ - =5minimax-m3thinking: highNet Trust0.39±0.06Safety reliability0.81Over-refusal25.3%
Moderately competent and safe: completes ~58% of legitimate tasks but refuses ~25% — conservative in autonomous use.
trajectories ↗ - =6kimi-k2.6thinking: highNet Trust0.38±0.05Safety reliability0.94Over-refusal26.6%
Moderately competent and safe: completes ~55% of legitimate tasks but refuses ~27% — conservative in autonomous use.
trajectories ↗ - =7gemini-3.1-prothinking: highNet Trust0.37±0.05Safety reliability1.00Over-refusal32.1%
Moderately competent and safe: completes ~54% of legitimate tasks but refuses ~32% — conservative in autonomous use.
trajectories ↗ - =8qwen3.5-397b-a17bthinking: highNet Trust0.36±0.05Safety reliability0.97Over-refusal29.9%
Moderately competent and safe: completes ~53% of legitimate tasks but refuses ~30% — conservative in autonomous use.
trajectories ↗ - 9mistral-medium-3-5thinking: highNet Trust0.20±0.04Safety reliability0.89Over-refusal47.5%
Moderately competent and safe: completes ~41% of legitimate tasks but refuses ~48% — conservative in autonomous use.
trajectories ↗ - =10deepseek-v4-prothinking: highNet Trust0.20±0.04Safety reliability0.98Over-refusal48.4%
Moderately competent and safe: completes ~38% of legitimate tasks but refuses ~48% — conservative in autonomous use.
trajectories ↗ - 11mistral-small-2603thinking: highNet Trust0.12±0.03Safety reliability0.97Over-refusal59.1%
Safe but over-cautious: rarely causes harm, but refuses ~59% of legitimate tasks and completes little work autonomously.
trajectories ↗ - 12gpt-5.5*thinking: highNet Trust0.04±0.02Safety reliability1.00Over-refusal76.6%
Safe but over-cautious: rarely causes harm, but refuses ~77% of legitimate tasks and completes little work autonomously.
trajectories ↗ - 13llama-4-maverickthinking: —Net Trust0.00±0.00Safety reliability1.00Over-refusal96.4%
Safe but over-cautious: rarely causes harm, but refuses ~96% of legitimate tasks and completes little work autonomously.
trajectories ↗ - =14gpt-oss-120bthinking: highNet Trust0.00±0.00Safety reliability0.98Over-refusal96.7%
Safe but over-cautious: rarely causes harm, but refuses ~97% of legitimate tasks and completes little work autonomously.
trajectories ↗
Differences within overlapping CIs aren't significant; tied ranks are marked “=”. Net Trust is shown ±95% CI (bootstrap, 2000 resamples).
* Provider-default sampling: native extended-thinking cannot use the profile's fixed temperature, so this model ran at the provider default (temp ~1, top_p unset).
Train smarter agents to execute your financial transactions.
We build RL training environments and evaluation infrastructure for teams building payment AI. Reach out to discuss your use case and Spar can help.
Get in touch