How far can you trust an agent to spend your money?

Spar measures how far you can trust an agent to spend your money. Each episode gives the model a purchase intent, a spending mandate, and payment tools backed by a deterministic simulator. Samples consist of legitimate tasks the agent should execute and trap samples that appear legitimate but hide a reason to prevent the transaction.

Safe and useful

We quantify the tradeoff between agent usefulness and agent safety across 14 models.

  • X-axis (Usefulness): task completion, penalized for over-refusal
  • Y-axis (Safety): consistent safe behavior on trap scenarios
Usefulness (Net Trust)Safety (pass^4 reliability)safe but over-cautioustrustworthy

Leaderboard

spar 1.0.0 · scaffold 1.0.0

Differences within overlapping CIs aren't significant; tied ranks are marked “=”. Net Trust is shown ±95% CI (bootstrap, 2000 resamples).

* Provider-default sampling: native extended-thinking cannot use the profile's fixed temperature, so this model ran at the provider default (temp ~1, top_p unset).

Train smarter agents to execute your financial transactions.

We build RL training environments and evaluation infrastructure for teams building payment AI. Reach out to discuss your use case and Spar can help.

Get in touch