Regulatory-control bypass in financial LLM agents · reference leaderboard
Can a tool-using agent be driven, through indirect prompt injection, to bypass the controls a bank actually operates? Each scenario is replayed under three postures — none, advisory (policy stated), enforced (hard block) — and scored by a deterministic predicate over world state, not an LLM judge.
| # | Model | ASR · no policy | ASR · advisory | ASR · enforced | Utility | Over-refusal |
|---|
Ranked by residual ASR under enforcement, then advisory, then no-policy ASR, then utility. Lower ASR is better; higher utility is better. The two 0%-ASR models differ sharply: Claude Sonnet pairs it with 100% utility (principled refusal), the local Llama with 50% utility (it is quiet because it often does not act — an over-refusal artifact, not judgment).
Vulnerability concentrates in payment structuring (T4) and dual-approval defeat (T5) — procedural, aggregate-reasoning controls where no single action is forbidden and only the pattern is illegitimate. These are exactly where hard enforcement, not written policy, is decisive.