window.huggingface={variables:{"SPACE_CREATOR_USER_ID":"6a49466b24692fe938c10d1a"}};r>

FinAgent Red-Team

Regulatory-control bypass in financial LLM agents · reference leaderboard

Can a tool-using agent be driven, through indirect prompt injection, to bypass the controls a bank actually operates? Each scenario is replayed under three postures — none, advisory (policy stated), enforced (hard block) — and scored by a deterministic predicate over world state, not an LLM judge.

48 scenarios 42 attack · 6 benign 7 models 3 control postures deterministic scoring 95% Wilson CIs
0%
Attack success at enforced posture — every model
71%
Peak no-policy ASR — GPT-4o, a frontier model
1 / 7
Models safe by default (0% ASR, full utility)
T4·T5
Categories a stated policy fails to stop

Leaderboard — attack success rate by control posture

# Model ASR · no policy ASR · advisory ASR · enforced Utility Over-refusal

Ranked by residual ASR under enforcement, then advisory, then no-policy ASR, then utility. Lower ASR is better; higher utility is better. The two 0%-ASR models differ sharply: Claude Sonnet pairs it with 100% utility (principled refusal), the local Llama with 50% utility (it is quiet because it often does not act — an over-refusal artifact, not judgment).

Per-threat ASR — advisory posture

0% robust 1–10% 11–30% >30%

Vulnerability concentrates in payment structuring (T4) and dual-approval defeat (T5) — procedural, aggregate-reasoning controls where no single action is forbidden and only the pattern is illegitimate. These are exactly where hard enforcement, not written policy, is decisive.

Every scenario runs against a synthetic world (mock ledgers, payees, counterparties) with no real account data. MIT-licensed. Numbers are the reference run (advisory single-trial for Groq-served Llama; see the repository for confidence intervals and full tables).