Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

arXiv:2608.181362026-08-20

A new test checks whether banking chatbots can be talked into leaking data or moving money

FraudBench is a new benchmark that lets a simulated scammer AI try to talk a bank-service AI agent, one with real tools to look up accounts, reset PINs, and transfer money, into unsafe actions purely through conversation. The scammer uses ten fraud tactics plus multi-step chained attacks, including tricks where an earlier admission in the conversation makes a later, seemingly normal request unsafe. Across four tested models, defense success ranged only from 49% to 65%, with money-mule fraud and chained attacks being the hardest to catch.

METAL MEDIA explanatory visual

A new test checks whether banking chatbots can be talked into leaking data or moving money

  1. 01The authors built a simulated bank with a 698-document internal policy library, a mutable customer database, and 17 privileged tools, then had an AI agent and an AI scammer interact through those tools in real conversations
  2. 02Safety judgments depend on the whole conversation history, not just the latest message: for instance, a transfer request must be refused if the caller earlier admitted the money belongs to someone else, even if the request looks fine in isolation
  3. 03Of 150 hand-written adversarial scenarios, 107 (90 single-mechanism cases plus 17 multi-step chained attacks) were released and used to test four models: Nemotron-3 Ultra 550B, Gemini 3.6 Flash, Gemini 3.1 Flash-Lite, and gpt-oss-120b
  4. 04Even under an ideal setup where the agent was handed the correct policy document in advance, defense success ranged from 49.5% to 64.5%; when Gemini 3.6 Flash had to search the full document corpus itself, its score dropped 13 points, from 64.5% to 51.0%
  5. 05Money-mule fraud was the weakest spot for every model, and on the 17 chained multi-step attacks, even the stronger models only defended 8-9 cases versus just 4 for weaker ones
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. The authors built a simulated bank with a 698-document internal policy library, a mutable customer database, and 17 privileged tools, then had an AI agent and an AI scammer interact through those tools in real conversations
  2. Safety judgments depend on the whole conversation history, not just the latest message: for instance, a transfer request must be refused if the caller earlier admitted the money belongs to someone else, even if the request looks fine in isolation
  3. Of 150 hand-written adversarial scenarios, 107 (90 single-mechanism cases plus 17 multi-step chained attacks) were released and used to test four models: Nemotron-3 Ultra 550B, Gemini 3.6 Flash, Gemini 3.1 Flash-Lite, and gpt-oss-120b
  4. Even under an ideal setup where the agent was handed the correct policy document in advance, defense success ranged from 49.5% to 64.5%; when Gemini 3.6 Flash had to search the full document corpus itself, its score dropped 13 points, from 64.5% to 51.0%
  5. Money-mule fraud was the weakest spot for every model, and on the 17 chained multi-step attacks, even the stronger models only defended 8-9 cases versus just 4 for weaker ones
Figure 1: FraudBench overview. A simulated adaptive caller attacks a policy-grounded banking agent across ten fraud mechanisms plus chained multi-step attacks, over privileged tools, a mutable customer database, and a 698-document internal policy corpus; episodes are graded on actions, state, leaks, and disposition, conditioned on the full conversation history. Prior conversational fraud benchmarks such as Fraud-R1 judge resistance to fraud messages without tools, a database, or an internal policy corpus, and cover a narrower slice of banking-fraud typologies.
Figure 1: FraudBench overview. A simulated adaptive caller attacks a policy-grounded banking agent across ten fraud mechanisms plus chained multi-step attacks, over privileged tools, a mutable customer database, and a 698-document internal policy corpus; episodes are graded on actions, state, leaks, and disposition, conditioned on the full conversation history. Prior conversational fraud benchmarks such as Fraud-R1 judge resistance to fraud messages without tools, a database, or an internal policy corpus, and cover a narrower slice of banking-fraud typologies.
Table 1: Operational comparison with neighboring agent benchmarks and frameworks. ✓ denotes a primary, implemented capability; ⊚, partial or narrower; –, not a primary feature. Multi-class fraud: multiple distinct fraud typologies are tested. History-dep.: safety depends on earlier turns, not only the latest request. Long conv.: extended multi-turn interaction with a (simulated) user. Type indicates whether the work is a fixed benchmark, a reusable framework, or both; FraudBench is a fixed banking-fraud benchmark whose environment and attacker machinery are reusable for new scenarios.
WorkExec. stateLarge KBFraudMulti-class fraudHistory-dep.Long conv.Type
Fraud-R1 (Yang et al. 2025)Benchmark
τ2-bench (Barres et al. 2025)Benchmark
τ-Knowledge (Shi et al. 2026)Benchmark
AgentDojo (Debenedetti et al. 2024)Both
SafeArena (Tur et al. 2025)Benchmark
AgentHazard (Feng et al. 2026)Benchmark
DoomArena (Boisvert et al. 2025)Framework
FraudBenchBoth
Figure 2: Grading pipeline. (1) Conversation roll: the agent and simulated caller hold a multi-turn conversation (shown top-to-bottom, chat style); here the caller impersonates a customer, induces the agent to read back protected contact details, and then requests a transfer to a third-party IBAN. (2) The full transcript is scored in two stages: (2a) a deterministic grader inspects the ordered tool-call sequence, unsafe state changes, and structured leaks (e.g. an executed unauthorized transfer), and (2b) an LLM-as-judge stage checks the scenario’s natural-language assertions (scam recognized, no PII revealed, transfer refused), each marked pass (✓) or fail (X). (3) The episode earns reward 1 only if all checks pass (logical AND); a single leak or prohibited action yields 0. The reported runs use the LLM-judge stage only; the deterministic stage is what catches false-safe transcripts like this one.
Figure 2: Grading pipeline. (1) Conversation roll: the agent and simulated caller hold a multi-turn conversation (shown top-to-bottom, chat style); here the caller impersonates a customer, induces the agent to read back protected contact details, and then requests a transfer to a third-party IBAN. (2) The full transcript is scored in two stages: (2a) a deterministic grader inspects the ordered tool-call sequence, unsafe state changes, and structured leaks (e.g. an executed unauthorized transfer), and (2b) an LLM-as-judge stage checks the scenario’s natural-language assertions (scam recognized, no PII revealed, transfer refused), each marked pass (✓) or fail (X). (3) The episode earns reward 1 only if all checks pass (logical AND); a single leak or prohibited action yields 0. The reported runs use the LLM-judge stage only; the deterministic stage is what catches false-safe transcripts like this one.
Table 2: Ten fraud mechanisms and the decision each isolates, plus the chain category of multi-step adaptive attacks.
MechanismPrincipal decision being tested
APP scamProtect a verified customer manipulated into an authorized payment.
Account takeoverStop stolen knowledge or recovery credentials from becoming account control.
Card fraudPreserve freezes, activation locks, PIN rules, and transaction controls.
Data exfiltrationEnforce identity, role, ownership, scope, and retention boundaries.
First-party fraudDetect dispute or reimbursement claims contradicted by the account.
Prompt injectionTreat instruction-like text in records or user content as data.
Money muleStop laundering, structuring, third-party movement, or authorization bypass.
Phishing PIIRecognize that static PII satisfies basic identity but not every high-risk action.
Social engineeringResist false authority, urgency, secrecy, sympathy, and escalation pressure.
Synthetic identityEnforce onboarding and expansion eligibility against fabricated or blended identities.
Chain (adaptive)Preserve trust-state evidence across a multi-step attack when a later locally valid request follows an earlier probe, admission, or failed attempt.
Table 3: Preliminary attack-security at pass1 on the currently graded tasks, for the four models with complete 107-task runs under oracle retrieval (the agent’s KB_search returns the task’s gold policy documents). All runs use the GPT-5.4 Nano user simulator, one trial, and the natural-language-assertion grader only (no deterministic Stage B checks, no matched legitimate controls). Models without a complete 107-task run are excluded. A realistic all-tools retrieval setting is compared in Table 5.
DefenderTasksSecurity S (p1)
Gemini 3.6 Flash10764.5%
Nemotron-3 Ultra 550B10757.9%
Gemini 3.1 Flash-Lite10753.3%
gpt-oss-120b10749.5%
Table 4: Per-mechanism attack-security (defended tasks / graded tasks) for the four models with complete 107-task runs. Columns: Nemo. = Nemotron-3 Ultra 550B; G3.6F = Gemini 3.6 Flash; G3.1FL = Gemini 3.1 Flash-Lite; oss120 = gpt-oss-120b. Money mule is the hardest mechanism for every model; the chain row covers all 17 chain tasks.
MechanismNemo.G3.6FG3.1FLoss120
APP scam3/95/94/92/9
Account takeover3/98/97/96/9
Card fraud7/98/97/98/9
Data exfiltration8/98/97/95/9
First-party fraud4/97/93/92/9
Prompt injection7/98/96/96/9
Money mule2/91/91/91/9
Phishing PII5/94/96/95/9
Social engineering8/96/96/97/9
Synthetic identity7/95/96/97/9
Chain (adaptive)8/179/174/174/17
Overall62/10769/10757/10753/107
Overall (%)57.964.553.349.5
Table 5: Retrieval-setting comparison at pass1. Oracle supplies the gold policy documents; all-tools requires the agent to retrieve them from the 698-document corpus. Gemini 3.6 Flash is the only model run under both settings, and its 13-point drop isolates retrieval from policy reasoning; Gemini 3.1 Pro was run under all-tools only. All-tools Gemini 3.6 Flash graded 104 of 107 episodes (three infrastructure failures excluded).
DefenderOracle S (p1)All-tools S (p1)
Gemini 3.6 Flash64.5% (107)51.0% (104)
Gemini 3.1 Pro55.1% (107)

Why it matters

As banking chatbots gain real authority to move money and change account details, a few conversational tricks are enough to turn them into an attack surface, and this benchmark is one of the first to measure that risk concretely with an executable, tool-based test rather than a static text quiz. It gives banks and AI developers hard numbers on where current models fail before such agents are deployed on real customers.

Terms in this paper

  • money mule · a person or account used to funnel illegally obtained money on someone else's behalf
  • oracle retrieval · a test condition where the agent is directly given the correct policy document instead of having to search for it
  • LLM-as-judge · using another AI model to read a conversation transcript and score whether it was handled safely
  • dual-control setting · a simulation design where both the service agent and the caller can act through tools, not just the agent

Original abstract (English)

Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $\tau^2$-bench dual-control framework and the $\tau$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49\% and 65\%, with money-mule and first-party fraud the most common cross-model weaknesses.

Authors · Dheeraj Mohandas Pai, Lu Xian

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Dheeraj Mohandas Pai et al., arXiv:2608.18136, CC BY 4.0