Table 1: Main results on the base splits (GPT 5.4 agent, n=4; airline 50, retail/telecom 114 tasks). Cells report Pass1 and Pass4 overall and on the PV/Mut slices. The verifier is absent for ReAct, static code for ToolGuard, and GPT 5.4 for PolicyGuard and PolicyGuide.
Airline (50)
Retail (114)
Telecom (114)
System
Verifier
Overall
PV
Mut
Overall
PV
Mut
Overall
PV
Mut
Pass1
ReAct
—
0.640
0.865
0.433
0.800
0.900
0.791
0.384
0.721
0.180
ToolGuard
static code
0.575
0.969
0.212
—
—
—
—
—
—
PolicyGuard
GPT 5.4
0.710
1.000
0.442
0.645
0.975
0.613
0.406
0.733
0.208
PolicyGuide
GPT 5.4
0.775
0.979
0.587
0.809
0.975
0.793
0.866
0.895
0.849
Pass4
ReAct
—
0.460
0.750
0.192
0.596
0.700
0.587
0.193
0.442
0.042
ToolGuard
static code
0.520
0.875
0.192
—
—
—
—
—
—
PolicyGuard
GPT 5.4
0.580
1.000
0.192
0.360
0.900
0.308
0.202
0.488
0.028
PolicyGuide
GPT 5.4
0.620
0.917
0.346
0.614
0.900
0.587
0.614
0.721
0.549
Table 2: Workflow ablations (GPT 5.4 agent; Airline base split, Retail and Telecom benchmark test splits of 40 tasks). All cells report Pass4.
Domain
Metric
ReAct
PolicyGuide Self
PolicyGuide Raw
PolicyGuide
Airline
Overall
0.460
0.480
0.520
0.620
PV
0.750
0.833
0.875
0.917
Mut
0.192
0.154
0.192
0.346
Retail
Overall
0.575
0.350
0.575
0.725
PV
0.750
0.750
0.750
1.000
Mut
0.556
0.306
0.556
0.694
Telecom
Overall
0.250
0.325
0.350
0.675
PV
0.429
0.571
0.619
0.667
Mut
0.053
0.053
0.053
0.684
Table 3: Matched workflow-controller comparison on the 40-task Telecom benchmark test split. All values are Pass4.
System
Runtime control
Pass4
ReAct
actor only
0.250
PolicyGuard
action-local check
0.325
FlowAgent
PDL + API control
0.350
PolicyGuide
external graph verifier
0.675
Table 4: Agent-family generalization on Airline (50 tasks, n=4; verifier model paired to the agent). All metrics are Pass4. The GPT 5.4-authored workflow graph is reused without re-authoring.
Agent
Metric
ReAct
PolicyGuard
PolicyGuide
GPT 5.4
Overall
0.460
0.580
0.620
PV
0.750
1.000
0.917
Mut
0.192
0.192
0.346
Claude Sonnet 4.6
Overall
0.720
0.780
0.780
PV
0.958
1.000
1.000
Mut
0.500
0.577
0.577
Gemini 2.5 Pro
Overall
0.480
0.600
0.680
PV
0.750
1.000
0.917
Mut
0.231
0.231
0.462
Table 5: Author-designed Telecom ordered trace compliance (%; n=4). Step- and Trace-TCR condition on outcome-passing traces.
System
Step-TCR
Trace-TCR
Process-valid rate
ReAct
86.4
35.4
17.5
PolicyGuard
85.7
23.9
13.1
PolicyGuide
94.5
63.4
56.2
Table 6: Argument- (A), process- (P), and workflow-level (W) partition of the source policies (W splits PolicyGuard’s process-level class; P+W equals it).
Domain
A
P
W
Total
% P+W
% W
Airline
14
27
2
43
67.4%
4.7%
Retail
0
27
1
28
∼100%
3.6%
Telecom (main)
1
21
7
29
96.6%
24.1%
Telecom (manual)
0
1
20
21
100%
95.2%
Telecom (both)
1
22
27
50
98.0%
54.0%
Table 7: Hand-classified atomic requirements of the τ2-bench Retail policy document (28 requirements, 0 A / 27 P / 1 W; subtypes D-only 13, T-only 14, D+T 1). Line refers to retail/policy.md as released with τ2-bench; Type A = argument-level, P = process-level (flat), W = workflow-level (order-bound; Appendix B.1), with D = dialogue-dependent, T = requires a prior read-only tool call.
ID
Line
Requirement (paraphrased)
Type
Global rules
G1
10
Authenticate identity by locating the user id via email or name+zip—even when the user already provides the id
P (D+T)
G2
14
One user per conversation; deny any request about another user
P (D)
G3
16
List action details + obtain explicit “yes” before any DB-updating action
P (D)
G4
18
No fabricated information/knowledge/procedures; no subjective recommendations
P (D)
G5
20
At most one tool call per turn (not paired with a user-facing reply)
P (D)
G6
22
Deny user requests that are against the policy
P (D)
G7
24
Transfer iff unhandleable: call transfer_to_human_agents, then the literal handoff message
W (D)
Generic action rules
N1
82
Act only on orders with status pending or delivered
P (T)
N2
84
Exchange / modify-items tools callable only once per order
P (T)
N3
84
Collect all items to change into one list before making the call
P (D)
Cancel pending order
C1
88
Order status must be pending; check it before taking the action
P (T)
C2
90
User confirms order id + reason ∈ {‘no longer needed’, ‘ordered by mistake’}; no other reason
P (D)
Modify pending order
M1
96
Order status must be pending; check it before taking the action
P (T)
M2
98
Only shipping address, payment method, or item options may be modified—nothing else
P (D)
M3
102
New payment = a single method, different from the original
P (T)
M4
104
If the new payment is a gift card, its balance must cover the total amount
P (T)
M5
110
Modify-items is one-shot (order becomes unmodifiable): remind + confirm all items first
P (D)
M6
112
Each new item must be available
P (T)
M7
112
New item = same product, different option (no product-type change)
P (T)
M8
114
User provides a payment method for the price difference
P (D)
M9
114
If that payment is a gift card, its balance must cover the price difference
P (T)
Return delivered order
R1
118
Order status must be delivered; check it before taking the action
P (T)
R2
120
User confirms order id + the list of items to be returned
P (D)
R3
122–124
Refund method provided; must be the original payment method or an existing gift card
P (T)
Exchange delivered order
Table 8: Hand-classified atomic requirements of the τ2-bench Telecom main_policy.md (29 requirements, 1 A / 21 P / 7 W). Line refers to the document as released with τ2-bench; \raisebox{-0.4pt}{\scriptsize$n$}⃝ marks a step in an ordered procedure (“To do so you need to follow these steps”). Types as in Table 7.
ID
Line
Requirement (paraphrased)
Type
Global rules
G1
7
No fabricated information/knowledge/procedures; no subjective recommendations
P (D)
G2
9
At most one tool call per turn (not paired with a user-facing reply)
P (D)
G3
11
Deny user requests that are against the policy
P (D)
G4
13
Transfer iff unhandleable: call transfer_to_human_agents, then the literal handoff message
W (D)
G5
15
Try your best to resolve the issue before transferring
W (D)
Customer lookup
L1
94–97
Identify the customer via phone number, customer ID, or full name + date of birth
P (D+T)
L2
99
For name lookup, date of birth is required for verification
P (D)
Overdue bill payment (ordered procedure)
O1
105, 117
\raisebox{-0.4pt}{\scriptsize1}⃝Check the bill status is Overdue before acting (the API does not check it)
P (T)
O2
106
\raisebox{-0.4pt}{\scriptsize2}⃝Check the bill amount due
P (T)
O3
107–108
\raisebox{-0.4pt}{\scriptsize3}⃝Send the payment request (→ AWAITING PAYMENT); gated on O1
P (T)
O4
109–110
\raisebox{-0.4pt}{\scriptsize4}⃝Inform the user to check their payment requests
W (D)
O5
111
\raisebox{-0.4pt}{\scriptsize5}⃝Only after the user accepts, call make_payment
W (D+T)
O6
113
\raisebox{-0.4pt}{\scriptsize6}⃝Always verify the bill became PAID before telling the user
W (T)
O7
116
At most one bill in AWAITING PAYMENT at a time
P (T)
Line suspension
S1
125
Lift a suspension only after all overdue bills are paid
W (T)
S2
126
Do not lift if the contract end date is past—even if all bills are paid
P (T)
S3
128
After resuming, instruct the user to reboot the device
W (D)
Data refueling (ordered procedure)
F1
134
Refuel amount ≤2 GB
A
F2
136
\raisebox{-0.4pt}{\scriptsize1}⃝Ask how much data the user wants to refuel
P (D)
F3
137
\raisebox{-0.4pt}{\scriptsize2}⃝Confirm the price
P (D)
F4
138
\raisebox{-0.4pt}{\scriptsize3}⃝Apply the refuel to the line associated with the user’s phone number
P (D+T)
Change plan (ordered procedure)
P1
144
\raisebox{-0.4pt}{\scriptsize1}⃝Establish which line the plan change is for
P (D)
P2
145
\raisebox{-0.4pt}{\scriptsize2}⃝Gather the available plans
P (T)
P3
146
\raisebox{-0.4pt}{\scriptsize3}⃝Ask the user to select one
P (D)
Table 9: Hand-classified atomic requirements of the τ2-bench Telecom tech_support_manual.md (21 requirements, 1 P / 20 W). Line refers to the document as released with τ2-bench. Every rule is a diagnostic-gated (T) user-guidance (D) step; the three chapters form the prerequisite chain Service ⊂ Data ⊂ MMS; every row except TSS1 (the entry diagnostic) is workflow-level.
ID
Line
Requirement (diagnose → conditional fix → verify)
Type
Cellular service (ll. 55–99)
TSS1
69–72
Diagnose service via check_status_bar
P (T)
TSS2
74–78
If Airplane Mode ON → guide toggle_airplane_mode OFF
W (D+T)
TSS3
79–87
Check SIM: Missing → reseat; Locked → escalate; Active → ok (three-way branch)
W (D+T)
TSS4
88–92
If APN incorrect → guide reset_apn_settings, then reboot_device
W (D+T)
TSS5
93–99
If line suspended → handle per main policy, then verify service restored
W (T)
Mobile data (ll. 100–163)
TSD0
106–108
Prerequisite: the user must first have cellular service
W (T)
TSD1
122–127
Diagnose via run_speed_test
W (T)
TSD2
129–131
Airplane Mode (as in the Service chapter)
W (D+T)
TSD3
132–135
If mobile data disabled → guide toggle_data ON
W (D+T)
TSD4
136–141
If roaming abroad & data off → guide toggle_roaming + verify the line is roaming-enabled
W (D+T)
TSD5
142–145
If Data Saver ON → guide toggle_data_saver_mode OFF
W (D+T)
TSD6
146–150
If VPN ON & performance poor → guide disconnect_vpn
W (D+T)
TSD7
151–158
If usage exceeds the plan limit → offer change-plan or refuel
W (T)
TSD8
159–163
If network mode 2G/3G → guide set_network_mode_preference
W (D+T)
MMS (ll. 164–205)
TSM0
170–173
Prerequisite: the user must have cellular service and mobile data
W (T)
TSM1
181–183
Diagnose via can_send_mms
W (T)
TSM2
185–188
Ensure basic service + data connectivity first
W (T)
TSM3
189–193
If on 2G → guide set_network_mode_preference to 3G+
W (D+T)
TSM4
194–199
If MMSC URL unset → guide reset_apn_settings, then reboot_device
W (D+T)
TSM5
200–203
If Wi-Fi Calling ON → guide toggle_wifi_calling OFF
W (D+T)
TSM6
204–205
If the messaging app lacks storage/SMS permissions → guide grant_app_permission
W (D+T)
Table 10: Passk breakdown for the base-split results in Table 1 and Figure 4 (GPT 5.4, n=4). P4/P1 is the consistency ratio.
Domain
System
P1
P2
P3
P4
P4/P1
Airline
ReAct
0.640
0.530
0.485
0.460
0.72
ToolGuard
0.575
0.553
0.535
0.520
0.90
PolicyGuard
0.710
0.630
0.595
0.580
0.82
PolicyGuide
0.775
0.707
0.660
0.620
0.80
Retail
ReAct
0.800
0.700
0.638
0.596
0.75
PolicyGuard
0.645
0.506
0.421
0.360
0.56
PolicyGuide
0.809
0.715
0.654
0.614
0.76
Telecom
ReAct
0.384
0.273
0.226
0.193
0.50
PolicyGuard
0.406
0.292
0.237
0.202
0.50
PolicyGuide
0.866
0.763
0.682
0.614
0.71
Table 11: Pass1 in each of the four trials on the base splits. pstd is the population standard deviation across trial-level values.
Domain
System
T1
T2
T3
T4
pstd
Airline
ReAct
0.620
0.620
0.640
0.680
0.024
ToolGuard
0.560
0.580
0.580
0.580
0.009
PolicyGuard
0.700
0.740
0.700
0.700
0.017
PolicyGuide
0.800
0.720
0.820
0.760
0.038
Retail
ReAct
0.807
0.746
0.842
0.807
0.035
PolicyGuard
0.649
0.623
0.632
0.675
0.020
PolicyGuide
0.816
0.798
0.789
0.833
0.017
Telecom
ReAct
0.342
0.377
0.404
0.412
0.027
PolicyGuard
0.465
0.386
0.360
0.412
0.039
PolicyGuide
0.860
0.860
0.842
0.904
0.023
Table 12: Pooled stratified McNemar tests on per-task Pass4. D is the number of domain strata; a counts PolicyGuide-only passes and b the reverse, summed across strata.
Opponent
D
∑a
∑b
ndisc
Z
p
ReAct
3
77
19
96
+5.92
<10−8
PolicyGuard
3
97
19
116
+7.24
<10−12
Table 13: Per-domain paired-bootstrap differences in Pass4 on the base splits (10,000 task-level resamples). Positive values favor PolicyGuide.
Domain
Opponent
n
ΔP4 [95% CI]
Airline
ReAct
50
+0.160 [+0.020, +0.300]
ToolGuard
50
+0.100 [−0.040, +0.260]
PolicyGuard
50
+0.040 [−0.080, +0.160]
Retail
ReAct
114
+0.018 [−0.070, +0.105]
PolicyGuard
114
+0.254 [+0.149, +0.360]
Telecom
ReAct
114
+0.421 [+0.316, +0.526]
PolicyGuard
114
+0.412 [+0.298, +0.526]
Table 14: Guide-side model usage for the GPT 5.4 configuration (50 Airline and 40 Retail/Telecom tasks). Costs exclude the actor and user simulator.
Domain
Calls/ task
Prompt tok./call
Cached input
Output tok./call
Guide total $
Guide $/task
Airline
7.56
32,360
88.1%
2,478
20.10
0.40
Retail
7.42
22,803
85.8%
2,179
13.67
0.34
Telecom
11.47
28,518
86.5%
2,186
22.29
0.56
Table 15: Mean end-to-end wall-clock time per task.
Domain
ReAct (s/task)
PolicyGuide (s/task)
Ratio
Airline
36.4
210.1
5.78×
Retail
34.6
193.6
5.60×
Telecom
45.5
247.6
5.45×
Table 16: Programmatic validation rerun on the frozen workflow graphs.
Domain
Nodes
Auth. nodes
Validator flags
Airline
158
11
0
Retail
104
7
0
Telecom
127
5
1
Table 17: Call-NMR on passing Mut trajectories (n=4): percentage of successfully executed agent mutations missing a frozen guard-derived read prerequisite. †Telecom is an adapted, agent-side diagnostic whose read oracle saturates; its zeros do not establish equal procedural quality.
Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workflow-following systems support prescribed process execution, but primarily target workflow completion rather than safeguarding agent behavior. PolicyGuide instead compiles each domain policy into a workflow graph and invokes a proactive verifier at user-turn boundaries. From persisted graph state, the verifier reconciles open requests and returns step-specific remediation along a policy-compliant path. Across the $\tau^2$-bench airline, retail, and telecom domains with a GPT-5.4 agent and verifier, PolicyGuide raises mean $\mathrm{Pass}^4$ from $0.42$ to $0.62$, with the largest gain on telecom ($0.19$ to $0.61$), the most workflow-structured domain. The same workflows transfer to Claude Sonnet 4.6 and Gemini 2.5 Pro agents. Complementary evaluations find the lowest observed attack-success rate under adversarial users and the strongest procedural compliance in an author-designed workflow-level validation.