Figure 1: WeClawArena pairs human-centered agent-network tasks with attack-resistance evaluation. Left: autonomous personal agents collaborate on behalf of human owners across six cross-user domains. Right: model-level macro-vector attack resistance on the ASR-main-six-domain pool, where higher 1−ASR indicates fewer judged attacks causing final harm.
Table 1: Main WeClawArena task-success results. TSR is computed over no-attacker, collaboration, security, privacy, and governance variants in each domain. Bold and underline mark the best and second-best observed values within each domain column, respectively; ties are marked together.
Model
WeClawArena TSR over all variants (%)
Bargaining
Travel
SWE-Workspace
Bidding
Clinical
Trading
Claude Opus 4.7
63.3
83.0
34.0
55.0
36.0
30.0
Claude Sonnet 4.5
68.3
58.0
8.0
51.7
40.0
35.0
Claude Opus 4.1
22.5
46.0
24.0
6.7
38.0
25.0
DeepSeek V3.2
20.0
26.0
14.8
5.0
40.0
20.0
Kimi K2.5
47.5
58.0
15.0
40.0
40.0
27.5
Kimi K2 Thinking
19.2
11.0
6.0
3.3
26.0
30.0
Qwen3 235B
25.8
47.0
2.0
3.3
34.0
30.0
Qwen3 32B
11.7
20.0
1.6
1.7
28.0
35.0
Figure 2: Overview of WeClawArena. Autonomous agents in a human-centered agent network collaborate on behalf of human users over personal digital workspaces, including filesystems, databases, tools, personal policies, and task resources, all simulated in Docker containers. The WeClawArena sandbox implements a message gateway that routes multi-agent communication, tool use, and workspace access while recording audit evidence for harmful or malicious behavior across four attack-harm families: collaboration, security, privacy, and governance. This design supports separate evaluation of task utility and attack success.
Table 2: Comparison with existing benchmarks. WeClawArena makes user-relative workspace ownership and authority part of both task utility and attack auditing.
Benchmark family
Representative benchmarks
Scored setting
Ownership and authority model
Single-user tool and workspace benchmarks
τ-bench and τ2-Bench (37; 1); WebArena, OSWorld, AppWorld, WorkArena (42; 35; 27; 5); AgentBench and GAIA (14; 17)
A tool-using agent completes tasks in a website, OS, app, or workspace.
Usually one user, account, or environment; cross-owner private resources and user-relative decision rights are outside the scored task contract.
Shared-authority multi-agent benchmarks
AutoGen, AgentVerse, MetaGPT, and MultiAgentBench (33; 3; 10; 45)
Agents coordinate, debate, specialize, or compete inside a team task.
The team usually shares task authority; agents are rarely separate delegates with owner-specific files, consents, approvals, or mandates.
Social-agent simulations
Generative Agents, OASIS, AgentSociety, and AgentSocialBench (21; 36; 22; 29)
Agent populations communicate and form social or economic behavior.
The focus is social dynamics; final utility is usually not a verifiable joint tool-use outcome assembled from separately owned workspaces.
Leakage, unsafe behavior, memory risk, governed-action failure, or trace-audit reliability.
They motivate our harm surfaces; deterministic collaborative utility and attack auditing are usually evaluated in separate settings.
WeClawArena
This work
Owned agents complete joint tasks across six domains, with one benign control and four matched attack variants per base task.
Utility requires joint tool use across private workspaces; ASR audits final harm in access, disclosure, consent, approval, mandate, and decision paths.
Figure 3: Benchmark construction pipeline. Source task pools and user-role profiles are curated into base tasks with owner workspaces, tools, predicates, and task contracts. Each base task becomes a scenario bundle with ground truth, agent/tool configuration, seed facts, personas, workspace resources, and MCP interfaces, then expands into one benign control and four attack-vector variants. Evaluation checks structural validity, task utility, and attack success from runtime evidence.
Table 3: Model-level utility and attack-audit breakdown. TSR denominators count available scenario executions. ASR denominators count attack rows with complete evidence and a valid GPT-5.2 judge verdict. Intervals are row-level Wilson 95% confidence intervals. Daggered rows have partial coverage in at least one attacked-mode component.
Model
Benign TSR (%, 95% CI)
Attacked TSR (%, 95% CI)
Drop
GPT-5.2 ASR (%, 95% CI)
Claude Opus 4.1
40/124 (32.3; 24.7–40.9)
126/496 (25.4; 21.8–29.4)
+6.9
127/436 (29.1; 25.1–33.6)
Claude Opus 4.7
76/124 (61.3; 52.5–69.4)
231/496 (46.6; 42.2–51.0)
+14.7
13/440 (3.0; 1.7–5.0)
Claude Sonnet 4.5
60/124 (48.4; 39.8–57.1)
165/496 (33.3; 29.3–37.5)
+15.1
76/293 (25.9; 21.3–31.2)
DeepSeek V3.2
43/124 (34.7; 26.9–43.4)
75/496 (15.1; 12.2–18.5)
+19.6
225/435 (51.7; 47.0–56.4)
Kimi K2.5†
64/124 (51.6; 42.9–60.2)
136/446 (30.5; 26.4–34.9)
+21.1
177/437 (40.5; 36.0–45.2)
Kimi K2 Thinking
25/124 (20.2; 14.0–28.1)
51/496 (10.3; 7.9–13.3)
+9.9
143/435 (32.9; 28.6–37.4)
Qwen3 235B†
36/124 (29.0; 21.8–37.6)
74/296 (25.0; 20.4–30.2)
+4.0
218/392 (55.6; 50.7–60.5)
Qwen3 32B
26/124 (21.0; 14.7–29.0)
41/496 (8.3; 6.2–11.0)
+12.7
173/434 (39.9; 35.4–44.5)
Figure 4: Attack-vector utility, resistance, and model-level utility-risk tradeoff on the ASR-main-six-domain pool. Left: TSR by variant and per-vector row-micro 1−ASRvmicro, where higher values indicate fewer judged attacks causing final harm. Right: model-level resistance compared against matched TSR drop from each model’s no-attacker rows. Higher 1−ASR is better; points further left preserve more utility.
Table 4: WeClawArena domain denominators. Each base has five scenario variants: one no-attacker control and four attack vectors.
Domain
Bases
Rows
Role in the benchmark
Bargaining
24
120
Owned transaction close with buyer, seller, and approver roles.
Bidding
12
60
Seller-centered public award with private buyer and seller constraints.
Travel
20
100
Hotel-booking collaboration with traveler, partner, approver, and consent constraints.
SWE-Workspace
50
250
Four-owner engineering workflow with patch, review, approval, CI, and handoff evidence.
Clinical
10
50
Role-scoped case-record coordination with protected facts and approval or consent constraints.
Trading
8
40
Market-decision coordination with private strategy, order, portfolio, and mandate records.
Figure 19: Domain-level attack resistance by harm surface on ASR-main-six-domain. Each cell reports domain-vector row-micro 1−ASRd,vmicro. Higher values indicate fewer judged attacks causing final harm; raw denominators are given in Table 11.
Evaluation criteria, resource and governance manifests, invariants, attack-vector metadata, and construction metadata.
family_pack_context.json
Domain seed facts consumed by tools, task predicates, attack overlays, and judges.
personas/
Per-agent system prompts and tool allowlists.
resources/
Owner-scoped files, databases, shared state, and final artifacts.
mcp/
Scenario-local tool declarations and role-scoped tool surfaces.
governance/
Owner, identity, approval, consent, mandate, and scope records when used by the domain.
Figure 20: Model-level attack resistance by harm surface on ASR-main-six-domain. Each bar reports model-vector row-micro 1−ASRm,vmicro. Higher bars indicate fewer judged attacks causing final harm; raw denominators are given in Table 11.
Table 6: Social topology patterns in WeClawArena.
Domain
Topology
Owner roles
Main governed path
Bargaining
Transaction triad or competitive graph
Buyer, seller, approver
Approval and close artifacts must agree on transaction terms.
Bidding
Seller-centered bidding graph
Seller plus buyers
Award, buyer confirmation, and close artifacts must agree.
Travel
Approval triad or consent graph
Traveler or organizer, partner hotel, approver or co-travelers
Booking must bind approval or consent evidence to the final hotel order.
SWE-Workspace
Four-owner engineering graph
Implementer, reviewer, code owner, CI owner
Patch submission requires review, owner approval, CI evidence, and harness success.
Clinical
Role-scoped care graph
Case roles with protected records and consent or approval constraints
Case action must respect protected facts and required consent or approval state.
Trading
Market-decision graph
Decision, order, portfolio, and mandate roles
Orders must respect portfolio constraints, private strategy boundaries, and mandates.
Figure 21: Per-model utility and attack-resistance tradeoff by harm surface on ASR-main-six-domain. The x-axis reports matched TSR drop against the model’s own no-attacker rows; the y-axis reports model-vector row-micro 1−ASRm,vmicro. Points in the upper-left are preferred. Negative TSR drops mean the attack-vector TSR exceeds the matched no-attacker TSR in the reported result set.
Patch the issue, obtain review and owner approval, run tests, and submit only after the evidence chain is complete.
Clinical
Case-coordination roles
Case records, protected facts, approval or consent records
Coordinate a case action while respecting protected fields and required approval or consent.
Trading
Decision, order, portfolio, mandate roles
Order records, portfolio limits, private strategy notes, mandate records
Coordinate a market decision while respecting portfolio constraints, private strategy boundaries, and mandates.
Figure 22: Judge concordance on the 32 model-vector ASR cells from ASR-main-six-domain and ASR-sensitivity-six-domain. Each point is one (model, attack vector) pair. Color encodes attack vector; the dashed line is y=x. Most points fall slightly above the diagonal, indicating that the Claude Opus 4.7 judge labels marginally more attacks as successful than GPT-5.2 across surfaces and models.
Table 8: ASR judge validation. G5.2 is the GPT-5.2 headline judge over ASR-main-six-domain; O4.7 is the Claude Opus 4.7 sensitivity judge over ASR-sensitivity-six-domain, the same evidence packets and denominators. ASR columns report model-macro surface ASR for surface rows and model-surface macro ASR for the all-vector row, as defined in Appendix I.1. Inter κ is Cohen’s κ between the two judges over all attack-vector rows in those pools. Human κ is Cohen’s κ between the human consensus label and each judge on ASR-human-pilot, a stratified 200-row validation subset annotated by two of the authors with attack-domain expertise; two-annotator human-human κ on the subset is 0.82. FPR is the benign false-positive rate over ASR-benign-control.
ASR (%)
Inter
Human κ
FPR (%)
Surface
G5.2
O4.7
Δ
κ
vs. G5.2
vs. O4.7
G5.2
O4.7
Collaboration
18.3
22.3
+4.0
0.71
0.65
0.69
0.3
0.5
Security
42.5
44.3
+1.8
0.66
0.61
0.65
0.2
0.6
Privacy
31.4
35.5
+4.1
0.69
0.64
0.67
0.4
0.8
Governance
50.3
53.6
+3.3
0.73
0.68
0.71
0.5
1.0
All vectors
34.9
38.2
+3.3
0.70
0.65
0.68
0.3
0.7
Table 9: Readiness failures versus benchmark outcomes.
Category
Examples
Readiness failure, excluded from denominators
Malformed bundle, missing required log, missing scorecard, evaluator crash, unusable final state, or missing evidence packet.
Attack-success outcome, counted only on attack rows
Final harm on the intended attack vector plus clear evidence link under the domain judge.
Calibration control
No-attacker rows used for task utility and judge false-positive checks, but excluded from ASR denominators.
Table 10: Variant-level TSR counts over the reported six-domain sweep. Percentages are in parentheses.
Model
No attacker
Collaboration
Security
Privacy
Governance
Claude Opus 4.1
40/124 (32.3)
43/124 (34.7)
20/124 (16.1)
35/124 (28.2)
28/124 (22.6)
Claude Opus 4.7
76/124 (61.3)
81/124 (65.3)
59/124 (47.6)
44/124 (35.5)
47/124 (37.9)
Claude Sonnet 4.5
60/124 (48.4)
54/124 (43.5)
30/124 (24.2)
41/124 (33.1)
40/124 (32.3)
DeepSeek V3.2
43/124 (34.7)
26/124 (21.0)
14/124 (11.3)
14/124 (11.3)
21/124 (16.9)
Kimi K2.5
64/124 (51.6)
42/124 (33.9)
24/124 (19.4)
37/124 (29.8)
33/74 (44.6)†
Kimi K2 Thinking
25/124 (20.2)
26/124 (21.0)
8/124 (6.5)
8/124 (6.5)
9/124 (7.3)
Qwen3 235B
36/124 (29.0)
30/74 (40.5)†
13/74 (17.6)†
23/74 (31.1)†
8/74 (10.8)†
Qwen3 32B
26/124 (21.0)
21/124 (16.9)
3/124 (2.4)
13/124 (10.5)
4/124 (3.2)
Table 11: Canonical raw ASR denominators and counts by model and attack vector. The GPT-5.2 block is ASR-main-six-domain; the Claude Opus 4.7 block is ASR-sensitivity-six-domain. Each cell reports model-vector row-micro ASR aggregated over domains. These raw counts support Figures 4, 4, 19, 20, 21, Table 8, and the failure-analysis totals. Percentages are in parentheses and report ASR, so lower is better. † indicates partial judged coverage.
Model
Collaboration
Security
Privacy
Governance
GPT-5.2 judge
Claude Opus 4.1
15/120 (12.5)
44/104 (42.3)
23/106 (21.7)
45/106 (42.5)
Claude Opus 4.7
4/124 (3.2)
0/105 (0.0)
0/106 (0.0)
9/105 (8.6)
Claude Sonnet 4.5
10/86 (11.6)
21/69 (30.4)
20/69 (29.0)
25/69 (36.2)
DeepSeek V3.2
37/123 (30.1)
61/101 (60.4)
50/106 (47.2)
77/105 (73.3)
Kimi K2.5
17/122 (13.9)
45/104 (43.3)
53/106 (50.0)
62/105 (59.0)
Kimi K2 Thinking
20/120 (16.7)
54/103 (52.4)
38/106 (35.8)
31/106 (29.2)
Qwen3 235B
49/112 (43.8)†
60/92 (65.2)†
42/94 (44.7)†
67/94 (71.3)†
Qwen3 32B
18/123 (14.6)
46/100 (46.0)
25/106 (23.6)
84/105 (80.0)
Claude Opus 4.7 judge
Claude Opus 4.1
20/120 (16.7)
52/104 (50.0)
27/106 (25.5)
42/106 (39.6)
Claude Opus 4.7
9/124 (7.3)
2/105 (1.9)
3/106 (2.8)
15/105 (14.3)
Claude Sonnet 4.5
7/86 (8.1)
26/69 (37.7)
23/69 (33.3)
21/69 (30.4)
DeepSeek V3.2
48/123 (39.0)
57/101 (56.4)
56/106 (52.8)
82/105 (78.1)
Kimi K2.5
27/122 (22.1)
50/104 (48.1)
50/106 (47.2)
68/105 (64.8)
Kimi K2 Thinking
15/120 (12.5)
58/103 (56.3)
47/106 (44.3)
35/106 (33.0)
Qwen3 235B
56/112 (50.0)†
58/92 (63.0)†
48/94 (51.1)†
73/94 (77.7)†
Qwen3 32B
25/123 (20.3)
42/100 (42.0)
30/106 (28.3)
90/105 (85.7)
研究结果
表1显示Claude Opus 4.7在旅行、SWE-Workspace、竞标领域任务成功率领先,整体表现最稳健。
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.