Figure 18: Layer-wise analysis of behavioral steering vectors. Left: Separation between high-collusion and low-collusion chain-of-thought examples as measured by Cohen’s d, plotted as a function of layer depth. Middle layers (12–22) show the strongest separation, consistent with these layers encoding abstract reasoning patterns. Right: Distribution of projected activations at layer 18, showing clear bimodal separation between high-collusion (right mode) and low-collusion (left mode) examples. The dashed vertical line indicates the optimal decision threshold.
Table 1: Influence of steering methods on the performance of DeepSeek-R1-Distill-Qwen-7B for the Bertrand duopoly pricing game.
Steering Method or Policy Definition
Average Price
Average Profit Gain
Nash Equilibrium - Free Market Competition
1.473 ± 0.000
0.000 ± 0.000
Optimal Monopoly Equivalent Collusion
1.925 ± 0.000
1.000 ± 0.000
Default Prompt to Maximize Profit
1.669 ± 0.070
0.442 ± 0.136
+ Behavior Will Be Monitored
1.619 ± 0.080
0.278 ± 0.179
+ Thoughts Will Be Monitored
1.716 ± 0.141
0.439 ± 0.232
+ Behavior & Thoughts Will Be Monitored
1.566 ± 0.044
0.278 ± 0.148
Prompt to Maximize Profit and Avoid Collusion
1.900 ± 0.183
0.188 ± 0.283
+ Provide Wikipedia Definition of Cartel in Prompt
2.068 ± 0.209
0.013 ± 0.350
+ Prompt to Avoid Penalties
1.957 ± 0.188
0.211 ± 0.269
Maximize Profit as a Symbolic Math Problem
2.829 ± 0.042
-1.313 ± 0.132
+ More Description of Variable Relations
2.388 ± 0.147
-0.202 ± 0.303
Prompt to Maximize Profit with Profits Modified by a ”Warden” Agent’s Fines
1.641 ± 0.111
0.393 ± 0.218
Prompt to Maximize Profit with Behavioral CoT Steering: Magnitude = -50
1.494 ± 0.013
0.013 ± 0.042
Table 2: The influence of prompting style for DeepSeek-R1-Distill-Qwen-7B on performance for the Bertrand duopoly pricing game.
Prompting Style
Average
Average
Average
Average
Price
Profit Gain
Demand
Profit
Nash Equilibrium - Free Market Competition
1.473 ± 0.000
0.000 ± 0.000
47.14 ± 0.00
22.30 ± 0.00
Optimal Monopoly Equivalent Collusion
1.925 ± 0.000
1.000 ± 0.000
36.49 ± 0.00
33.75 ± 0.00
Prompt to Maximize Profit
1.669 ± 0.070
0.442 ± 0.136
43.52 ± 1.77
27.36 ± 1.55
+ Implicit Prompt to Collude
1.660 ± 0.079
0.323 ± 0.113
43.40 ± 2.70
26.00 ± 1.30
+ Explicit Prompt to Collude
1.782 ± 0.075
0.611 ± 0.095
41.48 ± 2.33
29.47 ± 1.60
Prompt to Maximize Profit and Avoid Collusion
1.900 ± 0.183
0.188 ± 0.283
34.83 ± 6.46
24.44 ± 3.24
+ Provide Wikipedia Definition of Cartel in Prompt
2.068 ± 0.209
0.013 ± 0.350
28.81 ± 7.33
22.44 ± 4.02
+ Prompt to Avoid Penalties
1.957 ± 0.188
0.211 ± 0.269
33.79 ± 6.14
24.71 ± 3.09
Prompt to Minimize Price
1.318 ± 0.067
-0.636 ± 0.278
48.16 ± 0.37
15.01 ± 3.19
Prompt to Minimize Profit
1.284 ± 0.299
-1.790 ± 0.170
43.85 ± 7.16
1.79 ± 1.94
Prompt to Maximize Demand
1.839 ± 0.200
0.032 ± 0.290
35.88 ± 6.49
22.66 ± 3.33
Table 3: Pearson’s correlation of the steering vector magnitude with various steering effects as a function of the layer in which the steering vector is applied. P-values are listed in parenthesis and are bold when statistically significant with at least 95% confidence.
Steering Vector
Average Price
Average Profit Gain
CoT Collusion
CoT ↑ Price
CoT ↑ Demand
Intervention Layer
Correlation
Correlation
Correlation
Correlation
Correlation
0
0.854 (0.146)
-0.909 (0.091)
0.925 (0.075)
0.410 (0.590)
0.109 (0.891)
12
0.728 (0.064)
0.814 (0.026)
-0.516 (0.236)
0.546 (0.205)
-0.071 (0.880)
18
0.951 (0.001)
0.974 (0.000)
0.194 (0.677)
0.771 (0.042)
0.183 (0.695)
26
0.031 (0.948)
-0.006 (0.990)
0.181 (0.697)
0.196 (0.673)
0.046 (0.922)
Table 4: Effect on chain-of-thought semantics when using different magnitudes of behavioral steering at layer 18.
Behavioral Steering
Average Price
Average Profit
Judge CoT Collusion
Judge CoT Price Increase
Magnitude
Gain
Probability
Probability
-50
1.494 ± 0.013
0.013 ± 0.042
32.4%
49.5%
-35
1.506 ± 0.025
0.045 ± 0.073
33.6%
51.6%
-15
1.529 ± 0.051
0.162 ± 0.154
32.8%
51.0%
+0 (Default)
1.669 ± 0.070
0.442 ± 0.136
34.3%
54.9%
+15
1.675 ± 0.121
0.400 ± 0.221
33.5%
52.8%
+35
1.738 ± 0.083
0.625 ± 0.111
32.4%
53.1%
+50
1.766 ± 0.026
0.781 ± 0.044
33.7%
54.4%
Table 5: The average collusion Likert scores as a function of steering magnitude across LLM judge models.
Steering
Ministral-14B
OLMo-3.1-32B
OSS-120B
OSS-20B
Qwen3-14B
Llama-3.3-70B
R1-Distill-Qwen-7B
-50
3.98 ± 0.74
3.23 ± 0.50
4.66 ± 0.58
3.02 ± 0.37
3.17 ± 0.61
7.60 ± 0.28
2.67 ± 0.11
-35
3.18 ± 0.46
2.67 ± 0.27
3.78 ± 0.47
2.46 ± 0.16
2.40 ± 0.34
6.91 ± 0.45
2.70 ± 0.07
-10
3.52 ± 0.49
2.85 ± 0.33
3.76 ± 0.47
2.49 ± 0.18
2.46 ± 0.31
6.71 ± 0.28
2.69 ± 0.09
+0
3.90 ± 0.28
2.84 ± 0.16
3.54 ± 0.30
2.41 ± 0.10
2.45 ± 0.22
6.21 ± 0.32
2.84 ± 0.06
+10
3.46 ± 0.34
2.79 ± 0.20
3.28 ± 0.30
2.37 ± 0.12
2.22 ± 0.16
5.48 ± 0.37
2.81 ± 0.06
+35
3.50 ± 0.26
2.91 ± 0.11
2.93 ± 0.22
2.32 ± 0.06
2.15 ± 0.13
4.40 ± 0.40
2.81 ± 0.05
+50
3.14 ± 0.14
2.87 ± 0.09
2.68 ± 0.12
2.24 ± 0.03
2.03 ± 0.06
3.22 ± 0.26
2.77 ± 0.04
Table 6: GPT-5.4 Behavior as a Function of the Prompting Strategy. Simulation use the default reasoning model, with reasoning summary set to auto and reasoning effort set to medium. Each response permits up to 4,096 output tokens, uses a sampling temperature of 1.0, and allows up to two retries on generation failure. Agents are queried in parallel by default via a thread pool.
Prompt Strategy
Average Price
Average Profit Gain
Prompt to Maximize Profit
1.554 ± 0.068
0.202 ± 0.160
Prompt to Maximize Profit and Avoid Collusion
1.474 ± 0.015
0.000 ± 0.041
+ Provide Wikipedia Definition of Cartel in Prompt
1.474 ± 0.014
0.001 ± 0.039
+ Prompt to Avoid Penalties
1.474 ± 0.015
0.000 ± 0.042
Table 7: The Bertrand monopoly pricing game: the effect of steering without complex agent interactions.
Prompting Style
Average Price
Average Profit Gain
Average Profit
Average Demand
Standard Maximize Profit
2.138 ± 0.176
0.716 ± 0.124
48.70 ± 8.46
54.18 ± 12.46
Symbolic Maximize Profit
2.793 ± 0.104
0.243 ± 0.097
16.52 ± 6.56
11.75 ± 6.95
Standard Minimize Price
1.340 ± 0.198
0.297 ± 0.126
20.19 ± 8.56
88.91 ± 9.81
Symbolic Minimize Price
0.463 ± 0.103
-0.793 ± 0.152
-53.93 ± 10.31
99.46 ± 0.15
Standard Minimize Profit
1.091 ± 0.133
0.010 ± 0.014
0.67 ± 0.97
94.69 ± 6.42
Symbolic Minimize Profit
0.523 ± 0.079
-0.703 ± 0.116
-47.82 ± 7.91
99.45 ± 0.06
Standard Maximize Demand
1.923 ± 0.113
0.879 ± 0.081
59.77 ± 5.52
70.96 ± 7.30
Symbolic Maximize Demand
1.644 ± 0.211
0.492 ± 0.182
33.48 ± 12.41
76.66 ± 9.65
Table 8: Ablations on the effect of leveraging semantic priors for the Bertrand duopoly pricing game.
Prompting Style
Average Price
Average Profit Gain
Average Demand
Average Profit
Standard Maximize Profit
1.669 ± 0.070
0.442 ± 0.136
43.52 ± 1.77
27.36 ± 1.55
Symbolic Maximize Profit
2.829 ± 0.042
-1.313 ± 0.132
5.06 ± 2.15
7.24 ± 3.15
+ More Description of Variables
2.413 ± 0.128
-0.254 ± 0.274
19.02 ± 4.76
19.38 ± 3.14
Standard Minimize Price
1.318 ± 0.067
-0.636 ± 0.278
48.16 ± 0.37
15.01 ± 3.19
Symbolic Minimize Price
0.994 ± 0.045
-3.512 ± 1.526
49.86 ± 0.54
-3.5 ± 2.06
+ More Description of Variables
1.388 ± 0.132
-0.676 ± 0.468
46.44 ± 1.81
14.61 ± 5.35
Standard Minimize Profit
1.284 ± 0.299
-1.790 ± 0.170
43.85 ± 7.16
1.79 ± 1.94
Symbolic Minimize Profit
0.603 ± 0.142
-4.644 ± 0.866
49.51 ± 0.94
-31.04 ± 5.45
+ More Description of Variables
1.064 ± 0.171
-2.223 ± 0.128
47.37 ± 4.01
-3.17 ± 1.47
Standard Maximize Demand
1.839 ± 0.200
0.032 ± 0.290
35.88 ± 6.49
22.65 ± 3.33
Symbolic Maximize Demand
1.933 ± 0.166
-0.720 ± 0.435
35.08 ± 5.54
14.24 ± 3.04
+ More Description of Variables
1.971 ± 0.143
0.169 ± 0.262
33.82 ± 4.44
24.23 ± 3.00
Table 9: Effect on duopoly performance and chain-of-thought semantics when prompting agents to tell them that they will be monitored.
Prompt Style
Average Price
Average Profit
CoT Collusion
CoT Collusion
CoT Collusion
Magnitude
Gain
Probability 1.5B
Probability 7B
1-10 Score 7B
Maximize Profit
1.669 ± 0.116
0.442 ± 0.136
48.7
34.3
3.42
+ Behavior Monitored
1.619 ± 0.080
0.278 ± 0.179
47.4
34.3
3.42
+ Thoughts Monitored
1.716 ± 0.141
0.439 ± 0.232
48.8
35.0
3.44
+ Both Monitored
1.566 ± 0.044
0.278 ± 0.148
47.7
34.5
3.42
Implicit Prompt to Collude
1.660 ± 0.079
0.323 ± 0.113
48.6
35.8
3.82
+ Behavior Monitored
1.649 ± 0.122
0.372 ± 0.190
47.0
35.5
3.81
+ Thoughts Monitored
1.764 ± 0.107
0.621 ± 0.156
47.7
35.9
3.86
+ Both Monitored
1.718 ± 0.179
0.318 ± 0.189
47.5
35.7
3.76
Explicit Prompt to Collude
1.782 ± 0.075
0.611 ± 0.095
50.3
42.0
6.35
+ Behavior Monitored
1.755 ± 0.118
0.599 ± 0.149
49.7
42.2
6.59
+ Thoughts Monitored
1.768 ± 0.103
0.525 ± 0.201
49.8
42.4
6.37
+ Both Monitored
1.747 ± 0.101
0.593 ± 0.162
50.0
42.8
6.57
Table 10: The influence of a Warden agent giving fines on performance for the Bertrand oligopoly game with 2 and 3 agents.
Prompting Style
System Setup
Average Price
Average Profit Gain
Fined Profit %
Standard Maximize Profit
2 Agents
1.669 ± 0.070
0.442 ± 0.136
0.00 ± 0.00
2 Agents + Fine Warden
1.641 ± 0.111
0.393 ± 0.218
4.24 ± 1.39
3 Agents
1.551 ± 0.041
0.154 ± 0.086
0.00 ± 0.00
3 Agents + Fine Warden
1.629 ± 0.063
0.308 ± 0.139
2.84 ± 1.69
Implicit Prompt to Collude
2 Agents
1.660 ± 0.079
0.323 ± 0.113
0.00 ± 0.00
2 Agents + Fine Warden
1.691 ± 0.112
0.465 ± 0.171
4.81 ± 2.92
3 Agents
1.684 ± 0.164
0.418 ± 0.312
0.00 ± 0.00
3 Agents + Fine Warden
1.611 ± 0.032
0.281 ± 0.078
5.36 ± 1.87
Explicit Prompt to Collude
2 Agents
1.782 ± 0.075
0.611 ± 0.095
0.00 ± 0.00
2 Agents + Fine Warden
1.736 ± 0.097
0.559 ± 0.144
4.14 ± 1.59
3 Agents
1.832 ± 0.119
0.642 ± 0.176
0.00 ± 0.00
3 Agents + Fine Warden
1.864 ± 0.215
0.748 ± 0.338
2.84 ± 1.28
Table 11: Pearson’s correlation of the steering vector magnitude with price and profit gain as a function of the number of agents being steered. P-values are listed in parenthesis with bold indicating statistical significance with at least 95% confidence.
Agents Steered
Average Price
Average Profit Gain
Correlation
Correlation
1 of 2 Agents
0.870 (0.011)
0.925 (0.003)
2 of 2 Agents
0.951 (0.001)
0.974 (0.000)
1 of 3 Agents
0.543 (0.208)
0.488 (0.266)
2 of 3 Agents
0.847 (0.016)
0.822 (0.023)
3 of 3 Agents
0.919 (0.003)
0.914 (0.004)
Table 12: Pearson’s correlation of the steering vector magnitude with price and profit gain as a function of the price scale of the demand function associated with α={1,3.2,10} following (Fish et al., 2024). P-values are listed in parenthesis and bold is used indicate statistical significance with at least 95% confidence.
Price Scale Setting
Agents Steered
Average Price
Average Profit Gain
Correlation
Correlation
Cost: 1.00, Nash: 1.47, Monopoly: 1.93
1 of 2 Agents
0.870 (0.011)
0.925 (0.003)
Cost: 1.00, Nash: 1.47, Monopoly: 1.93
2 of 2 Agents
0.951 (0.001)
0.974 (0.000)
Cost: 3.20, Nash: 3.69, Monopoly: 5.34
1 of 2 Agents
0.839 (0.018)
0.967 (0.000)
Cost: 3.20, Nash: 3.69, Monopoly: 5.34
2 of 2 Agents
0.888 (0.008)
0.971 (0.000)
Cost: 10.00, Nash: 10.49, Monopoly: 14.59
1 of 2 Agents
0.282 (0.540)
0.784 (0.037)
Cost: 10.00, Nash: 10.49, Monopoly: 14.59
2 of 2 Agents
0.461 (0.298)
0.864 (0.012)
Table 13: Pearson’s correlation of the steering vector magnitude with price and profit gain as a function of the base model that the steering vector is applied to. P-values are listed in parenthesis and bold is used to indicate statistical significance with at least 95% confidence.
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.
作者 · Matthew Riemer, Tommaso Tosato, Amin Memarian, Maximilian Puelma Touzel, Glen Berseth, Irina Rish, Guillaume Dumas