AI 에이전트에게 진짜 온라인 무역회사를 통째로 맡겨보니, 15개 최신 모델 중 최고와 최악의 최종 자산 차이가 9배까지 벌어졌다
AI 에이전트에게 진짜 온라인 무역회사를 통째로 맡겨보니, 15개 최신 모델 중 최고와 최악의 최종 자산 차이가 9배까지 벌어졌다
Business Arena는 AI 에이전트가 알리바바닷컴의 실제 소싱 데이터와 실제 관세·시장 조건을 바탕으로 국경을 넘나드는 도매 상점을 처음부터 끝까지 운영하게 만드는 시험장이다. 15개 최신 언어모델을 같은 조건에서 평가한 결과 평균 최종 순자산이 최고 20만 달러대에서 최저 2만 달러대까지 9배 차이가 났고, 절반이 넘는 실행에서는 오히려 손실을 봤다. 가장 잘한 모델도 사람이 설계한 전략보다 한참 못 미쳤고, 연구팀은 단순 점수 대신 세부 역량 지표와 행동 단위 분석으로 왜 잘하거나 못했는지까지 추적했다.
METAL MEDIA 해설 도표
Business Arena의 평가 흐름
증거 상태측정 결과가 보고됨
- 실제 데이터로 시장 구성알리바바닷컴 실제 소싱 데이터와 실제 관세·수요 주기로 보정한 국경간 도매 시장을 만든다.
- 에이전트의 장기 운영15개 모델이 60개 이상 도구로 소싱, 가격결정, 광고, 고객응대, 컴플라이언스, 재무를 장기간 스스로 수행한다.
- 사람 설계 전략과 비교결정론적 전문가 전략을 기준선으로 삼아 시장에 실제로 존재했던 기회의 크기를 추정한다.
- 역량별·행동별 진단최종 자산 점수를 세부 역량 지표와 개별 행동 귀속으로 분해해 강점과 약점, 특정 이익·손실의 원인을 짚어낸다.
- 메커니즘 절제 검증방치나 시뮬레이터 허점 이용과 비교해 좋은 성적이 진짜 사업 판단력에서 나온 것인지 확인한다.
무엇을 했나
- 연구팀은 에이전트가 공급자에게서 사입해 구매자에게 판매하는 국경간(cross-border) 상점을 장기간 운영하도록 하는 통제된 환경 Business Arena를 만들었으며, 상품·공급자 조건·가격·최소주문량·리드타임은 알리바바닷컴 실제 매물, 수요 주기·관세 등은 공신력 있는 자료로 보정했다.
- 60개 이상의 도구(시장조사, 소싱, 재고, 가격결정, 광고, 고객응대, 컴플라이언스, 재무 등)를 제공해 에이전트가 사업의 전 과정을 스스로 판단하고 실행하도록 했고, 시장은 에이전트의 행동과 무관하게 경쟁자 재가격, 수요 변화, 공급 차질 등으로 계속 변화한다.
- GPT-5.6 Sol, GPT-5.5, Claude Fable 5, Opus 4.6/4.8, Gemini 3.1 Pro/3.5 Flash, Qwen 3.7 Max 등 상용 모델과 GLM-5.2, Kimi K2.6/K3, DeepSeek V4 Pro, MiniMax M2.5/M3, Qwen-3.8-Max-Preview 등 오픈웨이트 모델 총 15개를 동일한 세계 조건에서 각 10회씩 실행해 비교했다.
- 단순 최종 자산 점수만으로는 왜 성공하거나 실패했는지 알 수 없기 때문에, 사람이 설계한 전략과 비교해 시장에 존재한 기회 규모를 추정하고, 역량별 세부 지표로 강점과 약점을 드러내고, 실제로 발생한 이익과 손실을 그것을 만든 개별 행동까지 되짚어 추적했다.
- 에이전트가 시뮬레이터의 허점을 이용하거나 특정 업무를 방치해서 좋은 성적을 낸 것이 아님을 확인하기 위해, 의도된 행동과 방치·허점이용 정책을 비교하는 메커니즘 절제 실험(mechanism ablation)을 수행했다.

| Mechanism | Intended behavior | Intended policy | Neglect | Shortcut or misuse |
|---|---|---|---|---|
| Portfolio | Use demand to choose products | Evidence-guided +$63.6k | Blind bulk buying -$14.6k | Buy only cheap SKUs 0 |
| Market events | Check signals before investing | Evidence checked +$6.6k | Ignore events -$17.3k | Follow every rumor 0 |
| Pricing | Cover costs while sustaining sales | Full-cost pricing +$50.3k | Price near cost -$58.1k | Extreme markup 0 |
| Tariffs | Include tariffs when choosing markets | Tariff-aware routes +$25.9k | No active routing 0 | Tariff-blind U.S. focus -$3.7k |
| Customer service | Use buyer and product evidence | Evidence-based replies +$5.6k | Ignore inquiries -$0.1k | Generic replies 0 |
| Arena feature and purpose | Agent-facing tools |
|---|---|
| Decision-Making Under Uncertainty | |
| Demand and events. Form market beliefs from structural demand, trends, calendars, public events, and policy shocks. | get_base_demand_intel(), get_trends(), get_calendar(), get_world(), get_tariff_events(). |
| Market feedback and competition. Learn from realized outcomes and observe rival offers without accessing competitors’ private strategies. | get_orders(), get_competition(). |
| Strategic Planning Under Constraints | |
| Shop focus and initial portfolio. Select categories, compare opportunities, and decide how much capital to commit at opening. | get_catalog(), get_store_focus(), set_store_focus(), submit_setup(), skip_setup(). |
| Sourcing and supplier diligence. Search and rank offers, inspect supplier risk, and purchase inventory under cost, MOQ, quality, and lead-time constraints. | get_supplier_catalog(), get_supplier_flags(), buy_supplier(). |
| Capital, inventory, and recovery. Track deployed capital and obligations, finance expansion, and recover capital from weak positions. | get_state(), get_products(), get_payables(), get_loans(), borrow(), repay_loan(), get_factoring(), factor_ar(), liquidate_inventory(). |
| Insight-to-Action Alignment | |
| Pricing and listing. Translate market beliefs and cost calculations into concrete offers across products and segments. | get_listings(), list_product_on(), update_listing(), set_price_tiers(). |
| Tariffs, shipping, and route economics. Calculate the route-specific cost stack and choose viable destinations and commercial terms. | get_platforms(), get_countries(), get_shipping_rules(), get_tariff_table(), set_default_incoterm(). |
| Compliance. Identify market-entry requirements, apply early enough to clear approval lead times, and avoid unauthorized trading. | get_certifications(), get_compliance_status(), apply_certification(). |

| Arena feature and purpose | Agent-facing tools |
|---|---|
| Insight-to-Action Alignment (continued) | |
| Advertising. Allocate demand-generation spend and revise it using observed full-funnel performance. | get_ad_status(), set_ad_budget(). |
| Cooperation & Competition | |
| Customer service and buyer negotiation. Infer buyer needs, answer factual questions, and negotiate bulk transactions while protecting business value. | get_inquiries(), reply_inquiry(), get_rfqs(), respond_rfq(). |
| Returns and disputes. Respond to post-sale problems while managing refund, replacement, and escalation risk. | get_return_requests(), respond_to_return(), get_returns(), dispute_return(), get_disputes(), resolve_dispute(). |
| Supplier relationships and negotiation. Learn counterparty behavior, request better terms, and decide whether to accept supplier offers. | get_supplier_relations(), request_quote(), get_supplier_quotes(), respond_quote(). |
| Competitive response. Compare rival offers and adjust prices or demand-generation decisions as competitors change. | get_competition(), update_listing(), set_price_tiers(), set_ad_budget(). |
| Persistent Operation | |
| Persistent state and daily feedback. Preserve observations and plans, inspect previous decisions, and advance the market after completing the current operating cycle. | write_note(), read_notes(), delete_note(), end_round(). |
| Workspace and automation. Read and revise persistent files, construct reusable analyses, and execute model-authored workflows across business functions. | read(), write(), edit(), exec(), process(). |

| Archetype | Pricing Strategy | Behavior Summary |
|---|---|---|
| price_leader | undercut_median | Targets 95% of competitor median price; cuts further to 92% during peak festivals |
| follower | track_top_3 | Tracks average of 3 cheapest competitors with +2% offset |
| liquidator | aggressive_low | Targets 78% of competitor median; clearance pricing at 70% during festival endings |
| opportunist | dynamic_demand | Raises price when trailing demand exceeds 1.2× baseline; heavy pre-festival stocking |
| premium | cost-plus | Prices at 3× cost basis; EU suppliers; holds firm during festivals |
| cross_border | cost-plus | Prices at 2.5× cost basis; EU suppliers; stable across phases |
| wholesale | cost-plus | Prices at 1.3× cost basis; large inventory (35-day target), volume-driven |
| event_sniper | cost-plus | Prices at 1.65× cost basis normally; spikes to 2.5× during peak (narrow 5-SKU catalog) |
| new_entrant | undercut_until_orders | Extreme discounts (82% of median) until 50 orders, then switches to 1.05× cost basis |
| dormant | static | Never reprices; decays 5%/day after 10-day no-sale grace period |

| Financial factor | Current rule | Rationale |
|---|---|---|
| Operating drag | $200 fixed overhead per day, plus 0.5% of on-hand inventory value. | Penalizes passive operation and slow-moving stock; encourages sufficient throughput, disciplined purchasing, and inventory turnover. |
| Channel economics | Platform commission is generally 5%–12% of gross, with volume discounts; eligible export orders receive a 9% rebate. | Rewards pricing over the complete transaction-cost stack and selecting economically viable markets rather than maximizing gross revenue alone. |
| Short-term loans | Interest compounds at 0.15% per day and rises to 2.5× the normal rate when overdue. | Enables expansion when profitable opportunities exist, but penalizes borrowing without sufficiently fast and reliable capital recovery. |
| Supplier payables | Overdue balances accrue 0.10% per day, capped at 30% of the original invoice, and remain liabilities. | Rewards planning around payment deadlines; penalizes sourcing commitments that the agent cannot finance. |
| Invoice factoring | Eligible receivables can be converted to cash at a 5%–18% discount determined by maturity and buyer credit. | Lets agents accelerate cash recycling, while charging explicitly for liquidity obtained before customer payment. |
| Inventory liquidation | Inventory can be converted immediately to cash at 85% of cost basis. | Provides a controlled way to exit bad positions and redeploy capital, but preserves a meaningful loss so that poor sourcing is not costless. |
| Final settlement | Escrow and receivables recover at 97%, inventory at 85%, and liabilities remain at face value. | Penalizes unfinished operating cycles and rewards converting inventory and receivables into cash before the episode ends. |

| Model | Mean | Within-model SD | 95% CI | Capital preserved |
|---|---|---|---|---|
| Gemini 3.1 Pro | $188,488 | $66,641 | [$140,816, $236,160] | 9/10 |
| GPT-5.6 Sol | $168,867 | $47,185 | [$135,113, $202,620] | 10/10 |
| Fable 5 | $164,204 | $34,141 | [$139,781, $188,627] | 10/10 |
| Gemini 3.5 Flash | $125,952 | $28,321 | [$105,692, $146,212] | 10/10 |
| GPT-5.5 | $117,481 | $18,816 | [$104,021, $130,941] | 10/10 |
| Kimi K3 | $112,278 | $42,644 | [$81,773, $142,784] | 8/10 |
| Opus 4.8 | $93,946 | $36,729 | [$67,672, $120,220] | 4/10 |
| Opus 4.6 | $93,066 | $31,759 | [$70,347, $115,785] | 6/10 |
| Qwen 3.8 Max | $89,423 | $50,974 | [$52,958, $125,887] | 3/10 |
| GLM 5.2 | $55,742 | $30,921 | [$33,623, $77,862] | 1/10 |
| Kimi K2.6 | $52,533 | $27,568 | [$32,812, $72,254] | 1/10 |
| Qwen 3.7 Max | $47,956 | $31,145 | [$25,676, $70,236] | 0/10 |
| MiniMax M3 | $43,064 | $35,218 | [$17,871, $68,258] | 1/10 |
| DeepSeek V4 Pro | $40,804 | $46,130 | [$7,805, $73,803] | 1/10 |
| MiniMax M2.5 | $20,856 | $27,897 | [$900, $40,813] | 0/10 |

| Component | Strategy coverage | Connection to the decision system |
|---|---|---|
| Evidence and memory | Market signals, events, competition, supplier offers, route costs, firm state, and realized operating history | Normalizes current public evidence and retains bounded observations and prior plans for the next cycle. |
| Opportunity beliefs | Market-depth, trend- and event-aware, realized-velocity, unit-economics, inventory-risk, and customer-value signals | Converts heterogeneous evidence into demand, route, inventory, and customer beliefs consumed by the portfolio planner. |
| Capital and portfolio | Diversified probing, conviction-weighted deployment, broad velocity portfolios, compact capital-efficient books, adaptive focus, and recovery phases | Sets the cash reserve, risk budget, portfolio scope, and per-opportunity allocation that constrain all downstream spending. |
| Sourcing | Cost-balanced, fast-turn, quality-led, risk-adjusted, and relationship-aware supplier selection | Selects supplier, quantity, timing, and terms within the portfolio allocation; fulfillment and realized quality revise supplier eligibility. |
| Pricing and listing | Margin-preserving, competition-aware, volume-oriented, and premium offer policies with route-specific cost floors | Translates sourced inventory and route economics into viable offers; conversion, margin, and inventory age update future prices. |
| Compliance | Route checks, permit application, temporary listing pauses, and reopening after approval | Gates sourcing and listing plans before trade; permit status, violations, and fines feed operational reliability. |
| Ads and promotion | Bounded experimentation, test–scale–stop rules, event-timed promotion, and recovery shutdown | Operates only on viable, inventory-backed offers; full-funnel returns affect advertising and replenishment decisions. |
| Customer operations | Inquiry and RFQ handling, fulfillment-aware responses, conservative negotiation, and, where enabled, retention-oriented CRM | Converts incoming demand while protecting feasibility and contribution; service outcomes update customer-value and demand beliefs. |
| Finance and recovery | Cash reserves, evidence-gated borrowing and repayment, position limits, and stale-inventory liquidation | Expands deployment only when supported by visible economics and returns capital when continued ownership is no longer justified. |
| Feedback control | Sell-through learning, winner scaling, loser pauses, route adaptation, focus revision, and portfolio rebalancing | Routes realized orders, margins, stock, advertising, service, and failures back to both beliefs and capital allocation. |

실제로 확인된 결과
- 15개 모델의 평균 최종 순자산은 최고 188,488달러(Gemini 3.1 Pro)에서 최저 20,856달러(MiniMax M2.5)로 9.0배 차이가 났고, 시작 자본 8만 달러를 기준으로 전체 실행의 51%가 손실로 끝났으며 시작 자본을 모든 실행에서 지켜낸 모델은 4개뿐이었다.
- 가장 강한 사람 설계 전략은 같은 시장 조건에서 436,195달러를 달성해 최고 성적 모델 평균의 두 배를 넘었으며, 이는 시장 근거와 소싱·가격·컴플라이언스·자본 배분을 함께 조율한 결과였다.
- 10회 실행 평균의 신뢰도는 ICC=0.944로 안정적이었고, 5회씩 나눈 부분집합도 순위를 대체로 보존(ρ=0.898)하며 98.7%의 경우 같은 상위 그룹을 재현했다.
- Gemini 3.1 Pro, GPT-5.6 Sol, Fable 5는 각각 199%, 167%, 150%의 누적 자본가동률을 보이며 재고회전율도 1.0에 근접했지만, Qwen 3.7 Max는 시작 자본의 22.3%만 투입했고 MiniMax M3는 재고회전율이 0.36에 그쳐 자본이 재고에 묶였다.
- 전문가 설계 전략들은 46~61%의 판매완료율(sell-through)에서도 58~73%의 주문마진을 지킨 반면, GPT-5.6 Sol과 Opus 4.6은 94~95%의 판매완료율을 보였지만 마진은 35~39%에 그쳐 최종 자산에서 최상위 전문가 전략들에 미치지 못했고, 모든 전문가 전략은 컴플라이언스 위반으로 인한 벌금이 0이었다.

어디에 쓸 수 있나
- 전자상거래·B2B 소싱 자동화 도구를 개발할 때 에이전트가 어느 단계(소싱, 가격결정, 고객응대, 컴플라이언스)에서 취약한지 미리 점검하는 진단 프레임워크로 참고할 수 있다.
- 장기간 지연된 피드백과 변화하는 환경을 다루는 다른 에이전트 평가(재고관리, 광고 예산 배분 등)를 설계할 때 방법론적 참고 사례로 활용할 수 있다.
- 체크포인트-분기-재개(save-fork-load) 방식의 상태 저장형 평가 기법은 특정 의사결정 지점에서 여러 대안을 비교하는 테스트 설계에 응용할 수 있다.

한계와 남은 검증
- 평가된 것은 알리바바닷컴 데이터와 미중 관세 변화 등으로 보정된 특정 시뮬레이션 시장 조건이며, 실제 살아있는 시장이나 다른 업종·지역에 그대로 일반화된다는 근거는 제시되지 않았다.
- 가장 잘한 모델도 사람이 설계한 최상위 전략의 절반 수준에 그쳐, 현재 에이전트가 실제 사업을 안정적으로 운영할 수 있다고 보기는 어렵다.
- 적용처로 제시한 항목들은 논문이 실제로 검증한 성능이 아니라 이 평가 체계가 시사하는 가능성이며, 실제 자동화 도구에 적용했을 때의 성과는 별도로 측정되지 않았다.
- 트레이스 서치(trace search)를 통한 테스트 시점 연산 배분 실험은 Qwen 3.8 Max Preview 한 모델에 대해서만 수행되어 다른 모델로의 일반화는 확인되지 않았다.

왜 중요한가
실제 사업을 통째로 운영시켜 보기 전에 이런 통제된 가상 시장에서 먼저 시험해 보면, 실제 자본을 위험에 빠뜨리지 않으면서 에이전트가 어디에서 강하고 어디에서 무너지는지 미리 파악할 수 있다. 자동화된 매장 운영, 소싱 대행, 가격 결정 도구 등을 만드는 사람에게는 어떤 역량이 아직 부족한지 구체적으로 보여주는 참고자료가 된다.

이 논문의 용어
- Business Arena · AI 에이전트가 국경간 도매상점을 장기간 스스로 운영하도록 만든 통제된 시뮬레이션 평가 환경
- 메커니즘 절제 실험(mechanism ablation) · 특정 요소를 방치하거나 악용했을 때와 비교해 좋은 성적이 진짜 실력 때문인지 확인하는 실험 방법
- 역량별 세부 지표(skill-level metrics) · 최종 성과 점수 하나로는 안 보이는, 재고관리·가격결정·고객응대 등 개별 업무 능력을 따로 측정하는 지표
- 행동 단위 귀속(action-level attribution) · 발생한 이익이나 손실을 어떤 구체적 행동(소싱, 가격결정 등)이 만들었는지 되짚어 연결하는 분석
- MCP 호출 · 모델이 도구를 사용할 때 추론과 행동을 번갈아 하도록 정형화된 방식으로 도구를 부르는 규격
최신 논문
- AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다
- 논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- 고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법
- 로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다
- 에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법
- AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다
- AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다
- AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다
METAL MEDIA 최신 기사
그림 출처: Yijun Pan et al., arXiv:2608.08621, CC BY-SA 4.0