컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI 에이전트에게 진짜 온라인 무역회사를 통째로 맡겨보니, 15개 최신 모델 중 최고와 최악의 최종 자산 차이가 9배까지 벌어졌다

arXiv:2608.086212026-08-08

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

AI 에이전트에게 진짜 온라인 무역회사를 통째로 맡겨보니, 15개 최신 모델 중 최고와 최악의 최종 자산 차이가 9배까지 벌어졌다

Business Arena는 AI 에이전트가 알리바바닷컴의 실제 소싱 데이터와 실제 관세·시장 조건을 바탕으로 국경을 넘나드는 도매 상점을 처음부터 끝까지 운영하게 만드는 시험장이다. 15개 최신 언어모델을 같은 조건에서 평가한 결과 평균 최종 순자산이 최고 20만 달러대에서 최저 2만 달러대까지 9배 차이가 났고, 절반이 넘는 실행에서는 오히려 손실을 봤다. 가장 잘한 모델도 사람이 설계한 전략보다 한참 못 미쳤고, 연구팀은 단순 점수 대신 세부 역량 지표와 행동 단위 분석으로 왜 잘하거나 못했는지까지 추적했다.

METAL MEDIA 해설 도표

Business Arena의 평가 흐름

증거 상태측정 결과가 보고됨

  1. 실제 데이터로 시장 구성알리바바닷컴 실제 소싱 데이터와 실제 관세·수요 주기로 보정한 국경간 도매 시장을 만든다.
  2. 에이전트의 장기 운영15개 모델이 60개 이상 도구로 소싱, 가격결정, 광고, 고객응대, 컴플라이언스, 재무를 장기간 스스로 수행한다.
  3. 사람 설계 전략과 비교결정론적 전문가 전략을 기준선으로 삼아 시장에 실제로 존재했던 기회의 크기를 추정한다.
  4. 역량별·행동별 진단최종 자산 점수를 세부 역량 지표와 개별 행동 귀속으로 분해해 강점과 약점, 특정 이익·손실의 원인을 짚어낸다.
  5. 메커니즘 절제 검증방치나 시뮬레이터 허점 이용과 비교해 좋은 성적이 진짜 사업 판단력에서 나온 것인지 확인한다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 연구팀은 에이전트가 공급자에게서 사입해 구매자에게 판매하는 국경간(cross-border) 상점을 장기간 운영하도록 하는 통제된 환경 Business Arena를 만들었으며, 상품·공급자 조건·가격·최소주문량·리드타임은 알리바바닷컴 실제 매물, 수요 주기·관세 등은 공신력 있는 자료로 보정했다.
  2. 60개 이상의 도구(시장조사, 소싱, 재고, 가격결정, 광고, 고객응대, 컴플라이언스, 재무 등)를 제공해 에이전트가 사업의 전 과정을 스스로 판단하고 실행하도록 했고, 시장은 에이전트의 행동과 무관하게 경쟁자 재가격, 수요 변화, 공급 차질 등으로 계속 변화한다.
  3. GPT-5.6 Sol, GPT-5.5, Claude Fable 5, Opus 4.6/4.8, Gemini 3.1 Pro/3.5 Flash, Qwen 3.7 Max 등 상용 모델과 GLM-5.2, Kimi K2.6/K3, DeepSeek V4 Pro, MiniMax M2.5/M3, Qwen-3.8-Max-Preview 등 오픈웨이트 모델 총 15개를 동일한 세계 조건에서 각 10회씩 실행해 비교했다.
  4. 단순 최종 자산 점수만으로는 왜 성공하거나 실패했는지 알 수 없기 때문에, 사람이 설계한 전략과 비교해 시장에 존재한 기회 규모를 추정하고, 역량별 세부 지표로 강점과 약점을 드러내고, 실제로 발생한 이익과 손실을 그것을 만든 개별 행동까지 되짚어 추적했다.
  5. 에이전트가 시뮬레이터의 허점을 이용하거나 특정 업무를 방치해서 좋은 성적을 낸 것이 아님을 확인하기 위해, 의도된 행동과 방치·허점이용 정책을 비교하는 메커니즘 절제 실험(mechanism ablation)을 수행했다.
Figure 1: Model performance in Business Arena. Over the same long horizon, the strongest models more than double their capital, a middle group earns modest returns, and the weakest ones finish with less than they started.
Figure 1: Model performance in Business Arena. Over the same long horizon, the strongest models more than double their capital, a middle group earns modest returns, and the weakest ones finish with less than they started.
Table 1: Mechanism ablations. For each mechanism, the stronger of the neglect and shortcut policies is set as the baseline. Values report the mean change in final net worth relative to this baseline. The customer-service comparison uses three matched worlds; the remaining financial ladders use ten seeds.
MechanismIntended behaviorIntended policyNeglectShortcut or misuse
PortfolioUse demand to choose productsEvidence-guided +$63.6kBlind bulk buying -$14.6kBuy only cheap SKUs 0
Market eventsCheck signals before investingEvidence checked +$6.6kIgnore events -$17.3kFollow every rumor 0
PricingCover costs while sustaining salesFull-cost pricing +$50.3kPrice near cost -$58.1kExtreme markup 0
TariffsInclude tariffs when choosing marketsTariff-aware routes +$25.9kNo active routing 0Tariff-blind U.S. focus -$3.7k
Customer serviceUse buyer and product evidenceEvidence-based replies +$5.6kIgnore inquiries -$0.1kGeneric replies 0
Figure 2: Overview of arena design. The agent selects markets, purchases inventory, prices and lists products, learns from sales, and adapts its operation. It acts on partial observations while supplier disruptions, competitor repricing, and demand shifts create a changing market. Meanwhile, operational obligations persist and economic feedback remains delayed.
Figure 2: Overview of arena design. The agent selects markets, purchases inventory, prices and lists products, learns from sales, and adapts its operation. It acts on partial observations while supplier disruptions, competitor repricing, and demand shifts create a changing market. Meanwhile, operational obligations persist and economic feedback remains delayed.
Table 2: Mapping from business capabilities to arena features and agent-facing tools. Capability headers form the first layer, arena features describe the corresponding business problems, and the final column lists the tools through which agents gather evidence and act.
Arena feature and purposeAgent-facing tools
Decision-Making Under Uncertainty
Demand and events. Form market beliefs from structural demand, trends, calendars, public events, and policy shocks.get_base_demand_intel(), get_trends(), get_calendar(), get_world(), get_tariff_events().
Market feedback and competition. Learn from realized outcomes and observe rival offers without accessing competitors’ private strategies.get_orders(), get_competition().
Strategic Planning Under Constraints
Shop focus and initial portfolio. Select categories, compare opportunities, and decide how much capital to commit at opening.get_catalog(), get_store_focus(), set_store_focus(), submit_setup(), skip_setup().
Sourcing and supplier diligence. Search and rank offers, inspect supplier risk, and purchase inventory under cost, MOQ, quality, and lead-time constraints.get_supplier_catalog(), get_supplier_flags(), buy_supplier().
Capital, inventory, and recovery. Track deployed capital and obligations, finance expansion, and recover capital from weak positions.get_state(), get_products(), get_payables(), get_loans(), borrow(), repay_loan(), get_factoring(), factor_ar(), liquidate_inventory().
Insight-to-Action Alignment
Pricing and listing. Translate market beliefs and cost calculations into concrete offers across products and segments.get_listings(), list_product_on(), update_listing(), set_price_tiers().
Tariffs, shipping, and route economics. Calculate the route-specific cost stack and choose viable destinations and commercial terms.get_platforms(), get_countries(), get_shipping_rules(), get_tariff_table(), set_default_incoterm().
Compliance. Identify market-entry requirements, apply early enough to clear approval lead times, and avoid unauthorized trading.get_certifications(), get_compliance_status(), apply_certification().
Figure 3: Main leaderboard over 15 model families, averaged across ten runs under the same world condition. Dashed lines denote expert-designed strategies that use only agent-visible information.
Figure 3: Main leaderboard over 15 model families, averaged across ten runs under the same world condition. Dashed lines denote expert-designed strategies that use only agent-visible information.
Table 3: Mapping from business capabilities to arena features and agent-facing tools (continued).
Arena feature and purposeAgent-facing tools
Insight-to-Action Alignment (continued)
Advertising. Allocate demand-generation spend and revise it using observed full-funnel performance.get_ad_status(), set_ad_budget().
Cooperation & Competition
Customer service and buyer negotiation. Infer buyer needs, answer factual questions, and negotiate bulk transactions while protecting business value.get_inquiries(), reply_inquiry(), get_rfqs(), respond_rfq().
Returns and disputes. Respond to post-sale problems while managing refund, replacement, and escalation risk.get_return_requests(), respond_to_return(), get_returns(), dispute_return(), get_disputes(), resolve_dispute().
Supplier relationships and negotiation. Learn counterparty behavior, request better terms, and decide whether to accept supplier offers.get_supplier_relations(), request_quote(), get_supplier_quotes(), respond_quote().
Competitive response. Compare rival offers and adjust prices or demand-generation decisions as competitors change.get_competition(), update_listing(), set_price_tiers(), set_ad_budget().
Persistent Operation
Persistent state and daily feedback. Preserve observations and plans, inspect previous decisions, and advance the market after completing the current operating cycle.write_note(), read_notes(), delete_note(), end_round().
Workspace and automation. Read and revise persistent files, construct reusable analyses, and execute model-authored workflows across business functions.read(), write(), edit(), exec(), process().
Figure 4: Model diagnostic profiles. Models exhibit different strengths across operating fluency, capital deployment, selling, customer interaction, and compliance. Colors indicate cohort-relative performance from weaker to stronger.
Figure 4: Model diagnostic profiles. Models exhibit different strengths across operating fluency, capital deployment, selling, customer interaction, and compliance. Colors indicate cohort-relative performance from weaker to stronger.
Table 4: NPC seller archetypes grouped by pricing family. Top: competitor-aware strategies. Middle: cost-anchored strategies. Bottom: phase-switching and passive. The population is weighted toward high-liquidity archetypes (price_leader, follower, liquidator) that supply everyday buyer demand, with niche archetypes (event_sniper, dormant) appearing rarely. The distribution is fixed across all regions and seeds.
ArchetypePricing StrategyBehavior Summary
price_leaderundercut_medianTargets 95% of competitor median price; cuts further to 92% during peak festivals
followertrack_top_3Tracks average of 3 cheapest competitors with +2% offset
liquidatoraggressive_lowTargets 78% of competitor median; clearance pricing at 70% during festival endings
opportunistdynamic_demandRaises price when trailing demand exceeds 1.2× baseline; heavy pre-festival stocking
premiumcost-plusPrices at 3× cost basis; EU suppliers; holds firm during festivals
cross_bordercost-plusPrices at 2.5× cost basis; EU suppliers; stable across phases
wholesalecost-plusPrices at 1.3× cost basis; large inventory (35-day target), volume-driven
event_snipercost-plusPrices at 1.65× cost basis normally; spikes to 2.5× during peak (narrow 5-SKU catalog)
new_entrantundercut_until_ordersExtreme discounts (82% of median) until 50 orders, then switches to 1.05× cost basis
dormantstaticNever reprices; decays 5%/day after 10-day no-sale grace period
Figure 5: Core operating trade-offs. Successful agents identify opportunities and deploy capital into inventory that will sell (left), then preserve margin without pricing themselves out of the market (right). Crosses denote model-family means and lighter points individual runs.
Figure 5: Core operating trade-offs. Successful agents identify opportunities and deploy capital into inventory that will sell (left), then preserve margin without pricing themselves out of the market (right). Crosses denote model-family means and lighter points individual runs.
Table 5: Financial frictions and capital tools in Business Arena. The paper condition uses a mid-scale shop with $200 daily overhead. The mechanisms penalize idle capital, excess inventory, unfinished transactions, and poorly timed leverage, while allowing agents to pay explicit costs to recover or accelerate capital.
Financial factorCurrent ruleRationale
Operating drag$200 fixed overhead per day, plus 0.5% of on-hand inventory value.Penalizes passive operation and slow-moving stock; encourages sufficient throughput, disciplined purchasing, and inventory turnover.
Channel economicsPlatform commission is generally 5%–12% of gross, with volume discounts; eligible export orders receive a 9% rebate.Rewards pricing over the complete transaction-cost stack and selecting economically viable markets rather than maximizing gross revenue alone.
Short-term loansInterest compounds at 0.15% per day and rises to 2.5× the normal rate when overdue.Enables expansion when profitable opportunities exist, but penalizes borrowing without sufficiently fast and reliable capital recovery.
Supplier payablesOverdue balances accrue 0.10% per day, capped at 30% of the original invoice, and remain liabilities.Rewards planning around payment deadlines; penalizes sourcing commitments that the agent cannot finance.
Invoice factoringEligible receivables can be converted to cash at a 5%–18% discount determined by maturity and buyer credit.Lets agents accelerate cash recycling, while charging explicitly for liquidity obtained before customer payment.
Inventory liquidationInventory can be converted immediately to cash at 85% of cost basis.Provides a controlled way to exit bad positions and redeploy capital, but preserves a meaningful loss so that poor sourcing is not costless.
Final settlementEscrow and receivables recover at 97%, inventory at 85%, and liabilities remain at face value.Penalizes unfinished operating cycles and rewards converting inventory and receivables into cash before the episode ends.
(b) Margin and sell-through.
(b) Margin and sell-through.
Table 6: Repeated-run variation across the 150-run cohort. Confidence intervals use two-sided t-intervals over ten runs. Capital preservation counts runs ending at or above the $80,000 starting capital.
ModelMeanWithin-model SD95% CICapital preserved
Gemini 3.1 Pro$188,488$66,641[$140,816, $236,160]9/10
GPT-5.6 Sol$168,867$47,185[$135,113, $202,620]10/10
Fable 5$164,204$34,141[$139,781, $188,627]10/10
Gemini 3.5 Flash$125,952$28,321[$105,692, $146,212]10/10
GPT-5.5$117,481$18,816[$104,021, $130,941]10/10
Kimi K3$112,278$42,644[$81,773, $142,784]8/10
Opus 4.8$93,946$36,729[$67,672, $120,220]4/10
Opus 4.6$93,066$31,759[$70,347, $115,785]6/10
Qwen 3.8 Max$89,423$50,974[$52,958, $125,887]3/10
GLM 5.2$55,742$30,921[$33,623, $77,862]1/10
Kimi K2.6$52,533$27,568[$32,812, $72,254]1/10
Qwen 3.7 Max$47,956$31,145[$25,676, $70,236]0/10
MiniMax M3$43,064$35,218[$17,871, $68,258]1/10
DeepSeek V4 Pro$40,804$46,130[$7,805, $73,803]1/10
MiniMax M2.5$20,856$27,897[$900, $40,813]0/10
Figure 6: Compliance exposes deployment-critical reliability failures (left), while customer service reveals specialized strengths that do not follow the aggregate leaderboard (right).
Figure 6: Compliance exposes deployment-critical reliability failures (left), while customer service reveals specialized strengths that do not follow the aggregate leaderboard (right).
Table 7: Coverage and coupling of strategy-reserve components. Each component declares the public evidence or upstream plan it consumes, the decision object it produces, and the realized feedback that can revise it.
ComponentStrategy coverageConnection to the decision system
Evidence and memoryMarket signals, events, competition, supplier offers, route costs, firm state, and realized operating historyNormalizes current public evidence and retains bounded observations and prior plans for the next cycle.
Opportunity beliefsMarket-depth, trend- and event-aware, realized-velocity, unit-economics, inventory-risk, and customer-value signalsConverts heterogeneous evidence into demand, route, inventory, and customer beliefs consumed by the portfolio planner.
Capital and portfolioDiversified probing, conviction-weighted deployment, broad velocity portfolios, compact capital-efficient books, adaptive focus, and recovery phasesSets the cash reserve, risk budget, portfolio scope, and per-opportunity allocation that constrain all downstream spending.
SourcingCost-balanced, fast-turn, quality-led, risk-adjusted, and relationship-aware supplier selectionSelects supplier, quantity, timing, and terms within the portfolio allocation; fulfillment and realized quality revise supplier eligibility.
Pricing and listingMargin-preserving, competition-aware, volume-oriented, and premium offer policies with route-specific cost floorsTranslates sourced inventory and route economics into viable offers; conversion, margin, and inventory age update future prices.
ComplianceRoute checks, permit application, temporary listing pauses, and reopening after approvalGates sourcing and listing plans before trade; permit status, violations, and fines feed operational reliability.
Ads and promotionBounded experimentation, test–scale–stop rules, event-timed promotion, and recovery shutdownOperates only on viable, inventory-backed offers; full-funnel returns affect advertising and replenishment decisions.
Customer operationsInquiry and RFQ handling, fulfillment-aware responses, conservative negotiation, and, where enabled, retention-oriented CRMConverts incoming demand while protecting feasibility and contribution; service outcomes update customer-value and demand beliefs.
Finance and recoveryCash reserves, evidence-gated borrowing and repayment, position limits, and stale-inventory liquidationExpands deployment only when supported by visible economics and returns capital when continued ownership is no longer justified.
Feedback controlSell-through learning, winner scaling, loser pauses, route adaptation, focus revision, and portfolio rebalancingRoutes realized orders, margins, stock, advertising, service, and failures back to both beliefs and capital allocation.
(b) Customer-service outcomes.
(b) Customer-service outcomes.

실제로 확인된 결과

  • 15개 모델의 평균 최종 순자산은 최고 188,488달러(Gemini 3.1 Pro)에서 최저 20,856달러(MiniMax M2.5)로 9.0배 차이가 났고, 시작 자본 8만 달러를 기준으로 전체 실행의 51%가 손실로 끝났으며 시작 자본을 모든 실행에서 지켜낸 모델은 4개뿐이었다.
  • 가장 강한 사람 설계 전략은 같은 시장 조건에서 436,195달러를 달성해 최고 성적 모델 평균의 두 배를 넘었으며, 이는 시장 근거와 소싱·가격·컴플라이언스·자본 배분을 함께 조율한 결과였다.
  • 10회 실행 평균의 신뢰도는 ICC=0.944로 안정적이었고, 5회씩 나눈 부분집합도 순위를 대체로 보존(ρ=0.898)하며 98.7%의 경우 같은 상위 그룹을 재현했다.
  • Gemini 3.1 Pro, GPT-5.6 Sol, Fable 5는 각각 199%, 167%, 150%의 누적 자본가동률을 보이며 재고회전율도 1.0에 근접했지만, Qwen 3.7 Max는 시작 자본의 22.3%만 투입했고 MiniMax M3는 재고회전율이 0.36에 그쳐 자본이 재고에 묶였다.
  • 전문가 설계 전략들은 46~61%의 판매완료율(sell-through)에서도 58~73%의 주문마진을 지킨 반면, GPT-5.6 Sol과 Opus 4.6은 94~95%의 판매완료율을 보였지만 마진은 35~39%에 그쳐 최종 자산에서 최상위 전문가 전략들에 미치지 못했고, 모든 전문가 전략은 컴플라이언스 위반으로 인한 벌금이 0이었다.
Figure 7: GPT-5.6 Sol behaves like a Volume Wholesaler. Its self-defined sourcing standards reject risky products, prioritize high-margin routes, and diversify inventory while retaining cash, supporting aggressive capital deployment and high sell-through.
Figure 7: GPT-5.6 Sol behaves like a Volume Wholesaler. Its self-defined sourcing standards reject risky products, prioritize high-margin routes, and diversify inventory while retaining cash, supporting aggressive capital deployment and high sell-through.

어디에 쓸 수 있나

  • 전자상거래·B2B 소싱 자동화 도구를 개발할 때 에이전트가 어느 단계(소싱, 가격결정, 고객응대, 컴플라이언스)에서 취약한지 미리 점검하는 진단 프레임워크로 참고할 수 있다.
  • 장기간 지연된 피드백과 변화하는 환경을 다루는 다른 에이전트 평가(재고관리, 광고 예산 배분 등)를 설계할 때 방법론적 참고 사례로 활용할 수 있다.
  • 체크포인트-분기-재개(save-fork-load) 방식의 상태 저장형 평가 기법은 특정 의사결정 지점에서 여러 대안을 비교하는 테스트 설계에 응용할 수 있다.
Figure 8: Gemini 3.1 Pro resembles a Premium House. Its route-aware repricer incorporates supplier cost, freight, tariffs, and competition while enforcing a 15% margin floor, producing higher margins at the cost of lower sell-through.
Figure 8: Gemini 3.1 Pro resembles a Premium House. Its route-aware repricer incorporates supplier cost, freight, tariffs, and competition while enforcing a 15% margin floor, producing higher margins at the cost of lower sell-through.

한계와 남은 검증

  • 평가된 것은 알리바바닷컴 데이터와 미중 관세 변화 등으로 보정된 특정 시뮬레이션 시장 조건이며, 실제 살아있는 시장이나 다른 업종·지역에 그대로 일반화된다는 근거는 제시되지 않았다.
  • 가장 잘한 모델도 사람이 설계한 최상위 전략의 절반 수준에 그쳐, 현재 에이전트가 실제 사업을 안정적으로 운영할 수 있다고 보기는 어렵다.
  • 적용처로 제시한 항목들은 논문이 실제로 검증한 성능이 아니라 이 평가 체계가 시사하는 가능성이며, 실제 자동화 도구에 적용했을 때의 성과는 별도로 측정되지 않았다.
  • 트레이스 서치(trace search)를 통한 테스트 시점 연산 배분 실험은 Qwen 3.8 Max Preview 한 모델에 대해서만 수행되어 다른 모델로의 일반화는 확인되지 않았다.
Figure 9: GPT-5.5 demonstrates adaptive recovery. After detecting zero cash and slow-moving inventory, it liquidates stock, reduces advertising, adopts FOB terms, and resets prices to release trapped capital and continue operating.
Figure 9: GPT-5.5 demonstrates adaptive recovery. After detecting zero cash and slow-moving inventory, it liquidates stock, reduces advertising, adopts FOB terms, and resets prices to release trapped capital and continue operating.

왜 중요한가

실제 사업을 통째로 운영시켜 보기 전에 이런 통제된 가상 시장에서 먼저 시험해 보면, 실제 자본을 위험에 빠뜨리지 않으면서 에이전트가 어디에서 강하고 어디에서 무너지는지 미리 파악할 수 있다. 자동화된 매장 운영, 소싱 대행, 가격 결정 도구 등을 만드는 사람에게는 어떤 역량이 아직 부족한지 구체적으로 보여주는 참고자료가 된다.

Figure 10: Realized value attribution within one trajectory. Two decisions by Gemini 3.5 Flash produce opposite outcomes: route-specific landed-cost reasoning preserves margin for SH-04, while underestimated delivery cost makes a TB-03 order loss-making.
Figure 10: Realized value attribution within one trajectory. Two decisions by Gemini 3.5 Flash produce opposite outcomes: route-specific landed-cost reasoning preserves margin for SH-04, while underestimated delivery cost makes a TB-03 order loss-making.

이 논문의 용어

  • Business Arena · AI 에이전트가 국경간 도매상점을 장기간 스스로 운영하도록 만든 통제된 시뮬레이션 평가 환경
  • 메커니즘 절제 실험(mechanism ablation) · 특정 요소를 방치하거나 악용했을 때와 비교해 좋은 성적이 진짜 실력 때문인지 확인하는 실험 방법
  • 역량별 세부 지표(skill-level metrics) · 최종 성과 점수 하나로는 안 보이는, 재고관리·가격결정·고객응대 등 개별 업무 능력을 따로 측정하는 지표
  • 행동 단위 귀속(action-level attribution) · 발생한 이익이나 손실을 어떤 구체적 행동(소싱, 가격결정 등)이 만들었는지 되짚어 연결하는 분석
  • MCP 호출 · 모델이 도구를 사용할 때 추론과 행동을 번갈아 하도록 정형화된 방식으로 도구를 부르는 규격

저자 · Yijun Pan

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Yijun Pan et al., arXiv:2608.08621, CC BY-SA 4.0