AI 투자 에이전트에게 '업무 매뉴얼'을 쥐어주면 성적이 확 오른다, 스스로 쓰게 하면 별 소용없다
arXiv:2608.180992026-08-20
FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management
AI 투자 에이전트에게 '업무 매뉴얼'을 쥐어주면 성적이 확 오른다, 스스로 쓰게 하면 별 소용없다
FinSkillBench는 포트폴리오 구성, 리스크 관리, 재무제표 분석 3개 영역 12개 과제 2,603건으로 AI 에이전트의 금융 실무 능력을 테스트하는 벤치마크다. 사람이 만든 절차 문서와 검증된 계산 코드를 묶은 '큐레이티드 스킬'을 주면 평균 점수가 0.366에서 0.528로 올랐지만, 에이전트가 스스로 절차를 써서 쓰게 한 '자체 생성 스킬'은 계산 비용만 늘고 효과가 거의 없었다. 9개 모델, 17,820건의 평가와 별도의 Hermes Agent 프레임워크로 8개 모델, 5,280건을 추가 검증해 같은 패턴을 재확인했다.
METAL MEDIA 해설 도표
AI 투자 에이전트에게 '업무 매뉴얼'을 쥐어주면 성적이 확 오른다, 스스로 쓰게 하면 별 소용없다
01포트폴리오 최적화, 마감 위반 감시, 스트레스 테스트, 재무제표 정규화 등 실제 투자 업무에 가까운 12개 과제를 만들고 각 과제마다 정답을 자동 채점하는 검증기를 붙였다.
02같은 과제와 도구, 대화 턴 수를 고정한 채 스킬 문서 없음, 사람이 만든 스킬 제공, 에이전트가 직접 스킬을 작성하는 3가지 조건만 바꿔 비교했다.
03gpt-5.4, claude-sonnet-4.6, gemini-2.5-pro 등 9개 모델로 17,820건을 돌린 결과 큐레이티드 스킬은 평균 점수를 16.2퍼센트포인트 올렸고, 특히 포트폴리오 구성과 리스크 관리에서 효과가 컸다.
04자체 생성 스킬은 토큰과 턴을 더 많이 써서 비용은 늘었지만 정확도 개선은 0.5퍼센트포인트에 그쳐 사실상 효과가 없었다.
05다른 에이전트 실행 환경인 Hermes Agent로 재검증해도 같은 방향의 결과가 나와, 스킬 제공 없이는 못 풀던 문제를 스킬을 주면 풀 수 있게 된다는 결론이 프레임워크에 관계없이 성립함을 확인했다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
포트폴리오 최적화, 마감 위반 감시, 스트레스 테스트, 재무제표 정규화 등 실제 투자 업무에 가까운 12개 과제를 만들고 각 과제마다 정답을 자동 채점하는 검증기를 붙였다.
같은 과제와 도구, 대화 턴 수를 고정한 채 스킬 문서 없음, 사람이 만든 스킬 제공, 에이전트가 직접 스킬을 작성하는 3가지 조건만 바꿔 비교했다.
gpt-5.4, claude-sonnet-4.6, gemini-2.5-pro 등 9개 모델로 17,820건을 돌린 결과 큐레이티드 스킬은 평균 점수를 16.2퍼센트포인트 올렸고, 특히 포트폴리오 구성과 리스크 관리에서 효과가 컸다.
자체 생성 스킬은 토큰과 턴을 더 많이 써서 비용은 늘었지만 정확도 개선은 0.5퍼센트포인트에 그쳐 사실상 효과가 없었다.
다른 에이전트 실행 환경인 Hermes Agent로 재검증해도 같은 방향의 결과가 나와, 스킬 제공 없이는 못 풀던 문제를 스킬을 주면 풀 수 있게 된다는 결론이 프레임워크에 관계없이 성립함을 확인했다.
Table 1. Scorer-to-subtask mapping. Soft-vs-hard grading is calibrated to operational severity.
Table 2. Aggregate condition means, excluding Phi-4 (5,280 evaluations per condition). CIs are bootstrap 95% intervals (10,000 resamples) on paired deltas.
Condition
Mean
Δ vs. No-Skill
95% CI
No-Skill
0.366
—
—
Curated
0.528
+0.162
[+0.152, +0.171]
Self-Generated
0.371
+0.005
[−0.002, +0.011]
Table 3. Domain-level results (excl. Phi-4). PC and RM benefit most because curated skills contain validated solvers and precise input-passing instructions. CI shown for the curated − no-skill paired delta.
Table 7. Hermes harness: aggregate, by-domain, and by-model results. “†” indicates the curated and no-skill 95% CIs do not overlap. Numbers are unweighted means over 1,920 (overall, no-skill) and 5,280 (overall, curated skills) overall episodes. No-skill was a sample run of the first 20 episodes per each 12 subtask (240 total per model).
Group
N
No-Skill (95% CI)
N
Curated (95% CI)
Δ
Overall
All
1920
0.354 [0.338, 0.370]
5280
0.679 [0.669, 0.688]
+0.325
†
By domain
Fundamental analysis
480
0.439 [0.408, 0.470]
2400
0.567 [0.553, 0.582]
+0.129
†
Portfolio construction
800
0.325 [0.301, 0.350]
1600
0.760 [0.744, 0.776]
+0.435
†
Risk management
640
0.325 [0.297, 0.354]
1280
0.785 [0.767, 0.802]
+0.459
†
By model
claude-sonnet-4.6
240
0.304 [0.257, 0.353]
660
0.795 [0.775, 0.815]
+0.490
†
DeepSeek-V3.2
240
0.371 [0.325, 0.419]
660
0.686 [0.659, 0.712]
+0.315
†
gemini-2.5-pro
240
0.333 [0.291, 0.375]
660
0.477 [0.441, 0.512]
+0.144
†
gemma-4-31b-it
240
0.375 [0.332, 0.418]
660
0.642 [0.614, 0.670]
+0.267
†
glm-5.1
240
0.397 [0.350, 0.444]
660
0.779 [0.758, 0.801]
+0.383
†
gpt-4.1
240
0.285 [0.244, 0.327]
660
0.699 [0.672, 0.726]
+0.414
†
gpt-5.4
240
0.408 [0.365, 0.451]
660
0.660 [0.631, 0.687]
+0.252
†
grok-4.20
240
0.357 [0.314, 0.401]
660
0.691 [0.664, 0.719]
+0.334
†
Table 8. Tool availability by condition. “Always” means available in all three conditions; “Curated only” means available only when the curated skill package is mounted; “Self-gen only” means available only in the self-generated condition.
Tool
Availability
Description
submit_answer
Always
Submit final JSON answer (session-ending)
load_skill
Always
Load a SKILL.md into context (returns “not found” when empty)
load_references
Always
Load supplementary reference documents for a loaded skill
get_task_data
Always
Retrieve specific fields from the task input at full precision
query_xbrl
Always (FA)
Query the XBRL financial data panel by ticker, period, and metrics
run_skill_ script
Curated only
Execute a validated Python script from the skill directory
save_skill
Self-gen only
Write a SKILL.md to the per-episode scratch directory
Table 9. Cognitive and procedural demands by subtask. ✓= primary demand; ∘ = secondary demand.
Subtask
Numerical optim.
Tool param.
Multi-step comp.
Domain knowledge
Data retrieval
Precision mgmt.
Multi-output
Unconstrained optim.
✓
✓
∘
∘
∘
✓
Constrained optim.
✓
✓
∘
✓
∘
✓
✓
Tool-use param.
✓
✓
Rebalancing
✓
✓
✓
∘
∘
✓
✓
Black–Litterman
✓
✓
✓
✓
∘
✓
✓
Constraint monitoring
✓
✓
∘
✓
Risk identification
✓
✓
✓
✓
Stress testing
∘
✓
✓
✓
✓
✓
Risk remediation
✓
∘
✓
✓
∘
✓
Normalization
∘
✓
✓
Earnings quality
✓
✓
✓
✓
Driver decomposition
✓
✓
✓
✓
Table 10. Complete scorer specifications. Most component scores are clipped or averaged into [0,1]; the exp05 ranked-list scorer does not final-clip NDCG or the composite and can slightly exceed 1 when duplicate predicted risk types receive repeated relevance credit.
Scorer
Subtask
Formula
L2 distance
Unconstrained optim.
max(0,min(1, 1−∥wp−we∥2/(4θ))), θ=0.05
Constraint gate + L2
Constrained optim.
All expected constraint_satisfaction keys must be present and equal; extra predicted keys are ignored. If the gate passes, weight L2 score as above; else 0
Parameter match
Tool-use param.
Fraction of union leaf parameters matching (numeric: relative error < 0.2; categorical: exact; extra/missing leaves count against the denominator)
Turnover compliance
Rebalancing
0.5⋅sL2+0.25⋅sturn+0.25⋅strade, where sL2=max(0,min(1,1−L2/0.2)), sturn=max(0,1−min(1,|T^−T|/max(|T|,10−9))), and strade is mean expected-ticker trade consistency
View + weights
Black–Litterman
Mean of active components: posterior-returns MAE score max(0, 1−MAE/0.05) and weight L2 score max(0,min(1,1−L2/(4θ))) with θ=0.10
Exact match
Constraint monitoring
0.4⋅ 1[compliant match]+0.6⋅constraint accuracy, where accuracy is status match over expected normalized constraint types
Ranked recall
Risk identification
0.4⋅recall@k+0.3⋅precision@k+0.3⋅NDCG@k; recall/precision use multiset type overlap; exp05 NDCG uses max magnitude per expected type and is not final-clipped
Fraction of scored non-metadata metrics with relative error ≤ per-metric tolerance (default 5%); unwraps metrics wrapper when present on both sides
EQ composite
Earnings quality
Weighted mean over active terms: component accuracy (episode weight 0.4), Beneish-flag match (0.3), flag F1 (0.3); numeric-ratio accuracy is included only if numeric_ratio_weight>0 in exp05/zqbok_experiment05
Driver F1
Driver decomposition
Mean of active terms: revenue-driver recall/precision, matched revenue-driver direction, matched revenue numeric accuracy, margin recall/direction, matched margin numeric accuracy, and revenue-Δ accuracy
왜 중요한가
실무에서 AI 에이전트를 금융 분석에 투입할 때 모델 자체의 성능보다 검증된 절차 문서와 실행 가능한 도구를 함께 제공하는지가 더 중요할 수 있음을 보여준다. 반대로 에이전트가 스스로 절차를 만들어 쓰는 방식은 아직 신뢰하기 어렵다는 경고이기도 하다.
이 논문의 용어
에이전틱 AI · 단순히 답을 생성하는 것을 넘어 도구를 호출하고 여러 단계를 거쳐 작업을 수행하는 AI 시스템
스킬(skill) · 에이전트가 재사용할 수 있도록 만든 절차 문서와 실행 코드 묶음
큐레이티드 스킬 · 사람이 직접 검증해 만든 절차 문서와 코드 패키지
자체 생성 스킬 · 에이전트가 문제를 풀기 전 스스로 절차 문서를 작성하고 그것을 다시 불러와 사용하는 방식
포인트인타임(point-in-time) 데이터 · 특정 시점 이후의 정보가 섞이지 않도록 그 시점까지만 알 수 있는 데이터로 제한한 것