83억 개의 가상 인물로 사람 대신 AI 제품을 테스트하는 인프라, MatrAIx
83억 개의 가상 인물로 사람 대신 AI 제품을 테스트하는 인프라, MatrAIx
MatrAIx는 사람을 직접 모아 설문·테스트하기 어려운 문제를 풀기 위해, 1,290개 항목으로 정의된 83억 개의 가상 페르소나 데이터베이스와 이들이 설문·챗봇·웹·앱에서 행동하게 만드는 실행 환경, 그리고 1,010개의 평가용 과업을 하나로 묶은 인프라다. 저자들은 400건짜리 통제 실험으로 페르소나가 지시받은 성향을 실제로 따르는지 확인했고, 사람 심사자와 LLM 심사자를 동원해 추출된 페르소나의 품질도 점검했다. 다만 저자는 페르소나 결과가 실제 사람 행동을 그대로 대변한다고 주장하지 않으며, 오히려 사용 시 주의할 한계를 상세히 밝혔다.
METAL MEDIA 해설 도표
MatrAIx 파이프라인: 페르소나 생성부터 결과 보고까지
증거 상태측정 결과와 예정된 검증이 함께 있음
- Persona 8B1,290개 속성으로 정의된 83억 개의 가상 인물, 방향성 그래프 샘플링(합성)과 실제 자료 매핑(인간 기반)으로 구성되며 약 100만 명 코어셋이 공개된다
- 코호트 선택평가자가 원하는 대상(예: 연령대, 소득, 지역)을 지정하면 해당 조건에 맞는 페르소나 집단을 Persona 8B에서 추출한다
- 네 가지 실행 환경Survey(설문), AI Chatbot(대화형 AI), Web(웹 브라우징), App(데스크톱·모바일 앱) 환경에서 페르소나 에이전트가 상호작용하며 대화, 행동, 화면 상태 등을 기록한다
- 과업 및 검증자1,010개 과업 명세가 목표와 성공 조건을 정의하고, 프로그램형 검증자 또는 사람·LLM 심사자가 결과가 조건을 만족했는지 확인한다
- 집계 보고서개별 시행 결과를 과업·코호트·하위집단 단위로 모아 대표 지표(예: 유료 전환율)와 근거 데이터를 함께 보여주는 보고서를 만든다
무엇을 했나
- 사람 평가는 비싸고 느려서 규모를 키우기 어렵고, 오프라인 벤치마크는 빠르지만 사용자마다 다른 요구와 반응을 반영하지 못한다는 문제에서 출발했다.
- Persona 8B라는 1,290개 속성(나이, 언어, 위험 선호 등)으로 정의된 83억 개의 가상 인물 데이터베이스를 만들었는데, 속성 간 상관관계를 보존하는 방향성 그래프(부모 속성이 자식 속성에 영향을 주는 방식)로 합성 인물을 생성하고, 위키백과·아마존 리뷰·개발자 설문 등 실제 자료에서 뽑아낸 인물도 같은 틀에 맞춰 결합했다.
- 이 중 검증을 거친 약 100만 명 규모의 공개 서브셋(실제 자료 기반 599,847명, 합성 400,000명)을 배포했고, 페르소나가 설문(Survey)·AI챗봇(Chatbot)·웹(Web)·앱(App) 네 가지 환경에서 실제로 행동하도록 하는 실행 플랫폼과 1,010개 과업 목록을 함께 제공한다.
- 8개 대표 과업에 대해 Claude Opus 4.8, GPT 5.5, Claude Haiku 4.5 세 모델로 총 18,189회 시뮬레이션을 돌려 가격 인상 후 구매 망설임, AI 실패 후 재시도 의향, 응답 지연 인내심 등 사람마다 다른 반응 패턴을 확인했다.
- 400회 통제 실험에서 지시된 성향이 실제 행동으로 표현되거나 반대 성향이 올바르게 억제된 비율이 91.5%(366/400건)로 나타났고, 별도로 사람 평가자와 LLM 평가자가 추출된 페르소나의 품질을 점수화했다.

| Top-level group | Dims. | Representative attributes | Representative grounding |
|---|---|---|---|
| Background | 238 | Age, region, language, education, family, career, industry | Population statistics, household surveys, education and labor taxonomies |
| Psychology | 210 | Personality, values, worldview, motivation, risk | Validated instruments, values surveys, schema design priors |
| Capability | 331 | Domain expertise, general skills, tools, programming, developer context | Occupational taxonomies, technology/developer surveys |
| Behavior and Interaction | 124 | Preferences, habits, interaction state, work practices, technology adoption | Time-use, consumer, workplace, and technology-use evidence |
| Lifestyle | 387 | Interests, media, culture, hobbies, sports, food, health, fitness | Health statistics, consumption surveys, cultural sources |
| Total | 1,290 |
| Source | Released records |
|---|---|
| Wikipedia extraction | 323,438 |
| Amazon Review extraction | 97,915 |
| Stack Overflow survey extraction | 113,120 |
| PRISM Alignment | 1,487 |
| General Social Survey | 63,532 |
| MatrAIx volunteer survey | 355 |
| Human-grounded subtotal | 599,847 |
| Full-DAG synthetic | 400,000 |
| Total | 999,847 |
| Commerce | Software | Finance | Healthcare | Other | Total | |
|---|---|---|---|---|---|---|
| Survey | 202 | 138 | 141 | 139 | 1 | 621 |
| AI Chatbot | 3 | 11 | 17 | 29 | 311 | 371 |
| Web | 2 | 2 | 2 | 0 | 6 | 12 |
| App | 0 | 5 | 1 | 0 | 0 | 6 |
| Total | 207 | 156 | 161 | 168 | 318 | 1,010 |
| Group | Subgroup | Schema category | Count | Representative attributes |
|---|---|---|---|---|
| Background | Demographics | Demographic: Core | 25 | Age bracket; region; gender identity |
| Background | Demographics | Demographic: Cultural | 2 | Cultural background; attitude toward immigration |
| Background | Demographics | Demographic: Family | 1 | Household size |
| Background | Demographics | Demographic: Life Events | 24 | Life stage; major life events; childhood environment |
| Background | Language | Linguistic: Language | 53 | Primary language; English proficiency; multilingualism |
| Background | Language | Linguistic: Communication | 37 | Expected tone; verbosity; communication preferences |
| Background | Education | Learning: Academic | 34 | Highest education; academic field; institution tier |
| Background | Education | Learning: Style | 1 | Learning style |
| Background | Career | Professional: Career | 4 | Research output; seniority; years of experience |
| Background | Career | Professional: Industry | 51 | Company size; role function; industry |
| Background | Career | Developer: Professional Context | 6 | Professional status; role archetype; contribution context |
| Psychology | Personality | Personality: Character | 34 | Domain stance; dominant trait; curiosity |
| Psychology | Personality | Personality: Big Five | 50 | Imagination; artistic interest; emotionality |
| Psychology | Personality | Personality: MBTI | 2 | Neurotype; Myers-Briggs type |
| Psychology | Personality | Personality: Relationships | 4 | Attachment anxiety; attachment avoidance; interpersonal agency |
| Psychology | Worldview | Values & Motivation | 46 | Core value; religiosity; economic motivation |
| Psychology | Worldview | Worldview: Beliefs | 67 | Political leaning; trust level; safety sensitivity |
| Psychology | Decision-Making | Risk & Decision | 7 | Risk tolerance; decision style; need for closure |
| Capability | Domains | Expertise: Domains | 144 | Domain; subject specialty; technology savviness |
| Capability | Skills | Expertise: Skills | 64 | Writing; copywriting; editing |
| Capability | Skills | Skills: Tools | 69 | Excel; Google Sheets; Python |
| Capability | Skills | Skills: Programming | 44 | Comment style; summary documentation; naming verbosity |
| Capability | Skills | Developer: Code Maintenance | 10 | Complexity tolerance; modularity preference; type-system orientation |
| Behavior and Interaction | Personal Behavior | Behavior: Preferences | 34 | Modality preference; accessibility needs; media diet |
| Behavior and Interaction | Personal Behavior | Behavior: Habits | 30 | Journaling; meditation; use of to-do lists |
| Behavior and Interaction | Personal Behavior | Behavior: Time | 3 | Time pressure; sleep schedule; micromanagement aversion |
| Behavior and Interaction | Interaction State | State: Emotional | 5 | Emotional state; intent; query complexity |
| Behavior and Interaction | Work Practices | Behavior: Work | 2 | Work schedule; office versus remote work |
| Behavior and Interaction | Work Practices | Developer: Open Source Behavior | 7 | Open-source activity; GitHub contribution mode; pull-request style |
| Behavior and Interaction | Work Practices | Developer: Community Behavior | 4 | Stack Overflow use; participation style; help-seeking preference |
| Schema group | Facets | Grounding roles | Sources |
|---|---|---|---|
| Background | Demographics (52); language (90); education (35); career (61) | Category definitions; population priors; household, language, education, and labor dependencies | UN World Population Prospects and Population Data [87, 86]; World Bank WDI and WorldPop [94, 97]; Eurostat, ACS PUMS, and IPUMS [21, 82, 38]; DHS, UNICEF MICS, and OECD Family Database [77, 85, 62]; Pew and World Values Survey [69, 96]; UNESCO UIS, World Bank Education Statistics, OECD INES, and OECD PISA [84, 93, 61, 64]; ILOSTAT, BLS OEWS, and O*NET [34, 81, 60]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]. |
| Psychology | Personality (90); worldview (113); decision-making (7) | Instrument and value-set design; selected prevalence estimates; validation | IPIP and MIDUS [35, 56]; Pew and World Values Survey [69, 96]; GSS, European Social Survey, ISSP, and Gallup World Poll [58, 20, 36, 23]; Afrobarometer, Arab Barometer, Asian Barometer, Latinobarometro, and Eurobarometer [1, 3, 5, 46, 19]; ARDA [6]. |
| Capability | Domain expertise (144); general skills (64); tools (69); programming (44); developer context (10) | Occupational and skill taxonomies; technology access and adoption; developer-tool prevalence | ITU Statistics, World Bank WDI, DataReportal, and Pew Internet [37, 94, 15, 70]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]; O*NET [60]. |
| Behavior and Interaction | Personal behavior (67); interaction state (5); work practices (13); technology use (39) | Time-use and consumer priors; workplace behavior; technology and AI adoption | American Time Use Survey, Consumer Expenditure Surveys, and OECD Time Use [79, 80, 63]; ITU, DataReportal, and Pew Internet [37, 15, 70]; Stack Overflow Survey, GitHub Octoverse, JetBrains Developer Ecosystem, and O*NET [75, 25, 39, 60]. |
| Lifestyle | Interests (358); physical health (25); fitness (2); health lifestyle (2) | Health and disability priors; consumption and time use; cultural and interest category design | WHO GHO and IHME GBD [95, 33]; ACS PUMS, NHIS, BRFSS, DHS, and UNICEF MICS [82, 11, 10, 77, 85]; ATUS, CEX, and OECD Time Use [79, 80, 63]; FAOSTAT and UNESCO Culture Statistics [22, 83]; Eurobarometer, Pew, DataReportal, WVS, and Gallup World Poll [19, 69, 70, 15, 96, 23]. |
| v | πi(v) | qlang | qregion | ri(v) | mi(v) | πirimi | pθ(v∣xPa(i)) |
|---|---|---|---|---|---|---|---|
| None | 0.30 | 0.02 | 0.10 | 0.022 | 0 | 0.000 | 0.000 |
| Basic | 0.30 | 0.08 | 0.20 | 0.178 | 1 | 0.053 | 0.028 |
| Fluent | 0.25 | 0.30 | 0.35 | 1.680 | 1 | 0.420 | 0.224 |
| Native | 0.15 | 0.60 | 0.35 | 9.333 | 1 | 1.400 | 0.747 |
| sum | 1.00 | 1.00 | 1.00 | 1.873 | 1.000 |

| Stage | Rejected | Remaining |
|---|---|---|
| Original corpus | – | 10,002,288,277 |
| Contradiction filter | 239,310 | 10,002,048,967 |
| Human exact/MinHash deduplication | 41,597 | 2,222,496 human |
| Synthetic projection deduplication | 252,936,392 | 9,746,848,482 synthetic |
| Synthetic deterministic cutoff | 1,349,070,978 | 8,397,777,504 synthetic |
| Audited baseline | 8,400,000,000 |

| Dimension | Answered | Share of those answering |
|---|---|---|
| Age bracket | 321 | 25–34 22.1%, 55–64 19.0%, 35–44 15.9%, 18–24 13.7%, 45–54 13.4%, 65–74 8.4%, 75–84 4.0%, 85+ 3.4% |
| Gender identity | 322 | Woman 47.5%, Man 42.9%, Non-binary 3.7%, Self-described 3.4%, Prefer not to say 2.5% |
| Region | 329 | South Asia 20.4%, Sub-Saharan Africa 19.1%, East Asia 13.4%, Southeast Asia 10.9%, Latin America 9.7%, Western Europe 7.0%, MENA 6.7%, North America 5.8%, Eastern Europe 4.6%, Oceania 2.4% |
| Urbanicity | 328 | Rural 33.2%, Dense urban 22.3%, Small town 21.0%, Suburban 20.4%, Nomadic / remote 3.0% |
| Socioeconomic band | 340 | Low 31.8%, Lower-middle 30.0%, Middle 20.3%, Upper-middle 12.6%, High 5.3% |
| Employment | 330 | Full-time 25.8%, Retired 16.1%, Homemaker 14.2%, Self-employed 11.8%, Part-time 9.1%, Student 8.2%, Unemployed 7.6%, Gig / freelance 7.3% |

| Collection | Tasks | Environment | Composition |
|---|---|---|---|
| Synthetic persona surveys | 405 | Survey | 135 Finance, 135 Healthcare, and 135 Software questionnaires, generally sharing a common survey and reporting template. |
| Product surveys | 200 | Survey | Twenty product-research archetypes, including purchase intent, price sensitivity, recommendation, and retention, crossed with ten retail products. |
| Synthetic chatbots | 351 | AI Chatbot | Scenario families across 27 domains, including travel, legal, healthcare, telecommunications, real estate, insurance, finance, and education. |

| Contract component | Declared content | Audit purpose |
|---|---|---|
| Task metadata | Stable name and version, environment type, product domain, difficulty, tags, artifact paths, and runtime budgets | Identifies the recipe and prevents results from silently moving between task versions. |
| Persona-facing scenario | Context, user goal, constraints, disclosure policy, and required submission | Defines what every sampled persona is asked to do without exposing verifier internals. |
| Cohort strategy | Persona-schema filters, sampling mode, stratification fields, sample size, weights, and seed policy | Makes the target audience explicit and permits the cohort to be redrawn. |
| Product attachment | Questionnaire or stimulus, chat endpoint or sidecar, website target, or native application backend | Identifies the system under test and how the runtime reaches it. |
| Verifier | Required artifacts, structured finding schema, objective checks, timeouts, and failure conditions | Converts a trial into reproducible outcomes with supporting evidence. |
| Reporting policy | Aggregations, subgroup facets, summaries, optional judge directives, and disclosure rules | Defines how trial findings become a cohort-level report. |

| Environment | Task-owned inputs | Primary trial artifacts |
|---|---|---|
| Survey | Stimulus, questionnaire schema, response constraints, and optional rationale prompts | Typed responses, missing or invalid items, rationales, confidence, and completion summary. |
| AI Chatbot | Product endpoint or sidecar, user goal, staged-disclosure policy, turn and safety limits | Full transcript, tool or service events, resolution state, termination reason, and post-run feedback. |
| Web | URL or hosted site, browser backend, exploration requirements, and submission schema | Page and action trace, considered options, screenshots where enabled, final submission, and terminal page or site state. |
| App | Native target, platform backend, initial state, permissions, and terminal-state checks | Screenshot and action trace, exported files, application state, permission changes, and cross-application side effects. |
| Task | Env. | Primary outcome | Share giving that answer | Range | q | ||
|---|---|---|---|---|---|---|---|
| GPT | Opus | Haiku | (pt) | ||||
| Annual checkup | Survey | q31 = e | 50.5% | 41.4% | 29.7% | 20.8 | <0.001∗∗∗ |
| Candy Land price | Survey | q_threshold = hesitate | 98.3% | 27.0% | 83.3% | 71.3 | <0.001∗∗∗ |
| OpenBB honesty | Chat | wouldStillContinueUse = unsure | 81.3% | 85.5% | 28.1% | 57.4 | <0.001∗∗∗ |
| Meal planning† | Chat | adherenceLikelihood = 7 | 50.6% | 0.2% | 40.0% | 50.4 | <0.001∗∗∗ |
| Notion plans | Web | decision_subject_id = plus | 63.5% | 21.5% | 75.7% | 54.2 | <0.001∗∗∗ |
| MIT OCW course | Web | task_course_level = Graduate | 40.0% | 34.5% | 47.5% | 13.0 | <0.001∗∗∗ |
| News+ subscription | App | clicked_get_started = true | 4.2% | 20.8% | 0.0% | 20.8 | 0.025∗ |
| Stocks sentiment | App | sentiment = hold | 60.0% | 30.0% | 50.0% | 30.0 | 0.153 ns |
| Task | Product-level measure | GPT 5.5 | Opus 4.8 | Haiku 4.5 |
|---|---|---|---|---|
| Annual checkup | Very likely to schedule in time | 50.5% (505/1,000) | 41.4% (414/1,000) | 29.7% (297/1,000) |
| Candy Land price | Hesitates or worse at new price | 98.3% (983/1,000) | 27.0% (270/1,000) | 83.3% (833/1,000) |
| OpenBB honesty | Would not continue using | 18.5% (185/1,000) | 14.5% (132/909) | 71.9% (719/1,000) |
| Meal planning† | Stated need fully satisfied | 46.2% (462/1,000) | 0.2% (2/1,000) | 28.1% (281/1,000) |
| Notion plans | Selected a paid plan | 75.8% (776/1,024) | 23.2% (237/1,022) | 93.9% (958/1,020) |
| MIT OCW course | Chose a graduate-level course | 40.0% (403/1,008) | 34.5% (347/1,007) | 47.5% (473/996) |
| News+ subscription | Subscribed | 4.2% (1/24) | 20.8% (5/24) | 0.0% (0/24) |
| Stocks sentiment | Buy opinion | 40.0% (8/20) | 70.0% (14/20) | 47.4% (9/19) |
| Task | Persona dimension | Groups | Sig. | Spearman ρ, exact p | ||
|---|---|---|---|---|---|---|
| GPT/Opus | GPT/Haiku | Opus/Haiku | ||||
| Annual checkup | age bracket | 6 | ∘∘∙ | +0.26 p=0.329 | −0.43 p=0.822 | −0.77 p=0.971 |
| Candy Land price | economic motivation | 4 | ∘∘∘ | −0.40 p=0.792 | −0.20 p=0.625 | +0.80 p=0.167 |
| OpenBB honesty | trust level | 4 | ∙∙∙ | +1.00 p=0.042 | +1.00 p=0.042 | +1.00 p=0.042 |
| Meal planning† | life stage | 4 | ∘∘∘ | +0.95 p=0.083 | +0.32 p=0.500 | +0.33 p=0.417 |
| Notion plans | company size | 8 | ∘∘∘ | +0.44 p=0.138 | +0.60 p=0.059 | −0.24 p=0.725 |
| MIT OCW course | academic field | 8 | ∘∘∘ | +0.93 p=0.001 | +0.09 p=0.420 | +0.06 p=0.452 |
| News+ subscr.‡ | economic motivation | 3 | ∘∘∘ | +0.00 p=0.667 | flat | flat |
| Stocks sentiment‡ | risk tolerance | 5 | ∘∘∘ | +0.47 p=0.267 | −0.92 p=1.000 | −0.67 p=0.933 |
| GPT × Opus | GPT × Haiku | Opus × Haiku | |
|---|---|---|---|
| Paired agreement over 88 joinable fields | |||
| Median Cohen’s κ | 0.000 | 0.000 | +0.001 |
| Fields at κ≤0 | 59 of 88 | 50 of 88 | 40 of 88 |
| Fields reaching κ≥0.2 | 7 of 88 | 8 of 88 | 7 of 88 |
| Fields at ≥50% agreement, κ<0.1 | 48 of 56 | 25 of 34 | 24 of 34 |
| Self-report fidelity: age band matches the persona | |||
| GPT 5.5 16.3% (ns) Opus 4.8 16.9% (ns) Haiku 4.5 100.0% | chance 16.7% |
| Survey | Chat | Web | OS-App | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Attribute | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | |||
| code-comment-style | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| code-naming-verbosity | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| code-summary-documentation | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-emoji-use | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-humor | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-politeness | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-storytelling | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-use-of-jargon | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-verbosity | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| register | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 |
| Attribute | Positive (declared) | Negative (opposite) |
|---|---|---|
| code-comment-style | Nearly every line carries an inline comment (# Iterate over every integer, # Check divisibility) | Code contains no comments whatsoever |
| code-naming-verbosity | calculate_average_of_passing_scores, number_of_passing_scores | Variables s, t, a, c, x, all single-letter |
| code-summary-documentation | Every function opens with a tldr: docstring | Prose plus code, no TLDR/summary header |
| cog-emoji-use | Many emoji across a short paragraph | No emoji despite a casual app-store prompt |
| cog-humor | Playful, witty asides (“the boxes may yet win”) | Measured, earnest tone, no jokes |
| cog-politeness | “Might I kindly ask that you resend the document at your earliest convenience” | “Hey, you forgot the attachment. Again.”: blunt and sarcastic |
| cog-storytelling | A narrated scene (“a storm came through, half the lodge went dark”) | Abstract, value-driven, no concrete scene |
| cog-use-of-jargon | Dense technical jargon (“TCP three-way handshake, SYN-ACK”) | Plain terms (“translate the name into a numerical address”) |
| cog-verbosity | Rambling, multi-paragraph, tangential answer | Short clipped fragments, no elaboration |
| register | Standard/formal phrasing | Colloquial (“cuppa and the papers”, “the wife’s doing a roast”) |
| Metric | Name | Definition |
|---|---|---|
| M1 | Claim validity | Does each populated claim remain defensible when its value, cited evidence, and free-text description are considered jointly? A field is problematic if its value is indefensible, its evidence is plainly unrelated or non-supporting, or its description adds implausible unsupported specifics. |
| M2 | No over-claiming | Has an attribute been populated without any plausible basis? Clearly unsupported assignments count as problems, while a plausible, appropriately labeled inference does not. |
| M3 | Coverage | Did the extraction omit an important attribute that was clearly available in the source? |
| M4 | Internal consistency | Do extracted fields contradict one another, especially on core identity, role, region, life stage, language, or technology-use information? |
| M5 | Overall fidelity and plausibility | Is the extracted record usable for faithfully role-playing the source person, and are its inferred psychological, interest, and communication traits plausible and coherent? |
| Metric | GPT | Claude | Δ | ≤1 |
|---|---|---|---|---|
| M1 Claim validity | 3.433 | 3.937 | −0.504 | 81.4% |
| M2 No over-claiming | 3.532 | 3.646 | −0.114 | 82.3% |
| M3 Coverage | 4.465 | 4.031 | +0.434 | 99.2% |
| M4 Internal consistency | 3.606 | 3.932 | −0.326 | 85.0% |
| M5 Fidelity and plausibility | 3.829 | 4.109 | −0.280 | 97.8% |
| Overall | 3.773 | 3.931 | −0.158 | 89.1% |
| Metric | Human mean | H–H ≤1 | GPT–H ≤1 | Claude–H ≤1 |
|---|---|---|---|---|
| M1 | 4.105 | 99.1% | 69.0% | 92.0% |
| M2 | 3.770 | 92.2% | 81.0% | 88.0% |
| M3 | 4.223 | 97.1% | 95.0% | 100.0% |
| M4 | 4.537 | 98.6% | 55.0% | 92.0% |
| M5 | 4.040 | 99.1% | 96.0% | 97.0% |
| Overall | 4.135 | 97.2% | 79.2% | 93.8% |
| Agent Bench | GAIA | Web Arena | Web Shop | Mind2 Web | App World | Tool LLM | Bench Flow | τ-bench | Generative Agents | SOTOPIA | OASIS | Silicon sampling | MatrAIx | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Product-as-SUT evaluation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Persona conditioning | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Behavioral grounding | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ |
| Native-device execution | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Reproducible reporting | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Trajectory inspection | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| End-to-end validation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
| Silicon sampling | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Reproducible cohorts | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ |
| Multi-agent simulation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ |
실제로 확인된 결과
- 400건 통제 실험에서 지시된 행동 성향이 표현되거나 올바르게 억제된 비율이 91.5%(366/400건)로 나타났다.
- 1,000건의 추출 페르소나에 대해 GPT-5.5와 Claude 두 LLM 심사자가 평가한 결과 전체 평균 점수는 각각 3.773, 3.931이었고, 5,000개의 짝지은 점수 중 89.1%가 1점 이내로 일치했다.
- 100명 서브셋에 대한 6인 사람 평가 평균은 5점 만점에 4.135였고, 3,000개 점수 중 84.7%가 4점 또는 5점이었다.
- 사람 평균 점수와 비교했을 때 GPT는 79.2%, Claude는 93.8%의 경우에서 1점 이내로 일치했다.
- 동일 코호트에 대해 세 모델을 바꿔 실행했을 때 한 제품 페이지의 유료 요금제 선택 비율이 23.2%에서 93.9%까지 벌어졌고, 88개의 공통 필드에서 모델 간 짝별 일치도(코헨의 카파)는 거의 0에 가까웠다.
어디에 쓸 수 있나
- 신제품이나 앱 기능을 실제 사용자에게 공개하기 전에 다양한 배경의 가상 사용자로 초기 반응(가격 민감도, 사용성 문제 등)을 사전 스크리닝하는 데 활용할 수 있다.
- 챗봇이나 AI 어시스턴트가 실패했을 때 사용자가 계속 이용할지, 응답 지연을 얼마나 견디는지 같은 상호작용 패턴을 미리 점검하는 데 쓸 수 있다.
- 제품 버전을 바꾼 뒤 동일한 페르소나 집단과 과업으로 재실행해 변경 전후를 비교하는 반복 실험에 적용할 수 있다.
- 특정 하위 집단(연령, 소득, 지역 등)별로 제품이 다르게 작동하는지 세분화해 살펴보는 코호트 단위 분석에 쓸 수 있다.
한계와 남은 검증
- 페르소나 집단은 실제 인구의 확률 표본이 아니며, 페르소나 결과는 실제 사람이 어떻게 행동할지에 대한 직접적 증거가 아니라 가설을 만드는 용도로만 다뤄야 한다고 저자가 명시했다.
- 페르소나를 연기하는 모델과 평가 대상 시스템이 같은 모델 계열을 공유할 경우, 좋은 결과가 실제 만족인지 모델이 자기 출력을 선호하는 편향인지 구분할 수 없는데, 이번 실험은 이 둘을 분리하는 실험(같은 과업에서 페르소나 모델과 시스템 모델을 교차시키는 실험)을 아직 수행하지 않았다.
- 실제 사용자 상호작용에서 나타나는 정보 은닉, 반박, 포기 같은 행동을 페르소나가 실제 사람처럼 보이는지는 아직 직접 측정되지 않았으며, 이를 확인할 후속 실험(실제 대화 로그와의 비교, 개인 단위 대화 이어쓰기 검증 등)이 우선 과제로 남아 있다.
- 공개된 코어셋은 내부 전체 인구가 아니라 필터링과 중복 제거를 거친 약 100만 명 규모이며, 합성 부분은 연령대·지역·성별·도시화 정도 등 4개 변수에만 맞춰 보정되어 나머지 1,290개 속성 전체의 결합분포를 대표하지는 않는다.
- 건강, 금융, 고용 등 실제 사람에게 영향을 미치는 규제 대상 의사결정에는 시뮬레이션 결과만으로 결론을 내려서는 안 되며, 같은 과업과 도구로 사람 검증을 거쳐야 한다고 저자가 강조했다.
왜 중요한가
AI 제품이나 챗봇을 실제 사용자에게 테스트하기 전에, 다양한 배경을 가진 수십억 명 규모의 가상 사용자로 값싸고 빠르게 사전 점검할 수 있는 길을 제시한다. 다만 저자 스스로 페르소나 결과가 실제 사람 행동의 직접적 증거가 아니라 '가설을 만드는 도구'에 가깝다고 강조하므로, 실제 배포 전 사람 검증을 대체할 수는 없다.
이 논문의 용어
- 페르소나(Persona) · 나이, 언어, 성향 등 여러 속성으로 정의된 가상의 사용자 프로필
- 방향성 그래프(DAG) 샘플링 · 속성들을 서로 영향을 주는 순서대로 하나씩 뽑아, 나이와 학력처럼 서로 관련 있는 값들이 어긋나지 않게 인물을 생성하는 방법
- 인간 기반(human-grounded) 레코드 · 위키백과, 리뷰, 설문 응답 등 실제 사람 데이터에서 뽑아 만든 페르소나 항목
- 검증자(verifier) · 페르소나 행동 결과가 과업이 요구한 조건을 만족했는지 자동 또는 사람이 확인하는 절차
- 코호트(cohort) · 특정 조건(연령대, 지역 등)으로 선택된 평가용 페르소나 집단
최신 논문
- AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다
- 논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- 고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법
- 로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다
- 에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법
- AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다
- AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다
- AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다
METAL MEDIA 최신 기사
그림 출처: Xiaomin Li et al., arXiv:2608.04205, arxiv-nonexclusive