컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

83억 개의 가상 인물로 사람 대신 AI 제품을 테스트하는 인프라, MatrAIx

arXiv:2608.042052026-08-03

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

83억 개의 가상 인물로 사람 대신 AI 제품을 테스트하는 인프라, MatrAIx

MatrAIx는 사람을 직접 모아 설문·테스트하기 어려운 문제를 풀기 위해, 1,290개 항목으로 정의된 83억 개의 가상 페르소나 데이터베이스와 이들이 설문·챗봇·웹·앱에서 행동하게 만드는 실행 환경, 그리고 1,010개의 평가용 과업을 하나로 묶은 인프라다. 저자들은 400건짜리 통제 실험으로 페르소나가 지시받은 성향을 실제로 따르는지 확인했고, 사람 심사자와 LLM 심사자를 동원해 추출된 페르소나의 품질도 점검했다. 다만 저자는 페르소나 결과가 실제 사람 행동을 그대로 대변한다고 주장하지 않으며, 오히려 사용 시 주의할 한계를 상세히 밝혔다.

METAL MEDIA 해설 도표

MatrAIx 파이프라인: 페르소나 생성부터 결과 보고까지

증거 상태측정 결과와 예정된 검증이 함께 있음

  1. Persona 8B1,290개 속성으로 정의된 83억 개의 가상 인물, 방향성 그래프 샘플링(합성)과 실제 자료 매핑(인간 기반)으로 구성되며 약 100만 명 코어셋이 공개된다
  2. 코호트 선택평가자가 원하는 대상(예: 연령대, 소득, 지역)을 지정하면 해당 조건에 맞는 페르소나 집단을 Persona 8B에서 추출한다
  3. 네 가지 실행 환경Survey(설문), AI Chatbot(대화형 AI), Web(웹 브라우징), App(데스크톱·모바일 앱) 환경에서 페르소나 에이전트가 상호작용하며 대화, 행동, 화면 상태 등을 기록한다
  4. 과업 및 검증자1,010개 과업 명세가 목표와 성공 조건을 정의하고, 프로그램형 검증자 또는 사람·LLM 심사자가 결과가 조건을 만족했는지 확인한다
  5. 집계 보고서개별 시행 결과를 과업·코호트·하위집단 단위로 모아 대표 지표(예: 유료 전환율)와 근거 데이터를 함께 보여주는 보고서를 만든다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 사람 평가는 비싸고 느려서 규모를 키우기 어렵고, 오프라인 벤치마크는 빠르지만 사용자마다 다른 요구와 반응을 반영하지 못한다는 문제에서 출발했다.
  2. Persona 8B라는 1,290개 속성(나이, 언어, 위험 선호 등)으로 정의된 83억 개의 가상 인물 데이터베이스를 만들었는데, 속성 간 상관관계를 보존하는 방향성 그래프(부모 속성이 자식 속성에 영향을 주는 방식)로 합성 인물을 생성하고, 위키백과·아마존 리뷰·개발자 설문 등 실제 자료에서 뽑아낸 인물도 같은 틀에 맞춰 결합했다.
  3. 이 중 검증을 거친 약 100만 명 규모의 공개 서브셋(실제 자료 기반 599,847명, 합성 400,000명)을 배포했고, 페르소나가 설문(Survey)·AI챗봇(Chatbot)·웹(Web)·앱(App) 네 가지 환경에서 실제로 행동하도록 하는 실행 플랫폼과 1,010개 과업 목록을 함께 제공한다.
  4. 8개 대표 과업에 대해 Claude Opus 4.8, GPT 5.5, Claude Haiku 4.5 세 모델로 총 18,189회 시뮬레이션을 돌려 가격 인상 후 구매 망설임, AI 실패 후 재시도 의향, 응답 지연 인내심 등 사람마다 다른 반응 패턴을 확인했다.
  5. 400회 통제 실험에서 지시된 성향이 실제 행동으로 표현되거나 반대 성향이 올바르게 억제된 비율이 91.5%(366/400건)로 나타났고, 별도로 사람 평가자와 LLM 평가자가 추출된 페르소나의 품질을 점수화했다.
Figure 1: Overview of the MatrAIx simulated-user evaluation framework. Evaluators describe the target audience, and MatrAIx retrieves matching personas from Persona 8B to form an evaluation cohort. The four environments define how these persona agents interact, while MatrAIx Applications provides the task specifications they execute. Shared telemetry and task-owned verification preserve the evidence needed for subgroup and population-level reporting.
Figure 1: Overview of the MatrAIx simulated-user evaluation framework. Evaluators describe the target audience, and MatrAIx retrieves matching personas from Persona 8B to form an evaluation cohort. The four environments define how these persona agents interact, while MatrAIx Applications provides the task specifications they execute. Shared telemetry and task-owned verification preserve the evidence needed for subgroup and population-level reporting.
Table 1: Persona 8B schema overview. Dimension counts refer to emitted categorical attributes. Sources provide different kinds and strengths of grounding; they do not imply direct population estimates for every value.
Top-level groupDims.Representative attributesRepresentative grounding
Background238Age, region, language, education, family, career, industryPopulation statistics, household surveys, education and labor taxonomies
Psychology210Personality, values, worldview, motivation, riskValidated instruments, values surveys, schema design priors
Capability331Domain expertise, general skills, tools, programming, developer contextOccupational taxonomies, technology/developer surveys
Behavior and Interaction124Preferences, habits, interaction state, work practices, technology adoptionTime-use, consumer, workplace, and technology-use evidence
Lifestyle387Interests, media, culture, hobbies, sports, food, health, fitnessHealth statistics, consumption surveys, cultural sources
Total1,290
Figure 3: Controlled behavioral adherence across four environments. (a) Each cell pools five personas declaring one pole of an attribute and five declaring the opposite pole; 10/10 means that all five positive personas expressed the target and all five negative personas expressed the opposite value. (b) Overall successful-trial share by environment and the number of strong attributes, defined as at least four of five successes in both arms. An LLM judge evaluates the recorded trajectory or artifact. Overall, 366 of 400 trials succeed.
Figure 3: Controlled behavioral adherence across four environments. (a) Each cell pools five personas declaring one pole of an attribute and five declaring the opposite pole; 10/10 means that all five positive personas expressed the target and all five negative personas expressed the opposite value. (b) Overall successful-trial share by environment and the number of strong attributes, defined as at least four of five successes in both arms. An LLM judge evaluates the recorded trajectory or artifact. Overall, 366 of 400 trials succeed.
Table 2: Composition of the public Persona 1M coreset. “Human-grounded” identifies the origin of a record, not a guarantee that every extracted field is a verified fact. These counts mirror the Composition table of the dataset card (footnote * ‣ 1), which is the authoritative record for the release.
SourceReleased records
Wikipedia extraction323,438
Amazon Review extraction97,915
Stack Overflow survey extraction113,120
PRISM Alignment1,487
General Social Survey63,532
MatrAIx volunteer survey355
Human-grounded subtotal599,847
Full-DAG synthetic400,000
Total999,847
(b) Environment summary.
(b) Environment summary.
Table 3: Application-task coverage. Counts are unique specifications on the repository’s main branch together with the batch collections on its synthetic-task branches; individually contributed tasks still under review on open pull requests are not counted. “Other” aggregates more than 25 additional domains.
CommerceSoftwareFinanceHealthcareOtherTotal
Survey2021381411391621
AI Chatbot3111729311371
Web2220612
App051006
Total2071561611683181,010
Figure 4: Complete three-layer taxonomy of the Persona 8B schema. The left column gives the five top-level groups and their attribute totals; the middle column gives 16 subgroups; and the right column gives all 55 conceptual categories with counts. The hierarchy covers exactly 1,290 attributes and excludes latent or helper variables used only by the generation graph.
Figure 4: Complete three-layer taxonomy of the Persona 8B schema. The left column gives the five top-level groups and their attribute totals; the middle column gives 16 subgroups; and the right column gives all 55 conceptual categories with counts. The hierarchy covers exactly 1,290 attributes and excludes latent or helper variables used only by the generation graph.
Table 4: Complete index of the 43 schema categories. Counts sum to 1,290 attributes. Examples are labels from the authoritative dimension catalog; the complete attribute-level mapping is supplied as supp_material/persona_taxonomy_mapping.csv.
GroupSubgroupSchema categoryCountRepresentative attributes
BackgroundDemographicsDemographic: Core25Age bracket; region; gender identity
BackgroundDemographicsDemographic: Cultural2Cultural background; attitude toward immigration
BackgroundDemographicsDemographic: Family1Household size
BackgroundDemographicsDemographic: Life Events24Life stage; major life events; childhood environment
BackgroundLanguageLinguistic: Language53Primary language; English proficiency; multilingualism
BackgroundLanguageLinguistic: Communication37Expected tone; verbosity; communication preferences
BackgroundEducationLearning: Academic34Highest education; academic field; institution tier
BackgroundEducationLearning: Style1Learning style
BackgroundCareerProfessional: Career4Research output; seniority; years of experience
BackgroundCareerProfessional: Industry51Company size; role function; industry
BackgroundCareerDeveloper: Professional Context6Professional status; role archetype; contribution context
PsychologyPersonalityPersonality: Character34Domain stance; dominant trait; curiosity
PsychologyPersonalityPersonality: Big Five50Imagination; artistic interest; emotionality
PsychologyPersonalityPersonality: MBTI2Neurotype; Myers-Briggs type
PsychologyPersonalityPersonality: Relationships4Attachment anxiety; attachment avoidance; interpersonal agency
PsychologyWorldviewValues & Motivation46Core value; religiosity; economic motivation
PsychologyWorldviewWorldview: Beliefs67Political leaning; trust level; safety sensitivity
PsychologyDecision-MakingRisk & Decision7Risk tolerance; decision style; need for closure
CapabilityDomainsExpertise: Domains144Domain; subject specialty; technology savviness
CapabilitySkillsExpertise: Skills64Writing; copywriting; editing
CapabilitySkillsSkills: Tools69Excel; Google Sheets; Python
CapabilitySkillsSkills: Programming44Comment style; summary documentation; naming verbosity
CapabilitySkillsDeveloper: Code Maintenance10Complexity tolerance; modularity preference; type-system orientation
Behavior and InteractionPersonal BehaviorBehavior: Preferences34Modality preference; accessibility needs; media diet
Behavior and InteractionPersonal BehaviorBehavior: Habits30Journaling; meditation; use of to-do lists
Behavior and InteractionPersonal BehaviorBehavior: Time3Time pressure; sleep schedule; micromanagement aversion
Behavior and InteractionInteraction StateState: Emotional5Emotional state; intent; query complexity
Behavior and InteractionWork PracticesBehavior: Work2Work schedule; office versus remote work
Behavior and InteractionWork PracticesDeveloper: Open Source Behavior7Open-source activity; GitHub contribution mode; pull-request style
Behavior and InteractionWork PracticesDeveloper: Community Behavior4Stack Overflow use; participation style; help-seeking preference
Figure 5: The full persona DAG. All 1,308 graph nodes and 6,999 directed edges. Nodes are placed left to right by the topological order used in Equation 5 and grouped vertically into one lane per schema category, with lanes ordered by their mean topological position; marker area scales with a node’s directed degree. Edges are drawn as translucent curves, so darker bands mark the dependency bundles that connect demographic and educational roots to downstream expertise, interest, and behavior dimensions.
Figure 5: The full persona DAG. All 1,308 graph nodes and 6,999 directed edges. Nodes are placed left to right by the topological order used in Equation 5 and grouped vertically into one lane per schema category, with lanes ordered by their mean topological position; marker area scales with a node’s directed degree. Edges are drawn as translucent curves, so darker bands mark the dependency bundles that connect demographic and educational roots to downstream expertise, interest, and behavior dimensions.
Table 5: Detailed schema-to-source grounding map. The mapping records source families and their roles in schema design, prior estimation, dependency construction, compatibility rules, or downstream validation.
Schema groupFacetsGrounding rolesSources
BackgroundDemographics (52); language (90); education (35); career (61)Category definitions; population priors; household, language, education, and labor dependenciesUN World Population Prospects and Population Data [87, 86]; World Bank WDI and WorldPop [94, 97]; Eurostat, ACS PUMS, and IPUMS [21, 82, 38]; DHS, UNICEF MICS, and OECD Family Database [77, 85, 62]; Pew and World Values Survey [69, 96]; UNESCO UIS, World Bank Education Statistics, OECD INES, and OECD PISA [84, 93, 61, 64]; ILOSTAT, BLS OEWS, and O*NET [34, 81, 60]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39].
PsychologyPersonality (90); worldview (113); decision-making (7)Instrument and value-set design; selected prevalence estimates; validationIPIP and MIDUS [35, 56]; Pew and World Values Survey [69, 96]; GSS, European Social Survey, ISSP, and Gallup World Poll [58, 20, 36, 23]; Afrobarometer, Arab Barometer, Asian Barometer, Latinobarometro, and Eurobarometer [1, 3, 5, 46, 19]; ARDA [6].
CapabilityDomain expertise (144); general skills (64); tools (69); programming (44); developer context (10)Occupational and skill taxonomies; technology access and adoption; developer-tool prevalenceITU Statistics, World Bank WDI, DataReportal, and Pew Internet [37, 94, 15, 70]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]; O*NET [60].
Behavior and InteractionPersonal behavior (67); interaction state (5); work practices (13); technology use (39)Time-use and consumer priors; workplace behavior; technology and AI adoptionAmerican Time Use Survey, Consumer Expenditure Surveys, and OECD Time Use [79, 80, 63]; ITU, DataReportal, and Pew Internet [37, 15, 70]; Stack Overflow Survey, GitHub Octoverse, JetBrains Developer Ecosystem, and O*NET [75, 25, 39, 60].
LifestyleInterests (358); physical health (25); fitness (2); health lifestyle (2)Health and disability priors; consumption and time use; cultural and interest category designWHO GHO and IHME GBD [95, 33]; ACS PUMS, NHIS, BRFSS, DHS, and UNICEF MICS [82, 11, 10, 77, 85]; ATUS, CEX, and OECD Time Use [79, 80, 63]; FAOSTAT and UNESCO Culture Statistics [22, 83]; Eurobarometer, Pew, DataReportal, WVS, and Gallup World Poll [19, 69, 70, 15, 96, 23].
Figure 6: How much of the 1,290-item instrument the volunteer subset exercises. Each vertical bar is the median coverage of one interface category’s items, sorted, for all 43 categories. Labels carry the category’s leaf name, dropping the group prefix (Interests, Behavior, and so on) that would otherwise repeat under every bar. The vertical axis starts at 88.5% rather than zero because the finding is how narrow the range is, from 89.3% (Time, under Behavior) to 92.4% (Agent Adoption, under Developer; both extremes are named on their bars), which a zero-based axis would render as 43 nearly identical bars. The dashed line marks the 91.3% category median. The narrow range indicates that missingness is not concentrated in a small set of categories. Separately, 272 of the 355 records answer every one of the 1,290 items and the remaining 83 answer between 520 and 1,186.
Figure 6: How much of the 1,290-item instrument the volunteer subset exercises. Each vertical bar is the median coverage of one interface category’s items, sorted, for all 43 categories. Labels carry the category’s leaf name, dropping the group prefix (Interests, Behavior, and so on) that would otherwise repeat under every bar. The vertical axis starts at 88.5% rather than zero because the finding is how narrow the range is, from 89.3% (Time, under Behavior) to 92.4% (Agent Adoption, under Developer; both extremes are named on their bars), which a zero-based axis would render as 43 nearly identical bars. The dashed line marks the 91.3% category median. The narrow range indicates that missingness is not concentrated in a small set of categories. Separately, 272 of the 355 records answer every one of the 1,290 items and the remaining 83 answer between 520 and 1,186.
Table 6: One child dimension through the local CPD (illustrative). english_proficiency conditioned on primary_language=English and region=North America. The prior alone would give Native a 15% share; the two dependency factors raise it to 75%, and the compatibility mask removes None outright rather than merely making it unlikely. Numbers illustrate Equation 4 and are not the deployed parameter values.
vπi​(v)qlangqregionri​(v)mi​(v)πi​ri​mipθ​(v∣xPa⁡(i))
None0.300.020.100.02200.0000.000
Basic0.300.080.200.17810.0530.028
Fluent0.250.300.351.68010.4200.224
Native0.150.600.359.33311.4000.747
sum1.001.001.001.8731.000
Figure 8: The MatrAIx Playground. A user browses the persona population 8(a), inspects and filters individual persona records 8(b), configures a study over a persona cohort 8(c), runs an interactive evaluation of the system under test 8(d), and reads the aggregated population-level report 8(e), all without writing code.
Figure 8: The MatrAIx Playground. A user browses the persona population 8(a), inspects and filters individual persona records 8(b), configures a study over a persona cohort 8(c), runs an interactive evaluation of the system under test 8(d), and reads the aggregated population-level report 8(e), all without writing code.
Table 7: Detailed post-processing accounting for the audited baseline. Human and synthetic remaining counts have different scopes until the final row.
StageRejectedRemaining
Original corpus10,002,288,277
Contradiction filter239,31010,002,048,967
Human exact/MinHash deduplication41,5972,222,496 human
Synthetic projection deduplication252,936,3929,746,848,482 synthetic
Synthetic deterministic cutoff1,349,070,9788,397,777,504 synthetic
Audited baseline8,400,000,000
(b) Persona browser. Individual records are shown as cards carrying their grounded attributes (age, career, region, and further schema dimensions), and can be filtered to assemble an evaluation cohort.
(b) Persona browser. Individual records are shown as cards carrying their grounded attributes (age, career, region, and further schema dimensions), and can be filtered to assemble an evaluation cohort.
Table 8: Declared composition of the 355 released volunteer records. Shares are of the records answering each dimension, so the denominator differs by row. Computed from the public release, so the table matches what a reader downloading the dataset obtains.
DimensionAnsweredShare of those answering
Age bracket32125–34 22.1%, 55–64 19.0%, 35–44 15.9%, 18–24 13.7%, 45–54 13.4%, 65–74 8.4%, 75–84 4.0%, 85+ 3.4%
Gender identity322Woman 47.5%, Man 42.9%, Non-binary 3.7%, Self-described 3.4%, Prefer not to say 2.5%
Region329South Asia 20.4%, Sub-Saharan Africa 19.1%, East Asia 13.4%, Southeast Asia 10.9%, Latin America 9.7%, Western Europe 7.0%, MENA 6.7%, North America 5.8%, Eastern Europe 4.6%, Oceania 2.4%
Urbanicity328Rural 33.2%, Dense urban 22.3%, Small town 21.0%, Suburban 20.4%, Nomadic / remote 3.0%
Socioeconomic band340Low 31.8%, Lower-middle 30.0%, Middle 20.3%, Upper-middle 12.6%, High 5.3%
Employment330Full-time 25.8%, Retired 16.1%, Homemaker 14.2%, Self-employed 11.8%, Part-time 9.1%, Student 8.2%, Unemployed 7.6%, Gig / freelance 7.3%
(c) Study setup. A study is specified as a decision to test (here, a $2 price increase), with the persona cohort on the left and the task catalog on the right, before launching the simulation.
(c) Study setup. A study is specified as a decision to test (here, a $2 price increase), with the persona cohort on the left and the task catalog on the right, before launching the simulation.
Table 9: Large template-based task collections. Members of a collection are separate task specifications but reuse a common contract and verifier pattern.
CollectionTasksEnvironmentComposition
Synthetic persona surveys405Survey135 Finance, 135 Healthcare, and 135 Software questionnaires, generally sharing a common survey and reporting template.
Product surveys200SurveyTwenty product-research archetypes, including purchase intent, price sensitivity, recommendation, and retention, crossed with ten retail products.
Synthetic chatbots351AI ChatbotScenario families across 27 domains, including travel, legal, healthcare, telecommunications, real estate, insurance, finance, and education.
(d) Interactive evaluation. A persona-conditioned agent converses with the system under test in a live transcript, scored in real time against a task-owned rubric (here, an AI counselor).
(d) Interactive evaluation. A persona-conditioned agent converses with the system under test in a live transcript, scored in real time against a task-owned rubric (here, an AI counselor).
Table 10: Logical components of an application-task contract.
Contract componentDeclared contentAudit purpose
Task metadataStable name and version, environment type, product domain, difficulty, tags, artifact paths, and runtime budgetsIdentifies the recipe and prevents results from silently moving between task versions.
Persona-facing scenarioContext, user goal, constraints, disclosure policy, and required submissionDefines what every sampled persona is asked to do without exposing verifier internals.
Cohort strategyPersona-schema filters, sampling mode, stratification fields, sample size, weights, and seed policyMakes the target audience explicit and permits the cohort to be redrawn.
Product attachmentQuestionnaire or stimulus, chat endpoint or sidecar, website target, or native application backendIdentifies the system under test and how the runtime reaches it.
VerifierRequired artifacts, structured finding schema, objective checks, timeouts, and failure conditionsConverts a trial into reproducible outcomes with supporting evidence.
Reporting policyAggregations, subgroup facets, summaries, optional judge directives, and disclosure rulesDefines how trial findings become a cohort-level report.
(e) Population report. Trials aggregate into a population-level outcome (here, 61% projected retention) with response distributions, decision reasons, and the per-persona response-level dataset backing the headline number.
(e) Population report. Trials aggregate into a population-level outcome (here, 61% projected retention) with response distributions, decision reasons, and the per-persona response-level dataset backing the headline number.
Table 11: Environment-specific inputs and artifacts. App covers desktop and mobile native applications, including Linux, macOS, and iOS.
EnvironmentTask-owned inputsPrimary trial artifacts
SurveyStimulus, questionnaire schema, response constraints, and optional rationale promptsTyped responses, missing or invalid items, rationales, confidence, and completion summary.
AI ChatbotProduct endpoint or sidecar, user goal, staged-disclosure policy, turn and safety limitsFull transcript, tool or service events, resolution state, termination reason, and post-run feedback.
WebURL or hosted site, browser backend, exploration requirements, and submission schemaPage and action trace, considered options, screenshots where enabled, final submission, and terminal page or site state.
AppNative target, platform backend, initial state, permissions, and terminal-state checksScreenshot and action trace, exported files, application state, permission changes, and cross-application side effects.
Figure 9: Move-transition graphs for the meal-planning chat study, by economic motivation. 1,000 GPT-5.5 persona conversations, 500 per stratum, one fixed assistant; 13,120 move transitions in total. A conversation averages six exchanges and fourteen moves, so it traverses the graph in loops rather than one pass; edges curving back up the page are those returns. Each panel draws only its own moves and transitions and lays itself out, so shape differences reflect the data. Black edges trace each step’s most frequent successor; every drawn edge carries its transition count, and a pair of moves with no edge between them co-occurred fewer than the panel’s count floor of about 90 transitions.
Figure 9: Move-transition graphs for the meal-planning chat study, by economic motivation. 1,000 GPT-5.5 persona conversations, 500 per stratum, one fixed assistant; 13,120 move transitions in total. A conversation averages six exchanges and fourteen moves, so it traverses the graph in loops rather than one pass; edges curving back up the page are those returns. Each panel draws only its own moves and transitions and lays itself out, so shape differences reflect the data. Black edges trace each step’s most frequent successor; every drawn edge carries its transition count, and a pair of moves with no edge between them co-occurred fewer than the panel’s count floor of about 90 transitions.
Table 12: Complete primary-outcome results for the eight validation tasks. Range is the largest minus smallest model share in percentage points. The q column reports a three-arm χ2 test after Benjamini–Hochberg correction across the eight tasks (∗∗∗q<0.001, ∗q<0.05; ns otherwise).
TaskEnv.Primary outcomeShare giving that answerRangeq
GPTOpusHaiku(pt)
Annual checkupSurveyq31 = e50.5%41.4%29.7%20.8<0.001∗∗∗
Candy Land priceSurveyq_threshold = hesitate98.3%27.0%83.3%71.3<0.001∗∗∗
OpenBB honestyChatwouldStillContinueUse = unsure81.3%85.5%28.1%57.4<0.001∗∗∗
Meal planning†ChatadherenceLikelihood = 750.6%0.2%40.0%50.4<0.001∗∗∗
Notion plansWebdecision_subject_id = plus63.5%21.5%75.7%54.2<0.001∗∗∗
MIT OCW courseWebtask_course_level = Graduate40.0%34.5%47.5%13.0<0.001∗∗∗
News+ subscriptionAppclicked_get_started = true4.2%20.8%0.0%20.80.025∗
Stocks sentimentAppsentiment = hold60.0%30.0%50.0%30.00.153 ns
Table 13: Product-level conclusions under three agent models on identical cohorts. Denominators are trials in which the field was answered. Notion aggregates its three paid tiers against the free tier. †The meal-planning GPT 5.5 arm did not run the declared cohort and is not comparable with the other two arms.
TaskProduct-level measureGPT 5.5Opus 4.8Haiku 4.5
Annual checkupVery likely to schedule in time50.5% (505/1,000)41.4% (414/1,000)29.7% (297/1,000)
Candy Land priceHesitates or worse at new price98.3% (983/1,000)27.0% (270/1,000)83.3% (833/1,000)
OpenBB honestyWould not continue using18.5% (185/1,000)14.5% (132/909)71.9% (719/1,000)
Meal planning†Stated need fully satisfied46.2% (462/1,000)0.2% (2/1,000)28.1% (281/1,000)
Notion plansSelected a paid plan75.8% (776/1,024)23.2% (237/1,022)93.9% (958/1,020)
MIT OCW courseChose a graduate-level course40.0% (403/1,008)34.5% (347/1,007)47.5% (473/996)
News+ subscriptionSubscribed4.2% (1/24)20.8% (5/24)0.0% (0/24)
Stocks sentimentBuy opinion40.0% (8/20)70.0% (14/20)47.4% (9/19)
Table 14: Rank correlation between model arms’ orderings of the same persona subgroups. Significance marks are ordered GPT / Opus / Haiku (∙ significant after correction, ∘ not). A flat arm has no ordering to compare. †The GPT arm has a cohort-integrity exception. ‡App rows contain only three to eight personas per subgroup.
TaskPersona dimensionGroupsSig.Spearman ρ, exact p
GPT/OpusGPT/HaikuOpus/Haiku
Annual checkupage bracket6∘∘∙+0.26 p=0.329−0.43 p=0.822−0.77 p=0.971
Candy Land priceeconomic motivation4∘∘∘−0.40 p=0.792−0.20 p=0.625+0.80 p=0.167
OpenBB honestytrust level4∙∙∙+1.00 p=0.042+1.00 p=0.042+1.00 p=0.042
Meal planning†life stage4∘∘∘+0.95 p=0.083+0.32 p=0.500+0.33 p=0.417
Notion planscompany size8∘∘∘+0.44 p=0.138+0.60 p=0.059−0.24 p=0.725
MIT OCW courseacademic field8∘∘∘+0.93 p=0.001+0.09 p=0.420+0.06 p=0.452
News+ subscr.‡economic motivation3∘∘∘+0.00 p=0.667flatflat
Stocks sentiment‡risk tolerance5∘∘∘+0.47 p=0.267−0.92 p=1.000−0.67 p=0.933
Table 15: Persona fidelity across all three model pairs. Cohen’s κ corrects raw agreement for each model’s answer distribution. GPT and Opus age-band matches are indistinguishable from uniform guessing (q=0.84 and q=0.85); Haiku matches 1,000 of 1,000 trials.
GPT × OpusGPT × HaikuOpus × Haiku
Paired agreement over 88 joinable fields
Median Cohen’s κ0.0000.000+0.001
Fields at κ≤059 of 8850 of 8840 of 88
Fields reaching κ≥0.27 of 888 of 887 of 88
Fields at ≥50% agreement, κ<0.148 of 5625 of 3424 of 34
Self-report fidelity: age band matches the persona
GPT 5.5 16.3% (ns) Opus 4.8 16.9% (ns) Haiku 4.5 100.0%chance 16.7%
Table 16: Persona adherence per attribute and environment, both acting agents. Each environment is split into two sub-columns: Opus 4.8 (left) and GPT-5.6-sol (right), under the same Opus 4.8 judge. Each cell shows two groups of five persona icons—the left group the positive cohort, the right group the negative cohort. A filled icon (🚹) marks a persona whose behavior expressed the declared value (for the negative cohort, correctly expressed the opposite value—target suppressed); a faint icon (🚹) marks one that did not. More filled icons is better in both groups. Overall Opus 366/400=91.5% (33/40 cells strong, ≥4 filled per group) vs. GPT-5.6-sol 317/400=79.2%.
SurveyChatWebOS-App
AttributeOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-sol
code-comment-style🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
code-naming-verbosity🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
code-summary-documentation🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-emoji-use🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-humor🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-politeness🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-storytelling🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-use-of-jargon🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-verbosity🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
register🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
Table 17: Cited judge evidence, Survey environment. For each attribute, the behavior the judge identified in a positive persona (declared value) versus its matched negative persona (opposite value). Excerpts are quoted from the trial trajectories.
AttributePositive (declared)Negative (opposite)
code-comment-styleNearly every line carries an inline comment (# Iterate over every integer, # Check divisibility)Code contains no comments whatsoever
code-naming-verbositycalculate_average_of_passing_scores, number_of_passing_scoresVariables s, t, a, c, x, all single-letter
code-summary-documentationEvery function opens with a tldr: docstringProse plus code, no TLDR/summary header
cog-emoji-useMany emoji across a short paragraphNo emoji despite a casual app-store prompt
cog-humorPlayful, witty asides (“the boxes may yet win”)Measured, earnest tone, no jokes
cog-politeness“Might I kindly ask that you resend the document at your earliest convenience”“Hey, you forgot the attachment. Again.”: blunt and sarcastic
cog-storytellingA narrated scene (“a storm came through, half the lodge went dark”)Abstract, value-driven, no concrete scene
cog-use-of-jargonDense technical jargon (“TCP three-way handshake, SYN-ACK”)Plain terms (“translate the name into a numerical address”)
cog-verbosityRambling, multi-paragraph, tangential answerShort clipped fragments, no elaboration
registerStandard/formal phrasingColloquial (“cuppa and the papers”, “the wife’s doing a roast”)
Table 18: Extraction-quality metrics. Higher scores indicate better extraction quality.
MetricNameDefinition
M1Claim validityDoes each populated claim remain defensible when its value, cited evidence, and free-text description are considered jointly? A field is problematic if its value is indefensible, its evidence is plainly unrelated or non-supporting, or its description adds implausible unsupported specifics.
M2No over-claimingHas an attribute been populated without any plausible basis? Clearly unsupported assignments count as problems, while a plausible, appropriately labeled inference does not.
M3CoverageDid the extraction omit an important attribute that was clearly available in the source?
M4Internal consistencyDo extracted fields contradict one another, especially on core identity, role, region, life stage, language, or technology-use information?
M5Overall fidelity and plausibilityIs the extracted record usable for faithfully role-playing the source person, and are its inferred psychological, interest, and communication traits plausible and coherent?
Table 19: Paired LLM-judge results on all 1,000 extractions. Δ is GPT minus Claude; ≤1 is the proportion of paired scores that differ by at most one point.
MetricGPTClaudeΔ≤1
M1 Claim validity3.4333.937−0.50481.4%
M2 No over-claiming3.5323.646−0.11482.3%
M3 Coverage4.4654.031+0.43499.2%
M4 Internal consistency3.6063.932−0.32685.0%
M5 Fidelity and plausibility3.8294.109−0.28097.8%
Overall3.7733.931−0.15889.1%
Table 20: Human evaluation and within-one agreement on the 100-persona subset. H–H compares all 15 pairs of human raters; GPT–H and Claude–H compare each LLM judge with the six-rater human mean.
MetricHuman meanH–H ≤1GPT–H ≤1Claude–H ≤1
M14.10599.1%69.0%92.0%
M23.77092.2%81.0%88.0%
M34.22397.1%95.0%100.0%
M44.53798.6%55.0%92.0%
M54.04099.1%96.0%97.0%
Overall4.13597.2%79.2%93.8%
Table 21: Capability comparison across representative agent benchmarks, simulation systems, and synthetic-sampling methods. Systems, left to right: AgentBench [52], GAIA [55], WebArena [111], WebShop [102], Mind2Web [17], AppWorld [78], ToolLLM [72], BenchFlow [7], τ-bench [103], Generative Agents [66], SOTOPIA [113], OASIS [101], and Silicon sampling [4]. A check (✓) indicates explicit support in the published system; a cross (✗) indicates the capability is not stated. Dimensions read as follows: product-as-SUT evaluation treats the product, not the agent, as the subject under test; persona conditioning instantiates a user from a sampled population profile; behavioral grounding verifies that behavior tracks a target attribute rather than confounders; native-device execution covers real mobile/desktop apps beyond API demos; reproducible reporting separates verifier facts from reporting policy; trajectory inspection exposes per-trial traces; end-to-end validation closes the persona–task–product loop; silicon sampling substitutes simulated respondents for survey/UX pre-studies; reproducible cohorts deterministically re-instantiate the same cohort; and multi-agent simulation denotes persistent multi-agent social worlds (a deliberate non-goal for MatrAIx). The matrix compares capability coverage rather than providing an overall ranking.
Agent BenchGAIAWeb ArenaWeb ShopMind2 WebApp WorldTool LLMBench Flowτ-benchGenerative AgentsSOTOPIAOASISSilicon samplingMatrAIx
Product-as-SUT evaluation
Persona conditioning
Behavioral grounding
Native-device execution
Reproducible reporting
Trajectory inspection
End-to-end validation
Silicon sampling
Reproducible cohorts
Multi-agent simulation

실제로 확인된 결과

  • 400건 통제 실험에서 지시된 행동 성향이 표현되거나 올바르게 억제된 비율이 91.5%(366/400건)로 나타났다.
  • 1,000건의 추출 페르소나에 대해 GPT-5.5와 Claude 두 LLM 심사자가 평가한 결과 전체 평균 점수는 각각 3.773, 3.931이었고, 5,000개의 짝지은 점수 중 89.1%가 1점 이내로 일치했다.
  • 100명 서브셋에 대한 6인 사람 평가 평균은 5점 만점에 4.135였고, 3,000개 점수 중 84.7%가 4점 또는 5점이었다.
  • 사람 평균 점수와 비교했을 때 GPT는 79.2%, Claude는 93.8%의 경우에서 1점 이내로 일치했다.
  • 동일 코호트에 대해 세 모델을 바꿔 실행했을 때 한 제품 페이지의 유료 요금제 선택 비율이 23.2%에서 93.9%까지 벌어졌고, 88개의 공통 필드에서 모델 간 짝별 일치도(코헨의 카파)는 거의 0에 가까웠다.

어디에 쓸 수 있나

  • 신제품이나 앱 기능을 실제 사용자에게 공개하기 전에 다양한 배경의 가상 사용자로 초기 반응(가격 민감도, 사용성 문제 등)을 사전 스크리닝하는 데 활용할 수 있다.
  • 챗봇이나 AI 어시스턴트가 실패했을 때 사용자가 계속 이용할지, 응답 지연을 얼마나 견디는지 같은 상호작용 패턴을 미리 점검하는 데 쓸 수 있다.
  • 제품 버전을 바꾼 뒤 동일한 페르소나 집단과 과업으로 재실행해 변경 전후를 비교하는 반복 실험에 적용할 수 있다.
  • 특정 하위 집단(연령, 소득, 지역 등)별로 제품이 다르게 작동하는지 세분화해 살펴보는 코호트 단위 분석에 쓸 수 있다.

한계와 남은 검증

  • 페르소나 집단은 실제 인구의 확률 표본이 아니며, 페르소나 결과는 실제 사람이 어떻게 행동할지에 대한 직접적 증거가 아니라 가설을 만드는 용도로만 다뤄야 한다고 저자가 명시했다.
  • 페르소나를 연기하는 모델과 평가 대상 시스템이 같은 모델 계열을 공유할 경우, 좋은 결과가 실제 만족인지 모델이 자기 출력을 선호하는 편향인지 구분할 수 없는데, 이번 실험은 이 둘을 분리하는 실험(같은 과업에서 페르소나 모델과 시스템 모델을 교차시키는 실험)을 아직 수행하지 않았다.
  • 실제 사용자 상호작용에서 나타나는 정보 은닉, 반박, 포기 같은 행동을 페르소나가 실제 사람처럼 보이는지는 아직 직접 측정되지 않았으며, 이를 확인할 후속 실험(실제 대화 로그와의 비교, 개인 단위 대화 이어쓰기 검증 등)이 우선 과제로 남아 있다.
  • 공개된 코어셋은 내부 전체 인구가 아니라 필터링과 중복 제거를 거친 약 100만 명 규모이며, 합성 부분은 연령대·지역·성별·도시화 정도 등 4개 변수에만 맞춰 보정되어 나머지 1,290개 속성 전체의 결합분포를 대표하지는 않는다.
  • 건강, 금융, 고용 등 실제 사람에게 영향을 미치는 규제 대상 의사결정에는 시뮬레이션 결과만으로 결론을 내려서는 안 되며, 같은 과업과 도구로 사람 검증을 거쳐야 한다고 저자가 강조했다.

왜 중요한가

AI 제품이나 챗봇을 실제 사용자에게 테스트하기 전에, 다양한 배경을 가진 수십억 명 규모의 가상 사용자로 값싸고 빠르게 사전 점검할 수 있는 길을 제시한다. 다만 저자 스스로 페르소나 결과가 실제 사람 행동의 직접적 증거가 아니라 '가설을 만드는 도구'에 가깝다고 강조하므로, 실제 배포 전 사람 검증을 대체할 수는 없다.

이 논문의 용어

  • 페르소나(Persona) · 나이, 언어, 성향 등 여러 속성으로 정의된 가상의 사용자 프로필
  • 방향성 그래프(DAG) 샘플링 · 속성들을 서로 영향을 주는 순서대로 하나씩 뽑아, 나이와 학력처럼 서로 관련 있는 값들이 어긋나지 않게 인물을 생성하는 방법
  • 인간 기반(human-grounded) 레코드 · 위키백과, 리뷰, 설문 응답 등 실제 사람 데이터에서 뽑아 만든 페르소나 항목
  • 검증자(verifier) · 페르소나 행동 결과가 과업이 요구한 조건을 만족했는지 자동 또는 사람이 확인하는 절차
  • 코호트(cohort) · 특정 조건(연령대, 지역 등)으로 선택된 평가용 페르소나 집단

저자 · Xiaomin Li

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Xiaomin Li et al., arXiv:2608.04205, arxiv-nonexclusive