Operationalizing Narrative Entropy (Sn): A Two-Scene Registered Pilot Report and Pre-Validation Protocol
试图给故事带给读者的'信息负荷'打分,结果却出乎作者意料
作者首次将名为叙事熵(Narrative Entropy, Sn)的构想套用到两个真实场景上,人工计算出分数。塔伦蒂诺电影中九人快速对话的餐厅场景得分18.8,反而低于雷蒙德·卡佛短篇小说中单人内心独白段落的30.0分,与作者原本的直觉相反。作者没有为了迎合直觉去修改公式,而是如实记录这一矛盾,并提前公开了后续验证的实验计划。
METAL MEDIA 解读图
试图给故事带给读者的'信息负荷'打分,结果却出乎作者意料
- 01单独一名评分者按照公式Sn=If×Cb×t,对塔伦蒂诺电影对话场景和卡佛小说独白段落进行了人工计数打分
- 02结果:单人独白场景得分30.0,高于九人快速对话场景的18.8分,与作者最初的预期相反
- 03作者提出三种可能解释但未下结论:公式本身不完整、作者的直觉本身有误、或计数与假设的阅读速度等测量环节出了差错
- 04这个结果其实与论文所依据的理论框架一致——该框架认为需要读者自行推断的'隐藏式'表达比直接说明的'明示式'表达负荷更高,但作者拒绝将其作为事后自我辩护,而是设计了正式检验步骤
- 05后续计划已提前登记公开:扩展到四个场景、检验多名评分者之间的一致性,甚至测量读者阅读时的心率变异性和皮肤电反应等生理信号
他们做了什么
- 单独一名评分者按照公式Sn=If×Cb×t,对塔伦蒂诺电影对话场景和卡佛小说独白段落进行了人工计数打分
- 结果:单人独白场景得分30.0,高于九人快速对话场景的18.8分,与作者最初的预期相反
- 作者提出三种可能解释但未下结论:公式本身不完整、作者的直觉本身有误、或计数与假设的阅读速度等测量环节出了差错
- 这个结果其实与论文所依据的理论框架一致——该框架认为需要读者自行推断的'隐藏式'表达比直接说明的'明示式'表达负荷更高,但作者拒绝将其作为事后自我辩护,而是设计了正式检验步骤
- 后续计划已提前登记公开:扩展到四个场景、检验多名评分者之间的一致性,甚至测量读者阅读时的心率变异性和皮肤电反应等生理信号
| Metric | Reservoir Dogs (diner) | Cathedral (first block) |
|---|---|---|
| Duration (t) | 420 s (7.0 min) | 450 s (7.5 min @ 200 WPM) |
| Speaking / mentioned characters | 9 / 9 | 1 / 7 |
| Speaker turns | 92 | 1 (interior voice) |
| Spatial Matrix (Mp / Mn) | 4 / 4 | 4 / 1 |
| Time references | 8 | 15 |
| Topic shifts (total) | 11 | 19 |
| Causal Branching (Cb) | 1.57 / min | 2.53 / min |
| New information units | 20 | 27 |
| Uncertainty ratio | 0.60 | 0.44 |
| Information Friction (If) | 1.71 | 1.58 |
| Narrative Entropy (Sn) | 18.8 | 30.0 |
为什么重要
这是把一个模糊概念——故事给读者带来多少认知负荷——变成可实际计数的东西的尝试,并且在结果与预期不符时如实报告,而不是事后调整方法去迎合直觉。它为叙事与内容分析提供了一个讲究可证伪性的严谨测量设计案例。
本文术语
- 叙事熵 (Narrative Entropy, Sn) · 一种衡量故事给读者施加处理负荷速度的指标构想
- 信息摩擦 (Information Friction, If) · 衡量新引入信息在多大程度上未被充分说明、带有不确定性的数值
- 因果分支 (Causal Branching, Cb) · 统计叙事在话题之间切换频率的数值
- 预注册 (pre-registration) · 在收集数据之前就公开固定研究设计和判定标准,防止事后按结果调整
- 明示模式/隐藏模式 (Told mode / Shown mode) · 信息直接在表面说明(明示)与信息被隐藏、需读者自行推断(隐藏)两种表达方式
论文原文摘要(英文)
Narrative Entropy ($S_n$) is a proposed quantitative descriptor within the Bulut Doctrine, intended to capture the rate at which a narrative text imposes processing load on a reader. To date the construct has been defined theoretically but not operationalized against real texts. This report documents the first such operationalization (the v2.0 pilot): two narrative scenes -- the opening restaurant scene of Tarantino's Reservoir Dogs and the opening interior-monologue block of Carver's Cathedral -- were coded manually by a single rater and scored with the candidate formula $S_n = I_f \times C_b \times t$. The result was a divergence from the author's naive intuition: the single-voice monologue ($S_n = 30.0$) scored higher than the nine-character dialogue scene ($S_n = 18.8$). We treat this not as a result to be explained away but as the central finding, and we refuse post-hoc adjustment of the formula. Three competing interpretations are presented -- formula incompleteness, genuine high-load prose, and measurement error -- and the design that would discriminate among them is pre-registered. This v2.1 revision adds: (i) explicit acknowledgement that the divergence is consistent with the pre-existing architectural framework which privileges inferential reconstruction over surface declaration, and that what was called "contrary to expectation" in v2.0 reflected the author's anticipatory intuition rather than the methodology's own predictions; (ii) a pre-registered construct validity test for $I_f$, motivated by the observation that $I_f$ values were nearly equal across the two scenes (1.71 vs 1.58) despite the headline $S_n$ divergence. The document functions simultaneously as a pilot report ($n=2$) and as a pre-registration of the next-stage protocol. It does not claim that $S_n$ has been validated.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调