K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

arXiv:2608.150082026-08-14

把AI智能体的11种记忆存储方式放在同一条件下比拼,发现没有一种能通吃所有场景

研究者对长时间运行的大模型智能体常用的11种记忆存储方式做了统一对照评测,涵盖向量检索、知识图谱、层级摘要、参数微调等七大类。结果显示知识图谱式记忆在对话问答上表现最好,却在具身行动规划任务上拖后腿;而把经验提炼压缩的记忆方式则正好相反。结论是智能体记忆系统需要根据任务动态切换存储方式,而不是固定用一种。

METAL MEDIA 解读图

把AI智能体的11种记忆存储方式放在同一条件下比拼,发现没有一种能通吃所有场景

  1. 01对52个现有记忆增强系统的调查发现,62%的评测集中在LoCoMo和LongMemEval两个对话类基准上,只有21%的系统报告了延迟、成本等效率指标
  2. 02研究在3个模型骨干、4个基准、26项指标下,统一对比了7大类共11种记忆方式,包括稠密/稀疏向量索引、文本记录、知识图谱、层级树、经验蒸馏记忆、LoRA参数微调、KV缓存压缩等
  3. 03在长对话问答任务上,图谱加向量的混合方式(M5)准确率最高,但速度慢10到100倍;在具身行动规划任务(ALFWorld)上,把经验提炼成精简策略的记忆方式获胜,而单纯检索更多原始记忆反而降低任务成功率
  4. 04调节检索条数k发现:在长对话问答中准确率随k增大持续上升,但在行动规划任务中任务成功率随k增大而下降,原因是模型的注意力被检索内容抢走,忽略了当前需要关注的场景
  5. 05把对话长度从6千个token压力测试到26万个token时,知识图谱类记忆的处理速度大幅下降,而经验蒸馏类记忆在速度和质量上都保持稳定扩展
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 对52个现有记忆增强系统的调查发现,62%的评测集中在LoCoMo和LongMemEval两个对话类基准上,只有21%的系统报告了延迟、成本等效率指标
  2. 研究在3个模型骨干、4个基准、26项指标下,统一对比了7大类共11种记忆方式,包括稠密/稀疏向量索引、文本记录、知识图谱、层级树、经验蒸馏记忆、LoRA参数微调、KV缓存压缩等
  3. 在长对话问答任务上,图谱加向量的混合方式(M5)准确率最高,但速度慢10到100倍;在具身行动规划任务(ALFWorld)上,把经验提炼成精简策略的记忆方式获胜,而单纯检索更多原始记忆反而降低任务成功率
  4. 调节检索条数k发现:在长对话问答中准确率随k增大持续上升,但在行动规划任务中任务成功率随k增大而下降,原因是模型的注意力被检索内容抢走,忽略了当前需要关注的场景
  5. 把对话长度从6千个token压力测试到26万个token时,知识图谱类记忆的处理速度大幅下降,而经验蒸馏类记忆在速度和质量上都保持稳定扩展
Table 1: Memory substrates on LoCoMo, LME-S, and MAB (LRU, TTL, CR). P4 ↑: LLM-judge score; E15 ↓: per-query latency (s). Within each model band, green and red mark the per-column best and worst; bold marks the per-model P4 winner on each benchmark.
LoCoMoLME-SMAB LRUMAB TTLMAB CR
ModelSub.FamilyP4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓
Qwen3-8BExternal
M1Flat0.5400.770.4787.300.5789.940.5600.920.2701.65
M2Flat0.4700.330.4801.750.6203.230.4700.330.2701.39
M3Text0.5624.180.49555.580.563146.650.64027.150.2402.30
M4Struct0.4356.360.430225.670.493274.060.48012.650.2701.50
M5Struct0.64827.840.537187.210.578313.180.54023.700.4406.01
M6Hier0.5561.390.50719.960.49359.980.2903.240.2903.30
M8Refine0.5093.030.490102.850.592197.090.5601.020.3402.07
Internal
M9Weight0.3791.760.25051.500.30085.000.2208.500.2601.95
M10Act0.5894.550.1206.610.6766.580.4802.280.2701.64
M11Act0.3072.030.09014.610.63421.300.5600.560.2401.76
Qwen3-32B-AWQExternal
M1Flat0.5822.590.54812.000.67632.390.8101.810.4604.33
M2Flat0.4872.370.4977.310.66223.140.7801.200.4104.03
M3Text0.5866.470.56067.420.775477.970.85094.290.3403.94
M4Struct0.4528.700.490234.000.535278.930.71017.180.1901.73
M5Struct0.68329.220.627194.080.578333.950.63026.020.3809.38
M6Hier0.5735.200.59320.270.57773.160.3303.770.4108.70
M8Refine0.5335.530.533106.780.592228.660.8101.930.4403.39
Internal
M9Weight0.3994.730.290171.700.422253.630.25022.150.2805.50
M10Act0.66916.250.19022.980.62027.430.7602.430.3803.74
M11Act0.2744.170.11352.390.62025.000.8101.450.3002.65
Gemma-4-26B-A4B-ITExternal
M1Flat0.5740.930.5158.490.67616.500.77014.500.4604.09
M2Flat0.5550.490.5133.370.6629.180.7300.200.4603.88
M3Text0.5994.810.53031.780.720238.780.75036.210.4303.67
M4Struct0.4766.610.523230.390.535267.000.72016.120.4603.97
Table 2: Memory substrates evaluated on ALFWorld-unseen and BigCodeBench-Hard. Within each model band, green marks the best value on a metric, red marks the worst. Pavg is the mean per-task partial-success score. Bold marks the per-model performance winner (TSR for ALFWorld, Pass@1 for BCB). M10 is omitted from both agent-centric benchmarks because the cumulative context exceeds the tested context windows; M9 is omitted from the Gemma-4 band because it is incompatible with the A4B MoE routing.
ALFWorld-unseenBigCodeBench-Hard
ModelSub.FamilyTSR ↑𝐒𝐭𝐞𝐩𝐬|𝓢 ↓𝑷𝐚𝐯𝐠 ↑𝑬𝟐𝐩𝟗𝟎 ↓𝑬𝟏𝟓 ↓Pass@1 ↑𝑷𝐚𝐯𝐠 ↑𝑬𝟐𝐚𝐯𝐠 ↓𝑬𝟏𝟓 ↓
Qwen3-8B-NoMem5.712.213.513.24888.112.43.84
M1Flat5.25.96.616.16729.520.15.05
M2Flat6.76.09.415.764710.111.94.95
M3Text5.28.38.416.36538.112.45.35
M4Struct7.57.98.514.860510.120.44.65
M5Struct9.017.412.515.762015.517.128.328
M6Hier4.515.08.315.665612.222.24.87
M7Refine7.513.79.616.062912.216.94.617
M8Refine8.218.59.515.362014.924.05.119
M9Weight3.019.53.222.09217.44.612.713
M11Act11.914.214.513.555313.523.616.817
Qwen3-32B-AWQ-NoMem22.411.323.720.358317.625.63.94
M1Flat27.69.616.231.11.08k18.224.25.05
M2Flat21.610.815.430.81.09k19.621.85.05
M3Text26.99.921.830.395018.224.55.48
M4Struct26.912.424.328.681617.622.44.920
M5Struct23.111.423.027.485116.216.631.231
M6Hier23.99.922.826.881716.216.84.47
M7Refine32.111.129.927.271617.613.54.321
M8Refine22.412.421.328.088613.523.04.622
M9Weight29.912.127.421.461214.215.514.615
M11Act26.912.324.722.078313.513.821.321
Gemma-4-26B-A4B-IT-NoMem7.511.317.316.467114.219.519.820
M1Flat10.415.016.321.678618.226.421.922
M2Flat9.711.517.221.475319.627.020.621
M3Text10.414.814.920.578114.219.523.423
M4Struct8.214.513.922.594218.926.922.423
M5Struct4.58.59.419.974920.928.445.646
M6Hier10.411.418.121.174118.926.024.328
Table 3: Landscape of 52 memory-augmented LLM systems (2023–2026). Benchmark codes: LC = LoCoMo, LME = LongMemEval, MH = Multi-hop QA, DS = DialSim, DMR = DMR. Metric codes: A = accuracy, P = per-type, T = token, L = latency, R = runtime, M = memory size, K = API calls. #B: distinct benchmarks used. #E: efficiency dimensions reported (bold ≥2). Base: G = GPT-series, O = open-source. Summary: 81% GPT-family; 62% of benchmark pairs on LC + LME; only 21% report any efficiency metric; no system reaches ≥3 benchmarks × ≥3 efficiency dimensions.
SystemYearBench.Metrics#B#EBaseSystemYearBench.Metrics#B#EBase
Generative Agents [39]2023OtherA10GMem0 [3]2025LCA, P, T, L12G
MemGPT [36]2023DMRA10GMem-α [64]2025OtherA, P, M11O
MemoChat [29]2023OtherA10GMemAgent [75]2025OtherA10O
MemoryBank [82]2023OtherA10GMemOS [27]2025LC, LMEA, P20G
RSum [59]2023OtherA10GMemoria [44]2025LMEA, T, L12G
SCM [55]2023OtherA10GMemoryOS [19]2025LCA, P10G
AI PERSONA [61]2024OtherA10GMemory-R1 [72]2025LC, LME, OtherA, P30O
EM-LLM [7]2024OtherA, P10OMemory-T1 [5]2025LC, OtherA, P20O
HippoRAG [9]2024MHA, P10OMIRIX [63]2025LC, OtherA, P, M21G
LD-Agent [24]2024OtherA, P10GMMS [77]2025LCA, P10G
MemTree [42]2024OtherA10GNemori [32]2025LC, LMEA, P20G
RAPTOR [45]2024OtherA10GO-Mem [58]2025LC, LMEA, P20G
ReadAgent [23]2024OtherA10GPRINCIPLES [21]2025OtherA10G
THEANINE [34]2024OtherA10GPropRAG [56]2025MH, OtherA, P20O
A-Mem [70]2025LC, DSA, P, T21GR3Mem [62]2025OtherA10O
ComoRAG [57]2025OtherA10GRGMem [53]2025LC, OtherA, P20G
GAM [71]2025LCA10GRMM [51]2025LCA, P10G
H-Mem [74]2025LCA, P, R11OSeCom [38]2025LCA, P10G
Hindsight [22]2025LC, LMEA, P20GSGMem [66]2025LC, LMEA, P20G
HippoRAG 2 [10]2025MHA, P10OEMem [84]2025LC, LMEA, P20G
LightMem [6]2025LC, LMEA, P, T, L, R, K24GZep [40]2025DMR, LMEA, P, L21G
LiCoMemory [16]2025LC, LMEA, P, L21G
Agentic Memory [76]2026LCA, P10GMemMachine [60]2026LC, LME, MHA, P, T31G
EverMemOS [11]2026LC, LMEA, P20GMemori [1]2026LCA, P10G
Cognis [4]2026LC, LMEA, P20GMMM [13]2026LC, LMEA, P20G
MAGMA [17]2026LC, LMEA, P20GSwiftMem [52]2026LC, LMEA, P, L21G
TSM [48]2026LC, LMEA, P20G
Table 4: Memory substrate configurations. TF: training-free (✓) or requires weight modification (✗). LLM: auxiliary LLM required (✓) beyond final-answer generation. ‡ attention-guided utterance selection with re-prefill (Qwen3 hybrid-attention adaptation). § zero-shot LLM controller in place of PPO.
FamilyMethodTFLLMWriteReadMgmt
External Memory
Flat IndexM1 Dense Vector [68]Encode; add to indexDense top-k
M2 Sparse Vector [43]Tokenize; append to BM25Sparse top-k
Text RecordM3 Gist Index [23]LLM gist per page; index gistsLLM selects pages to expandText de-dup
StructuralM4 Evolving Notes [70]LLM note per write; build linksANN + link expansionEvolution-triggered neighbour rewrites
M5 Dual-Level Graph [8]LLM extracts entities + relations; build KGKG + vector hybrid (mix mode)Lazy batch indexing; cached graph
HierarchicalM6 Hierarchical Tree [45, 49]LLM assigns four levelsCollapsed-tree top-k; all levelsTop-level de-dup; decay
RefinementM7 Distilled Strategies [35]LLM judges; distills strategies; MaTTSDense over strategies; top-1Trajectory judge; strategy de-dup
M8 Skill Bundles§ [79]Cluster experiences into bundlesZero-shot LLM; all bundles prefilledMerge and prune on overflow
Internal Memory
WeightM9 Adapter Tuning [80]LLM QA pairs; LoRA answer-only lossDirect generation on base + adapter
ActivationM10 Full ContextAppend to bufferConcatenate; generateTruncate at window
M11 Episode-Clustered Re-prefill‡ [20]Cluster into episodes; keep 25%Match to centroid; re-prefill episodeEpisodic clustering; budgeting
Table 5: Auxiliary-LLM cost on a full LoCoMo run (ten conversations, 1,986 queries; GPT-4o-mini aux model, embedding excluded). Costs are one to two orders of magnitude above the lightest retrieval baselines (e.g. M2 BM25 issues zero aux-LLM calls and finishes in roughly ten minutes).
SubstrateMethodAuxiliary LLM callsExecution time
Text RecordMemGPT [36]2,73916.3 h
Text RecordMem0 [3]8,98411.8 h
StructuralZep [40]7,62423.1 h
Table 6: Metric taxonomy for evaluating memory substrates. Two families: Performance (P1–P6, P8–P12) captures answer accuracy and retrieval quality; Efficiency (E1–E15) captures latency, storage, LLM call costs, and aggregate totals. The Benchmark column records which benchmarks officially report each metric. Abbreviations: LC = LoCoMo, MA = MemoryAgentBench, ALF = ALFWorld, BC = BigCodeBench-Hard; All denotes substrate-level metrics applicable across all four. Arrows indicate the preferred direction.
IDMetricDefinitionBenchmark
Performance: answer accuracy and retrieval quality
P1 ↑Exact Match𝟙[𝚗𝚘𝚛𝚖(y^)=𝚗𝚘𝚛𝚖(y)]; lowercase, strip articles, punctuation, whitespaceLC
P2 ↑Token F1Harmonic mean of token level precision and recall between 𝚗𝚘𝚛𝚖⁡(y^) and 𝚗𝚘𝚛𝚖⁡(y)LC
P3 ↑BLEU 1Clipped unigram precision with brevity penalty: BP⋅∑wmin⁡(cwy^,cwy)/|y^|LC
P4 ↑LLM JudgeLLM labels the answer CORRECT/WRONG given (q,y,y^); fraction correct in [0,1]LC, MA
P5Compression Ratio|input tokens|/|stored tokens|; larger values indicate more lossy compression (unsigned)All
P6 ↑Recall@k|𝒢∩ℛ|/|𝒢|; coverage of gold evidence dia_idsLC, MA
P8 ↑Task SuccessFraction of episodes where all goal conditions are satisfied (binary per episode); ALFWorld primary metricALF
P9 ↑GC SuccessFraction of individual sub goal conditions satisfied (ALFRED inherited); reported in parentheses alongside P8ALF
P10 ↓Steps to GoalEnvironment actions per successful episode; auxiliary efficiency signal, not officially reported by ALFWorldALF
P11 ↑Pass@1Unit test pass rate under greedy decoding; BigCodeBench primary metricBC
P12 ↑SubEM𝟙[𝚗𝚘𝚛𝚖(y)⊆𝚗𝚘𝚛𝚖(y^)]; substring exact match. MA primary metric for Accurate Retrieval and Conflict ResolutionMA
Efficiency: latency, storage, LLM call costs, and aggregate totals
E1 ↓Memory SizeMemory unit’s data structures in bytes; for parametric methods (M9), includes edited weights or adapter sizeAll
E2 ↓Inference TimeTotal per query wall time: read() + generate(); mean, p50, p90 (ms)All
E3 ↓Write LatencyPer message write() wall time; mean, p50, p90 (ms)All
E4 ↓Retrieval LatencyPer query read() wall time; mean, p50, p90 (ms)All
E5 ↓Write TokensLLM tokens consumed during the write phase (input + output)All
E6 ↓Write CallsNumber of LLM calls during the write phaseAll
E7Retrieved TokensTokens surfaced by memory to support answer generationAll
E8 ↓Read TokensLLM tokens consumed during read phase calls (input + output)All
E9 ↓Read CallsNumber of LLM calls during the read phaseAll
E10 ↓Mgmt. TokensLLM tokens for memory management (dedup, update, delete, conflict resolution)All
E11 ↓Mgmt. CallsNumber of LLM calls for memory managementAll
E12 ↓Total TokensAggregate auxiliary-LLM tokens across phases: E5+E8+E10 (excludes retrieved tokens E7)All
E13 ↓Total CallsAggregate LLM calls across phases: E6+E9+E11All
E14 ↓Total Wall-clockEnd-to-end run time across the benchmark (s); includes write, read, and management phasesAll
E15 ↓Per-query LatencyE14/nq, the principal cost reported in the main tables (s/query)All
Table 7: Table for different substrate and model LoCoMo Performance.
ModelSub.FamilyP1 EMP2 F1P3 BLEUP4 JudgeP6 Recall@k
Qwen3-8BM1Flat0.2730.4660.2850.5400.696
M2Flat0.2830.4360.2410.4700.595
M3Text0.1900.3660.2690.5620.620
M4Struct0.2260.3550.2160.4350.465
M5Struct0.1960.3810.2520.6480.455
M6Hier0.2510.4360.2760.5560.797
M8Refine0.3010.4560.2520.5090.615
M9Weight0.0190.0990.0710.379
M10Act0.1590.3420.2850.589
M11Act0.0660.1610.1380.307
Qwen3-32B-AWQM1Flat0.2080.4080.3280.5820.696
M2Flat0.1870.3520.2720.4870.595
M3Text0.1730.3680.2890.5860.620
M4Struct0.1650.3200.2400.4520.465
M5Struct0.1340.3290.2860.6830.455
M6Hier0.1920.3870.3070.5730.797
M8Refine0.1900.3650.2810.5330.615
M9Weight0.0240.1010.0820.399
M10Act0.1900.4120.3370.669
M11Act0.0350.1180.0950.274
Gemma-4-26BM1Flat0.3290.5130.3110.5740.696
M2Flat0.3020.4510.2570.5550.595
M3Text0.3030.4840.3000.5990.620
M4Struct0.3100.4050.2320.4760.465
M5Struct0.2030.4510.2710.7190.455
M6Hier0.3100.4830.2920.5920.797
M8Refine0.3350.4820.2410.5540.615
M10Act0.3000.5150.3080.688
M11Act0.0850.1810.1490.330
Table 8: Table for different substrate and model LoCoMo Latency & Quality.
ModelSub.Family𝑬𝟐 (ms/q)𝑬𝟑 (ms/wr)𝑬𝟒 (ms/rd)𝑬𝟏𝟒 (s)𝑬𝟏𝟓 (s/q)𝑷𝟓
Qwen3-8BM1Flat5221992261,5340.771.00
M2Flat302016600.331.00
M3Text2,2921,6637818,2934.185.29
M4Struct1696,210912,6386.360.71
M5Struct20,6805,1169,61055,29027.840.44
M6Hier6836411952,7581.390.62
M8Refine4312,5701796,0193.031.09
M9Weight3351,4003,4981.76
M10Act4,1329,0274.55
M11Act1,8484,0372.03
Qwen3-32B-AWQM1Flat2,1741946485,1422.591.00
M2Flat2,153014,7032.371.00
M3Text4,3791,6231,10012,8526.475.19
M4Struct2,2896,3581217,2698.700.72
M5Struct21,9495,2219,58658,02529.220.43
M6Hier2,65065361010,3275.200.60
M8Refine2,7072,60016610,9915.531.11
M9Weight8603,8009,3864.73
M10Act14,77432,27516.25
M11Act3,7908,2804.17
Gemma-4-26BM1Flat6672001881,8500.931.00
M2Flat445019720.491.00
M3Text2,8701,6641,0189,5554.815.37
M4Struct3946,1451113,1296.610.73
M5Struct18,9825,1359,44451,54425.950.43
M6Hier7126261742,8221.420.61
M8Refine6002,5721666,3883.221.06
M10Act4,69810,2635.17
M11Act1,9674,2972.16
Table 9: Table for different substrate LoCoMo Storage, Tokens & Calls.
Sub.Family𝑬𝟏𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟖𝑬𝟗𝑬𝟏𝟎𝑬𝟏𝟏𝑬𝟏𝟐𝑬𝟏𝟑
M1Flat3.59M00889.8K000000
M2Flat173.7K00806.6K000000
M3Text257.4K311.8K39011.52M8.23M2,040008.54M2,430
M4Struct177.6K4.59M3,26896.2K003.16M2,3127.75M5,580
M5Struct167.8K33.0M9,8606.00M17.8M3,9800050.8M13,840
M6Hier4.33M381.8K9101.87M00127.8K190509.6K1,100
M8Refine3.40M1.46M481806.2K2.94M2,037242.9K514.64M2,569
M9Weighthost-dep
M10Act101.2K60.9M
M11Act100.7K15.3M
Table 10: P12 (SubEM), P2 (F1), P3 (BLEU-1) per (substrate, generator, sweep) for LongMemEval-S and MemoryAgentBench.
LME-SMAB LRUMAB TTLMAB CR
ModelIDFamilyP12P2P3P12P2P3P12P2P3P12P2P3
Qwen3-8BM1Flat0.3900.1700.1130.5920.4910.4170.5700.4760.4050.3300.2110.171
M2Flat0.3400.1700.1140.6060.5270.4480.4800.3990.3400.2800.1880.155
M3Text0.1230.0470.0290.5490.4790.4070.8400.5440.4620.2700.1920.163
M4Struct0.3030.1450.0940.4650.4190.3560.4900.4680.3970.2900.1210.103
M5Struct0.3800.1600.1030.5490.4910.4170.5800.4680.3970.5000.3620.308
M6Hier0.3570.1620.1110.5210.4190.3560.2900.2460.2100.2700.1710.138
M8Refine0.3730.1590.1010.5920.5030.4270.9600.5700.4840.3100.2720.231
M9Weight0.0830.0820.0580.2390.1920.1630.1700.1670.1420.1600.1080.089
M10Act0.0900.0750.0530.6480.5750.4880.1400.1300.1110.2700.1610.124
M11Act0.2020.1110.0650.5920.5390.4580.5600.4680.3970.0500.0370.032
Qwen3-32B-AWQM1Flat0.4430.1800.1220.6340.5750.4880.8200.6890.5850.2800.1340.093
M2Flat0.3770.1630.1090.6200.5630.4780.7900.6630.5640.2700.1360.091
M3Text0.2250.1240.0720.7040.6580.5600.8600.7220.6140.2800.1580.120
M4Struct0.3570.1600.1060.5210.4550.3870.6800.6660.5660.1400.0720.049
M5Struct0.4600.1590.1040.5920.4910.4170.6200.5100.4330.4400.1910.127
M6Hier0.4130.1760.1170.5770.6230.5290.3400.2810.2380.2800.1350.090
M8Refine0.4030.1660.1110.5770.5030.4270.9400.7900.6720.3100.1620.115
M9Weight0.1000.0980.0650.4080.3590.3050.3240.3180.2700.1600.0860.059
M10Act0.1370.0870.0610.6060.5270.4480.7600.6460.5490.2700.1280.085
M11Act0.0800.0800.0560.5920.5150.4380.8200.6890.5850.2300.1220.086
Gemma-4-26BM1Flat0.4330.2000.1350.6620.5750.4880.7700.6290.5350.5200.2670.223
M2Flat0.3570.1800.1220.6620.6730.5720.7300.6200.5270.5200.2480.202
M3Text0.2230.1230.0710.3520.2750.2340.7980.6160.5240.5200.3040.259
M4Struct0.3530.1660.1110.5210.4550.3870.6900.6160.5240.0800.0680.058
M5Struct0.4530.1920.1290.5350.3830.3260.7200.6120.5200.6600.2940.214
M6Hier0.3970.1750.1180.4930.4310.3660.2500.2120.1810.4800.2280.182
M8Refine0.3730.1640.1080.5490.4550.3870.9500.8070.6860.4700.3270.278
M10Act0.0630.0450.0390.5770.5790.4920.4600.3910.3320.5000.2630.222
M11Act0.2340.1290.0750.5720.4430.3760.7020.6160.5240.0500.0500.043
Table 11: LongMemEval-S storage and latency per substrate.
ModelIDFamily𝑬𝟏𝑬𝟐,𝐦𝐞𝐚𝐧𝑬𝟐,𝐩𝟗𝟎𝑬𝟏𝟒𝑬𝟏𝟓
Qwen3-8BM1Flat9.7M2,0754,6022,1917.30
M2Flat2.7M1,7473,5805251.75
M3Text4.5M4,5009,00016,67455.58
M4Struct1.6M1,8284,23367,702225.67
M5Struct32.0M8,50018,00056,162187.21
M6Hier27.5M19,9565,1955,98719.96
M8Refine22.7M1,8694,39230,855102.85
M9Weight80.0M2,4005,00015,45151.50
M10Act60K3,7445,5391,9836.61
M11Act80K1,9004,2004,38214.61
Qwen3-32B-AWQM1Flat9.7M7,41318,1003,60012.00
M2Flat2.7M7,30718,8242,1937.31
M3Text4.5M16,07632,15320,22667.42
M4Struct1.6M7,07018,01970,200234.00
M5Struct32.0M18,00035,00058,225194.08
M6Hier27.5M20,27117,8166,08220.27
M8Refine23.4M6,14915,19232,034106.78
M9Weight320M8,50019,00051,511171.70
M10Act60K14,65028,0226,89422.98
M11Act80K7,00017,00015,71752.39
Gemma-4-26BM1Flat9.7M4,1498,5502,5468.49
M2Flat2.7M3,3726,5011,0123.37
M3Text4.5M8,99817,9969,53431.78
M4Struct1.6M3,8059,02769,116230.39
M5Struct32.0M12,00023,00057,280190.93
M6Hier27.5M17,1899,0075,15717.19
M8Refine21.7M3,1246,54031,140103.80
M10Act60K3,3403,9452,0736.91
M11Act80K5,55018.50
Table 12: LongMemEval-S token usage (E5–E13).
ModelIDFamily𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟏𝟐𝑬𝟏𝟑𝑷𝟓
Qwen3-8BM1Flat00101K001.00
M2Flat00153K001.00
M3Text437K6715.0M5.3M7314.50
M4Struct4.4M1310108K4.4M13106.20
M5Struct10.3M2371375K10.4M24313.80
M6Hier280K8046K280K802.80
M8Refine1.8M175221K2.78M350412.50
M9Weight
M10Act1.9M
M11Act250K
Qwen3-32B-AWQM1Flat00101K001.00
M2Flat00153K001.00
M3Text437K6715.0M5.3M7314.50
M4Struct4.4M1310108K4.4M13106.20
M5Struct10.4M2369385K10.4M24293.80
M6Hier280K8046K280K802.80
M8Refine1.8M175220K2.78M350412.50
M9Weight
M10Act1.9M
M11Act250K
Gemma-4-26BM1Flat00104K001.00
M2Flat00158K001.00
M3Text437K6715.0M5.3M7314.50
M4Struct4.4M1310111K4.4M13106.20
M5Struct10.4M2377379K10.4M24373.80
M6Hier280K8046K280K802.80
M8Refine1.9M175223K2.98M350412.50
M9Weight
M10Act422K
M11Act250K
Table 13: Table for different substrate and model ALFWorld-unseen Performance.
ModelSub.FamilyP8 TSRStepsP10 Steps∣𝒮Pavg
Qwen3-8BNoMem5.747.012.213.5
M1Flat5.247.75.96.6
M2Flat6.747.06.09.4
M3Text5.247.88.38.4
M4Struct7.546.97.98.5
M5Struct9.047.117.412.5
M6Hier4.548.415.08.3
M7Refine7.547.313.79.6
M8Refine8.247.418.59.5
M9Weight3.049.119.53.2
M11Act11.945.714.214.5
Qwen3-32B-AWQNoMem22.441.611.323.7
M1Flat27.639.69.616.2
M2Flat21.641.410.815.4
M3Text26.939.49.921.8
M4Struct26.939.512.424.3
M5Struct23.141.111.423.0
M6Hier23.940.29.922.8
M7Refine32.137.511.129.9
M8Refine22.441.612.421.3
M9Weight29.938.712.127.4
M11Act26.939.512.324.7
Gemma-4-26BNoMem7.547.111.317.3
M1Flat10.446.315.016.3
M2Flat9.746.311.517.2
M3Text10.446.314.814.9
M4Struct8.247.114.513.9
M5Struct4.548.18.59.4
M6Hier10.446.011.418.1
M7Refine8.246.710.117.8
Table 14: Table for different substrate and model ALFWorld-unseen Latency & Storage.
ModelSub.Family𝑬𝟏𝑬𝟐𝑬𝟐,p90𝑬𝟑𝑬𝟒𝑬𝟏𝟒𝑬𝟏𝟓
Qwen3-8BNoMem012,72813,20065,392488
M1Flat2.17M14,08616,10418822090,048672
M2Flat1.49M13,74715,6920186,698647
M3Text1.43M13,65516,3181,62075487,502653
M4Struct942K12,90314,8436,134981,070605
M5Struct532K13,16715,7025,0219,41283,080620
M6Hier25.1M13,53915,64162419187,904656
M7Refine3.90M13,29715,9882,51917584,286629
M8Refine5.27M13,06915,3002,52117383,080620
M9Weight15.3M18,75722,0021,400123,414921
M11Act6.9K12,72813,50074,102553
Qwen3-32B-AWQNoMem020,10020,30078,122583
M1Flat2.17M20,50031,100192644144,7201,080
M2Flat1.49M20,30030,80001146,0601,090
M3Text1.43M20,90030,3001,6031,087127,300950
M4Struct942K20,60028,6006,28412109,344816
M5Struct532K20,73627,4005,1549,477114,034851
M6Hier25.1M20,40026,800642600109,478817
M7Refine3.90M20,20027,2002,57516595,944716
M8Refine5.27M20,50028,0002,571164118,724886
M9Weight41.9M15,82521,4003,80082,008612
M11Act6.9K20,10022,000104,922783
Gemma-4-26BNoMem016,06016,40089,914671
M1Flat2.17M16,80021,600195184105,324786
M2Flat1.49M16,50021,40001100,902753
M3Text1.43M16,85720,4661,6491,009104,654781
M4Struct942K20,00722,5176,08911126,228942
M5Struct532K15,56419,9065,0899,358100,366749
M6Hier25.1M16,11421,09961917399,294741
M7Refine3.90M16,86621,1072,542161105,592788
Table 15: Table for different substrate ALFWorld-unseen Token Accounting.
Sub.Family𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟖𝑬𝟗𝑬𝟏𝟎𝑬𝟏𝟏𝑬𝟏𝟐𝑬𝟏𝟑
M1Flat00391K000000
M2Flat003.66M000000
M3Text372K2008.14M1.84M1,340002.21M1,540
M4Struct1.62M1995.05M00480K3802.10M579
M5Struct1.42M3852611.21M770002.63M1,155
M6Hier357K460235K0084.6K92442K552
M7Refine398K200183K0000398K200
M8Refine931K57713.5K00277K3801.21M957
M9Weight8.7K
M11Act118K
Table 16: Table for different substrate and model BigCodeBench-Hard Performance & Quality.
ModelSub.FamilyPass@1 ↑𝑷𝐚𝐯𝐠 ↑𝑷𝟓
Qwen3-8BNoMem8.112.4
M1Flat9.520.11.000
M2Flat10.111.91.000
M3Text8.112.41.964
M4Struct10.120.40.982
M5Struct15.517.10.817
M6Hier12.222.20.820
M7Refine12.216.95.705
M8Refine14.924.04.432
M9Weight7.44.6
M11Act13.523.6
Qwen3-32B-AWQNoMem17.625.6
M1Flat18.224.21.000
M2Flat19.621.81.000
M3Text18.224.51.995
M4Struct17.622.40.910
M5Struct16.216.60.786
M6Hier16.216.80.799
M7Refine17.613.52.261
M8Refine13.523.03.575
M9Weight14.215.5
M11Act13.513.8
Gemma-4-26BNoMem14.219.5
M1Flat18.226.41.000
M2Flat19.627.01.000
M3Text14.219.52.444
M4Struct18.926.90.982
M5Struct20.928.40.813
M6Hier18.926.00.813
M7Refine16.925.11.102
Table 17: Table for different substrate and model BigCodeBench-Hard Latency & Storage.
ModelSub.Family𝑬𝟏𝑬𝟐𝑬𝟐,p90𝑬𝟑𝑬𝟒𝑬𝟏𝟒𝑬𝟏𝟓
Qwen3-8BNoMem150K3,8004,9405924
M1Flat1.32M4,9766,6311991837765.2
M2Flat245K4,8536,6000.137184.9
M3Text305K5,3006,8906,853467845.3
M4Struct150K4,6426,4321371387475.0
M5Struct239K28,27336,7008,77912,4324,18428.3
M6Hier3.42M4,7646,2341,4441849946.7
M7Refine990K4,6386,2569,0291832,51917.0
M8Refine1.22M5,0536,82510,4401972,83619.2
M9Weight14.6M12,70016,5101,92413
M11Act151K16,80021,8402,51617
Qwen3-32B-AWQNoMem150K3,9005,0705924
M1Flat1.32M5,0006,5002562257405
M2Flat245K5,0006,5000.137405
M3Text296K5,4007,0209,3091,8431,1848
M4Struct150K4,9006,37012,4603762,96020
M5Struct249K31,22128,2397,88814,7234,62131.2
M6Hier3.50M4,4005,7201,9952201,0367
M7Refine1.93M4,3005,59012,9941953,10821
M8Refine1.14M4,6005,98012,4471863,25622
M9Weight38.0M14,60018,9802,22015
M11Act161K21,30027,6903,15221.3
Gemma-4-26BNoMem150K19,80025,7402,96020
M1Flat1.32M21,91628,5332132893,28622.2
M2Flat245K20,60227,8600.023,04920.6
M3Text272K23,40030,42018,378303,46323.4
M4Struct150K22,37531,073173923,37222.8
M5Struct240K45,59547,4368,19413,0876,74845.6
M6Hier3.47M24,31033,6142,7172514,14128.0
M7Refine3.13M24,14931,5138,9213175,46036.9
Table 18: Table for different substrate and model BigCodeBench-Hard Token Accounting.
ModelSub.Family𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟖𝑬𝟗𝑬𝟏𝟎𝑬𝟏𝟏𝑬𝟏𝟐𝑬𝟏𝟑
Qwen3-8BNoMem4.75M
M1Flat002.29M000000
M2Flat002.26M000000
M3Text55.9K362.80M2.84M148002.90M184
M4Struct002.16M000000
M5Struct0040.7K116K14800116K148
M6Hier44.7K821.14M0011.5K1656.2K98
M7Refine244K200593K0000244K200
M8Refine308K200550K00146K200454K400
M9Weight0
M11Act1.44M
Qwen3-32B-AWQNoMem4.75M
M1Flat002.29M000000
M2Flat002.26M000000
M3Text55.6K362.76M2.81M148002.86M184
M4Struct524K199286K0000524K199
M5Struct0041.7K121K14800121K148
M6Hier45.1K86575K0013.5K2458.6K110
M7Refine245K20091.6K0000245K200
M8Refine293K20084.8K00148K200441K400
M9Weight0
M11Act1.48M
Gemma-4-26BNoMem4.75M
M1Flat002.29M000000
M2Flat002.26M000000
M3Text52.2K362.26M2.30M148002.35M184
M4Struct002.16M000000
M5Struct0040.1K177K14800177K148
M6Hier44.4K821.14M0012.6K2457.0K106
M7Refine161K200901K0000161K200

为什么重要

这为构建长期运行的编程助手、个人AI伴侣或机器人智能体提供了选择记忆架构的实证依据。它表明通用的智能体记忆系统必须根据任务类型和对话历史长度动态切换存储策略,而不能固定使用某一种方式。

本文术语

  • 记忆载体(memory substrate) · 智能体存放记忆所用的底层介质,如向量数据库、知识图谱或摘要树
  • 稠密/稀疏向量索引 · 把文本转成数值向量做相似度检索(稠密),或基于关键词匹配检索(稀疏)
  • 知识图谱(结构化存储) · 把实体和关系以节点边形式连接存储,便于多跳推理
  • top-k检索 · 只取与问题最相关的k条存储记忆来生成答案
  • LLM裁判评分 · 用另一个大模型(如GPT-4o-mini)自动给答案质量打分,替代人工评审

无法转载的图表

  • Figure 1: Evaluation landscape of 52 memory-augmented LLM systems (2023–2026). (a) Benchmark adoption: 62% of usage concentrates on LoCoMo and LongMemEval (b) Metric coverage: every system reports accuracy but only 21% report any efficiency metric; 81% of systems use GPT-family backbones. (c) Breadth and efficiency: half of systems use a single benchmark with no efficiency metrics; no system simultaneously achieves broad coverage and comprehensive efficiency reporting.
  • Figure 2: Overview of the harness evaluation. 11 memory methods (M1–M11) spanning seven substrate families are evaluated on identical interaction histories across four benchmarks, instrumented with 26 metrics in two families (performance and efficiency).
  • Figure 3: Performance vs. latency across four benchmarks on Qwen3-32B-AWQ. (a) LoCoMo and (b) LongMemEval-S: absolute P4. (c) ALFWorld-unseen and (d) BigCodeBench-Hard: Δ-performance vs. latency overhead anchored at NoMem (No memory baseline).
  • Figure 4: Retrieval breadth k scales oppositely on long-context QA versus sequential decision-making. Top: LoCoMo (Qwen3-8B, n=1986, k∈{1,2,5,10,20}). Bottom: ALFWorld-unseen (Qwen3-32B-AWQ, n=134, k∈{1,2,3,4,5}). green when the trend is in the preferred direction, red otherwise. No Mem stands for no memory baseline.
  • Figure 5: Response-token attention under M1 retrieval for (a) LoCoMo (b) ALFWorld.
  • Figure 6: MAB CR scalability on Qwen3-32B-AWQ: P4 (bars, left) and per-query latency (lines, right, log) at 6K, 32K, 262K tokens.
在原文中查看图表 →

论文原文摘要(英文)

Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms. Across three backbone models and four benchmark suites spanning user-centric question answering and agent-centric decision-making, we instrument 26 performance and efficiency metrics under a unified harness. Our results show that no single substrate consistently dominates: broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context. Scalability introduces a further routing axis, as substrates that perform well at moderate history lengths can become costly or brittle at longer horizons. These findings motivate substrate routing as a necessary component of adaptive agent memory systems and provide empirical guidance for designing efficient, reliable, and regime-aware long-term memory for LLM agents. Code will be made available upon acceptance.

作者 · Wei-Chieh Huang

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道