컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI 에이전트의 '기억 저장 방식' 11가지를 똑같은 조건에서 겨뤄보니, 만능은 없었다

arXiv:2608.150082026-08-14

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

AI 에이전트의 '기억 저장 방식' 11가지를 똑같은 조건에서 겨뤄보니, 만능은 없었다

장기간 작동하는 AI 에이전트가 과거 경험을 저장하고 꺼내 쓰는 다양한 방식(메모리 서브스트레이트) 11가지를 벡터 검색부터 지식그래프, 계층형 요약, 파라미터 학습까지 같은 조건에서 비교했다. 대화형 질의응답에서는 지식그래프 방식이 앞섰지만 로봇 행동 계획 같은 과제에서는 오히려 성능을 깎아먹었고, 반대로 경험을 압축하는 방식은 그 반대였다. 즉 상황에 맞춰 저장 방식을 바꿔 쓰는 '라우팅'이 필요하다는 결론이다.

METAL MEDIA 해설 도표

AI 에이전트의 '기억 저장 방식' 11가지를 똑같은 조건에서 겨뤄보니, 만능은 없었다

  1. 01기존 52개 메모리 시스템 조사 결과 62%가 LoCoMo, LongMemEval 두 대화 벤치마크에만 몰려 있고, 정작 속도·비용 같은 효율성 지표는 21%만 보고했다
  2. 02덴스/스파스 벡터 검색, 텍스트 기록, 지식그래프, 계층형 트리, 경험 정제, 파라미터 미세조정, KV캐시 압축 등 7개 계열 11개 방식을 3개 모델·4개 벤치마크·26개 지표로 통제 비교했다
  3. 03대화 질의응답에서는 지식그래프+벡터 결합 방식(M5)이 최고 성능을 냈지만 속도는 10~100배 느렸고, 로봇 행동계획(ALFWorld)에서는 경험을 요약·정제한 방식이 이기고 단순 검색량을 늘리면 오히려 성공률이 떨어졌다
  4. 04검색 범위를 늘릴수록 질의응답 성능은 계속 오르지만 로봇 행동계획 성공률은 떨어지는 정반대 패턴을 확인했고, 모델의 주의력이 검색된 내용 쪽으로 쏠리면서 정작 봐야 할 현재 상황을 놓치는 게 원인으로 나타났다
  5. 05대화 길이를 6천에서 26만 토큰까지 늘려보니 지식그래프 방식은 처리 속도가 급격히 느려진 반면 경험 압축 방식은 속도 저하 없이 안정적으로 확장됐다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 52개 메모리 시스템 조사 결과 62%가 LoCoMo, LongMemEval 두 대화 벤치마크에만 몰려 있고, 정작 속도·비용 같은 효율성 지표는 21%만 보고했다
  2. 덴스/스파스 벡터 검색, 텍스트 기록, 지식그래프, 계층형 트리, 경험 정제, 파라미터 미세조정, KV캐시 압축 등 7개 계열 11개 방식을 3개 모델·4개 벤치마크·26개 지표로 통제 비교했다
  3. 대화 질의응답에서는 지식그래프+벡터 결합 방식(M5)이 최고 성능을 냈지만 속도는 10~100배 느렸고, 로봇 행동계획(ALFWorld)에서는 경험을 요약·정제한 방식이 이기고 단순 검색량을 늘리면 오히려 성공률이 떨어졌다
  4. 검색 범위를 늘릴수록 질의응답 성능은 계속 오르지만 로봇 행동계획 성공률은 떨어지는 정반대 패턴을 확인했고, 모델의 주의력이 검색된 내용 쪽으로 쏠리면서 정작 봐야 할 현재 상황을 놓치는 게 원인으로 나타났다
  5. 대화 길이를 6천에서 26만 토큰까지 늘려보니 지식그래프 방식은 처리 속도가 급격히 느려진 반면 경험 압축 방식은 속도 저하 없이 안정적으로 확장됐다
Table 1: Memory substrates on LoCoMo, LME-S, and MAB (LRU, TTL, CR). P4 ↑: LLM-judge score; E15 ↓: per-query latency (s). Within each model band, green and red mark the per-column best and worst; bold marks the per-model P4 winner on each benchmark.
LoCoMoLME-SMAB LRUMAB TTLMAB CR
ModelSub.FamilyP4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓P4 ↑𝑬𝟏𝟓 ↓
Qwen3-8BExternal
M1Flat0.5400.770.4787.300.5789.940.5600.920.2701.65
M2Flat0.4700.330.4801.750.6203.230.4700.330.2701.39
M3Text0.5624.180.49555.580.563146.650.64027.150.2402.30
M4Struct0.4356.360.430225.670.493274.060.48012.650.2701.50
M5Struct0.64827.840.537187.210.578313.180.54023.700.4406.01
M6Hier0.5561.390.50719.960.49359.980.2903.240.2903.30
M8Refine0.5093.030.490102.850.592197.090.5601.020.3402.07
Internal
M9Weight0.3791.760.25051.500.30085.000.2208.500.2601.95
M10Act0.5894.550.1206.610.6766.580.4802.280.2701.64
M11Act0.3072.030.09014.610.63421.300.5600.560.2401.76
Qwen3-32B-AWQExternal
M1Flat0.5822.590.54812.000.67632.390.8101.810.4604.33
M2Flat0.4872.370.4977.310.66223.140.7801.200.4104.03
M3Text0.5866.470.56067.420.775477.970.85094.290.3403.94
M4Struct0.4528.700.490234.000.535278.930.71017.180.1901.73
M5Struct0.68329.220.627194.080.578333.950.63026.020.3809.38
M6Hier0.5735.200.59320.270.57773.160.3303.770.4108.70
M8Refine0.5335.530.533106.780.592228.660.8101.930.4403.39
Internal
M9Weight0.3994.730.290171.700.422253.630.25022.150.2805.50
M10Act0.66916.250.19022.980.62027.430.7602.430.3803.74
M11Act0.2744.170.11352.390.62025.000.8101.450.3002.65
Gemma-4-26B-A4B-ITExternal
M1Flat0.5740.930.5158.490.67616.500.77014.500.4604.09
M2Flat0.5550.490.5133.370.6629.180.7300.200.4603.88
M3Text0.5994.810.53031.780.720238.780.75036.210.4303.67
M4Struct0.4766.610.523230.390.535267.000.72016.120.4603.97
Table 2: Memory substrates evaluated on ALFWorld-unseen and BigCodeBench-Hard. Within each model band, green marks the best value on a metric, red marks the worst. Pavg is the mean per-task partial-success score. Bold marks the per-model performance winner (TSR for ALFWorld, Pass@1 for BCB). M10 is omitted from both agent-centric benchmarks because the cumulative context exceeds the tested context windows; M9 is omitted from the Gemma-4 band because it is incompatible with the A4B MoE routing.
ALFWorld-unseenBigCodeBench-Hard
ModelSub.FamilyTSR ↑𝐒𝐭𝐞𝐩𝐬|𝓢 ↓𝑷𝐚𝐯𝐠 ↑𝑬𝟐𝐩𝟗𝟎 ↓𝑬𝟏𝟓 ↓Pass@1 ↑𝑷𝐚𝐯𝐠 ↑𝑬𝟐𝐚𝐯𝐠 ↓𝑬𝟏𝟓 ↓
Qwen3-8B-NoMem5.712.213.513.24888.112.43.84
M1Flat5.25.96.616.16729.520.15.05
M2Flat6.76.09.415.764710.111.94.95
M3Text5.28.38.416.36538.112.45.35
M4Struct7.57.98.514.860510.120.44.65
M5Struct9.017.412.515.762015.517.128.328
M6Hier4.515.08.315.665612.222.24.87
M7Refine7.513.79.616.062912.216.94.617
M8Refine8.218.59.515.362014.924.05.119
M9Weight3.019.53.222.09217.44.612.713
M11Act11.914.214.513.555313.523.616.817
Qwen3-32B-AWQ-NoMem22.411.323.720.358317.625.63.94
M1Flat27.69.616.231.11.08k18.224.25.05
M2Flat21.610.815.430.81.09k19.621.85.05
M3Text26.99.921.830.395018.224.55.48
M4Struct26.912.424.328.681617.622.44.920
M5Struct23.111.423.027.485116.216.631.231
M6Hier23.99.922.826.881716.216.84.47
M7Refine32.111.129.927.271617.613.54.321
M8Refine22.412.421.328.088613.523.04.622
M9Weight29.912.127.421.461214.215.514.615
M11Act26.912.324.722.078313.513.821.321
Gemma-4-26B-A4B-IT-NoMem7.511.317.316.467114.219.519.820
M1Flat10.415.016.321.678618.226.421.922
M2Flat9.711.517.221.475319.627.020.621
M3Text10.414.814.920.578114.219.523.423
M4Struct8.214.513.922.594218.926.922.423
M5Struct4.58.59.419.974920.928.445.646
M6Hier10.411.418.121.174118.926.024.328
Table 3: Landscape of 52 memory-augmented LLM systems (2023–2026). Benchmark codes: LC = LoCoMo, LME = LongMemEval, MH = Multi-hop QA, DS = DialSim, DMR = DMR. Metric codes: A = accuracy, P = per-type, T = token, L = latency, R = runtime, M = memory size, K = API calls. #B: distinct benchmarks used. #E: efficiency dimensions reported (bold ≥2). Base: G = GPT-series, O = open-source. Summary: 81% GPT-family; 62% of benchmark pairs on LC + LME; only 21% report any efficiency metric; no system reaches ≥3 benchmarks × ≥3 efficiency dimensions.
SystemYearBench.Metrics#B#EBaseSystemYearBench.Metrics#B#EBase
Generative Agents [39]2023OtherA10GMem0 [3]2025LCA, P, T, L12G
MemGPT [36]2023DMRA10GMem-α [64]2025OtherA, P, M11O
MemoChat [29]2023OtherA10GMemAgent [75]2025OtherA10O
MemoryBank [82]2023OtherA10GMemOS [27]2025LC, LMEA, P20G
RSum [59]2023OtherA10GMemoria [44]2025LMEA, T, L12G
SCM [55]2023OtherA10GMemoryOS [19]2025LCA, P10G
AI PERSONA [61]2024OtherA10GMemory-R1 [72]2025LC, LME, OtherA, P30O
EM-LLM [7]2024OtherA, P10OMemory-T1 [5]2025LC, OtherA, P20O
HippoRAG [9]2024MHA, P10OMIRIX [63]2025LC, OtherA, P, M21G
LD-Agent [24]2024OtherA, P10GMMS [77]2025LCA, P10G
MemTree [42]2024OtherA10GNemori [32]2025LC, LMEA, P20G
RAPTOR [45]2024OtherA10GO-Mem [58]2025LC, LMEA, P20G
ReadAgent [23]2024OtherA10GPRINCIPLES [21]2025OtherA10G
THEANINE [34]2024OtherA10GPropRAG [56]2025MH, OtherA, P20O
A-Mem [70]2025LC, DSA, P, T21GR3Mem [62]2025OtherA10O
ComoRAG [57]2025OtherA10GRGMem [53]2025LC, OtherA, P20G
GAM [71]2025LCA10GRMM [51]2025LCA, P10G
H-Mem [74]2025LCA, P, R11OSeCom [38]2025LCA, P10G
Hindsight [22]2025LC, LMEA, P20GSGMem [66]2025LC, LMEA, P20G
HippoRAG 2 [10]2025MHA, P10OEMem [84]2025LC, LMEA, P20G
LightMem [6]2025LC, LMEA, P, T, L, R, K24GZep [40]2025DMR, LMEA, P, L21G
LiCoMemory [16]2025LC, LMEA, P, L21G
Agentic Memory [76]2026LCA, P10GMemMachine [60]2026LC, LME, MHA, P, T31G
EverMemOS [11]2026LC, LMEA, P20GMemori [1]2026LCA, P10G
Cognis [4]2026LC, LMEA, P20GMMM [13]2026LC, LMEA, P20G
MAGMA [17]2026LC, LMEA, P20GSwiftMem [52]2026LC, LMEA, P, L21G
TSM [48]2026LC, LMEA, P20G
Table 4: Memory substrate configurations. TF: training-free (✓) or requires weight modification (✗). LLM: auxiliary LLM required (✓) beyond final-answer generation. ‡ attention-guided utterance selection with re-prefill (Qwen3 hybrid-attention adaptation). § zero-shot LLM controller in place of PPO.
FamilyMethodTFLLMWriteReadMgmt
External Memory
Flat IndexM1 Dense Vector [68]Encode; add to indexDense top-k
M2 Sparse Vector [43]Tokenize; append to BM25Sparse top-k
Text RecordM3 Gist Index [23]LLM gist per page; index gistsLLM selects pages to expandText de-dup
StructuralM4 Evolving Notes [70]LLM note per write; build linksANN + link expansionEvolution-triggered neighbour rewrites
M5 Dual-Level Graph [8]LLM extracts entities + relations; build KGKG + vector hybrid (mix mode)Lazy batch indexing; cached graph
HierarchicalM6 Hierarchical Tree [45, 49]LLM assigns four levelsCollapsed-tree top-k; all levelsTop-level de-dup; decay
RefinementM7 Distilled Strategies [35]LLM judges; distills strategies; MaTTSDense over strategies; top-1Trajectory judge; strategy de-dup
M8 Skill Bundles§ [79]Cluster experiences into bundlesZero-shot LLM; all bundles prefilledMerge and prune on overflow
Internal Memory
WeightM9 Adapter Tuning [80]LLM QA pairs; LoRA answer-only lossDirect generation on base + adapter
ActivationM10 Full ContextAppend to bufferConcatenate; generateTruncate at window
M11 Episode-Clustered Re-prefill‡ [20]Cluster into episodes; keep 25%Match to centroid; re-prefill episodeEpisodic clustering; budgeting
Table 5: Auxiliary-LLM cost on a full LoCoMo run (ten conversations, 1,986 queries; GPT-4o-mini aux model, embedding excluded). Costs are one to two orders of magnitude above the lightest retrieval baselines (e.g. M2 BM25 issues zero aux-LLM calls and finishes in roughly ten minutes).
SubstrateMethodAuxiliary LLM callsExecution time
Text RecordMemGPT [36]2,73916.3 h
Text RecordMem0 [3]8,98411.8 h
StructuralZep [40]7,62423.1 h
Table 6: Metric taxonomy for evaluating memory substrates. Two families: Performance (P1–P6, P8–P12) captures answer accuracy and retrieval quality; Efficiency (E1–E15) captures latency, storage, LLM call costs, and aggregate totals. The Benchmark column records which benchmarks officially report each metric. Abbreviations: LC = LoCoMo, MA = MemoryAgentBench, ALF = ALFWorld, BC = BigCodeBench-Hard; All denotes substrate-level metrics applicable across all four. Arrows indicate the preferred direction.
IDMetricDefinitionBenchmark
Performance: answer accuracy and retrieval quality
P1 ↑Exact Match𝟙[𝚗𝚘𝚛𝚖(y^)=𝚗𝚘𝚛𝚖(y)]; lowercase, strip articles, punctuation, whitespaceLC
P2 ↑Token F1Harmonic mean of token level precision and recall between 𝚗𝚘𝚛𝚖⁡(y^) and 𝚗𝚘𝚛𝚖⁡(y)LC
P3 ↑BLEU 1Clipped unigram precision with brevity penalty: BP⋅∑wmin⁡(cwy^,cwy)/|y^|LC
P4 ↑LLM JudgeLLM labels the answer CORRECT/WRONG given (q,y,y^); fraction correct in [0,1]LC, MA
P5Compression Ratio|input tokens|/|stored tokens|; larger values indicate more lossy compression (unsigned)All
P6 ↑Recall@k|𝒢∩ℛ|/|𝒢|; coverage of gold evidence dia_idsLC, MA
P8 ↑Task SuccessFraction of episodes where all goal conditions are satisfied (binary per episode); ALFWorld primary metricALF
P9 ↑GC SuccessFraction of individual sub goal conditions satisfied (ALFRED inherited); reported in parentheses alongside P8ALF
P10 ↓Steps to GoalEnvironment actions per successful episode; auxiliary efficiency signal, not officially reported by ALFWorldALF
P11 ↑Pass@1Unit test pass rate under greedy decoding; BigCodeBench primary metricBC
P12 ↑SubEM𝟙[𝚗𝚘𝚛𝚖(y)⊆𝚗𝚘𝚛𝚖(y^)]; substring exact match. MA primary metric for Accurate Retrieval and Conflict ResolutionMA
Efficiency: latency, storage, LLM call costs, and aggregate totals
E1 ↓Memory SizeMemory unit’s data structures in bytes; for parametric methods (M9), includes edited weights or adapter sizeAll
E2 ↓Inference TimeTotal per query wall time: read() + generate(); mean, p50, p90 (ms)All
E3 ↓Write LatencyPer message write() wall time; mean, p50, p90 (ms)All
E4 ↓Retrieval LatencyPer query read() wall time; mean, p50, p90 (ms)All
E5 ↓Write TokensLLM tokens consumed during the write phase (input + output)All
E6 ↓Write CallsNumber of LLM calls during the write phaseAll
E7Retrieved TokensTokens surfaced by memory to support answer generationAll
E8 ↓Read TokensLLM tokens consumed during read phase calls (input + output)All
E9 ↓Read CallsNumber of LLM calls during the read phaseAll
E10 ↓Mgmt. TokensLLM tokens for memory management (dedup, update, delete, conflict resolution)All
E11 ↓Mgmt. CallsNumber of LLM calls for memory managementAll
E12 ↓Total TokensAggregate auxiliary-LLM tokens across phases: E5+E8+E10 (excludes retrieved tokens E7)All
E13 ↓Total CallsAggregate LLM calls across phases: E6+E9+E11All
E14 ↓Total Wall-clockEnd-to-end run time across the benchmark (s); includes write, read, and management phasesAll
E15 ↓Per-query LatencyE14/nq, the principal cost reported in the main tables (s/query)All
Table 7: Table for different substrate and model LoCoMo Performance.
ModelSub.FamilyP1 EMP2 F1P3 BLEUP4 JudgeP6 Recall@k
Qwen3-8BM1Flat0.2730.4660.2850.5400.696
M2Flat0.2830.4360.2410.4700.595
M3Text0.1900.3660.2690.5620.620
M4Struct0.2260.3550.2160.4350.465
M5Struct0.1960.3810.2520.6480.455
M6Hier0.2510.4360.2760.5560.797
M8Refine0.3010.4560.2520.5090.615
M9Weight0.0190.0990.0710.379
M10Act0.1590.3420.2850.589
M11Act0.0660.1610.1380.307
Qwen3-32B-AWQM1Flat0.2080.4080.3280.5820.696
M2Flat0.1870.3520.2720.4870.595
M3Text0.1730.3680.2890.5860.620
M4Struct0.1650.3200.2400.4520.465
M5Struct0.1340.3290.2860.6830.455
M6Hier0.1920.3870.3070.5730.797
M8Refine0.1900.3650.2810.5330.615
M9Weight0.0240.1010.0820.399
M10Act0.1900.4120.3370.669
M11Act0.0350.1180.0950.274
Gemma-4-26BM1Flat0.3290.5130.3110.5740.696
M2Flat0.3020.4510.2570.5550.595
M3Text0.3030.4840.3000.5990.620
M4Struct0.3100.4050.2320.4760.465
M5Struct0.2030.4510.2710.7190.455
M6Hier0.3100.4830.2920.5920.797
M8Refine0.3350.4820.2410.5540.615
M10Act0.3000.5150.3080.688
M11Act0.0850.1810.1490.330
Table 8: Table for different substrate and model LoCoMo Latency & Quality.
ModelSub.Family𝑬𝟐 (ms/q)𝑬𝟑 (ms/wr)𝑬𝟒 (ms/rd)𝑬𝟏𝟒 (s)𝑬𝟏𝟓 (s/q)𝑷𝟓
Qwen3-8BM1Flat5221992261,5340.771.00
M2Flat302016600.331.00
M3Text2,2921,6637818,2934.185.29
M4Struct1696,210912,6386.360.71
M5Struct20,6805,1169,61055,29027.840.44
M6Hier6836411952,7581.390.62
M8Refine4312,5701796,0193.031.09
M9Weight3351,4003,4981.76
M10Act4,1329,0274.55
M11Act1,8484,0372.03
Qwen3-32B-AWQM1Flat2,1741946485,1422.591.00
M2Flat2,153014,7032.371.00
M3Text4,3791,6231,10012,8526.475.19
M4Struct2,2896,3581217,2698.700.72
M5Struct21,9495,2219,58658,02529.220.43
M6Hier2,65065361010,3275.200.60
M8Refine2,7072,60016610,9915.531.11
M9Weight8603,8009,3864.73
M10Act14,77432,27516.25
M11Act3,7908,2804.17
Gemma-4-26BM1Flat6672001881,8500.931.00
M2Flat445019720.491.00
M3Text2,8701,6641,0189,5554.815.37
M4Struct3946,1451113,1296.610.73
M5Struct18,9825,1359,44451,54425.950.43
M6Hier7126261742,8221.420.61
M8Refine6002,5721666,3883.221.06
M10Act4,69810,2635.17
M11Act1,9674,2972.16
Table 9: Table for different substrate LoCoMo Storage, Tokens & Calls.
Sub.Family𝑬𝟏𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟖𝑬𝟗𝑬𝟏𝟎𝑬𝟏𝟏𝑬𝟏𝟐𝑬𝟏𝟑
M1Flat3.59M00889.8K000000
M2Flat173.7K00806.6K000000
M3Text257.4K311.8K39011.52M8.23M2,040008.54M2,430
M4Struct177.6K4.59M3,26896.2K003.16M2,3127.75M5,580
M5Struct167.8K33.0M9,8606.00M17.8M3,9800050.8M13,840
M6Hier4.33M381.8K9101.87M00127.8K190509.6K1,100
M8Refine3.40M1.46M481806.2K2.94M2,037242.9K514.64M2,569
M9Weighthost-dep
M10Act101.2K60.9M
M11Act100.7K15.3M
Table 10: P12 (SubEM), P2 (F1), P3 (BLEU-1) per (substrate, generator, sweep) for LongMemEval-S and MemoryAgentBench.
LME-SMAB LRUMAB TTLMAB CR
ModelIDFamilyP12P2P3P12P2P3P12P2P3P12P2P3
Qwen3-8BM1Flat0.3900.1700.1130.5920.4910.4170.5700.4760.4050.3300.2110.171
M2Flat0.3400.1700.1140.6060.5270.4480.4800.3990.3400.2800.1880.155
M3Text0.1230.0470.0290.5490.4790.4070.8400.5440.4620.2700.1920.163
M4Struct0.3030.1450.0940.4650.4190.3560.4900.4680.3970.2900.1210.103
M5Struct0.3800.1600.1030.5490.4910.4170.5800.4680.3970.5000.3620.308
M6Hier0.3570.1620.1110.5210.4190.3560.2900.2460.2100.2700.1710.138
M8Refine0.3730.1590.1010.5920.5030.4270.9600.5700.4840.3100.2720.231
M9Weight0.0830.0820.0580.2390.1920.1630.1700.1670.1420.1600.1080.089
M10Act0.0900.0750.0530.6480.5750.4880.1400.1300.1110.2700.1610.124
M11Act0.2020.1110.0650.5920.5390.4580.5600.4680.3970.0500.0370.032
Qwen3-32B-AWQM1Flat0.4430.1800.1220.6340.5750.4880.8200.6890.5850.2800.1340.093
M2Flat0.3770.1630.1090.6200.5630.4780.7900.6630.5640.2700.1360.091
M3Text0.2250.1240.0720.7040.6580.5600.8600.7220.6140.2800.1580.120
M4Struct0.3570.1600.1060.5210.4550.3870.6800.6660.5660.1400.0720.049
M5Struct0.4600.1590.1040.5920.4910.4170.6200.5100.4330.4400.1910.127
M6Hier0.4130.1760.1170.5770.6230.5290.3400.2810.2380.2800.1350.090
M8Refine0.4030.1660.1110.5770.5030.4270.9400.7900.6720.3100.1620.115
M9Weight0.1000.0980.0650.4080.3590.3050.3240.3180.2700.1600.0860.059
M10Act0.1370.0870.0610.6060.5270.4480.7600.6460.5490.2700.1280.085
M11Act0.0800.0800.0560.5920.5150.4380.8200.6890.5850.2300.1220.086
Gemma-4-26BM1Flat0.4330.2000.1350.6620.5750.4880.7700.6290.5350.5200.2670.223
M2Flat0.3570.1800.1220.6620.6730.5720.7300.6200.5270.5200.2480.202
M3Text0.2230.1230.0710.3520.2750.2340.7980.6160.5240.5200.3040.259
M4Struct0.3530.1660.1110.5210.4550.3870.6900.6160.5240.0800.0680.058
M5Struct0.4530.1920.1290.5350.3830.3260.7200.6120.5200.6600.2940.214
M6Hier0.3970.1750.1180.4930.4310.3660.2500.2120.1810.4800.2280.182
M8Refine0.3730.1640.1080.5490.4550.3870.9500.8070.6860.4700.3270.278
M10Act0.0630.0450.0390.5770.5790.4920.4600.3910.3320.5000.2630.222
M11Act0.2340.1290.0750.5720.4430.3760.7020.6160.5240.0500.0500.043
Table 11: LongMemEval-S storage and latency per substrate.
ModelIDFamily𝑬𝟏𝑬𝟐,𝐦𝐞𝐚𝐧𝑬𝟐,𝐩𝟗𝟎𝑬𝟏𝟒𝑬𝟏𝟓
Qwen3-8BM1Flat9.7M2,0754,6022,1917.30
M2Flat2.7M1,7473,5805251.75
M3Text4.5M4,5009,00016,67455.58
M4Struct1.6M1,8284,23367,702225.67
M5Struct32.0M8,50018,00056,162187.21
M6Hier27.5M19,9565,1955,98719.96
M8Refine22.7M1,8694,39230,855102.85
M9Weight80.0M2,4005,00015,45151.50
M10Act60K3,7445,5391,9836.61
M11Act80K1,9004,2004,38214.61
Qwen3-32B-AWQM1Flat9.7M7,41318,1003,60012.00
M2Flat2.7M7,30718,8242,1937.31
M3Text4.5M16,07632,15320,22667.42
M4Struct1.6M7,07018,01970,200234.00
M5Struct32.0M18,00035,00058,225194.08
M6Hier27.5M20,27117,8166,08220.27
M8Refine23.4M6,14915,19232,034106.78
M9Weight320M8,50019,00051,511171.70
M10Act60K14,65028,0226,89422.98
M11Act80K7,00017,00015,71752.39
Gemma-4-26BM1Flat9.7M4,1498,5502,5468.49
M2Flat2.7M3,3726,5011,0123.37
M3Text4.5M8,99817,9969,53431.78
M4Struct1.6M3,8059,02769,116230.39
M5Struct32.0M12,00023,00057,280190.93
M6Hier27.5M17,1899,0075,15717.19
M8Refine21.7M3,1246,54031,140103.80
M10Act60K3,3403,9452,0736.91
M11Act80K5,55018.50
Table 12: LongMemEval-S token usage (E5–E13).
ModelIDFamily𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟏𝟐𝑬𝟏𝟑𝑷𝟓
Qwen3-8BM1Flat00101K001.00
M2Flat00153K001.00
M3Text437K6715.0M5.3M7314.50
M4Struct4.4M1310108K4.4M13106.20
M5Struct10.3M2371375K10.4M24313.80
M6Hier280K8046K280K802.80
M8Refine1.8M175221K2.78M350412.50
M9Weight
M10Act1.9M
M11Act250K
Qwen3-32B-AWQM1Flat00101K001.00
M2Flat00153K001.00
M3Text437K6715.0M5.3M7314.50
M4Struct4.4M1310108K4.4M13106.20
M5Struct10.4M2369385K10.4M24293.80
M6Hier280K8046K280K802.80
M8Refine1.8M175220K2.78M350412.50
M9Weight
M10Act1.9M
M11Act250K
Gemma-4-26BM1Flat00104K001.00
M2Flat00158K001.00
M3Text437K6715.0M5.3M7314.50
M4Struct4.4M1310111K4.4M13106.20
M5Struct10.4M2377379K10.4M24373.80
M6Hier280K8046K280K802.80
M8Refine1.9M175223K2.98M350412.50
M9Weight
M10Act422K
M11Act250K
Table 13: Table for different substrate and model ALFWorld-unseen Performance.
ModelSub.FamilyP8 TSRStepsP10 Steps∣𝒮Pavg
Qwen3-8BNoMem5.747.012.213.5
M1Flat5.247.75.96.6
M2Flat6.747.06.09.4
M3Text5.247.88.38.4
M4Struct7.546.97.98.5
M5Struct9.047.117.412.5
M6Hier4.548.415.08.3
M7Refine7.547.313.79.6
M8Refine8.247.418.59.5
M9Weight3.049.119.53.2
M11Act11.945.714.214.5
Qwen3-32B-AWQNoMem22.441.611.323.7
M1Flat27.639.69.616.2
M2Flat21.641.410.815.4
M3Text26.939.49.921.8
M4Struct26.939.512.424.3
M5Struct23.141.111.423.0
M6Hier23.940.29.922.8
M7Refine32.137.511.129.9
M8Refine22.441.612.421.3
M9Weight29.938.712.127.4
M11Act26.939.512.324.7
Gemma-4-26BNoMem7.547.111.317.3
M1Flat10.446.315.016.3
M2Flat9.746.311.517.2
M3Text10.446.314.814.9
M4Struct8.247.114.513.9
M5Struct4.548.18.59.4
M6Hier10.446.011.418.1
M7Refine8.246.710.117.8
Table 14: Table for different substrate and model ALFWorld-unseen Latency & Storage.
ModelSub.Family𝑬𝟏𝑬𝟐𝑬𝟐,p90𝑬𝟑𝑬𝟒𝑬𝟏𝟒𝑬𝟏𝟓
Qwen3-8BNoMem012,72813,20065,392488
M1Flat2.17M14,08616,10418822090,048672
M2Flat1.49M13,74715,6920186,698647
M3Text1.43M13,65516,3181,62075487,502653
M4Struct942K12,90314,8436,134981,070605
M5Struct532K13,16715,7025,0219,41283,080620
M6Hier25.1M13,53915,64162419187,904656
M7Refine3.90M13,29715,9882,51917584,286629
M8Refine5.27M13,06915,3002,52117383,080620
M9Weight15.3M18,75722,0021,400123,414921
M11Act6.9K12,72813,50074,102553
Qwen3-32B-AWQNoMem020,10020,30078,122583
M1Flat2.17M20,50031,100192644144,7201,080
M2Flat1.49M20,30030,80001146,0601,090
M3Text1.43M20,90030,3001,6031,087127,300950
M4Struct942K20,60028,6006,28412109,344816
M5Struct532K20,73627,4005,1549,477114,034851
M6Hier25.1M20,40026,800642600109,478817
M7Refine3.90M20,20027,2002,57516595,944716
M8Refine5.27M20,50028,0002,571164118,724886
M9Weight41.9M15,82521,4003,80082,008612
M11Act6.9K20,10022,000104,922783
Gemma-4-26BNoMem016,06016,40089,914671
M1Flat2.17M16,80021,600195184105,324786
M2Flat1.49M16,50021,40001100,902753
M3Text1.43M16,85720,4661,6491,009104,654781
M4Struct942K20,00722,5176,08911126,228942
M5Struct532K15,56419,9065,0899,358100,366749
M6Hier25.1M16,11421,09961917399,294741
M7Refine3.90M16,86621,1072,542161105,592788
Table 15: Table for different substrate ALFWorld-unseen Token Accounting.
Sub.Family𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟖𝑬𝟗𝑬𝟏𝟎𝑬𝟏𝟏𝑬𝟏𝟐𝑬𝟏𝟑
M1Flat00391K000000
M2Flat003.66M000000
M3Text372K2008.14M1.84M1,340002.21M1,540
M4Struct1.62M1995.05M00480K3802.10M579
M5Struct1.42M3852611.21M770002.63M1,155
M6Hier357K460235K0084.6K92442K552
M7Refine398K200183K0000398K200
M8Refine931K57713.5K00277K3801.21M957
M9Weight8.7K
M11Act118K
Table 16: Table for different substrate and model BigCodeBench-Hard Performance & Quality.
ModelSub.FamilyPass@1 ↑𝑷𝐚𝐯𝐠 ↑𝑷𝟓
Qwen3-8BNoMem8.112.4
M1Flat9.520.11.000
M2Flat10.111.91.000
M3Text8.112.41.964
M4Struct10.120.40.982
M5Struct15.517.10.817
M6Hier12.222.20.820
M7Refine12.216.95.705
M8Refine14.924.04.432
M9Weight7.44.6
M11Act13.523.6
Qwen3-32B-AWQNoMem17.625.6
M1Flat18.224.21.000
M2Flat19.621.81.000
M3Text18.224.51.995
M4Struct17.622.40.910
M5Struct16.216.60.786
M6Hier16.216.80.799
M7Refine17.613.52.261
M8Refine13.523.03.575
M9Weight14.215.5
M11Act13.513.8
Gemma-4-26BNoMem14.219.5
M1Flat18.226.41.000
M2Flat19.627.01.000
M3Text14.219.52.444
M4Struct18.926.90.982
M5Struct20.928.40.813
M6Hier18.926.00.813
M7Refine16.925.11.102
Table 17: Table for different substrate and model BigCodeBench-Hard Latency & Storage.
ModelSub.Family𝑬𝟏𝑬𝟐𝑬𝟐,p90𝑬𝟑𝑬𝟒𝑬𝟏𝟒𝑬𝟏𝟓
Qwen3-8BNoMem150K3,8004,9405924
M1Flat1.32M4,9766,6311991837765.2
M2Flat245K4,8536,6000.137184.9
M3Text305K5,3006,8906,853467845.3
M4Struct150K4,6426,4321371387475.0
M5Struct239K28,27336,7008,77912,4324,18428.3
M6Hier3.42M4,7646,2341,4441849946.7
M7Refine990K4,6386,2569,0291832,51917.0
M8Refine1.22M5,0536,82510,4401972,83619.2
M9Weight14.6M12,70016,5101,92413
M11Act151K16,80021,8402,51617
Qwen3-32B-AWQNoMem150K3,9005,0705924
M1Flat1.32M5,0006,5002562257405
M2Flat245K5,0006,5000.137405
M3Text296K5,4007,0209,3091,8431,1848
M4Struct150K4,9006,37012,4603762,96020
M5Struct249K31,22128,2397,88814,7234,62131.2
M6Hier3.50M4,4005,7201,9952201,0367
M7Refine1.93M4,3005,59012,9941953,10821
M8Refine1.14M4,6005,98012,4471863,25622
M9Weight38.0M14,60018,9802,22015
M11Act161K21,30027,6903,15221.3
Gemma-4-26BNoMem150K19,80025,7402,96020
M1Flat1.32M21,91628,5332132893,28622.2
M2Flat245K20,60227,8600.023,04920.6
M3Text272K23,40030,42018,378303,46323.4
M4Struct150K22,37531,073173923,37222.8
M5Struct240K45,59547,4368,19413,0876,74845.6
M6Hier3.47M24,31033,6142,7172514,14128.0
M7Refine3.13M24,14931,5138,9213175,46036.9
Table 18: Table for different substrate and model BigCodeBench-Hard Token Accounting.
ModelSub.Family𝑬𝟓𝑬𝟔𝑬𝟕𝑬𝟖𝑬𝟗𝑬𝟏𝟎𝑬𝟏𝟏𝑬𝟏𝟐𝑬𝟏𝟑
Qwen3-8BNoMem4.75M
M1Flat002.29M000000
M2Flat002.26M000000
M3Text55.9K362.80M2.84M148002.90M184
M4Struct002.16M000000
M5Struct0040.7K116K14800116K148
M6Hier44.7K821.14M0011.5K1656.2K98
M7Refine244K200593K0000244K200
M8Refine308K200550K00146K200454K400
M9Weight0
M11Act1.44M
Qwen3-32B-AWQNoMem4.75M
M1Flat002.29M000000
M2Flat002.26M000000
M3Text55.6K362.76M2.81M148002.86M184
M4Struct524K199286K0000524K199
M5Struct0041.7K121K14800121K148
M6Hier45.1K86575K0013.5K2458.6K110
M7Refine245K20091.6K0000245K200
M8Refine293K20084.8K00148K200441K400
M9Weight0
M11Act1.48M
Gemma-4-26BNoMem4.75M
M1Flat002.29M000000
M2Flat002.26M000000
M3Text52.2K362.26M2.30M148002.35M184
M4Struct002.16M000000
M5Struct0040.1K177K14800177K148
M6Hier44.4K821.14M0012.6K2457.0K106
M7Refine161K200901K0000161K200

왜 중요한가

긴 시간 작동하며 과거를 기억해야 하는 코딩 비서, 개인화 챗봇, 로봇 에이전트를 만들 때 어떤 메모리 구조를 골라야 하는지 실증 근거를 준다. 하나의 저장 방식을 고정해서 쓰기보다 과제 성격과 대화 길이에 따라 여러 방식을 섞어 전환하는 설계가 필요하다는 실무 지침을 제공한다.

이 논문의 용어

  • 메모리 서브스트레이트 · 에이전트가 기억을 담아두는 매체·구조. 벡터DB, 지식그래프, 요약 트리 등
  • 덴스/스파스 인덱스 · 문장을 숫자벡터로 바꿔 유사도 검색하는 방식(덴스)과 키워드 매칭 기반 검색(스파스)
  • 지식그래프(구조적 저장소) · 개체와 관계를 노드-엣지로 연결해 저장하는 방식으로 다단계 추론에 유리
  • top-k 검색 · 질문과 관련된 기억을 상위 k개만 꺼내와 답변에 활용하는 방식
  • LLM-as-a-judge · 정답 채점을 사람 대신 별도의 언어모델(GPT-4o-mini 등)이 수행하는 평가 방식

본문에 싣지 못한 그림

  • Figure 1: Evaluation landscape of 52 memory-augmented LLM systems (2023–2026). (a) Benchmark adoption: 62% of usage concentrates on LoCoMo and LongMemEval (b) Metric coverage: every system reports accuracy but only 21% report any efficiency metric; 81% of systems use GPT-family backbones. (c) Breadth and efficiency: half of systems use a single benchmark with no efficiency metrics; no system simultaneously achieves broad coverage and comprehensive efficiency reporting.
  • Figure 2: Overview of the harness evaluation. 11 memory methods (M1–M11) spanning seven substrate families are evaluated on identical interaction histories across four benchmarks, instrumented with 26 metrics in two families (performance and efficiency).
  • Figure 3: Performance vs. latency across four benchmarks on Qwen3-32B-AWQ. (a) LoCoMo and (b) LongMemEval-S: absolute P4. (c) ALFWorld-unseen and (d) BigCodeBench-Hard: Δ-performance vs. latency overhead anchored at NoMem (No memory baseline).
  • Figure 4: Retrieval breadth k scales oppositely on long-context QA versus sequential decision-making. Top: LoCoMo (Qwen3-8B, n=1986, k∈{1,2,5,10,20}). Bottom: ALFWorld-unseen (Qwen3-32B-AWQ, n=134, k∈{1,2,3,4,5}). green when the trend is in the preferred direction, red otherwise. No Mem stands for no memory baseline.
  • Figure 5: Response-token attention under M1 retrieval for (a) LoCoMo (b) ALFWorld.
  • Figure 6: MAB CR scalability on Qwen3-32B-AWQ: P4 (bars, left) and per-query latency (lines, right, log) at 6K, 32K, 262K tokens.
원문에서 그림 보기 →

저자 · Wei-Chieh Huang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사