컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

로봇이 미래를 예측하는 학습에서, 예측 목표를 행동을 만드는 바로 그 정책 자신에게서 뽑아내면 더 잘 작동한다

arXiv:2608.013972026-08-01

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

로봇이 미래를 예측하는 학습에서, 예측 목표를 행동을 만드는 바로 그 정책 자신에게서 뽑아내면 더 잘 작동한다

SG-WAM은 로봇 팔이 다음 행동을 만들어내는 동안 앞으로 장면이 어떻게 변할지 함께 예측하도록 학습하는 로봇 정책이다. 별도의 미래 예측기나 화면 복원 대신, 정책 자신의 과거 버전(EMA 사본)을 목표로 삼고, 3D 형태 인식 모델로 화면 토큰에 공간 정보를 주입한다. 시뮬레이션 LIBERO에서 평균 98.5%, 더 어려운 LIBERO-Plus에서 73.0% 성공률을 기록했고 실제 로봇 실험에서도 비교 대상보다 나은 성능을 보였다.

METAL MEDIA 해설 도표

SG-WAM 구조: 정책 자신을 목표로 삼는 미래예측 흐름

증거 상태측정 결과가 보고됨

  1. 입력 관찰 + 언어 + 다이내믹스 토큰현재 로봇 카메라 영상, 작업 지시문, 학습 가능한 8개 다이내믹스 토큰이 하나의 VLM(Qwen3.5-0.8B)에 함께 들어간다
  2. 기하 감독(VGGT 교사)고정된 3D 인식 모델 VGGT가 메인 화면 토큰에 공간 구조 정보를 주입하도록 코사인 유사도로 정렬시킨다
  3. 자기유도 미래예측기(SGWP)현재 다이내믹스 토큰 상태와 실제 로봇 행동 시퀀스를 받아, 미래 다이내믹스 토큰 상태를 예측한다
  4. EMA 목표 경로같은 정책의 천천히 갱신되는 복사본이 실제 미래 관찰을 보고 만든 다이내믹스 토큰 상태를 예측 목표로 삼는다(그래디언트 차단)
  5. 흐름매칭 액션 생성기다이내믹스 토큰을 포함한 전체 정책 문맥을 조건으로 받아 연속적인 로봇 행동을 생성하며, 추론 시에는 예측기·기하교사·EMA경로가 모두 제거된다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 로봇 미래예측모델(WAM)은 미래 이미지를 직접 그리거나, 별도의 잠재공간을 만들어 미래를 예측하는데, 둘 다 실제 행동 생성에 쓰이는 표현과 어긋나거나 불필요한 화면 정보에 신경을 뺏길 수 있다는 문제가 있었다.
  2. SG-WAM은 학습 가능한 '다이내믹스 토큰' 8개를 비전-언어 모델(VLM) 안에 끼워 넣고, 이 토큰들이 앞으로 어떻게 바뀔지 예측하게 한다. 예측 목표는 같은 정책의 EMA(지수이동평균) 사본이 미래 관찰을 보고 만든 값으로, 완전히 같은 표현 체계 안에서 목표를 만든다.
  3. 동시에 얼어붙은(고정된) 3D 형태 인식 모델 VGGT의 특징을 정책의 화면 토큰에 맞춰 학습시켜, 다이내믹스 토큰이 공간적 위치 정보까지 담도록 유도한다.
  4. 이 모든 예측·기하 학습은 실제 행동을 만드는 흐름매칭(flow-matching) 방식 액션 생성기와 한 번에 같이 학습되며, 추론(실제 사용) 시에는 예측기·기하학습·EMA 목표 경로가 모두 빠지고 순수 정책만 남는다.
  5. LIBERO 시뮬레이션 4개 과제 평균 98.5%, 시각·언어·조명 등 다양한 변화를 준 LIBERO-Plus에서 73.0%로 가장 높은 성능을 기록했고, 실제 UR5e 로봇 실험(픽앤플레이스, 수건 접기, 공구함 정리)에서도 VPP·VLA-JEPA보다 우수했다.
Figure 1: Conceptual comparison of existing WAMs and SG-WAM. Whereas explicit and auxiliary latent targets may introduce perceptual burden or target–policy mismatch, SG-WAM learns intervening-action-conditioned dynamics in geometry-structured, policy-derived representations.
Figure 1: Conceptual comparison of existing WAMs and SG-WAM. Whereas explicit and auxiliary latent targets may introduce perceptual burden or target–policy mismatch, SG-WAM learns intervening-action-conditioned dynamics in geometry-structured, policy-derived representations.
Table 1: Simulation results on LIBERO. Success rates are reported for the four standard LIBERO suites. Bold indicates the best result, and underlining indicates the second-best result.
MethodParamsEmbodied PT.SpatialObjectGoalLongAvg.
OpenVLA-OFT [21]7B97.698.497.994.597.1
π0 [6]3.3B98.096.894.488.494.4
π0-FAST [31]3.3B96.496.888.660.285.5
π0.5 [5]3.3B98.898.298.092.496.9
GR00T N1.6 [4]3B97.798.597.594.497.0
Spatial Forcing [24]7B99.499.698.896.098.5
WorldVLA [10]7B87.696.283.460.081.8
LAPA [44]7B55.458.874.673.865.7
RynnVLA-002 [9]7B99.099.896.494.497.4
Mantis [42]5.8B98.899.294.494.296.7
UniVLA [8]7B96.596.895.692.095.2
Fast-WAM [45]6B98.2100.097.095.297.6
VLA-JEPA [34]2B94.899.695.894.096.1
SG-WAM0.9B99.499.898.696.298.5
Figure 2: Overview of the SG-WAM framework. The VLM jointly contextualizes multi-view observations, language, and learnable dynamics tokens for action generation. During training, a frozen VGGT teacher shapes main-view image tokens, while SGWP predicts future dynamics-token states conditioned on intervening actions and aligns them with an EMA policy target. All auxiliary branches are removed at inference.
Figure 2: Overview of the SG-WAM framework. The VLM jointly contextualizes multi-view observations, language, and learnable dynamics tokens for action generation. During training, a frozen VGGT teacher shapes main-view image tokens, while SGWP predicts future dynamics-token states conditioned on intervening actions and aligns them with an EMA policy target. All auxiliary branches are removed at inference.
Table 2: Zero-shot transfer results on LIBERO-Plus. Success rates are reported under different perturbation settings. Bold indicates the best result, and underlining indicates the second-best result.
MethodParamsCameraRobotLanguageLightBackgroundNoiseLayoutOverall
WorldVLA [10]7B0.127.941.643.717.110.938.025.0
Spatial Forcing [24]7B20.113.440.929.133.425.739.329.1
Mantis [42]5.8B15.741.845.945.128.939.262.539.8
UniVLA [8]7B4.350.371.859.180.025.334.341.5
Fast-WAM [45]6B16.444.568.978.253.737.760.750.0
π0 [6]3.3B13.86.058.885.081.479.068.953.6
VLA-JEPA [34]2B40.355.772.988.270.538.274.662.9
OpenVLA-OFT [21]7B56.431.979.588.793.375.874.269.6
SG-WAM0.9B58.648.981.489.886.180.774.273.0
Figure 3: Left: Visualization of the real-world platform. We use a UR5e robot arm as the manipulation platform, the Kinect Azure camera as the main camera and the RealSense D405 as the gripper camera. Right: Visualization of 3 real-world tasks. 1) Top: Pick and Place. 2) Middle: Towel Folding. 3) Bottom: Toolbox Organization.
Figure 3: Left: Visualization of the real-world platform. We use a UR5e robot arm as the manipulation platform, the Kinect Azure camera as the main camera and the RealSense D405 as the gripper camera. Right: Visualization of 3 real-world tasks. 1) Top: Pick and Place. 2) Middle: Towel Folding. 3) Bottom: Toolbox Organization.
Table 3: Success rates under different visual perturbations. Bold indicates the best result.
ModelPick and PlaceTowel FoldingToolbox Organization
IDBackgroundLight ChangeNovel ObjectIDBackgroundLight ChangeNovel ObjectID
VLA-JEPA35%20%25%20%20%10%10%15%20%
VPP30%15%10%10%35%15%15%10%30%
SG-WAM75%55%60%40%45%25%35%25%50%
Figure 4: Representative middle-layer attention maps from dynamics tokens to main-view image tokens. The task is Pick up the black bowl between the plate and the ramekin and place it on the plate. With geometric supervision, the selected tokens attend more consistently to the robot–object interaction regions.
Figure 4: Representative middle-layer attention maps from dynamics tokens to main-view image tokens. The task is Pick up the black bowl between the plate and the ramekin and place it on the plate. With geometric supervision, the selected tokens attend more consistently to the robot–object interaction regions.
Table 4: Ablation study of geometric supervision and self-guided world modeling on LIBERO. Geo. denotes geometric supervision, and WM. denotes self-guided world modeling. Success rates (%) are reported. Bold indicates the best result.
Geo.WM.SpatialObjectGoalLongAvg.
97.097.695.691.095.3
98.297.898.092.296.6
97.899.898.294.497.6
99.499.898.696.298.5
Figure 5: Overview of the Self-Guided World Predictor (SGWP). The projected dynamics-token states are first contextualized through self-attention and then attend to the encoded intervening actions through cross-attention to predict future latent states.
Figure 5: Overview of the Self-Guided World Predictor (SGWP). The projected dynamics-token states are first contextualized through self-attention and then attend to the encoded intervening actions through cross-attention to predict future latent states.
Table 5: Ablation study on the number of learnable dynamics tokens on LIBERO. Success rates (%) are reported. Bold indicates the best result.
Number of TokenSpatialObjectGoalLongAvg.
197.298.898.290.296.1
499.098.497.495.297.5
899.499.898.696.298.5
1699.499.497.692.497.2
Figure 6: Complete visualization of the Pick and Place task.
Figure 6: Complete visualization of the Pick and Place task.
Table 6: Subtasks success rates of Towel Folding and Toolbox Organization tasks in ID settings. Bold indicates the best result.
ModelTowel FoldingToolbox Organization
First FoldingSecond FoldingPick the ScrewdriverPick the First GearPick the Second GearClose the Toolbox
VLA-JEPA50%20%50%40%20%20%
VPP60%35%70%50%40%30%
SG-WAM75%45%80%60%50%50%
Figure 7: Complete visualization of the Towel Folding task.
Figure 7: Complete visualization of the Towel Folding task.
Table 7: Ablation of the action information provided to the Self-Guided World Predictor on LIBERO. For the null-action variant, the ground-truth intervening action sequence is replaced with an all-zero sequence before being processed by the action encoder. The action-conditioning architecture, temporal positional embeddings, and all other model components and training settings remain unchanged. Success rates (%) are reported.
VariantSpatialObjectGoalLongAvg.
Null-Action Sequence98.499.298.094.697.6
SG-WAM99.499.898.696.298.5
Figure 8: Complete visualization of the Toolbox Organization task.
Figure 8: Complete visualization of the Toolbox Organization task.

실제로 확인된 결과

  • LIBERO 4개 표준 과제 평균 성공률 98.5%로, 훨씬 큰 모델이나 대규모 사전학습을 쓴 기존 최강 baseline과 대등했다.
  • 카메라 시점, 조명, 배경, 언어 등을 바꾼 LIBERO-Plus 제로샷 전이 평가에서 73.0%로 가장 높은 전체 성공률을 기록했으며, 특히 카메라·언어·조명·배치 변화에서 최고 성능을 보였다.
  • 실제 로봇 3개 과제(픽앤플레이스, 수건 접기, 공구함 정리)에서 기존 baseline인 VPP(명시적 미래영상 예측)와 VLA-JEPA(잠재공간 예측)보다 훈련 조건과 배경/조명/신규물체 변화 조건 모두에서 높은 성공률을 보였다.
  • 기하학습과 자기유도 미래예측 두 요소를 모두 제거하면 평균 95.3%로 떨어지고, 둘 다 넣으면 98.5%로 최고 성능을 냈다. 미래예측을 빼면 1.9%포인트, 기하학습을 빼면 0.9%포인트 하락했으며 특히 장기과제(LIBERO-Long)에서 차이가 컸다(96.2%→92.2%).
  • 다이내믹스 토큰 개수를 1개에서 8개로 늘리면 평균 성공률이 96.1%에서 98.5%로 올랐지만 16개로 더 늘리면 97.2%로 오히려 떨어졌다. 실제 행동 정보를 넣지 않은(모두 0으로 대체) 변형은 평균 성공률이 98.5%에서 97.6%로 낮아져, 실제 행동 정보가 예측에 도움을 준다는 것을 보여줬다.
Figure 10: Visualization of the attention weight matrix of all latent dynamics tokens attending to main-view image tokens. w/ Geo. and w/o Geo. denote with geometric supervision and without geometric supervision, respectively.
Figure 10: Visualization of the attention weight matrix of all latent dynamics tokens attending to main-view image tokens. w/ Geo. and w/o Geo. denote with geometric supervision and without geometric supervision, respectively.

어디에 쓸 수 있나

  • 작은 모델로도 로봇 조작 성능과 낯선 환경 대응력을 동시에 높이고 싶은 로봇 학습 파이프라인 설계
  • 미래 화면을 직접 생성하지 않고도 행동-환경 변화 관계를 학습에 반영하고 싶은 로봇 정책 연구
  • 장기간 여러 단계를 거치는 작업(예: 여러 물체를 순서대로 정리하는 작업)에서 중간 실패가 누적되는 문제를 줄이려는 연구
Figure 11: Complete visualization of the Pick and Place task under the background shift.
Figure 11: Complete visualization of the Pick and Place task under the background shift.

한계와 남은 검증

  • 실험은 UR5e 로봇 팔과 특정 카메라 구성, 3개의 실제 과제(픽앤플레이스, 수건 접기, 공구함 정리)에 한정되어 다른 로봇 형태나 훨씬 다양한 과제에 대한 검증은 아직 없다.
  • 0.9B 크기의 비교적 작은 모델로 대규모 사전학습 없이 얻은 결과이며, 저자들도 더 큰 정책 백본과 다양한 로봇 형태(cross-embodiment) 데이터로 확장하는 것을 향후 과제로 남겼다.
  • 다이내믹스 토큰 개수가 16개일 때 성능이 오히려 떨어지는 비단조적 경향이 관찰되어, 최적 토큰 수를 다른 과제나 모델 규모에 그대로 적용할 수 있는지는 불확실하다.
  • 실제 환경 OOD 평가는 배경 변화, 조명 변화, 신규 물체라는 세 가지 조건에 한정되어 있어, 더 극단적이거나 다른 종류의 환경 변화에 대한 강건성은 확인되지 않았다.
Figure 12: Complete visualization of the Pick and Place task under the light change.
Figure 12: Complete visualization of the Pick and Place task under the light change.

왜 중요한가

이 연구는 로봇이 '무엇을 할지'와 '그 결과 세상이 어떻게 바뀔지'를 따로 배우지 않고 하나의 표현 공간에서 함께 배우게 하는 설계 원칙을 제시한다. 별도의 거대한 화면 예측 모델이나 사전학습 없이도 작은 모델(0.9B)로 강한 성능과 낯선 환경 대응력을 얻을 수 있음을 보여준다.

이 논문의 용어

  • World Action Model(WAM) · 로봇의 행동 생성과 미래 상태 예측을 함께 학습하는 모델
  • 다이내믹스 토큰(dynamics token) · 장면이 행동에 따라 어떻게 변하는지를 담기 위해 모델에 추가한 학습 가능한 벡터들
  • EMA(지수이동평균) 사본 · 학습 중인 모델을 천천히 따라가는 안정된 복사본으로, 예측의 목표값을 만드는 데 쓰임
  • 흐름매칭(flow matching) · 노이즈에서 시작해 실제 행동 값으로 점차 변환해가며 연속적인 로봇 동작을 생성하는 학습 방식
  • VGGT · 이미지에서 3차원 기하 구조 정보를 뽑아내는, 이 연구에서는 학습 중 고정해 둔 3D 인식 모델

본문에 싣지 못한 그림

  • Figure 9: Visualization of 3 OOD settings. 1) Left: Background Shift. 2) Middle: Light Change. 3) Right: Novel Object.
원문에서 그림 보기 →

저자 · Ruiteng Zhao

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Ruiteng Zhao et al., arXiv:2608.01397, arxiv-nonexclusive