τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
로봇이 애매한 상황에서는 답을 바로 내지 않고 여러 수를 미리 상상해본 뒤 결정한다
τ0-VLA는 청소, 요리, 밀크티 만들기처럼 몇 분에서 12분까지 걸리는 긴 로봇 작업에서 다음에 할 하위 작업을 정할 때, 어려운 순간에만 여러 후보를 상상하고 결과 이미지를 예측해 점수를 매긴 뒤 최종 선택을 하는 계층형 시스템이다. 40,115시간 분량의 실제 로봇 데이터로 학습됐고, 실물 로봇 실험에서 추가 연산을 쓸수록 다음 행동 예측 정확도와 실제 작업 성공률이 함께 올라갔다. 저자는 Xiaowei Cai이며 논문은 arXiv 2608.16885에 공개됐다.
METAL MEDIA 해설 도표
로봇이 애매한 상황에서는 답을 바로 내지 않고 여러 수를 미리 상상해본 뒤 결정한다
01기존 계층형 로봇 AI는 다음에 할 일을 한 번의 계산만으로 정해서, 어려운 결정에 더 신경 쓸 방법이 없었다
02τ0-VLA는 확신이 낮을 때만 여러 후보 하위 작업을 만들고, 각 후보가 실행됐을 때의 최종 화면을 예측하는 세계모델과 그 결과를 채점하는 가치모델로 빔서치를 수행한 뒤 최종 결정을 내린다
03선택된 하위 작업은 40차원 통일 동작 공간을 쓰는 하위 실행 모델이 실제 로봇 팔다리 움직임으로 바꿔 여러 로봇 몸체에서 수행한다
04청소, 재료 준비, 볶음요리, 밀크티 제작, 빨래 수거, 책 정리 등 실물 로봇 과제에서 연산을 더 쓸수록 다음 작업 예측 정확도와 최종 성공률이 함께 향상됐다
05학습 데이터가 없던 낯선 책 배열 상황에서도 같은 경향이 유지돼 분포가 달라져도 방법이 견고함을 보였다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
기존 계층형 로봇 AI는 다음에 할 일을 한 번의 계산만으로 정해서, 어려운 결정에 더 신경 쓸 방법이 없었다
τ0-VLA는 확신이 낮을 때만 여러 후보 하위 작업을 만들고, 각 후보가 실행됐을 때의 최종 화면을 예측하는 세계모델과 그 결과를 채점하는 가치모델로 빔서치를 수행한 뒤 최종 결정을 내린다
선택된 하위 작업은 40차원 통일 동작 공간을 쓰는 하위 실행 모델이 실제 로봇 팔다리 움직임으로 바꿔 여러 로봇 몸체에서 수행한다
청소, 재료 준비, 볶음요리, 밀크티 제작, 빨래 수거, 책 정리 등 실물 로봇 과제에서 연산을 더 쓸수록 다음 작업 예측 정확도와 최종 성공률이 함께 향상됐다
학습 데이터가 없던 낯선 책 배열 상황에서도 같은 경향이 유지돼 분포가 달라져도 방법이 견고함을 보였다
Fig. 2: The hierarchical τ0-VLA architecture. (a) At a high-level inference step t, the proposal model P conditions on the latest multi-view observation ot, task instruction ℓ, carried execution memory ℳt−1, and previously generated subtask zt−1⋆. It produces observation-aligned memory ℳt and a direct proposal ztdir. (b) The low-level policy conditions on the generated subtask zt⋆, multi-view observation ot, proprioceptive state 𝐬t, and textual control metadata η. A vision-language backbone and Mixture-of-Transformers (MoT) action expert generate the action chunk 𝐚t:t+H−1 through conditional flow matching from a noisy action chunk. (c) On the TTC route, the proposal model is invoked N times for each retained branch to generate N candidates. Given the branch’s head-camera image and a candidate, the world model predicts the terminal head-camera image, and the value model assigns a candidate-quality score conditioned on the task instruction, candidate, and predicted image. Beam search globally retains the top-B branches by cumulative score and recursively expands them to depth D. The figure illustrates the root expansion with N=3 and B=2, where local and cumulative scores coincide. Deeper expansion is omitted for clarity. The reflective model then conditions on h¯t and the final branch summaries 𝒞t to generate the final subtask zt⋆. This output may coincide with a retained proposal but is not restricted to the retained set and is passed to the low-level policy for execution.
TABLE I: Long-horizon task performance. Each method–task setting uses 10 independently collected physical-robot trials. SR reports successful trials as x/10, and Progress is the normalized milestone-completion score. Avg. is the unweighted mean of the four task-level rates. The first four rows use direct execution. The final row uses the Hierarchical System with Plan Once and no beam search.
Method
Clean Room
Prepare Ingredients
Tomato and Egg Stir Fry
Make Milk Tea
Avg.
SR ↑
Progress ↑
SR ↑
Progress ↑
SR ↑
Progress ↑
SR ↑
Progress ↑
SR ↑
Progress ↑
GR00T N1.7 [26]
0/10
59.80%
1/10
68.57%
0/10
24.32%
0/10
28.46%
2.50%
45.29%
LingBot-VLA [40]
0/10
66.60%
0/10
35.00%
0/10
12.27%
0/10
63.85%
0.00%
44.43%
π0.5 [2]
4/10
86.20%
2/10
73.93%
0/10
49.77%
3/10
82.31%
22.50%
73.05%
τ0-VLA
4/10
92.80%
2/10
66.43%
0/10
65.00%
5/10
96.15%
27.50%
80.10%
τ0-VLA (Hierarchical System, Plan Once)
5/10
94.80%
4/10
82.86%
4/10
81.82%
5/10
91.92%
45.00%
87.85%
Fig. 3: Representative physical-robot evaluation tasks. (a) Clean Room requires collecting two dirty garments, placing them in a laundry basket, hanging a handbag, handing a blanket to a person, and disposing of table trash. (b) Prepare Ingredients requires retrieving a tomato and an egg, then cracking and stirring the egg while returning the tools. (c) Tomato and Egg Stir Fry requires chopping and transferring the tomatoes, cooking and seasoning the ingredients, plating the dish, and returning the cookware and utensils. (d) Make Milk Tea requires adding toppings, pouring milk and tea, sealing the cup, and inserting a straw. (e) Collect Laundry requires transferring a T-shirt from the bedside table to a laundry basket. (f) Tidy Makeup Table comprises three separately scored instruction-following groups that require different object selections and action sequences from matched visual states.
TABLE II: Direct-execution performance across embodiments. Tidy Makeup Table comprises three independently scored instruction-following groups. All methods execute the full task instruction without a high-level policy. Each task or task group is evaluated over 10 trials. SR reports successful trials as x/10, and Progress is the normalized milestone-completion score.
Method
Collect Laundry
Tidy Makeup Table
T-shirt
Cotton Pad
Eyelash Curler
Makeup Puff
SR ↑
Progress ↑
SR ↑
Progress ↑
SR ↑
Progress ↑
SR ↑
Progress ↑
GR00T N1.7 [26]
4/10
76.00%
10/10
87.50%
8/10
77.50%
7/10
52.50%
LingBot-VLA [40]
2/10
35.00%
9/10
67.50%
3/10
22.50%
3/10
33.75%
π0.5 [2]
9/10
88.00%
9/10
85.00%
8/10
85.00%
7/10
73.75%
τ0-VLA
10/10
97.00%
10/10
95.00%
9/10
92.50%
10/10
95.00%
Fig. 4: Next-subtask prediction accuracy under different high-level inference methods. We compare Plan Once, Best-of-N, and TTC (Ours) across four evaluation settings: Make Milk Tea, Book Organization (In-Domain), Book Organization (OOD), and Clean Room.
TABLE III: Closed-loop physical-robot performance with test-time computation. Each entry uses 10 independently collected trials. Book Organization uses shuffled initial arrangements and is reported without the in-domain and OOD split used in the open-loop evaluation. SR denotes task success rate, and Progress is the normalized milestone-completion score.
Method
Make Milk Tea
Book Organization
Clean Room
SR ↑
Progress ↑
SR ↑
Progress ↑
SR ↑
Progress ↑
Plan Once
5/10
91.92%
6/10
66.67%
5/10
94.80%
TTC
7/10
95.38%
9/10
93.33%
7/10
97.60%
Fig. 5: Relationship between computational cost and subtask-prediction accuracy. The accuracy increases with additional computation for the Make Milk Tea and Book Organization tasks, shown in panels (a) and (b), respectively. Each point represents an experimental result obtained at a different computational cost. The orange dashed curves show saturation fits to these results, and the gray dashed lines indicate the Plan Once baselines.
TABLE IV: Canonical 40-D state and action layout. Dimensions are one-indexed. For a rotation matrix 𝐑=[𝐫1,𝐫2,𝐫3], we use Rot6D(𝐑)=[𝐫1⊤,𝐫2⊤]⊤.
Coordinates
Dimensions
State representation
Left EEF position
1–3
Cartesian position in meters
Left EEF orientation
4–9
Rot6D(𝐑L)
Right EEF position
10–12
Cartesian position in meters
Right EEF orientation
13–18
Rot6D(𝐑R)
Left gripper
19
native opening coordinate
Right gripper
20
native opening coordinate
Waist
21–22
two native coordinates
Planar base velocity
23–24
two native coordinates
Left arm joints
25–32
q1L,…,q8L in radians
Right arm joints
33–40
q1R,…,q8R in radians
TABLE V: Maximum duration of each physical-robot trial.
Task
Maximum duration
Clean Room
20 min
Prepare Ingredients
20 min
Tomato and Egg Stir Fry
20 min
Make Milk Tea
10 min
Book Organization
5 min
Collect Laundry
5 min
Tidy Makeup Table (each group)
5 min
TABLE VI: High-level instance families, synthesized by perturbing only the input memory while reading the corrected target from the demonstration. ℳn is the memory upon entering segment n. A single unified <think>/<memory>/<subtask> format instantiates all of them at zero extra annotation.
Family
Sampling position
Input → target memory
Target subtask
Deployment failure countered
Mix
within-subtask
anywhere in seg. n
ℳn→ℳn
seg. n
— (aligned, normal progression)
58%
transition
tail of seg. n
ℳn→ℳn+1
seg. n+1
starting a new subtask after completion
15%
catch-up
head of seg. n
ℳn−1→ℳn
seg. n
memory lag (behind the visual state)
10%
rollback
late in seg. n
ℳn+1…n+3→ℳn
retry seg. n
memory run-ahead (over-optimistic)
12%
error-think
annotated failure frame
ℳn→ type-dependent
recovery step
unnoticed execution failure
5%
왜 중요한가
긴 작업을 하는 로봇은 잘못된 하위 결정을 내리면 아무리 손동작이 정확해도 실패하는데, 이 연구는 어려운 순간에만 더 오래 생각하게 만드는 방법을 실물 로봇으로 검증했다. 이는 언어모델의 테스트 타임 연산 확장 아이디어를 로봇 제어에 옮긴 사례로, 앞으로 범용 로봇 시스템 설계에 참고가 될 수 있다.
이 논문의 용어
VLA (Vision-Language-Action) 모델 · 카메라로 본 장면과 언어 지시를 받아 로봇 동작을 출력하는 인공지능 모델
테스트 타임 연산(Test-Time Computation) · 모델을 다시 학습시키지 않고, 실제 사용 시점에 더 많은 계산을 들여 답의 질을 높이는 방법
세계모델(World Model) · 어떤 행동을 하면 환경이 어떻게 바뀔지 미리 예측하는 모델
빔서치(Beam Search) · 여러 가능한 다음 수 중 점수가 높은 몇 개만 남기고 계속 확장해 나가는 탐색 방법
실행 메모리(Execution Memory) · 지금까지 로봇이 어디까지 작업을 끝냈는지 요약해 기억해두는 정보