컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

언어·이미지 이해·이미지 생성을 한 모델에 같이 학습시킬 때 무엇이 서로 돕고 무엇이 서로 방해하는지 실험으로 밝힌 연구

arXiv:2608.050002026-08-04

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

언어·이미지 이해·이미지 생성을 한 모델에 같이 학습시킬 때 무엇이 서로 돕고 무엇이 서로 방해하는지 실험으로 밝힌 연구

이 연구는 텍스트, 이미지 이해, 이미지 생성을 하나의 모델에서 함께 학습시키는 '통합 멀티모달 사전학습'에서 각 능력이 서로 지식을 주고받는 방식을 통제된 실험으로 분석했다. 실제 대규모 데이터와 CLEVR 기반 합성 데이터를 모두 사용해 언어→시각, 이해→생성, 생성→이해 방향의 전이가 각각 다르게 작동함을 보였다. 이 결과를 바탕으로 데이터 비율, 구조 설계, 학습 시점에 대한 실전 레시피를 만들고 135억 파라미터 MoE 모델을 2조 토큰으로 학습해 검증했다.

METAL MEDIA 해설 도표

통합 멀티모달 사전학습에서 지식이 흐르는 방향

증거 상태측정 결과가 보고됨

  1. 언어 데이터DCLM 언어 데이터를 늘리면 이미지 이해와 이미지 생성 성능이 함께 좋아진다
  2. 이미지 이해 데이터이해 데이터를 늘리면 이미지 생성은 크게 좋아지지만 순수 언어 성능은 약간 떨어진다
  3. 이미지 생성 데이터생성 데이터를 늘려도 언어와 이해 성능에는 뚜렷한 영향 없이 소폭 변동만 있다
  4. CLEVR 개념 전이 실험색깔·모양은 양방향 전이 실패, 공간관계·크기·개수는 이해→생성만 성공
  5. 구조·시점 설계(split_ffn, 조기 통합)어텐션·정규화 공유+FFN 분리, 시각 데이터의 조기·동시 도입이 시너지를 만든다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 언어 데이터 비율을 늘리면 이미지 이해와 이미지 생성 성능이 모두 꾸준히 좋아지는 반면, 이미지 이해 데이터를 늘리면 이미지 생성은 크게 좋아지지만 순수 언어 성능은 약간 떨어진다.
  2. 이미지 생성 데이터를 늘리는 것은 언어나 이미지 이해 성능에 뚜렷한 도움도 해도 주지 않고 소폭의 변동만 일으킨다.
  3. CLEVR 합성 데이터로 색깔·모양 같은 저수준 속성과 공간관계·크기·개수 같은 구조적 개념을 나눠 특정 개념을 한쪽 모달리티 학습에서 빼는 실험을 했더니, 색깔과 모양은 이해와 생성 사이에 전혀 전이되지 않고, 구조적 개념은 이해에서 생성 쪽으로는 전이되지만 반대 방향은 대체로 실패했다.
  4. 모델 구조에서 어텐션과 정규화는 공유하고 피드포워드 층만 모달리티별로 분리하면(split_ffn) 언어와 시각이 서로 방해하지 않으면서 시너지를 낼 수 있음을 확인했다.
  5. 시각 데이터를 처음부터 언어와 함께 학습시키는 '조기 통합'이 나중에 시각을 붙이는 방식보다 우수하며, 늦게 붙이면 모델이 시각 쪽을 충분히 최적화하지 않고 언어 지식에만 의존하는 '시각 게으름' 현상이 나타났다.
Figure 1: Impact of scaling language data on visual understanding and generation. Increasing the language ratio universally improves both vision capabilities.
Figure 1: Impact of scaling language data on visual understanding and generation. Increasing the language ratio universally improves both vision capabilities.
Figure 2: Impact of scaling visual understanding data. It significantly benefits visual generation but degrades pure language performance.
Figure 2: Impact of scaling visual understanding data. It significantly benefits visual generation but degrades pure language performance.
Table 1: Extensive grid search of data mixing ratios across three axes. We evaluate models on Language, Visual Understanding, and Visual Generation. The searches confirm that while language requires a dominant token share, visual capabilities peak at highly specific and asymmetrical ratios. The optimal configuration emerges in the "Next" sweep at a 70/25/5 split for Language, Understanding, and Generation respectively.
Mix %LanguageVisual UnderstandingVisual Generation
LUGPPL ↓Acc ↑Gen ↑Know ↑OCR ↑V-Ctr ↑Avg ↑DPG ↑GenEval ↑CLIP-Sim ↑DiffLoss ↓
Fix MM10454519.2741.8937.926.922.942.932.70.3260.1480.2560.2984
20404017.5144.4645.730.824.044.036.10.3950.1890.2730.2802
30353516.7345.0343.129.924.942.235.00.3610.1830.2730.2826
40303016.3145.7946.131.224.245.636.80.3310.1860.2690.2946
50252515.9846.5947.230.925.444.637.00.3850.2190.2740.2804
60202015.8146.7044.633.024.843.936.60.3870.2030.2750.2883
70151515.6846.9948.132.325.246.638.10.3990.2190.2730.2996
80101015.5746.8545.932.824.045.237.00.3880.2160.2710.2868
905515.4848.0843.731.821.446.335.80.3360.2040.2730.2894
Fix Lan5054516.0545.2643.829.423.743.135.00.3580.2060.2730.2785
50104016.0346.1845.931.024.245.636.70.4010.2210.2810.2801
50153516.0646.4246.632.123.445.436.90.3700.1930.2730.2834
50203016.0246.1746.831.223.545.236.70.3900.1990.2740.2909
50252515.9846.5947.230.925.444.637.00.3850.2190.2740.2804
50302016.0146.0545.531.125.244.736.60.3700.2080.2730.2893
50351516.0545.8547.032.523.846.037.30.3990.1990.2750.2821
50401016.0046.1447.732.926.545.738.20.4200.2160.2760.2874
5045516.0346.1446.932.325.846.137.80.3920.2000.2690.2931
Next7052515.6746.5547.032.923.943.536.80.3750.2040.2710.2823
70102015.7146.3446.830.622.744.336.10.3580.2060.2730.2855
70151515.6846.9948.132.325.246.638.10.3990.2190.2730.2996
70201015.6546.6548.031.926.346.338.10.4010.2210.2720.2934
7025515.6846.8648.332.725.847.138.50.4500.2370.2750.2868
Figure 3: Impact of scaling visual generation data. Adding visual generation causes minor fluctuations in language and most understanding tasks.
Figure 3: Impact of scaling visual generation data. Adding visual generation causes minor fluctuations in language and most understanding tasks.
Figure 4: Overview of the synthetic CLEVR testbed. We extend the standard CLEVR vocabulary across five conceptual axes: color, shape, spatial relation, size, and object count. To isolate knowledge flow, specific target concepts (highlighted in red) are systematically ablated from targeted modality training streams.
Figure 4: Overview of the synthetic CLEVR testbed. We extend the standard CLEVR vocabulary across five conceptual axes: color, shape, spatial relation, size, and object count. To isolate knowledge flow, specific target concepts (highlighted in red) are systematically ablated from targeted modality training streams.
Table 2: Scaling results and controlled baseline comparisons. We evaluate our model against three controlled baselines to validate our main design choices: data mixture recipes (Balanced Recipe), architecture design style (Dense Model), and vision alignment strategy (Late-Fusion).
ModelLanguageVisual UnderstandingVisual Generation
PPL ↓Acc ↑Gen ↑Know ↑OCR ↑V-Ctr ↑Avg ↑DPG ↑GenEval ↑CLIP-Sim ↑DiffLoss ↓
Balanced Recipe11.9752.8651.5038.9025.1550.1441.420.6760.4670.3100.261
Dense Model12.1452.0350.1236.6625.4349.7440.490.6670.4590.3080.266
Late-Fusion12.2551.7849.8937.0326.2249.5040.660.6720.4710.3080.269
Full11.6754.3153.6340.1127.2351.3343.080.6890.4820.3120.272
Figure 5: Zero-shot concept transfer results on CLEVR. Left (Color, Shape): Low-level (Color, Shape) attributes fail to transfer in either direction. Right (Relation, Size, Count): Structural concepts exhibit an asymmetric transfer. Understanding helps zero-shot generation, whereas generation largely fails to help understanding, with a minor exception for counting.
Figure 5: Zero-shot concept transfer results on CLEVR. Left (Color, Shape): Low-level (Color, Shape) attributes fail to transfer in either direction. Right (Relation, Size, Count): Structural concepts exhibit an asymmetric transfer. Understanding helps zero-shot generation, whereas generation largely fails to help understanding, with a minor exception for counting.
Figure 6: Concept recovery via fine-tuning. We measure how quickly models learn a missing low-level concept. Top row: Prior exposure via visual understanding provides no acceleration for color generation, but leaves a usable prior that accelerates shape generation. Bottom row: Prior exposure via visual generation acts as a booster, accelerating visual understanding learning across both color and shape.
Figure 6: Concept recovery via fine-tuning. We measure how quickly models learn a missing low-level concept. Top row: Prior exposure via visual understanding provides no acceleration for color generation, but leaves a usable prior that accelerates shape generation. Bottom row: Prior exposure via visual generation acts as a booster, accelerating visual understanding learning across both color and shape.

실제로 확인된 결과

  • 언어 비율을 0%에서 80%까지 늘리자 이미지 이해의 네 축(General, Knowledge, OCR&Chart, Vision-Centric) 성능이 모두 단조 증가했고, 이미지 생성의 조건부·비조건부 확산 손실도 함께 감소했다.
  • 이미지 이해 데이터 비율을 늘리면 이미지 생성 지표와 확산 손실이 크게 개선됐지만, 순수 언어 벤치마크 점수와 퍼플렉시티는 소폭 악화됐다.
  • 이미지 생성 데이터를 늘려도 언어 정확도·퍼플렉시티와 이미지 이해 네 축 성능은 뚜렷한 추세 없이 소폭 변동만 보였다.
  • CLEVR 실험에서 색깔·모양은 이해↔생성 어느 방향으로도 전이되지 않아 노출이 없을 때 수준으로 성능이 떨어졌고, 공간관계·크기·개수는 이해→생성 방향으로는 전이됐지만 생성→이해 방향은 개수를 제외하면 대체로 실패했다.
  • 모델 구조 실험에서 어텐션·정규화 공유+FFN만 분리(split_ffn)가 완전 공유(dense)의 성능 저하를 없애면서도 시너지를 냈고, 이 시너지는 RAE, Raw Pixels, CLIP+VAE, AR(UniTok) 등 네 가지 시각 토큰화 방식 모두에서 나타났으며, 순수 언어 사전학습 구간을 늘릴수록 시각 이해·생성 성능이 꾸준히 하락했고, 동시 학습(joint training)이 순차 학습(sequential training)보다 거의 모든 지표에서 우수했다.
Figure 7: Overview of data complexity progressions. Examples of visual (top) and language (bottom) data used to evaluate the impact of task complexity on modality interactions, ranging from simple synthetic patterns to complex real-world distributions.
Figure 7: Overview of data complexity progressions. Examples of visual (top) and language (bottom) data used to evaluate the impact of task complexity on modality interactions, ranging from simple synthetic patterns to complex real-world distributions.
Figure 8: Impact of task complexity on modality interaction. Left: Escalating visual task complexity gradually turns synergy into competition. Simple visual tasks (e.g., backgrounds, noise) improve language modeling, whereas complex visual distributions (SSTK, video) degrade text perplexity. Right: Introducing language universally aids visual generation, but the simplest linguistic distribution provides the maximum synergistic boost.
Figure 8: Impact of task complexity on modality interaction. Left: Escalating visual task complexity gradually turns synergy into competition. Simple visual tasks (e.g., backgrounds, noise) improve language modeling, whereas complex visual distributions (SSTK, video) degrade text perplexity. Right: Introducing language universally aids visual generation, but the simplest linguistic distribution provides the maximum synergistic boost.

어디에 쓸 수 있나

  • 언어, 이미지 이해, 이미지 생성을 함께 학습하는 통합 멀티모달 모델을 설계할 때 데이터 배합 비율(예: 언어 70%, 이해 25%, 생성 5%)을 정하는 참고 기준으로 활용
  • 트랜스포머 구조에서 어텐션·정규화는 공유하고 피드포워드 층만 모달리티별로 분리하는 설계를 검토하는 데 참고
  • 시각 데이터를 언제부터, 어떻게(순차 대 동시) 도입할지 결정하는 학습 커리큘럼 설계에 참고
Figure 9: Impact of parameter sharing on cross-modal performance. Fully shared (dense) parameters force modality competition, degrading both language and vision. Decoupling solely the FFNs (split_ffn) perfectly mitigates this competition while leveraging shared attention to foster strong synergy. Decoupling attention (split_ffn_attn) or normalization (split_ffn_norm) significantly diminishes these improvements, and fully isolating all parameters (split_all) yields identical results to baselines.
Figure 9: Impact of parameter sharing on cross-modal performance. Fully shared (dense) parameters force modality competition, degrading both language and vision. Decoupling solely the FFNs (split_ffn) perfectly mitigates this competition while leveraging shared attention to foster strong synergy. Decoupling attention (split_ffn_attn) or normalization (split_ffn_norm) significantly diminishes these improvements, and fully isolating all parameters (split_all) yields identical results to baselines.
Figure 10: Impact of vision encoder designs on modality synergy. Left: The impact of pairing pure background images with language across different encoder configurations on language perplexity (Δ PPL). Right: The relative change in diffusion loss (%) for conditional and unconditional generation when paired with simple language. Modality synergy consistently occurs across all four visual tokenization designs.
Figure 10: Impact of vision encoder designs on modality synergy. Left: The impact of pairing pure background images with language across different encoder configurations on language perplexity (Δ PPL). Right: The relative change in diffusion loss (%) for conditional and unconditional generation when paired with simple language. Modality synergy consistently occurs across all four visual tokenization designs.

한계와 남은 검증

  • 주요 통제 실험은 1.5B~2.3B 규모의 자체 백본 모델과 SSTK, DCLM 등 특정 데이터셋에서 수행되어, 다른 아키텍처·데이터 구성에 그대로 적용될지는 검증되지 않았다.
  • CLEVR 기반 결론은 합성 3D 장면이라는 단순화된 환경에서 나온 것이라 실제 복잡한 이미지·언어 분포에도 동일하게 적용되는지는 별도 확인이 필요하다.
  • 135B 규모 검증은 13.5B MoE 모델 2조 토큰 실험 한 건으로, 더 큰 규모나 다른 모델 계열에서의 재현 여부는 추가 실험이 필요하다.
  • 생성 품질 판정에 Qwen3-VL-8B-Instruct라는 또 다른 모델을 심판으로 썼기 때문에 그 판정 자체의 편향 가능성은 별도로 검증되지 않았다.
Figure 11: Timing of unification training. The x-axis represents the number of pure language tokens consumed before visual data is introduced to the training mix. While extending the initial pure language phase yields marginal improvements in unimodal text metrics like language accuracy and perplexity, it triggers a steep and consistent decline in performance across all visual understanding and visual generation benchmarks.
Figure 11: Timing of unification training. The x-axis represents the number of pure language tokens consumed before visual data is introduced to the training mix. While extending the initial pure language phase yields marginal improvements in unimodal text metrics like language accuracy and perplexity, it triggers a steep and consistent decline in performance across all visual understanding and visual generation benchmarks.
Figure 12: Impact of sequential versus joint pretraining across various modality orderings. The charts display the performance of six distinct sequential training paths. Solid bars denote strict sequential training, while patterned bars indicate training with a 12.5% replay buffer of previously seen modalities. The horizontal dashed line represents the simultaneous joint training baseline. The results clearly show that joint training dominates all sequential approaches across almost every metric. Although data replay slightly mitigates catastrophic forgetting, it fails to match the cross-modal synergies.
Figure 12: Impact of sequential versus joint pretraining across various modality orderings. The charts display the performance of six distinct sequential training paths. Solid bars denote strict sequential training, while patterned bars indicate training with a 12.5% replay buffer of previously seen modalities. The horizontal dashed line represents the simultaneous joint training baseline. The results clearly show that joint training dominates all sequential approaches across almost every metric. Although data replay slightly mitigates catastrophic forgetting, it fails to match the cross-modal synergies.

왜 중요한가

지금 여러 회사가 언어와 이미지를 한 모델에서 같이 학습시키는 통합 멀티모달 모델로 가고 있는데, 어떤 데이터 비율과 구조가 서로 도움이 되고 방해가 되는지는 그동안 경험칙에 의존해 왔다. 이 연구는 그 설계 선택을 실험으로 검증해 실전에서 쓸 수 있는 데이터 배합비와 구조 원칙을 제시한다.

이 논문의 용어

  • 통합 멀티모달 사전학습 · 텍스트 생성과 이미지 이해, 이미지 생성을 하나의 모델이 처음부터 함께 학습하는 방식
  • 조기 통합(Early Unification) · 학습 초반부터 시각 데이터를 언어 데이터와 함께 넣어 두 모달리티가 같이 성장하도록 하는 방식
  • 시각 게으름(Vision Laziness) · 시각 데이터를 늦게 넣으면 모델이 시각 부분을 충분히 학습하지 않고 언어 지식에 의존해버리는 현상
  • split_ffn · 어텐션과 정규화는 공유하되 피드포워드 신경망만 모달리티별로 따로 두는 구조 설계
  • CLEVR · 색깔, 모양, 개수, 위치 등을 정확히 통제해 만들 수 있는 합성 3D 장면 데이터셋

저자 · Junlin Han

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Junlin Han et al., arXiv:2608.05000, arxiv-nonexclusive