컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

얼굴과 목소리를 동시에, 스트리밍으로 바꿔치기하는 영상 합성 AI

arXiv:2608.117522026-08-12

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

얼굴과 목소리를 동시에, 스트리밍으로 바꿔치기하는 영상 합성 AI

UniSwap은 말하는 영상 속 인물의 얼굴과 목소리를 한 모델 안에서 동시에 바꾸면서 원본의 움직임, 배경, 발화 내용, 입 모양-음성 타이밍을 그대로 유지하는 프레임워크다. 기존에는 얼굴 교체와 음성 변환을 따로 처리해 두 모달리티가 어긋나는 문제가 있었는데, 하나의 오디오-비디오 디퓨전 트랜스포머로 통합해 실시간에 가까운 블록 단위 생성을 가능하게 했다. 짧은 영상과 1분짜리 긴 영상 실험에서 기존 조합 방식보다 입-음성 동기화가 우수하고, 정체성 보존과 긴 영상에서의 안정성도 경쟁력 있는 수준을 보였다.

METAL MEDIA 해설 도표

UniSwap 3단계 학습 및 추론 구조

증거 상태측정 결과가 보고됨

  1. 데이터 생성실제 영상에서 얼굴을 포즈 실루엣으로, 목소리를 다른 화자로 바꿔 '가짜 소스'를 만들고 원본을 복원 목표로 삼는 swap-and-reconstruct 파이프라인
  2. 1단계 문맥 내 사전학습소스·참조·목표 영상·음성 latent를 하나의 시퀀스로 묶어 전체를 한 번에 보며 얼굴·목소리 동시 교체를 학습
  3. 2단계 스트리밍 적응Decoupled Streaming Conditioning Mask로 블록 단위 인과적 주의 구조를 적용해 KV 캐시 기반 자동회귀 생성이 가능하게 전환
  4. 3단계 효율적 self-forcing DMD학생 모델이 자기 출력을 이어받아 학습하고, 교사-생성기-판별기를 LoRA 어댑터로 전환하며 디노이징을 30단계에서 3단계로 압축
  5. 추론 시 Feature-RoPE Decomposition캐시된 위치 좌표를 학습 범위 안으로 재매핑하고 첫 블록을 고정 앵커로 유지해 1분 분량 생성에서도 정체성 흐름을 안정화
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 말하는 영상에서 인물의 얼굴 생김새와 목소리 음색을 동시에 참조 이미지·참조 음성으로 바꾸되, 원본의 움직임, 배경, 말하는 내용, 입-소리 타이밍은 그대로 유지하는 과제를 다뤘다.
  2. 교차 정체성 학습 데이터가 부족한 문제를 풀기 위해 실제 영상에서 얼굴과 목소리 정체성을 지운 뒤 원본 영상을 복원 목표로 삼는 swap-and-reconstruct 데이터 생성 파이프라인을 만들었다.
  3. 양방향(전체 영상을 한 번에 보는) 백본을 시작으로, 문맥 내 사전학습 → 블록 단위 실시간 생성 적응 → 효율적 self-forcing DMD 증류의 3단계를 거쳐 디노이징 스텝을 30단계에서 3단계로 줄였다.
  4. Efficient Multi-LoRA Switching으로 교사·생성기·판별기 역할을 하나의 고정된 백본 위에서 어댑터만 바꿔가며 처리해 GPU 메모리를 80GB 초과(메모리 부족)에서 65.34GB로 줄였다.
  5. Feature-RoPE Decomposition으로 캐시된 위치 정보를 학습 범위 안에 묶어두어, 1분 분량의 긴 영상 생성에서도 정체성이 흐트러지지 않도록 안정화했다.
Figure 2: Overview of UniSwap. The three-stage pipeline comprises (a) In-Context Pretraining for joint audio-video replacement, (b) Conditional Streaming Adaptation with block-causal masking and KV-cached inference, and (c) Efficient Self-Forcing DMD, which reduces denoising to 3 steps per block. Feature-RoPE Decomposition bounds cached positions while preserving cross-modal physical-time alignment.
Figure 2: Overview of UniSwap. The three-stage pipeline comprises (a) In-Context Pretraining for joint audio-video replacement, (b) Conditional Streaming Adaptation with block-causal masking and KV-cached inference, and (c) Efficient Self-Forcing DMD, which reduces denoising to 3 steps per block. Feature-RoPE Decomposition bounds cached positions while preserving cross-modal physical-time alignment.
Table 1: Quantitative comparison on the short-video benchmark. Each video replacement method is paired with the same Seed-VC audio backend, which matches the converter used to synthesize our training sources; audio-visual synchronization is measured on the resulting cascade, and video quality on the generated video. Because all video replacement baselines share the same Seed-VC output, their repeated voice-quality values are omitted. Voice-conversion methods retain the source video. The best result in each column is in bold, and the second best is underlined. “–” denotes a metric that is not separately reported for that row.
MethodA–V SyncVideo QualityVoice Quality
Sync-C ↑Sync-D ↓ASE ↑IQA ↑DINO-S ↑SIG ↑BAK ↑OVRL ↑SECS ↑SSIM ↑
VideoVACE 180.832±0.32012.800±0.9332.059±0.3553.269±0.5470.400±0.110
Wan-Animate 42.874±1.65311.338±1.7132.098±0.3403.514±0.5200.580±0.140
SCAIL-2 393.289±1.59211.269±1.6232.409±0.3524.067±0.3990.630±0.150
MoCha 383.031±1.67811.198±1.8812.534±0.3404.249±0.3040.577±0.147
HunyuanCustom 140.894±0.40112.991±0.8642.319±0.4143.816±0.4890.624±0.142
VoiceOpenVoice 243.458±0.4153.438±0.6632.910±0.4860.755±0.0660.363±0.162
Seed-VC 223.489±0.3353.750±0.5713.074±0.4450.829±0.0470.212±0.115
CosyVoice 73.461±0.2493.738±0.4563.041±0.3660.802±0.0510.137±0.106
JointUniSwap3.633±1.23610.304±0.8492.097±0.2383.758±0.3180.629±0.1363.486±0.3273.563±0.5432.988±0.3660.730±0.0640.269±0.161
Figure 3: Swap-and-reconstruct paired data synthesis. Every real talking video serves as its own reconstruction target. ① The real clip provides the target video Vt and audio At. ② The visual identity is swapped out by replacing the person with a pose proxy composited onto the masked background plate, the vocal timbre is randomized towards a sampled speaker, and the reference identity is formed by a portrait frame Ir and a random 30% audio crop Ar. ③ The resulting identity-swapped source and reference identity condition the model, which is trained to reconstruct the original clip.
Figure 3: Swap-and-reconstruct paired data synthesis. Every real talking video serves as its own reconstruction target. ① The real clip provides the target video Vt and audio At. ② The visual identity is swapped out by replacing the person with a pose proxy composited onto the masked background plate, the vocal timbre is randomized towards a sampled speaker, and the reference identity is formed by a portrait frame Ir and a random 30% audio crop Ar. ③ The resulting identity-swapped source and reference identity condition the model, which is trained to reconstruct the original clip.
Table 2: Long-video comparison. Metrics are computed independently on three 20-second segments of 1-minute generated videos. The best result in each segment is in bold and the second best is underlined. UniSwap maintains stable quality and identity across the full duration, while the baselines fluctuate or degrade over time.
Method0–20 s20–40 s40–60 s
ASE ↑IQA ↑DINO-S ↑ASE ↑IQA ↑DINO-S ↑ASE ↑IQA ↑DINO-S ↑
SCAIL-2 392.628±0.2034.426±0.2650.566±0.1752.765±0.1974.568±0.2940.538±0.1542.656±0.3264.254±0.4010.517±0.135
Wan-Animate 42.241±0.3973.766±0.6350.554±0.1192.271±0.4133.741±0.7010.533±0.1152.238±0.3643.628±0.7140.528±0.114
UniSwap2.224±0.2263.966±0.3310.596±0.1222.236±0.1894.001±0.2760.590±0.1262.259±0.1914.032±0.2540.596±0.118
Figure 4: Qualitative comparison on the short-video benchmark. Given the reference image/audio and the source video/audio (top), video replacement methods (upper rows) transfer the appearance but keep the original voice, and voice conversion methods (middle rows) modify only the speech waveform, while UniSwap (bottom) jointly replaces the appearance and the voice, producing lip motion synchronized with the converted speech. Zoom in for details.
Figure 4: Qualitative comparison on the short-video benchmark. Given the reference image/audio and the source video/audio (top), video replacement methods (upper rows) transfer the appearance but keep the original voice, and voice conversion methods (middle rows) modify only the speech waveform, while UniSwap (bottom) jointly replaces the appearance and the voice, producing lip motion synchronized with the converted speech. Zoom in for details.
Table 3: Efficiency comparison on the short-video benchmark (241-frame clips). All measurements use one NVIDIA H100 GPU. Wall-clock FPS counts generated pixel frames per second. SCAIL-2 uses its accelerated 8-step LoRA configuration. Baseline times cover an entire clip, whereas UniSwap’s values marked by † are per block (3 latent frames or 24 pixel frames); per-step times are therefore not directly comparable across the two settings.
MethodStepsInfer. Time (s)Time/Step (s)FPS ↑
VACE 1850482.569.650.499
Wan-Animate 420176.348.821.367
HunyuanCustom 1450703.3414.070.343
SCAIL-2 398190.9823.871.262
MoCha 38301800.9460.030.134
UniSwap3†1.76†0.59†13.6
Figure 5: Qualitative comparison on the long-video benchmark. Frames are sampled every 10 seconds from 1-minute generations. The baselines exhibit identity drift and visual artifacts as generation progresses (highlighted in red), while UniSwap preserves the reference identity throughout the full duration. Zoom in for details.
Figure 5: Qualitative comparison on the long-video benchmark. Frames are sampled every 10 seconds from 1-minute generations. The baselines exhibit identity drift and visual artifacts as generation progresses (highlighted in red), while UniSwap preserves the reference identity throughout the full duration. Zoom in for details.
Table 4: Ablation on training stages and the condition positional encoding offset (short-video benchmark). The best result in each column is in bold and the second best is underlined.
SettingA–V SyncVideo QualityVoice Quality
Sync-C ↑Sync-D ↓ASE ↑IQA ↑DINO-S ↑SIG ↑BAK ↑OVRL ↑SECS ↑SSIM ↑
Stage 1 (In-context)5.272±1.5109.107±0.9412.253±0.2893.922±0.3250.635±0.1323.476±0.4223.619±0.6653.029±0.4980.782±0.0590.213±0.151
Stage 2 (Teacher forcing)4.620±1.1879.581±0.7522.233±0.3043.893±0.3320.623±0.1343.349±0.5783.059±0.7362.687±0.5490.681±0.0710.480±0.156
Stage 3 (Self-forcing DMD)3.633±1.23610.304±0.8492.097±0.2383.758±0.3180.629±0.1363.486±0.3273.563±0.5432.988±0.3660.730±0.0640.269±0.161
Stage 2 w/o condition PE offset1.738±1.73711.843±1.6062.152±0.3523.472±0.6000.463±0.1663.238±0.7283.280±0.9242.767±0.7020.624±0.0860.103±0.142
Figure 6: Qualitative ablation on Feature-RoPE Decomposition. Frames are sampled every 10 seconds from one-minute generations. Removing any component leads to visible identity drift and artifacts in later segments, whereas the full model remains more consistent throughout the sequence. Zoom in for details.
Figure 6: Qualitative ablation on Feature-RoPE Decomposition. Frames are sampled every 10 seconds from one-minute generations. Removing any component leads to visible identity drift and artifacts in later segments, whereas the full model remains more consistent throughout the sequence. Zoom in for details.
Table 5: Ablation on Feature-RoPE Decomposition (long-video benchmark, per-segment metrics). The best result in each segment is in bold and the second best is underlined.
Setting0–20 s20–40 s40–60 s
ASE ↑IQA ↑DINO-S ↑ASE ↑IQA ↑DINO-S ↑ASE ↑IQA ↑DINO-S ↑
UniSwap (full)2.224±0.2263.966±0.3310.596±0.1222.236±0.1894.001±0.2760.590±0.1262.259±0.1914.032±0.2540.596±0.118
w/o Window-Bounded RoPE2.115±0.2263.763±0.2280.599±0.1202.083±0.2273.500±0.1530.546±0.0932.083±0.2133.390±0.1410.517±0.071
w/o Reference Re-anchoring2.106±0.1973.741±0.3140.595±0.1142.111±0.2253.399±0.2070.522±0.0942.140±0.2273.208±0.2460.491±0.084
w/o Adaptive Sink Block2.083±0.1733.734±0.3130.589±0.1152.007±0.1433.336±0.1640.530±0.0872.099±0.2033.174±0.1820.499±0.070
Figure 7: Additional qualitative results on short videos (Part I). For each example, the left column provides the reference image and reference voice, while the right column shows sampled frames and the waveform of the source video/audio followed by the UniSwap output. Red waveforms indicate source audio and blue waveforms indicate generated audio. UniSwap changes the visual and vocal identity according to the references while preserving the source scene composition and motion.
Figure 7: Additional qualitative results on short videos (Part I). For each example, the left column provides the reference image and reference voice, while the right column shows sampled frames and the waveform of the source video/audio followed by the UniSwap output. Red waveforms indicate source audio and blue waveforms indicate generated audio. UniSwap changes the visual and vocal identity according to the references while preserving the source scene composition and motion.
Table 6: User-study results. Ratings are mean scores on a five-point Likert scale.
MethodAppearance ID ↑Voice ID ↑Lip Sync ↑Naturalness ↑
Wan-Animate + Seed-VC3.854.043.423.77
SCAIL-2 + Seed-VC4.034.053.673.85
MoCha + Seed-VC3.914.083.543.89
HunyuanCustom + Seed-VC3.643.983.283.61
UniSwap4.163.874.113.96
Figure 8: Additional qualitative results on short videos (Part II). The examples follow the layout of Fig. 7 and cover additional identities, viewpoints, gestures, and backgrounds. Red and blue denote the source- and generated-audio waveforms, respectively. The generated frames adopt the reference appearance while following the pose, expression, framing, and scene content of the source video.
Figure 8: Additional qualitative results on short videos (Part II). The examples follow the layout of Fig. 7 and cover additional identities, viewpoints, gestures, and backgrounds. Red and blue denote the source- and generated-audio waveforms, respectively. The generated frames adopt the reference appearance while following the pose, expression, framing, and scene content of the source video.

실제로 확인된 결과

  • 짧은 영상 벤치마크(100개 클립)에서 UniSwap은 평가된 교체 파이프라인 중 가장 높은 Sync-C(3.633)와 가장 낮은 Sync-D(10.304)를 기록해 입-음성 동기화가 가장 우수했다.
  • 얼굴 정체성 유사도(DINO-S)는 최고 성능 기준선과 0.001 차이(0.629 대 0.630)로 근접했지만, 화질·미적 점수는 MoCha와 SCAIL-2보다 낮았고, 음성 품질 일부 지표(BAK, OVRL, SECS, SSIM)는 최고 음성변환 기준선보다 낮았다.
  • 1분 길이 긴 영상 벤치마크에서 3개 구간 모두 UniSwap이 가장 높은 정체성 유사도(DINO-S)를 유지했으나, 미적·화질 점수는 SCAIL-2가 더 높았다.
  • 효율성 비교에서 24프레임 블록을 1.76초에 생성해 초당 13.6프레임을 냈고, 이는 가장 빠른 기준선(Wan-Animate, 1.367FPS)보다 약 10배, MoCha보다 약 100배 빠른 수치였다.
  • 30명이 참여한 사용자 평가에서 UniSwap이 외형 정체성, 입-음성 동기화, 자연스러움 항목에서 가장 높은 평점을 받았다.
Figure 9: Additional qualitative results on one-minute videos. Frames are sampled every 10 seconds from 0 to 60 seconds. Each row group shows the reference image and voice, the source video/audio (red waveform), and the UniSwap output (blue waveform). Across all three sequences, the generated character remains consistent with the reference identity throughout the minute while retaining the source scene and temporal performance.
Figure 9: Additional qualitative results on one-minute videos. Frames are sampled every 10 seconds from 0 to 60 seconds. Each row group shows the reference image and voice, the source video/audio (red waveform), and the UniSwap output (blue waveform). Across all three sequences, the generated character remains consistent with the reference identity throughout the minute while retaining the source scene and temporal performance.

어디에 쓸 수 있나

  • 영화·영상 후반작업에서 배우의 얼굴과 목소리를 동시에 다른 정체성으로 교체하는 작업
  • 다국어 더빙이나 콘텐츠 현지화 시 원본 입 모양과 어울리는 목소리를 함께 생성하는 작업
  • 저지연 인터랙티브 방송이나 실시간에 가까운 스트리밍 콘텐츠 제작 파이프라인 실험
  • 접근성 지원을 위한 영상 재더빙 및 인물 표현 변경 실험

한계와 남은 검증

  • 실시간 재생 기준인 초당 25프레임에는 아직 못 미치는 13.6FPS 수준이라 완전한 실시간 재생은 아니며 추가 시스템 최적화가 필요하다고 밝혔다.
  • 짧은 영상에서 화질·미적 점수, 음성 품질 일부 지표가 단일 모달리티 전문 기준선보다 낮아, 통합 처리와 개별 최적화 모델 사이 트레이드오프가 남아 있다.
  • 현재는 단일 화자가 등장하는 말하는 영상에 초점이 맞춰져 있고, 다중 화자 장면, 가림, 복잡한 상호작용은 아직 다루지 못한다고 명시했다.
  • 표정은 오디오 조건에 따라 자동으로 생성될 뿐 사용자가 별도로 표정을 편집하거나 지정하는 기능은 지원하지 않는다.
  • 안정성 개선 요소(윈도우 경계 RoPE, 참조 재정렬, 고정 앵커 블록)를 제거했을 때 시간이 지날수록 화질과 정체성이 저하되는 것이 확인되어, 각 요소가 긴 영상 안정성에 필수적임을 시사한다.

왜 중요한가

지금까지 얼굴 교체와 목소리 변환은 따로 만든 모델을 이어붙여 쓰다 보니 입 모양과 목소리가 어긋나는 문제가 흔했는데, 이를 한 모델에서 동시에 처리하고 스트리밍까지 가능하게 한 첫 사례라는 점에서 영화 후반작업, 다국어 더빙, 실시간 인터랙티브 콘텐츠 제작에 실질적인 방향을 제시한다. 다만 짧은 영상 벤치마크에서 화질·음질 일부 지표는 여전히 단일 모달리티 전문 모델에 못 미쳐, 통합과 전문화 사이의 트레이드오프가 남아 있다.

이 논문의 용어

  • 디퓨전 트랜스포머 · 노이즈를 점진적으로 제거해 영상·음성을 생성하는 트랜스포머 구조의 생성모델
  • KV 캐시 · 이전에 계산한 키·값을 저장해두고 재사용함으로써 매 단계 전체를 다시 계산하지 않게 하는 기법
  • RoPE(회전 위치 인코딩) · 토큰의 순서·시간 위치 정보를 회전 변환으로 표현하는 위치 인코딩 방식
  • DMD(분포 매칭 증류) · 여러 단계를 거치는 교사 모델의 출력 분포를 적은 단계의 학생 모델이 흉내 내도록 학습시키는 압축 기법
  • LoRA · 원본 모델은 고정한 채 작은 어댑터 파라미터만 학습해 적은 자원으로 미세조정하는 방법

저자 · Yuxuan Zhang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Yuxuan Zhang et al., arXiv:2608.11752, arxiv-nonexclusive