컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

말하는 얼굴 영상을 1스텝 추론으로 실시간 200프레임까지 뽑아내면서도 얼굴이 흐트러지지 않게 만든 기술

arXiv:2608.000792026-07-28

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

말하는 얼굴 영상을 1스텝 추론으로 실시간 200프레임까지 뽑아내면서도 얼굴이 흐트러지지 않게 만든 기술

음성과 얼굴 사진 한 장만 있으면 실시간으로 계속 이어지는 말하는 얼굴 영상을 만들 수 있는 LeapTalk을 소개한다. 기존 방식은 노이즈에서 매번 다시 그림을 그리다 보니 시간이 지날수록 사람 얼굴이 조금씩 다른 사람처럼 변하는 문제(아이덴티티 드리프트)가 있었는데, 이를 원본 사진을 기준점으로 고정하는 방식으로 해결했다. 한 대의 GPU에서 초당 최대 200프레임, 최대 15000배 빨라진 속도를 유지하면서도 입 모양 동기화와 화질을 지켰다고 보고한다.

METAL MEDIA 해설 도표

LeapTalk의 Bridge Forcing과 이질적 증류 구조

증거 상태측정 결과가 보고됨

  1. 참조 이미지 고정매 청크(짧은 영상 구간)마다 노이즈에서 다시 시작하지 않고, 원본 얼굴 사진을 고정된 출발점으로 사용해 정체성 드리프트를 막는다.
  2. Brownian Bridge 경로고정된 시작점(참조 이미지)과 목표 프레임을 잇는 다리 형태의 확률 경로를 따라 한 번의 순전파로 다음 프레임을 만든다.
  3. 시간 변환 Φ(τ) 기반 증류다단계 diffusion teacher와 1스텝 bridge student의 노이즈 수준을 맞춰주는 함수를 만들어, 서로 다른 방식의 두 모델 사이에서도 지식 전달이 가능하게 한다.
  4. 음성 기반 CFG음성 조건을 강조해 1스텝 생성에서도 입 모양 동기화와 움직임의 다양성을 유지하도록 결과를 조정한다.
  5. 실시간 스트리밍 청크 연결이전 청크의 마지막 K프레임을 다음 청크의 앞부분으로 이어붙여 정체성은 고정하면서도 움직임이 끊기지 않게 연결한다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 말하는 얼굴 생성 기술은 화질이 좋은 다단계 diffusion(여러 번 노이즈를 제거해 그림을 완성하는 방식)이면 너무 느려서 실시간 스트리밍이 안 되고, 빠른 자기회귀(직전 결과를 참고해 다음을 만드는 방식) 방식은 시간이 지나면 오류가 쌓여 얼굴이 다른 사람처럼 변하는 문제가 있었다.
  2. LeapTalk은 매번 무작위 노이즈에서 새로 그리는 대신, 원본 참조 사진을 고정된 시작점으로 두고 거기서 목표 프레임까지 이어지는 다리(Brownian Bridge)를 만드는 Bridge Forcing 방식을 제안해 얼굴 정체성이 흐트러지지 않게 했다.
  3. 느리지만 정교한 기존 diffusion teacher 모델의 지식을 1스텝짜리 student 모델로 옮기기 위해, 두 모델의 노이즈 수준을 맞춰주는 시간 변환 함수 Φ(τ)를 만들어 서로 다른 방식으로 학습된 모델 간에도 지식 전달(distillation)이 잘 되도록 했다.
  4. 1스텝만 쓰면 입 모양이 뭉개지거나 움직임이 어색해지는 문제를 막기 위해, 음성 신호를 활용한 classifier-free guidance(음성 조건을 강조해 결과를 조정하는 기법)를 추가해 입 모양 동기화와 움직임 다양성을 살렸다.
  5. HDTF, CelebV-HQ 등 두 데이터셋 80개 영상에서 여러 기존 방법과 같은 조건으로 비교 실험을 하고, DINO 유사도로 시간이 지나도 얼굴이 유지되는지, 사용자 설문으로 체감 품질까지 측정했다.
Figure 1: Overview of LeapTalk. Given audio and a reference image, our method enables open-ended streaming talking-head generation with consistent identity. It achieves 1-step inference per chunk at up to 200 FPS, delivering up to 𝟏𝟓𝟎𝟎𝟎× speedup while maintaining strong lip-sync accuracy.
Figure 1: Overview of LeapTalk. Given audio and a reference image, our method enables open-ended streaming talking-head generation with consistent identity. It achieves 1-step inference per chunk at up to 200 FPS, delivering up to 𝟏𝟓𝟎𝟎𝟎× speedup while maintaining strong lip-sync accuracy.
Table 1: Comparison on HDTF and CelebV-HQ datasets.
HDTFCelebV-HQ
ModelNFEFID↓FVD↓Sync-C↑Sync-D↓IQA↑ASE↑FPS↑FID↓FVD↓Sync-C↑Sync-D↓IQA↑ASE↑FPS↑
StableAvatar501763298.118.056.512.690.423184924.738.615.423.300.42
Echomimic307229815.3210.026.162.290.5264218851.0911.020.153.190.52
Hallo3518719727.149.236.242.600.3084211042.498.855.233.220.29
FantasyTalking304598846.339.416.132.450.1554414291.848.615.213.180.15
OmniAvatar501686233.1012.366.322.680.183519231.2510.035.393.270.16
Soulx-Flashhead4304528.078.246.602.9514.42716424.778.235.573.3514.42
OURS (Lite)1382858.147.896.222.70200474564.808.225.513.33200
OURS (Pro)1211978.387.696.532.7455423705.058.215.583.3455
Figure 2: Comparison between conventional noise-to-data flow matching and our Bridge Forcing paradigm. The identity consistency is preserved along bridge process, while error accumulates in the forward diffusion process.
Figure 2: Comparison between conventional noise-to-data flow matching and our Bridge Forcing paradigm. The identity consistency is preserved along bridge process, while error accumulates in the forward diffusion process.
Table 2: Ablation study results.
MethodFIDSync-CSync-D
LeapTalk218.387.69
w/o Brownian Bridge2177.1611.05
w/o Time Transformation3787.848.13
w/o Audio-Driven CFG1624.3410.21
Figure 3: Distillation pipeline of our approach. The one-step student generates frames conditioned on a static reference. To enable stable distillation across heterogeneous processes, we apply a time transformation t=Φ⁡(τ) to align noise levels between the flow-matching teacher and the bridge-based student, allowing consistent score supervision and a well-defined DMD objective.
Figure 3: Distillation pipeline of our approach. The one-step student generates frames conditioned on a static reference. To enable stable distillation across heterogeneous processes, we apply a time transformation t=Φ⁡(τ) to align noise levels between the flow-matching teacher and the bridge-based student, allowing consistent score supervision and a well-defined DMD objective.
Table 3: Perceptual loss weight sensitivity. Bold: best.
Metricλperc=0λperc=1.0λperc=2.0λperc=4.0λperc=8.0
PSNR↑18.6219.1819.3019.7017.31
SSIM↑0.6250.6830.6920.7040.349
LPIPS↓0.2030.1970.1940.1830.562
Figure 4: Qualitative comparison on long-video streaming generation. Frames are sampled from streaming rollouts at increasing time indices. The red boxes mark identity inconsistency and drift, and the blue boxes mark slight or incorrect lip-sync. LeapTalk maintains stable identity and lip motion as generation proceeds.
Figure 4: Qualitative comparison on long-video streaming generation. Frames are sampled from streaming rollouts at increasing time indices. The red boxes mark identity inconsistency and drift, and the blue boxes mark slight or incorrect lip-sync. LeapTalk maintains stable identity and lip motion as generation proceeds.
Table 4: User study results across different evaluation criteria. The table reports the percentage of participants who prefer our method to another method. A higher percentage indicates stronger user preference and better perceived performance.
MethodIdentity Cons.Lip-sync Acc.Visual QualityOverall Pref.
StableAvatar92.40%93.14%91.25%91.37%
Echomimic91.26%96.14%94.36%95.11%
SoulX-FlashHead95.23%92.16%91.89%95.32%
Hallo394.81%95.48%96.29%96.21%
FantasyTalking96.89%93.32%95.36%97.90%
OmniAvatar95.95%94.31%97.43%95.75%
Figure 5: DINO similarity over video time. We plot the similarity between each generated frame and the reference image as streaming generation progresses. A flatter and higher curve indicates less identity drift over long-duration generation.
Figure 5: DINO similarity over video time. We plot the similarity between each generated frame and the reference image as streaming generation progresses. A flatter and higher curve indicates less identity drift over long-duration generation.
Table 5: Comparison of autoencoder backbones used in Lite and Pro variants. Encoding and decoding speeds are measured on video clips of 81 frames under BF16 precision.
Autoenc.Arch.Enc. SpeedDec. SpeedEnc. Mem.Dec. Mem.
WanVAECausal Conv3D4.17s5.26s8.495GB10.128GB
TAEHVConv2D0.39s0.24s0.008GB0.411GB
Figure 6: Ablation results demonstrating the contribution of each component in our framework.
Figure 6: Ablation results demonstrating the contribution of each component in our framework.
Table 6: Motion diversity vs. baselines on the HDTF samples. Bold: best; uline: 2nd best.
MethodYaw Std↑Pitch Std↑Roll Std↑Avg Std↑BAS↑
EchoMimic1.7341.8670.7711.4570.650
OmniAvatar4.3325.8881.9094.0430.652
SoulX-Flashhead3.5942.9181.9592.8240.684
Ours6.1775.6322.1464.6520.696
Figure 7: Visualization of ablation effects of audio-driven CFG.
Figure 7: Visualization of ablation effects of audio-driven CFG.
Table 7: Audio CFG scale ablation on HDTF.
CFG ScaleYaw Std↑Pitch Std↑Roll Std↑Avg Std↑BAS↑
1.01.5532.6530.7571.6550.723
3.04.0214.4931.6683.3940.658
5.06.8405.7642.3965.0000.696
7.08.5337.6352.8016.3230.650
Figure 8: Sensitivity analysis of audio-driven CFG.
Figure 8: Sensitivity analysis of audio-driven CFG.
Table 8: Chunk size ablation at 512×512, 1-step inference, single A100 GPU. Tgen/Tchunk<1 indicates real-time. Bold represents chosen default, which achieves the best ratio.
Chunk SizeTchunk (s)Tgen (s)Tgen/Tchunk ↓FPS ↑
90.360.0880.2445.4
170.680.1460.2282.0
331.320.2680.20104.7
491.960.4320.22101.9
652.600.6290.2495.3
Figure 9: Visual comparison across different VAEs.
Figure 9: Visual comparison across different VAEs.
Table 9: Inference speed and average chunk generation latency under different resolutions on a single A100 GPU (1-step inference).
Resolution256×256384×384𝟓𝟏𝟐×𝟓𝟏𝟐768×7681024×1024
FPS ↑4761911043514
Latency (s) ↓0.0590.1460.2700.7931.908
Figure 10: Generated results under diverse and challenging input scenarios.
Figure 10: Generated results under diverse and challenging input scenarios.

실제로 확인된 결과

  • HDTF와 CelebV-HQ 80개 영상 비교에서 LeapTalk Pro가 FID/FVD 기준 가장 좋은 화질 점수(HDTF 21/197, CelebV-HQ 42/370)와 가장 강한 입-음성 동기화(Sync-C/D) 점수를 기록했고, Lite 모델은 최대 초당 200프레임까지 속도를 냈다.
  • DINO 유사도(생성된 프레임과 원본 사진의 닮음 정도) 측정에서 LeapTalk은 시간이 지나도 유사도가 높고 평탄하게 유지되었지만, OmniAvatar 같은 기존 방법은 시간이 지날수록 유사도가 눈에 띄게 떨어졌다.
  • 구성요소를 하나씩 제거하는 ablation 실험에서, Brownian Bridge를 빼면 FID가 21에서 217로 크게 나빠지고 Sync-C도 8.38에서 7.16으로 떨어졌으며, 시간 변환 Φ(τ)를 빼면 FID가 378로 나빠졌고, 음성 기반 CFG를 빼면 Sync-C 4.34, Sync-D 10.21로 입 동기화가 크게 나빠졌다.
  • 30명이 참여한 사용자 설문에서 정체성 일관성, 화질, 입 동기화, 전체 선호도 네 항목 모두에서 LeapTalk이 비교 대상보다 더 많이 선택되었다.
  • TAEHV라는 경량 오토인코더를 쓴 Lite 모델은 3D Conv 기반 WanVAE 대비 계산량과 메모리를 크게 줄이면서도 구조·정체성·움직임 품질이 비슷했고, 입술 등 미세 영역에서만 약간 흐릿해지는 정도였다.

어디에 쓸 수 있나

  • 화상 상담, 가상 비서, 라이브 방송 아바타처럼 음성에 실시간으로 반응해 얼굴 영상을 끊김 없이 계속 만들어야 하는 서비스
  • 짧은 클립이 아니라 임의로 긴 시간 동안 정체성이 유지되는 아바타 영상이 필요한 콘텐츠 제작이나 디지털 휴먼 응용
  • 고정형 캐릭터 사진뿐 아니라 만화, 조각상, 옆모습 등 특이한 입력 이미지로도 말하는 영상을 만들어야 하는 실험적 응용

한계와 남은 검증

  • 보고된 정량 비교는 HDTF, CelebV-HQ 두 데이터셋 80개 영상과 특정 GPU(H200, A100) 환경에 한정되어 있어, 다른 데이터나 하드웨어에서의 성능은 이 실험만으로 보장되지 않는다.
  • 본문에서 밝힌 200FPS는 H200 GPU 기준이고 A100 기준 속도는 별도 표로 제공되며, 저자들은 하드웨어 표기를 개정판에서 통일하겠다고 밝혔다.
  • 경량 오토인코더(TAEHV)를 쓰면 입술 등 세밀한 영역에서 약간의 흐릿함이 발생하며, 이는 추론 스텝을 1에서 2로 늘리면 완화된다고만 언급되어 있어 완전히 해결된 것은 아니다.
  • 음성 CFG 강도(α)가 커질수록 움직임 다양성은 늘지만 Beat Align Score(음성-동작 정렬 지표)는 낮아지는 경향이 있어, 값 설정에 따른 트레이드오프가 존재한다.
  • 동물, 그림, 조각상 등 다양한 입력에 대한 결과는 정성적 사례로만 제시되어 있고, 이런 비정형 입력에 대한 정량적 평가는 본문에 제공되지 않았다.

왜 중요한가

실시간 화상통화나 가상 아바타처럼 사람이 말하는 동안 지연 없이 얼굴 영상을 계속 만들어야 하는 서비스에서는, 화질과 속도를 동시에 만족시키는 방법이 그동안 부족했다. 이 연구는 그 둘을 절충하지 않고 둘 다 잡으려는 시도를 실제로 측정된 성능으로 보여준다는 점에서 실무적 참고가 된다.

이 논문의 용어

  • diffusion · 무작위 노이즈에서 여러 단계를 거쳐 점점 선명한 이미지나 영상으로 복원해가는 생성 방식
  • 자기회귀(autoregressive) · 직전에 생성한 결과를 다음 생성의 입력으로 다시 사용하는 방식
  • Brownian Bridge · 두 고정된 시작점과 끝점을 연결하면서 중간에만 무작위성을 허용하는 확률 과정
  • distillation(지식 증류) · 크고 느린 모델(teacher)의 판단을 작고 빠른 모델(student)이 따라 배우게 하는 학습 방법
  • classifier-free guidance(CFG) · 조건(여기서는 음성)을 강하게 반영하도록 생성 결과를 조정하는 기법

본문에 싣지 못한 그림

  • Figure 11: More qualitative results on challenging conditions including non-human faces (e.g., animals), artistic portraits (e.g., paintings and stylized illustrations), side-view photos, low-light conditions, partial occlusions, and even non-photorealistic objects such as sculptures.
원문에서 그림 보기 →

저자 · Rongxiang Zhang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Rongxiang Zhang et al., arXiv:2608.00079, CC BY 4.0