K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

arXiv:2608.180632026-08-17

先低分辨率编辑再对照高分辨率原图重新补全细节,让4K图像编辑仅需61秒完成

高分辨率图像编辑通常只能在1K以下运行,因为注意力计算量会随图像尺寸呈平方级增长,常见的折中方案是先在低分辨率下编辑再用超分辨率放大,但这种方式容易凭空捏造与原图不符的细节,或出现过度平滑、过度锐化的纹理问题。EditBridge不是从噪声重新生成图像,而是训练一个扩散桥模型,直接学习把低分辨率编辑结果转换为高分辨率结果,并以原始高分辨率原图作为引导;同时配合一种稀疏注意力机制,未编辑区域只参考原图中精确对应的位置,编辑区域只参考附近小范围,从而大幅降低计算量。最终在2K分辨率下比超分辨率方法快3.6到8.4倍,4K图像编辑仅需61秒完成。

METAL MEDIA 解读图

先低分辨率编辑再对照高分辨率原图重新补全细节,让4K图像编辑仅需61秒完成

  1. 01高分辨率图像编辑因注意力计算量随图像尺寸呈平方级增长,通常只能在1K以下分辨率运行
  2. 02常见的低分辨率编辑加超分辨率放大方案会出现信息发散(捏造出与原图不符的细节)和纹理退化(过度平滑或过度锐化)两大问题
  3. 03EditBridge采用扩散桥,直接学习从低分辨率编辑结果到高分辨率目标图像的转换路径,而不是从噪声重新生成图像,从而在整个过程中保留原始结构信息
  4. 04提出的先验引导块状稀疏注意力(PG-BSA)让未编辑区域只关注原图中精确对应的位置,编辑区域只关注周边小范围,显著减少计算开销
  5. 05在2K分辨率下比基于超分辨率的方法快3.6到8.4倍(耗时4.08秒),4K图像编辑仅需61秒完成,用户研究显示该方法在细节保留、真实感和美观度上均优于对比方法
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 高分辨率图像编辑因注意力计算量随图像尺寸呈平方级增长,通常只能在1K以下分辨率运行
  2. 常见的低分辨率编辑加超分辨率放大方案会出现信息发散(捏造出与原图不符的细节)和纹理退化(过度平滑或过度锐化)两大问题
  3. EditBridge采用扩散桥,直接学习从低分辨率编辑结果到高分辨率目标图像的转换路径,而不是从噪声重新生成图像,从而在整个过程中保留原始结构信息
  4. 提出的先验引导块状稀疏注意力(PG-BSA)让未编辑区域只关注原图中精确对应的位置,编辑区域只关注周边小范围,显著减少计算开销
  5. 在2K分辨率下比基于超分辨率的方法快3.6到8.4倍(耗时4.08秒),4K图像编辑仅需61秒完成,用户研究显示该方法在细节保留、真实感和美观度上均优于对比方法
Table 1: Quantitative comparison results. ↑/↓ indicate whether higher/lower values are better. Bold numbers represent the best results.
MethodHaar ↑mPSNR ↑mSSIM ↑mMSE ↓mLPIPS ↓Time (s) ↓
1K Resolution
Direct Inference0.397316.80780.75390.05490.411620.00
DiT4SR [14]0.412714.50310.70920.05540.481828.47
DiT-SR [9]0.452316.71520.79550.04300.44032.44
PiSA-SR [32]0.424715.21150.74810.04860.44920.30
TSD-SR [12]0.412715.38640.75310.05160.46280.42
HiFlow [7]0.301211.69640.64890.13230.60125.12
ScaleEdit [22]0.482416.84610.82060.04080.3784181.20
Ours0.499218.65360.86860.03690.38290.84
2K Resolution
Direct Inference0.225413.52620.74910.08800.670082.14
DiT4SR [14]0.380615.12980.74680.05160.5348113.32
DiT-SR [9]0.421017.24300.82650.04230.473814.82
PiSA-SR [32]0.404616.24330.80610.04460.470220.22
TSD-SR [12]0.404616.20210.79520.04670.491034.20
HiFlow [7]0.314413.93140.73770.07590.511333.11
ScaleEdit [22]0.458316.55100.68180.04850.5197662.50
Ours0.467318.63660.87120.03770.39034.08
4K Resolution
Direct Inference0.213410.70860.54110.12810.7774695.23
DiT4SR [14]0.400220.27710.75350.01330.3510533.83
DiT-SR [9]0.426716.86150.68180.04150.5436119.68
PiSA-SR [32]0.426721.62000.79740.01080.2687111.42
TSD-SR [12]0.412422.27660.80350.00740.2399108.72
HiFlow [7]0.330714.44000.67190.06640.6038129.30
ScaleEdit [22]0.411618.57140.76380.03900.60153601.20
Ours0.427924.13800.83470.00490.211261.10
Table 2: Performance comparison between our PG-BSA and full attention across 1K and 2K settings. ↑/↓ indicate whether higher/lower values are better. Bold numbers represent the best results within each resolution block.
SettingMethodHaarPSI↑mPSNR↑mSSIM↑mMSE↓mLPIPS↓Time (s)↓
1KOurs0.49918.6530.8680.0360.3820.84
FullAttention0.53919.0800.8570.0300.5990.73
2KOurs0.46718.6360.8710.0370.3904.08
FullAttention0.50517.9180.7050.0410.4254.71
Table 3: Quantitative evaluation of varying inference steps. ↑/↓ indicate whether higher/lower values are better. Bold numbers represent the best results.
MethodHaarPSI↑mPSNR↑mSSIM↑mMSE↓mLPIPS↓Time(s)↓
5 steps0.413115.72690.77630.04740.47394.03
10 steps0.402015.41900.76360.04940.46628.10
Ours (1 step)0.499118.65350.86850.03680.38290.84

为什么重要

高分辨率图像编辑是专业修图、设计等实际工作中的刚需,但此前速度与保真度往往难以兼得。这项研究证明二者可以同时实现,为将高分辨率编辑功能落地到实际产品降低了门槛。

本文术语

  • 扩散桥(Diffusion Bridge) · 一种直接学习从起始图像过渡到目标图像的生成模型,而非从随机噪声生成
  • 超分辨率(Super-Resolution) · 把低分辨率图像放大为高分辨率的技术,需要凭空补出缺失的细节
  • PG-BSA(先验引导块状稀疏注意力) · 一种只选择性关注原图中空间对应区域以降低计算量的注意力机制
  • DiT(扩散Transformer) · 一种基于Transformer架构的图像生成模型
  • LoRA(低秩适配) · 一种不重新训练整个模型,只训练少量新增参数即可完成微调的方法

无法转载的图表

  • Figure 1: Motivation of the proposed method. (a) and (b) compare the standard high-resolution image editing paradigm with our approach. (c) analyzes the inherent limitations of existing methods.
  • Figure 2: Overview of our proposed EditBridge. The upper section illustrates the inference process of the diffusion bridge, which transports the low-resolution edit to its high-resolution counterpart. The lower section details the construction of the correspondence prior and the proposed Prior-Guided Sparse Attention (PG-BSA) mechanism.
  • Figure 3: Qualitative comparison at 1K resolution. Our method better preserves fidelity while enhancing visual clarity.
  • Figure 4: Qualitative comparison at 2K resolution. Our method better preserves fidelity while enhancing visual clarity.
  • Figure 5: Ablation of Attention Sparsity. Full attention leads to noticeable source-induced artifacts, whereas our sparse attention maintains superior visual fidelity and alignment.
  • Figure 6: Impact of sampling steps. Perceptual quality and high-frequency details are maximized at T=1. Increasing the number of inference steps does not yield significant visual gains, further validating the efficiency of our refinement framework.
  • Figure 7: Results of the user study across three dimensions: Detail Preservation, Realism, and Aesthetics. Our method consistently outperforms all baselines in human preference.
  • Figure 8: Visual results at 1K resolution.
  • Figure 9: Visual results at 2K resolution.
在原文中查看图表 →

论文原文摘要(英文)

High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requirements. A prevalent workaround employs a two-stage pipeline: editing at low resolution followed by independent super-resolution. However, this approach suffers from two critical issues: information divergence, where hallucinated details contradict the original high-resolution (HR) source, and texture degradation, manifesting as over-smoothed or over-sharpened artifacts. We propose EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing. Unlike conventional diffusion that regenerates from noise, we formulate refinement as structured data-to-data translation from the low-resolution (LR) edited result to its HR counterpart, explicitly conditioned on the original HR source to preserve authentic details. To efficiently incorporate HR source guidance, we introduce a prior-guided block-wise sparse attention mechanism that exploits semantic correspondence from first-stage editing to constrain cross-image interactions to spatially aligned regions, significantly reducing computational overhead. Extensive experiments demonstrate that EditBridge achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4times speedup at 2K and enabling practical 4K editing in 61 seconds.

作者 · Jiayi Song

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道