Figure 1: PixRestore achieves the best overall performance in terms of restoration quality, model size, and inference speed. Top-left: radar charts comparing PSNR, LPIPS, and degradation removal performance across eight degradation types. Top-right: GFLOPs versus inference latency, where the bubble size denotes the number of parameters. Bottom: visual comparisons on eight restoration tasks. With only about 50M parameters and single-step inference, PixRestore achieves the best overall restoration quality while being the most efficient among diffusion-based methods.
Table 1: Latent diffusion vs. pixel diffusion under the same UIR training/test setting. The results are averaged over 8 restoration tasks. Pixel-space modeling provides a better overall trade-off in fidelity, perceptual quality, parameter count, and inference speed.
Model
VAE
Params(M)
Inf Time-DM (ms/step)
Inf Time-VAE (ms)
PSNR (dB) ↑
LPIPS ↓
MUSIQ ↑
Latent DiT-S
SD2VAE (f8c4, ps1)
106.41
25
78
22.10
0.2483
50.38
Latent DiT-S
FluxVAE (f8c16, ps1)
106.59
25
78
22.63
0.2109
50.86
Latent DiT-S
QwenVAE (f8c16, ps1)
67.37
25
41
22.80
0.2181
51.87
Pixel DiT-S
None (ps8)
23.41
25
0
26.62
0.1593
54.32
Figure 2: Overview of PixRestore. A frozen vision encoder extracts multi-layer features from the LQ image, and an adaptive layer router predicts per-layer weights to fuse them into a single conditioning feature. The LQ image and the noisy state xt are patchified, then processed by N DiT blocks, and finally decoded into the HQ output.
Table 2: Quantitative comparison on public benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained using the same training dataset as ours. Metrics: PSNR↑, SSIM↑, LPIPS↓, DISTS↓, DR-Score (degradation-removal score)↑.
Method
De-rainstreak
Denoise
Deblur
De-raindrop
PSNR
SSIM
LPIPS
DISTS
DR-Score
PSNR
SSIM
LPIPS
DISTS
DR-Score
PSNR
SSIM
LPIPS
DISTS
DR-Score
PSNR
SSIM
LPIPS
DISTS
PromptIR
24.50
0.7507
0.3529
0.2381
31.25
32.54
0.9029
0.2459
0.1604
55.35
23.82
0.7584
0.3131
0.2195
27.23
18.81
0.6962
0.3390
0.1885
PromptIR∗
28.43
0.8389
0.2659
0.1904
49.99
35.45
0.9377
0.1104
0.1065
79.33
29.10
0.8474
0.2051
0.1560
55.56
23.68
0.8005
0.2236
0.1326
DiffUIR
24.51
0.7618
0.3484
0.2248
40.13
26.61
0.8380
0.3103
0.2056
59.52
27.86
0.8200
0.2258
0.1699
44.09
18.85
0.7004
0.3455
0.1905
UniRestore
22.50
0.7291
0.4223
0.2689
34.31
31.75
0.8969
0.2336
0.1669
61.20
23.87
0.7205
0.2405
0.1716
46.94
18.50
0.6377
0.4114
0.2285
DA-CLIP
24.50
0.7560
0.3347
0.2176
43.09
27.35
0.8245
0.2459
0.1764
64.12
27.40
0.8113
0.1641
0.1303
55.00
20.12
0.6988
0.2715
0.1512
DA-CLIP∗
31.61
0.8606
0.1048
0.0885
78.72
33.97
0.8728
0.1462
0.1166
75.14
28.29
0.8228
0.1474
0.1148
58.38
23.58
0.7916
0.1325
0.0852
FoundIR
26.87
0.8263
0.2454
0.1799
52.37
32.46
0.7925
0.2794
0.1641
55.52
27.29
0.8046
0.2430
0.1781
39.12
18.87
0.6972
0.3586
0.2017
FoundIR∗
32.36
0.8888
0.1593
0.1203
71.12
36.16
0.9415
0.1051
0.1245
80.86
29.14
0.8466
0.1986
0.1503
49.86
24.40
0.8218
0.1942
0.1146
Flux-IR
20.98
0.6233
0.4625
0.2861
26.04
25.71
0.6842
0.4197
0.2367
54.86
23.51
0.6745
0.2746
0.1851
58.70
18.94
0.6459
0.3224
0.1774
Flux-IR∗
21.01
0.6344
0.4511
0.2822
31.04
26.80
0.7410
0.3789
0.2051
37.05
22.15
0.6446
0.2818
0.2036
73.18
17.80
0.5785
0.3266
0.2075
FoundIR-v2
23.17
0.6652
0.3659
0.2218
57.26
26.87
0.7481
0.2785
0.1946
74.47
24.40
0.7079
0.2203
0.1566
73.23
19.63
0.5459
0.2792
0.1563
FoundIR-v2∗
27.85
0.7598
0.1773
0.1322
81.03
28.17
0.7497
0.2889
0.1890
75.01
24.98
0.7351
0.1914
0.1323
77.12
20.81
0.5461
0.2300
0.1294
FAPE-IR
27.51
0.8226
0.2319
0.1679
71.09
32.99
0.9088
0.1238
0.1113
79.31
26.80
0.7895
0.2098
0.1556
50.72
21.39
0.6822
0.2219
0.1316
FAPE-IR∗
31.91
0.8695
0.0903
0.0762
84.64
34.22
0.9172
0.0741
0.0750
82.56
27.96
0.8207
0.1577
0.1094
66.75
23.88
0.7223
0.1555
0.0967
PixRestore
32.28
0.8847
0.0902
0.0905
82.75
34.87
0.9336
0.0624
0.0735
81.89
28.32
0.8284
0.1201
0.0940
70.19
24.48
0.7755
0.1258
0.0882
PixRestore-B
32.85
0.8898
0.0767
0.0817
84.72
34.62
0.9356
0.0564
0.0706
83.70
29.23
0.8511
0.1051
0.0835
72.69
25.21
0.7976
0.1087
0.0763
Method
Desnow
Dehaze
Low-light enhancement
Super-resolution
PSNR
SSIM
LPIPS
DISTS
DR-Score
PSNR
SSIM
LPIPS
DISTS
DR-Score
PSNR
SSIM
LPIPS
DISTS
DR-Score
PSNR
SSIM
LPIPS
DISTS
PromptIR
22.26
0.7939
0.2452
0.1672
24.35
21.34
0.8803
0.1426
0.1012
61.68
10.49
0.4781
0.5410
0.3765
27.87
24.26
0.7372
0.4394
0.2510
PromptIR∗
29.32
0.8532
0.1818
0.1402
72.39
21.25
0.8931
0.1348
0.0885
67.89
17.91
0.6881
0.3564
0.2512
58.58
27.59
0.7968
0.2839
0.2208
DiffUIR
22.95
0.7948
0.2392
0.1667
25.00
20.41
0.8656
0.1640
0.1175
57.19
21.72
0.7082
0.3819
0.2292
62.93
26.54
0.7582
0.3866
0.2378
UniRestore
22.33
0.7863
0.2464
0.1771
26.12
20.11
0.8469
0.2122
0.1359
67.33
10.94
0.5188
0.5006
0.3165
38.38
24.80
0.7543
0.3548
0.2284
DA-CLIP
23.60
0.7971
0.2221
0.1558
35.53
22.96
0.8751
0.1311
0.0885
70.33
22.25
0.7915
0.2469
0.1622
73.29
23.73
0.6789
0.3772
0.2367
DA-CLIP∗
28.31
0.8282
0.1360
0.1021
81.48
21.92
0.8779
0.1459
0.1077
58.58
18.01
0.8185
0.2129
0.1614
72.79
26.49
0.7585
0.2341
0.1763
FoundIR
23.03
0.7999
0.2406
0.1630
24.68
15.07
0.7901
0.2582
0.1919
33.17
15.34
0.7473
0.3034
0.2132
62.24
25.85
0.7399
0.4280
0.2492
FoundIR∗
29.82
0.8678
0.1524
0.1224
73.00
20.99
0.8903
0.1335
0.0962
64.81
23.34
0.9026
0.1792
0.1439
80.41
27.65
0.7952
0.2883
0.2251
Flux-IR
21.74
0.7231
0.3434
0.2204
26.80
14.73
0.7599
0.2976
0.1986
46.13
18.85
0.7022
0.3557
0.1998
62.16
22.49
0.6541
0.2903
0.2064
Flux-IR∗
21.79
0.6831
0.3023
0.1959
51.21
15.96
0.8061
0.2315
0.1578
52.88
18.40
0.7198
0.3386
0.2010
61.45
20.87
0.5853
0.3569
0.2519
FoundIR-v2
24.72
0.7347
0.2513
0.1671
67.47
19.06
0.7705
0.1928
0.1337
62.79
17.17
0.7448
0.3132
0.2027
72.49
23.85
0.6661
0.2959
0.1974
Figure 3: Motivation of adaptive hierarchical visual guidance. Left: Per-layer DINO feature visualizations for low-light enhancement and de-raindrop. Shallow layers preserve local structures and details, while deeper layers encode global semantics. Right: LQ–HQ feature similarity across DINOv2-B layers for eight types of degradations. We see that different layers are sensitive to different degradations.
Table 3: Complexity comparison of different methods. “NFE” denotes number of evaluations. All values are measured with input 1×3×512×512 on a single NVIDIA A800 GPU, with 5 warmup iterations and averaged over 100 runs.
Method
NFE
Params (M)
FLOPs (G)
Latency (ms)
PromptIR
-
35.59
1382
334
DiffUIR
4
36.26
3839
777
DA-CLIP
100
231.76
112927
18071
FoundIR
4
36.26
3839
777
UniRestore
1
1071.20
5255
158
FoundIR-v2
20
16910.09
119334
18293
Flux-IR
21
17698.23
795290
5790
FAPE-IR
1
21575.43
64758
1011
PixRestore
1
53.70
658
44
PixRestore-B
1
210.89
1842
79
Figure 4: No-reference quality metrics do not reliably reflect degradation removal. Here PixRestore removes the rainstreaks best, yet MUSIQ and AFINE-NR rank it worst, favoring the LQ input and the degradation-preserving output of FoundIR-v2, while our VLM-based DR-Score demonstrates strong alignment with human perceptual judgments.
Table 4: Ablation studies of PixRestore. “Avg.” denotes uniform averaging of selected DINO features for conditioning or supervision, while “Adap.” denotes adaptive averaging of selected DINO features for conditioning or supervision. “NFE” denotes the number of evaluations.
ID
Variant
NFE
Conditioning
Supervision
PSNR↑
SSIM↑
LPIPS↓
MUSIQ↑
A0
Pixel DiT-S
10
None
None
26.62
0.8454
0.1593
54.32
A1
+ single-layer conditioning
10
layer 2
None
27.07
0.8457
0.1561
54.45
A2
+ single-layer conditioning
10
layer 5
None
27.36
0.8487
0.1489
54.64
A3
+ single-layer conditioning
10
layer 11
None
27.12
0.8491
0.1508
54.60
A4
+ multi-layer conditioning
10
Avg. 2 layers
None
27.62
0.8540
0.1412
54.90
A5
+ multi-layer conditioning
10
Avg. 6 layers
None
27.72
0.8536
0.1407
55.01
A6
+ hierarchical loss
10
Avg. 6 layers
Avg. 6 layers
27.36
0.8444
0.1239
54.59
A7
+ adaptive hierarchical visual guidance
10
Adap. 6 layers
Adap. 6 layers
27.66
0.8500
0.1209
54.97
A8
+ adaptive hierarchical visual guidance
4
Adap. 6 layers
Adap. 6 layers
27.75
0.8584
0.1228
54.07
A9
+ adaptive hierarchical visual guidance
1
Adap. 6 layers
Adap. 6 layers
28.07
0.8640
0.1202
53.34
A10
+ single-step finetuning (PixRestore)
1
Adap. 6 layers
Adap. 6 layers
28.49
0.8589
0.1120
55.52
Figure 5: Visual comparisons on desnow (top) and low-light enhancement (bottom). PixRestore removes degradations effectively and recovers more faithful details and colors.
Table 5: Comparison between diffusion pretraining and finetuning with regression training under the same objective and total training iterations.
Scheme
Training iterations
PSNR↑
SSIM↑
LPIPS↓
MUSIQ↑
Regression Training
350k
27.00
0.8179
0.1494
52.00
Flow pretraining + one-step finetuning (ours)
250k + 100k
28.49
0.8589
0.1120
55.52
Figure 6: Scaling behavior of PixRestore under varying backbone GFLOPs and patch sizes.
Table 6: Quantitative comparison on real-world test set. The best and second-best results for each metric are highlighted in red bold and blue italic, respectively. Retrained methods are marked with ∗. ‘PR’ denotes the proposed PixRestore, the results of which are shaded in pink.
Degradation
Metric
PromptIR
PromptIR∗
DiffUIR
DA-CLIP
DA-CLIP∗
FoundIR
FoundIR∗
UniRestore
FoundIR-v2
FoundIR-v2∗
Flux-IR
Flux-IR∗
FAPE-IR
FAPE-IR∗
PR
PR-B
De-rainstreak
MUSIQ↑
59.90
60.06
61.02
62.23
59.94
61.00
60.87
62.58
63.06
62.57
62.98
60.23
61.07
59.53
62.62
62.78
Affine-NR↓
-0.90
-0.90
-0.92
-0.91
-0.91
-0.90
-0.92
-0.89
-0.96
-0.96
-0.94
-0.89
-1.00
-1.00
-1.02
-1.03
DR-Score ↑
30.85
30.05
33.77
33.62
51.68
31.48
38.83
32.87
49.07
71.72
25.50
28.37
64.73
76.32
72.12
74.69
Deblur
MUSIQ↑
34.28
32.57
41.64
46.61
43.92
32.85
34.10
49.26
69.99
72.78
55.26
65.73
44.89
45.84
53.37
56.56
Affine-NR↓
-0.73
-0.70
-0.78
-0.77
-0.76
-0.70
-0.72
-0.81
-1.04
-1.10
-0.87
-1.00
-0.80
-0.83
-0.88
-0.93
DR-Score ↑
26.56
31.08
28.12
48.78
45.22
26.58
29.35
46.77
66.66
74.94
38.09
72.18
63.35
65.41
65.43
73.08
De-raindrop
MUSIQ↑
64.18
63.47
64.04
66.26
54.46
61.85
60.62
63.87
65.81
57.74
65.60
66.28
52.71
39.61
47.99
54.54
Affine-NR↓
-0.83
-0.78
-0.82
-0.87
-0.71
-0.77
-0.71
-0.78
-0.83
-0.79
-0.92
-1.02
-0.85
-0.76
-0.72
-0.80
DR-Score↑
25.70
26.15
22.10
37.63
44.65
26.57
32.38
26.00
34.30
62.59
36.05
36.03
80.73
80.82
74.88
80.62
Desnow
MUSIQ↑
58.96
59.27
59.92
59.74
59.36
60.24
60.15
60.30
63.35
63.46
58.33
62.12
57.96
58.68
62.69
62.65
Affine-NR↓
-0.74
-0.75
-0.77
-0.78
-0.77
-0.75
-0.75
-0.73
-0.83
-0.86
-0.76
-0.89
-0.86
-0.88
-0.90
-0.91
DR-Score ↑
26.90
34.17
39.53
45.99
46.69
28.01
32.45
32.13
46.90
64.89
38.63
50.63
71.57
71.73
71.86
74.70
Dehaze
MUSIQ↑
59.83
60.40
59.66
61.00
60.21
60.26
60.55
60.85
63.38
61.32
63.39
60.32
59.77
60.33
61.93
61.47
Affine-NR↓
-0.88
-0.88
-0.87
-0.88
-0.87
-0.87
-0.88
-0.82
-0.90
-0.86
-0.89
-0.83
-0.88
-0.90
-0.93
-0.93
DR-Score ↑
32.90
37.68
24.55
32.36
28.86
31.02
32.43
47.43
42.61
37.52
40.37
40.74
32.67
35.94
44.51
42.06
Low-light
MUSIQ↑
47.67
55.03
54.21
64.66
53.18
49.61
57.10
48.47
63.33
61.31
54.10
52.75
49.91
57.60
58.17
58.63
Affine-NR↓
-0.88
-0.92
-0.71
-0.95
-0.90
-0.89
-0.98
-0.85
-1.00
-0.94
-0.93
-0.91
-0.89
-0.99
-0.98
-0.99
DR-Score↑
31.65
67.03
59.33
66.24
65.10
31.70
65.05
38.87
71.25
66.62
58.07
54.71
38.70
75.09
62.38
62.13
Average
MUSIQ↑
54.14
55.13
56.75
60.08
55.18
54.30
55.56
57.55
64.82
63.20
59.45
61.24
54.39
53.60
57.80
59.44
Affine-NR↓
-0.83
-0.82
-0.81
-0.86
-0.82
-0.81
-0.83
-0.81
-0.93
-0.92
-0.89
-0.92
-0.88
-0.88
-0.91
-0.93
DR-Score ↑
29.09
37.69
34.57
44.10
47.03
29.23
38.42
37.51
51.80
63.05
39.45
47.11
58.63
67.55
65.20
67.88
Figure 7: Visual comparisons on real-world desnow, dehaze, and deblur cases. PixRestore can remove degradations more effectively and recover cleaner details and sharper structures with fewer artifacts.
Table S.1: Training data sources for each degradation type.
RealESRGAN degradation from DF2K [34, 67], RealSR [42], ScreenSR [43]
Figure 8: Visual comparisons on real-world de-raindrop, low-light enhancement, and de-rainstreak cases. PixRestore can remove degradations more effectively and recover cleaner details and sharper structures with fewer artifacts.
Table S.2: Real-world test data sources for each degradation type.
Degradation
Real-world test data
Deblur
GyroBlur-Real [70]
Dehaze
RTTS and OpenReal-fog [44]
De-raindrop
OpenReal-raindrop [44]
De-rainstreak
OpenReal-rainstreak [44] and DiffUIR[10]
Desnow
Snow100K-realistic [65] and OpenReal-snow [44]
Low-light
enhancement
OpenReal-night [44] and ExDark [71]
Figure S.1: Illustration of DR-Score. Given the LQ input, the restored result, and the task description, the VLM judges whether the target degradation has been removed. A better restoration receives a higher DR-Score.
Table S.3: Pixel-space and latent-space comparison on 8 degradation types. We report comparisons on PSNR↑/LPIPS↓/MUSIQ↑.
Degradation
Latent DiT with FLUX-VAE
Latent DiT with Qwen-VAE
Latent DiT with SD2-VAE
Pixel DiT
SR
25.59/0.2038/58.68
25.62/0.1992/62.25
25.03/0.2586/58.53
26.48/0.2067/57.80
Deblur
26.66/0.1617/46.53
26.93/0.1545/46.69
26.00/0.1921/45.23
27.17/0.1850/41.76
Dehaze
15.31/0.2253/54.56
15.16/0.2391/52.50
15.08/0.2659/53.31
24.78/0.1020/58.17
Denoise
31.97/0.0952/47.22
32.33/0.1021/48.37
30.65/0.1252/47.90
32.58/0.1179/51.32
De-raindrop
20.43/0.2257/62.79
20.63/0.2463/64.71
20.01/0.2751/62.46
23.19/0.1707/66.66
De-rainstreak
25.07/0.1478/46.49
25.34/0.1821/47.63
24.53/0.1876/46.00
30.19/0.1350/49.43
Desnow
27.82/0.1304/49.00
28.41/0.1190/49.39
27.30/0.1483/49.45
28.96/0.1301/48.58
Low-light enhancement
10.82/0.4572/40.64
10.82/0.4530/42.13
10.81/0.4834/39.73
20.80/0.2122/57.98
Overall
22.63/0.2109/50.86
22.80/0.2181/51.87
22.10/0.2483/50.38
26.62/0.1593/54.32
Figure S.2: Human alignment analysis of different no-reference metrics with pairwise human preferences. Left: alignment ratios between metric-induced rankings and human judgments across real-world UIR tasks. Right: representative failure cases where two restored results are visually similar with subtle appearance differences but lead to disagreement between DR-Score and human judgment.
Table S.4: Comparison of frozen visual foundation priors under the same pixel DiT setting. All encoders share the ViT-B architecture and parameter count. Features of the same layer are selected. Metrics are averaged over 8 degradation types. Best results are highlighted in bold.
Visual Prior
PSNR (dB) ↑
SSIM ↑
LPIPS ↓
MUSIQ ↑
CLIP-B
26.61
0.8441
0.1598
54.30
MAE-B
26.70
0.8454
0.1593
54.20
SigLIP-B
26.66
0.8449
0.1592
54.19
DINOv2-B
27.25
0.8500
0.1531
54.21
Figure S.3: Visual comparisons on synthetic desnow, low-light enhancement, de-rainstreak, and denoise cases. Overall, PixRestore restores cleaner and more faithful results with better structural details and fewer artifacts.
Table S.5: Detailed quantitative comparison on Deblur benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
GoPro
UHD-blur
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
23.43
0.7903
0.3035
0.1972
25.50
-0.5874
24.21
0.7265
0.3228
0.2418
29.81
-0.6163
PromptIR∗
29.82
0.8775
0.2000
0.1440
36.18
-0.7123
28.39
0.8172
0.2101
0.1681
40.52
-0.7570
DiffUIR
29.32
0.8669
0.2041
0.1492
35.93
-0.7137
26.39
0.7731
0.2476
0.1907
38.01
-0.7075
UniRestore
24.11
0.7496
0.2216
0.1408
45.93
-0.7896
23.63
0.6914
0.2595
0.2024
42.99
-0.6870
DA-CLIP
28.57
0.8554
0.1279
0.0994
41.05
-0.7329
26.23
0.7671
0.2003
0.1612
42.83
-0.7479
DA-CLIP∗
28.87
0.8536
0.1317
0.1031
42.91
-0.7444
27.71
0.7920
0.1630
0.1264
47.05
-0.7656
FoundIR
27.02
0.8121
0.2610
0.1800
29.96
-0.6276
27.56
0.7971
0.2250
0.1762
38.80
-0.7544
FoundIR∗
29.95
0.8762
0.1887
0.1393
35.81
-0.7297
28.33
0.8170
0.2084
0.1613
39.98
-0.7696
FoundIR-v2
24.48
0.7179
0.2224
0.1451
57.73
-0.8954
24.32
0.6980
0.2183
0.1680
63.23
-1.0422
FoundIR-v2∗
24.98
0.7470
0.1965
0.1262
56.90
-0.8914
24.97
0.7232
0.1862
0.1385
58.39
-0.9774
Flux-IR
24.03
0.7012
0.2408
0.1550
55.90
-0.8500
22.98
0.6479
0.3084
0.2153
53.29
-0.8064
Flux-IR∗
22.59
0.6716
0.2717
0.1936
60.22
-0.9609
21.71
0.6176
0.2918
0.2137
58.22
-0.9313
FAPE-IR
28.02
0.8374
0.1527
0.1058
41.62
-0.7781
25.58
0.7417
0.2669
0.2054
35.00
-0.7267
FAPE-IR∗
28.68
0.8505
0.1347
0.0907
38.84
-0.7403
27.25
0.7908
0.1807
0.1282
41.57
-0.7694
PixRestore-S
28.96
0.8553
0.1067
0.0833
43.39
-0.8182
27.67
0.8016
0.1335
0.1046
49.73
-0.8105
PixRestore-B
30.00
0.8805
0.0895
0.0721
44.74
-0.8354
28.46
0.8218
0.1206
0.0949
49.54
-0.8200
PixRestore-L
30.79
0.8948
0.0787
0.0657
45.51
-0.8616
28.96
0.8346
0.1117
0.0891
50.45
-0.8515
PixRestore-XL
31.23
0.9025
0.0721
0.0621
46.34
-0.8712
29.07
0.8407
0.1090
0.0854
50.34
-0.8487
Figure S.4: Visual comparisons on synthetic dehaze, de-raindrop, deblur, and SR cases. Overall, PixRestore restores cleaner and more faithful results with better structural details and fewer artifacts.
Table S.6: Detailed quantitative comparison on Dehaze benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
RESIDE-6K
UHD-Haze
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
26.69
0.9572
0.0460
0.0418
51.24
-0.9422
15.98
0.8035
0.2391
0.1605
63.94
-0.9765
PromptIR∗
24.01
0.9417
0.0642
0.0553
51.58
-0.9362
18.49
0.8445
0.2054
0.1216
63.45
-1.0058
DiffUIR
24.66
0.9311
0.0708
0.0582
50.59
-0.9314
16.15
0.8001
0.2572
0.1768
62.57
-0.9683
UniRestore
23.63
0.9162
0.1171
0.0869
55.68
-0.9630
16.59
0.7776
0.3073
0.1849
62.14
-0.9155
DA-CLIP
28.66
0.9273
0.0559
0.0463
53.52
-0.9784
17.27
0.8230
0.2063
0.1307
65.88
-1.0698
DA-CLIP∗
28.15
0.9579
0.0431
0.0409
51.31
-0.9439
15.70
0.7979
0.2487
0.1746
63.98
-0.9540
FoundIR
16.52
0.8381
0.1665
0.1304
49.68
-0.8612
13.62
0.7421
0.3499
0.2533
59.75
-0.8267
FoundIR∗
25.07
0.9525
0.0523
0.0483
50.61
-0.9464
16.91
0.8282
0.2146
0.1441
64.37
-0.9976
FoundIR-v2
19.01
0.8101
0.1837
0.1318
57.83
-0.9396
19.11
0.7309
0.2018
0.1357
69.04
-1.0056
FoundIR-v2∗
18.57
0.7925
0.1997
0.1342
56.12
-0.9191
19.62
0.7365
0.1870
0.1222
68.08
-1.0543
Flux-IR
15.64
0.7796
0.2591
0.1639
65.46
-1.0185
13.83
0.7402
0.3361
0.2333
61.74
-0.8161
Flux-IR∗
17.27
0.8416
0.1578
0.1074
50.54
-0.8348
14.65
0.7706
0.3052
0.2083
60.48
-0.8251
FAPE-IR
31.36
0.9628
0.0378
0.0364
50.35
-0.9573
19.20
0.8386
0.1542
0.0937
66.12
-1.1555
FAPE-IR∗
28.90
0.9575
0.0411
0.0385
50.32
-0.9488
21.78
0.8537
0.1442
0.0905
65.45
-1.1361
PixRestore-S
28.41
0.9555
0.0473
0.0444
52.81
-0.9777
22.51
0.8729
0.1318
0.0855
66.98
-1.1836
PixRestore-B
29.87
0.9636
0.0390
0.0388
51.79
-0.9737
23.41
0.8820
0.1188
0.0760
67.26
-1.2451
PixRestore-L
30.63
0.9660
0.0365
0.0366
52.11
-0.9772
23.23
0.8836
0.1227
0.0771
66.81
-1.2381
PixRestore-XL
31.73
0.9677
0.0340
0.0355
51.83
-0.9741
23.14
0.8827
0.1223
0.0776
66.92
-1.2237
Table S.7: Detailed quantitative comparison on Denoise benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
DIV2K (Gaussian)
PolyU
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
34.57
0.9079
0.1411
0.1384
64.24
-0.9098
30.50
0.8978
0.3506
0.1825
30.88
-0.6068
PromptIR∗
33.91
0.8975
0.1463
0.1388
61.00
-0.8517
37.00
0.9780
0.0744
0.0742
32.70
-0.6156
DiffUIR
21.24
0.7540
0.3437
0.2523
55.21
-0.7013
31.98
0.9221
0.2769
0.1588
32.67
-0.6270
UniRestore
30.38
0.8728
0.1659
0.1420
64.25
-0.8868
33.11
0.9209
0.3012
0.1918
31.60
-0.5629
DA-CLIP
28.91
0.7532
0.2412
0.1764
57.33
-0.8155
25.79
0.8957
0.2506
0.1765
33.00
-0.6137
DA-CLIP∗
30.76
0.7704
0.2448
0.1570
58.12
-0.7929
37.19
0.9753
0.0477
0.0761
33.67
-0.5932
FoundIR
27.14
0.6061
0.4920
0.2574
45.98
-0.6097
37.77
0.9789
0.0668
0.0707
33.55
-0.6303
FoundIR∗
34.06
0.8992
0.1572
0.1469
62.82
-0.9074
38.25
0.9837
0.0529
0.1022
34.24
-0.6195
FoundIR-v2
25.43
0.6544
0.2658
0.1825
62.76
-0.9197
28.31
0.8418
0.2913
0.2066
52.61
-0.8119
FoundIR-v2∗
25.96
0.6584
0.2488
0.1837
62.60
-0.9080
30.39
0.8410
0.3290
0.1943
41.44
-0.6478
Flux-IR
21.33
0.4780
0.5758
0.2660
50.37
-0.7322
30.08
0.8903
0.2637
0.2073
43.26
-0.7832
Flux-IR∗
24.34
0.5837
0.4296
0.2386
49.22
-0.7338
29.26
0.8984
0.3281
0.1716
32.85
-0.6348
FAPE-IR
31.09
0.8540
0.1147
0.1013
60.49
-0.8739
34.89
0.9636
0.1329
0.1212
35.29
-0.6526
FAPE-IR∗
31.32
0.8572
0.1080
0.0978
60.52
-0.8830
37.11
0.9772
0.0401
0.0522
33.52
-0.6135
PixRestore-S
33.04
0.8923
0.0783
0.0863
63.62
-0.9410
36.71
0.9749
0.0465
0.0607
34.54
-0.6139
PixRestore-B
33.40
0.8968
0.0725
0.0777
63.84
-0.9591
35.84
0.9745
0.0402
0.0636
34.15
-0.6030
PixRestore-L
33.35
0.8962
0.0730
0.0768
64.12
-0.9671
35.91
0.9736
0.0385
0.0709
34.39
-0.6022
PixRestore-XL
33.46
0.8979
0.0729
0.0774
64.06
-0.9648
36.50
0.9745
0.0341
0.0753
33.95
-0.5994
Table S.8: Detailed quantitative comparison on De-rainstreak benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
RainDS-real
RealRain-1K
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
25.15
0.7483
0.2047
0.1392
60.54
-0.8602
23.85
0.7531
0.5011
0.3370
42.82
-0.6698
PromptIR∗
26.62
0.7933
0.1977
0.1254
61.79
-0.8963
30.23
0.8845
0.3341
0.2554
36.49
-0.5757
DiffUIR
26.11
0.7888
0.1885
0.1168
63.88
-0.9157
22.91
0.7348
0.5084
0.3329
45.21
-0.7440
UniRestore
23.47
0.7143
0.3220
0.1838
65.71
-0.8797
21.52
0.7439
0.5226
0.3540
45.97
-0.7192
DA-CLIP
24.66
0.7467
0.1834
0.1213
63.20
-0.9041
24.35
0.7654
0.4861
0.3139
46.39
-0.7354
DA-CLIP∗
25.72
0.7516
0.1432
0.0859
63.31
-0.9261
37.50
0.9695
0.0665
0.0910
32.27
-0.6034
FoundIR
26.78
0.7890
0.1630
0.1075
63.26
-0.9045
26.97
0.8635
0.3278
0.2523
39.15
-0.6486
FoundIR∗
27.27
0.8120
0.1993
0.1191
67.29
-0.9812
37.44
0.9655
0.1193
0.1216
35.47
-0.6340
FoundIR-v2
23.65
0.6078
0.1956
0.1111
63.58
-0.9722
22.69
0.7226
0.5362
0.3326
48.23
-0.7375
FoundIR-v2∗
23.74
0.6043
0.1906
0.1088
64.82
-1.0001
31.96
0.9153
0.1641
0.1556
36.17
-0.6297
Flux-IR
22.34
0.6665
0.2695
0.1769
64.41
-0.9676
19.63
0.5801
0.6554
0.3952
50.43
-0.7374
Flux-IR∗
21.66
0.6410
0.2912
0.1803
62.00
-0.9306
20.37
0.6278
0.6110
0.3840
45.93
-0.6811
FAPE-IR
26.53
0.7636
0.1407
0.0835
62.03
-0.9748
28.50
0.8817
0.3232
0.2523
42.51
-0.7340
FAPE-IR∗
26.67
0.7642
0.1270
0.0746
61.14
-0.9816
37.14
0.9748
0.0535
0.0778
31.15
-0.6092
PixRestore-S
27.06
0.7950
0.1180
0.0773
66.07
-1.0433
37.50
0.9744
0.0625
0.1038
32.07
-0.6152
PixRestore-B
27.41
0.8038
0.1078
0.0693
65.65
-1.0461
38.28
0.9758
0.0456
0.0941
32.27
-0.6153
PixRestore-L
27.54
0.8065
0.0990
0.0664
65.60
-1.0532
38.51
0.9760
0.0404
0.0933
32.65
-0.6170
PixRestore-XL
27.58
0.8079
0.1000
0.0650
65.18
-1.0463
38.60
0.9753
0.0401
0.1028
32.81
-0.6160
Table S.9: Detailed quantitative comparison on De-raindrop benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
RainDS-real
UAV-Rain1k
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
20.70
0.7069
0.2773
0.1560
55.31
-0.8088
16.92
0.6855
0.4007
0.2211
66.57
-0.8022
PromptIR∗
24.53
0.7529
0.2637
0.1347
59.14
-0.8560
22.84
0.8482
0.1835
0.1305
68.19
-0.8475
DiffUIR
20.56
0.6996
0.3119
0.1678
56.28
-0.8034
17.13
0.7011
0.3790
0.2132
66.67
-0.7942
UniRestore
20.25
0.6906
0.3482
0.1888
58.44
-0.8111
16.75
0.5847
0.4746
0.2683
65.96
-0.7044
DA-CLIP
22.99
0.7018
0.1763
0.0977
61.66
-0.8832
17.26
0.6958
0.3667
0.2047
67.27
-0.8080
DA-CLIP∗
24.14
0.7077
0.1570
0.0905
60.92
-0.8855
23.01
0.8755
0.1079
0.0799
69.43
-0.9228
FoundIR
20.67
0.7171
0.3145
0.1749
59.34
-0.8041
17.07
0.6773
0.4026
0.2285
67.57
-0.7823
FoundIR∗
25.41
0.7708
0.2487
0.1319
65.20
-0.9316
23.40
0.8729
0.1397
0.0973
69.58
-0.9111
FoundIR-v2
20.22
0.5725
0.3079
0.1566
59.84
-0.8704
19.03
0.5194
0.2505
0.1560
69.47
-0.8975
FoundIR-v2∗
21.72
0.5641
0.2435
0.1255
64.75
-0.9876
19.91
0.5280
0.2165
0.1332
69.30
-0.9315
Flux-IR
21.01
0.6540
0.2335
0.1221
63.08
-1.0046
16.87
0.6377
0.4113
0.2328
66.83
-0.8472
Flux-IR∗
18.85
0.5672
0.3313
0.1964
68.66
-1.1729
16.76
0.5897
0.3218
0.2186
72.29
-1.0842
FAPE-IR
24.71
0.7255
0.1841
0.0930
59.93
-0.9452
18.06
0.6389
0.2598
0.1702
65.88
-0.8070
FAPE-IR∗
25.27
0.7263
0.1502
0.0790
58.40
-0.9358
22.49
0.7183
0.1607
0.1145
68.35
-0.9177
PixRestore-S
25.42
0.7394
0.1326
0.0761
61.54
-0.9911
23.54
0.8117
0.1190
0.1002
70.15
-0.9816
PixRestore-B
25.76
0.7461
0.1245
0.0720
61.41
-0.9928
24.65
0.8490
0.0928
0.0806
70.35
-1.0176
PixRestore-L
25.89
0.7504
0.1195
0.0686
61.30
-0.9963
25.16
0.8631
0.0810
0.0719
70.63
-1.0379
PixRestore-XL
26.01
0.7567
0.1171
0.0677
61.36
-0.9939
25.69
0.8774
0.0715
0.0655
70.55
-1.0392
Table S.10: Detailed quantitative comparison on Low-light Enhancement benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
UHD-LL
LOL
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
11.82
0.5660
0.5088
0.3116
35.06
-0.6577
9.17
0.3902
0.5732
0.4415
38.49
-0.7859
PromptIR∗
25.33
0.8846
0.2303
0.1757
45.85
-0.6794
10.50
0.4916
0.4824
0.3268
44.53
-0.8607
DiffUIR
17.67
0.5064
0.6057
0.3320
37.19
-0.4779
25.76
0.9099
0.1581
0.1264
68.79
-0.9938
UniRestore
12.41
0.6102
0.4665
0.2903
36.98
-0.6417
9.48
0.4274
0.5347
0.3426
43.85
-0.7569
DA-CLIP
20.51
0.7436
0.3675
0.2201
48.51
-0.6172
23.99
0.8395
0.1263
0.1043
74.18
-1.0050
DA-CLIP∗
16.64
0.7776
0.2632
0.2016
47.98
-0.7589
19.38
0.8594
0.1626
0.1212
66.85
-0.9450
FoundIR
14.22
0.6984
0.3548
0.2386
42.38
-0.7037
16.47
0.7963
0.2519
0.1877
65.83
-1.0337
FoundIR∗
24.28
0.8885
0.2113
0.1687
51.48
-0.8552
22.40
0.9168
0.1471
0.1191
71.35
-1.0455
FoundIR-v2
16.40
0.7120
0.3643
0.2381
60.71
-0.9313
17.95
0.7776
0.2620
0.1674
67.72
-1.0373
FoundIR-v2∗
17.72
0.7272
0.3541
0.2273
55.85
-0.8751
16.36
0.7341
0.2834
0.1929
62.16
-0.9620
Flux-IR
15.35
0.5519
0.5455
0.2965
38.47
-0.5140
22.36
0.8525
0.1658
0.1032
72.10
-1.1667
Flux-IR∗
14.33
0.6028
0.4866
0.2852
37.25
-0.5275
22.47
0.8369
0.1906
0.1167
70.91
-1.2086
FAPE-IR
12.57
0.6315
0.3899
0.2669
39.07
-0.7489
26.34
0.8942
0.1261
0.1006
66.74
-1.0265
FAPE-IR∗
26.55
0.8980
0.1445
0.1096
51.54
-0.8156
25.53
0.9049
0.1296
0.1040
65.50
-1.0535
PixRestore-S
26.28
0.8889
0.1420
0.1060
53.31
-0.8461
24.96
0.8899
0.1300
0.0957
66.07
-1.0386
PixRestore-B
26.36
0.8886
0.1384
0.1010
53.45
-0.8534
25.08
0.8983
0.1212
0.0899
65.88
-1.0329
PixRestore-L
26.89
0.8928
0.1326
0.0979
54.53
-0.8613
25.30
0.8960
0.1273
0.0932
64.39
-1.0055
PixRestore-XL
26.37
0.8934
0.1324
0.0982
54.62
-0.8686
26.20
0.8955
0.1286
0.0949
63.64
-1.0015
Table S.11: Detailed quantitative comparison on Desnow benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Method
WeatherBench
PSNR↑
SSIM↑
LPIPS↓
DISTS↓
MUSIQ↑
AFINE-NR↓
PromptIR
22.26
0.7939
0.2452
0.1672
45.60
-0.6023
PromptIR∗
29.32
0.8532
0.1818
0.1402
46.30
-0.6161
DiffUIR
22.95
0.7948
0.2392
0.1667
48.10
-0.6185
UniRestore
22.33
0.7863
0.2464
0.1771
49.39
-0.6451
DA-CLIP
23.60
0.7971
0.2221
0.1558
46.48
-0.6127
DA-CLIP∗
28.31
0.8282
0.1360
0.1021
48.82
-0.6424
FoundIR
23.03
0.7999
0.2406
0.1630
45.77
-0.6029
FoundIR∗
29.82
0.8678
0.1524
0.1224
47.98
-0.6908
FoundIR-v2
24.72
0.7347
0.2513
0.1671
59.98
-0.8032
FoundIR-v2∗
26.28
0.7690
0.1824
0.1311
54.44
-0.7440
Flux-IR
21.74
0.7231
0.3434
0.2204
56.08
-0.6879
Flux-IR∗
21.79
0.6831
0.3023
0.1959
57.03
-0.7504
FAPE-IR
26.02
0.8191
0.1759
0.1189
46.29
-0.6280
FAPE-IR∗
30.19
0.8676
0.1136
0.0849
47.84
-0.6595
PixRestore-S
31.26
0.8859
0.0853
0.0669
49.86
-0.6829
PixRestore-B
31.93
0.8959
0.0688
0.0584
50.10
-0.6825
PixRestore-L
32.35
0.9039
0.0656
0.0569
50.36
-0.6859
PixRestore-XL
32.57
0.9077
0.0623
0.0553
50.25
-0.6864
Table S.12: Detailed quantitative comparison on Super-resolution benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with ∗ are retrained under the same training setting as ours.
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models for their strong capacity and generative priors. However, the variational autoencoder (VAE) in latent T2I models may discard restoration-sensitive details, while the open-ended synthesis prior can introduce content-inconsistent artifacts. We present PixRestore, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining. PixRestore performs flow matching directly on patchified pixels, preserving fine-grained details while keeping the token sequence tractable. To adapt to different degradations, PixRestore learns to predict the reliability of layer features using LQ--HQ DINO feature similarity. Features from more reliable layers are fused as dense conditioning, while less reliable layers receive stronger HQ-feature supervision to encourage degradation removal. We train PixRestore on a large-scale corpus of diverse scenes and degradations, and further finetune it into a one-step generator using DINO-based adversarial objectives for efficient inference. Experiments on public benchmarks and real-world test sets show that, with only about 50M parameters and single-step inference, PixRestore achieves the best overall fidelity, perceptual quality, and robustness to degradations among competing UIR models while being far more efficient. Larger PixRestore variants can further boost performance, demonstrating the scalability of our pixel-space design. Code and the curated benchmark can be found at https://github.com/csslc/PixRestore.