Table 1: Cross-modal retrieval performance of CLIP-based fine-tuning methods on four dense-caption benchmarks. We report Recall@K (%) on both Text→Image and Image→Text settings. All fine-tuned methods start from the same Long-CLIP-L backbone with an identical training budget; Long-CLIP denotes the released checkpoint. Best results in bold; second best underlined. Δ denotes the margin over the best competitor per column, with gain in ↑ green.
DOCCI
DCI
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP[ECCV’24]
78.78
95.24
98.02
66.75
91.92
96.31
67.83
83.19
87.69
64.13
84.84
89.74
FineLIP[CVPR’25]
77.51
96.02
98.41
69.90
93.43
97.45
72.69
87.14
90.65
65.48
86.84
91.00
GOAL[CVPR’25]
81.53
97.02
98.80
80.86
96.24
98.63
77.29
90.25
93.30
74.84
89.94
93.25
StructXLIP[CVPR’26]
84.73
97.69
99.00
82.61
97.08
98.71
75.84
89.94
93.65
74.49
90.05
93.40
HN-CLIP
88.25
98.45
99.43
86.24
98.12
99.22
80.69
92.40
95.10
78.84
91.90
94.65
Δ
↑ 3.52
↑ 0.76
↑ 0.43
↑ 3.63
↑ 1.04
↑ 0.51
↑ 3.40
↑ 2.15
↑ 1.45
↑ 4.00
↑ 1.85
↑ 1.25
Table 2: Plug-and-play enhancement of our ℒHN on CLIP-based fine-tuning. Results on DOCCI and Long-DCI for Text→Image and Image→Text retrieval. Upper: full-parameter fine-tuning; lower: parameter-efficient tuning. Our method consistently boosts diverse CLIP variants, and the gains grow with caption length. Best in bold, with gain in ↑ green.
DOCCI
Long-DCI
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP[ECCV’24]
86.45
98.00
99.31
84.10
97.84
99.04
70.24
89.62
94.00
68.44
89.35
94.67
+ our ℒHN
88.31
98.75
99.45
86.39
98.08
99.20
78.43
92.52
95.27
76.02
91.78
94.76
Δ
↑ 1.86
↑ 0.75
↑ 0.14
↑ 2.29
↑ 0.24
↑ 0.16
↑ 8.19
↑ 2.90
↑ 1.27
↑ 7.58
↑ 2.43
↑ 0.09
FineLIP[CVPR’25]
77.51
96.02
98.41
69.90
93.43
97.45
59.24
77.86
83.19
49.52
75.08
82.39
+ our ℒHN
85.47
97.71
99.04
83.88
97.41
98.82
74.92
90.52
93.62
73.39
89.87
93.23
Δ
↑ 7.96
↑ 1.69
↑ 0.63
↑ 13.98
↑ 3.98
↑ 1.37
↑ 15.68
↑ 12.66
↑ 10.43
↑ 23.87
↑ 14.79
↑ 10.84
GOAL[CVPR’25]
81.53
97.02
98.80
80.86
96.24
98.63
74.29
92.77
95.58
73.31
92.14
95.77
+ our ℒHN
86.24
98.14
99.35
84.86
97.69
99.20
84.27
94.99
96.57
82.28
94.26
96.18
Δ
↑ 4.71
↑ 1.12
↑ 0.55
↑ 4.00
↑ 1.45
↑ 0.57
↑ 9.98
↑ 2.22
↑ 0.99
↑ 8.97
↑ 2.12
↑ 0.41
StructXLIP[CVPR’26]
84.73
97.69
99.00
82.61
97.08
98.71
75.34
93.15
95.84
72.30
92.79
95.78
+ our ℒHN
85.45
97.92
99.12
83.31
97.45
98.92
82.04
94.18
96.23
78.98
93.22
95.65
Δ
↑ 0.72
↑ 0.23
↑ 0.12
↑ 0.70
↑ 0.37
↑ 0.21
↑ 6.70
↑ 1.03
↑ 0.39
↑ 6.68
↑ 0.43
↓ 0.13
LoRA[ICLR’22]
80.49
96.16
98.55
77.96
95.80
98.04
60.08
80.68
86.86
58.97
80.02
86.53
+ our ℒHN
82.82
97.14
98.86
80.84
96.49
98.63
64.87
83.09
88.35
62.79
81.77
87.31
Δ
↑ 2.33
↑ 0.98
↑ 0.31
↑ 2.88
↑ 0.69
↑ 0.59
↑ 4.79
↑ 2.41
↑ 1.49
↑ 3.82
↑ 1.75
↑ 0.78
DoRA[ICML’24]
80.76
96.25
98.61
78.20
95.90
98.10
60.76
81.19
87.48
59.65
80.31
87.13
+ our ℒHN
83.41
97.33
99.00
81.41
96.65
98.67
65.46
83.57
88.61
63.38
82.15
87.56
Δ
↑ 2.65
↑ 1.08
↑ 0.39
↑ 3.21
↑ 0.75
↑ 0.57
↑ 4.70
↑ 2.38
↑ 1.13
↑ 3.73
↑ 1.84
↑ 0.43
Table 3: Ablations. (a) Sensitivity to the boost strength γ (Recall@1): every γ∈[0.25,1] beats γ=0 in-domain, while under transfer (Urban-1K) milder boosts generalize better. (b) Loss components under an identical recipe; only the loss changes. (c) G¯ recomputed from the live encoder (default) vs. frozen at initialization; R@1/5/10 in Table A2. Best per column in bold.
DOCCI
DCI
Long-DCI
Urban-1K
γ
R@1
R@1
R@1
R@1
R@1
R@1
R@1
R@1
Avg
γ=0
86.45
84.29
79.24
77.54
70.92
68.00
91.30
93.00
81.34
γ=0.25
87.82
86.63
80.74
79.29
76.38
74.80
91.70
93.00
83.80
γ=0.5 (default)
88.25
86.24
80.69
78.84
79.03
76.44
91.10
90.30
83.86
γ=0.75
88.20
85.73
81.04
78.59
79.96
76.72
89.10
89.50
83.61
γ=1.0
87.98
85.02
80.34
78.19
79.94
76.59
87.90
88.40
83.04
Table A1: Benchmark statistics. Mean caption length is measured on the evaluation split in words; similarities use the pre-trained Long-CLIP-L text encoder.
Benchmark
#train
#test
words
pairwise sim
hardest sim
DOCCI
9450
5100
123
0.84
0.93
DCI
5445
1999
133
0.85
0.92
Long-DCI
5445 (=DCI)
7444
134
0.85
0.93
Urban-1K
14579 (VG)
1000
107
0.88
0.94
Table A2: Recomputed vs. frozen G¯ (Recall@K, Text→Image and Image→Text). Building the boost from a frozen copy of the pre-trained text encoder makes the margins exactly stationary but removes their implicit annealing: accuracy peaks at epoch 1 on DCI and Urban-1K and then declines, and the default wins 22 of 24 columns. Best per column in bold. Upper band: DOCCI and DCI; lower band: Long-DCI and the transfer setting Urban-1K.
DOCCI
DCI
G¯ source
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Recomputed G¯ (default)
88.25
98.45
99.43
86.24
98.12
99.22
80.69
92.40
95.10
78.84
91.90
94.65
Frozen G¯
85.86
97.61
99.18
82.43
96.82
98.73
77.79
90.80
93.35
75.84
89.49
92.95
Δ (default − frozen)
↑ 2.39
↑ 0.84
↑ 0.25
↑ 3.81
↑ 1.30
↑ 0.49
↑ 2.90
↑ 1.60
↑ 1.75
↑ 3.00
↑ 2.41
↑ 1.70
best epoch (default / frozen)
6 / 5
5 / 1
Long-DCI
Urban-1K
G¯ source
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Recomputed G¯ (default)
79.03
93.76
96.34
76.44
92.79
95.91
91.10
98.10
99.40
90.30
98.20
99.30
Frozen G¯
79.93
90.84
93.82
76.75
89.54
92.62
87.30
97.70
98.90
83.50
95.90
97.90
Δ (default − frozen)
↓ 0.90
↑ 2.92
↑ 2.52
↓ 0.31
↑ 3.25
↑ 3.29
↑ 3.80
↑ 0.40
↑ 0.50
↑ 6.80
↑ 2.30
↑ 1.40
best epoch (default / frozen)
10 / 10
4 / 1
Table A3: Training-efficiency comparison on DCI fine-tuning (10 epochs, 5.4k images, single Ascend 910B, effective batch 128). Wall-clock for GOAL/StructXLIP is measured from same-device sequential runs minus evaluation overhead; FineLIP ran in parallel and is not attributable. Offline preprocessing time (segmentation, edge extraction, LLM filtering) is not included in the wall-clock.
Method
Aux. training inputs
Offline prep.
Wall-clock
Throughput
Long-CLIP (plain FT)
none
none
≈16 min
55.5 img/s
FineLIP
none
none
—
—
GOAL
SAM segments
required
≈41 min
≈22 img/s
StructXLIP
edges + LLM lexicon
required
≈92 min
≈10 img/s
HN-CLIP (ours)
none
none
17 min
53 img/s
Table A4: Plug-and-play enhancement of ℒHN on CLIP-based fine-tuning, all four benchmarks. Results on DOCCI, DCI, Long-DCI, and Urban-1K for Text→Image and Image→Text retrieval; per framework we report the official baseline, the same recipe with ℒHN replacing its global InfoNCE term, and the per-column difference Δ (gain in ↑ green, drop in ↓ gray; differences within ±0.15 are shown as ≈0). Bold marks the +ℒHN value where it is the better of the pair. Upper band: DOCCI and DCI; lower band: Long-DCI and the transfer setting Urban-1K.
DOCCI
DCI
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP
86.45
98.00
99.31
84.10
97.84
99.04
79.39
91.60
94.65
78.34
92.85
95.25
+our ℒHN
88.31
98.75
99.45
86.39
98.08
99.20
81.04
92.20
95.00
78.54
91.95
94.60
Δ
↑ 1.86
↑ 0.75
≈0
↑ 2.29
↑ 0.24
↑ 0.16
↑ 1.65
↑ 0.60
↑ 0.35
↑ 0.20
↓ 0.90
↓ 0.65
FineLIP†
77.12
95.94
98.29
70.16
93.14
97.29
72.69
87.14
90.65
65.48
86.84
91.00
+our ℒHN
85.47
97.71
99.04
83.88
97.41
98.82
80.74
91.90
94.55
79.99
91.80
94.35
Δ
↑ 8.35
↑ 1.77
↑ 0.75
↑ 13.72
↑ 4.27
↑ 1.53
↑ 8.05
↑ 4.76
↑ 3.90
↑ 14.51
↑ 4.96
↑ 3.35
GOAL†
81.96
96.94
98.78
80.84
96.33
98.61
77.29
90.25
93.30
74.84
89.94
93.25
+our ℒHN
86.24
98.14
99.35
84.86
97.69
99.20
79.59
90.80
93.30
77.09
90.10
93.00
Δ
↑ 4.28
↑ 1.20
↑ 0.57
↑ 4.02
↑ 1.36
↑ 0.59
↑ 2.30
↑ 0.55
≈0
↑ 2.25
↑ 0.16
↓ 0.25
StructXLIP
84.73
97.69
99.00
82.61
97.08
98.71
75.84
89.94
93.65
74.49
90.05
93.40
+our ℒHN
85.45
97.92
99.12
83.31
97.45
98.92
77.34
90.75
93.10
74.14
89.89
92.30
Δ
↑ 0.72
↑ 0.23
≈0
↑ 0.70
↑ 0.37
↑ 0.21
↑ 1.50
↑ 0.81
↓ 0.55
↓ 0.35
↓ 0.16
↓ 1.10
LoRA
80.49
96.16
98.55
77.96
95.80
98.04
73.84
89.44
92.85
71.89
88.74
92.85
+our ℒHN
82.82
97.14
98.86
80.84
96.49
98.63
75.84
90.20
93.00
73.44
88.54
92.25
Δ
↑ 2.33
↑ 0.98
↑ 0.31
↑ 2.88
↑ 0.69
↑ 0.59
↑ 2.00
↑ 0.76
↑ 0.15
↑ 1.55
↓ 0.20
↓ 0.60
DoRA
80.76
96.25
98.61
78.20
95.90
98.10
74.14
89.74
93.00
72.44
89.24
92.95
+our ℒHN
83.41
97.33
99.00
81.41
96.65
98.67
76.24
90.45
93.15
73.89
88.59
92.30
Δ
↑ 2.65
↑ 1.08
↑ 0.39
↑ 3.21
↑ 0.75
↑ 0.57
↑ 2.10
↑ 0.71
↑ 0.15
↑ 1.45
↓ 0.65
↓ 0.65
Long-DCI
Urban-1K
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP
70.24
89.62
94.00
68.44
89.35
94.67
91.70
99.10
99.50
93.20
99.00
99.50
+our ℒHN
78.43
92.52
95.27
76.02
91.78
94.76
90.60
98.30
99.40
90.20
98.30
99.40
Δ
↑ 8.19
↑ 2.90
↑ 1.27
↑ 7.58
↑ 2.43
≈0
↓ 1.10
↓ 0.80
≈0
↓ 3.00
↓ 0.70
≈0
FineLIP†
59.24
77.86
83.19
49.52
75.08
82.39
81.10
94.50
97.40
77.50
94.80
97.50
+our ℒHN
74.92
90.52
93.62
73.39
89.87
93.23
91.80
98.10
99.30
93.20
98.80
99.30
Δ
↑ 15.68
↑ 12.66
↑ 10.43
↑ 23.87
↑ 14.79
↑ 10.84
↑ 10.70
↑ 3.60
↑ 1.90
↑ 15.70
↑ 4.00
↑ 1.80
GOAL†
74.34
92.89
95.77
72.03
92.37
95.69
86.30
97.00
98.90
86.10
97.40
98.90
+our ℒHN
84.27
94.99
96.57
82.28
94.26
96.18
85.70
96.70
98.20
86.80
96.60
98.00
Δ
↑ 9.93
↑ 2.10
↑ 0.80
↑ 10.25
↑ 1.89
↑ 0.49
↓ 0.60
↓ 0.30
↓ 0.70
↑ 0.70
↓ 0.80
↓ 0.90
Table A5: Loss ablation, full resolution (all four benchmarks, all ranks). Best per column in bold.
DOCCI
DCI
ℒtok
ℒHN
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
✗
✗
86.45
98.00
99.31
99.86
99.98
84.10
97.84
99.04
99.78
99.98
79.39
91.60
94.65
97.25
97.95
78.34
92.85
95.25
✓
✗
86.45
97.98
99.43
99.86
99.98
84.29
97.90
98.94
99.78
99.94
79.24
91.55
94.25
97.00
98.00
77.54
92.95
95.35
✗
✓
88.31
98.75
99.45
99.84
99.98
86.39
98.08
99.20
99.82
99.96
81.04
92.20
95.00
97.40
98.10
78.54
91.95
94.60
✓
✓
88.25
98.45
99.43
99.86
99.98
86.24
98.12
99.22
99.84
99.94
80.69
92.40
95.10
97.35
98.15
78.84
91.90
94.65
Long-DCI
Urban-1K
ℒtok
ℒHN
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
✗
✗
70.24
89.62
94.00
97.66
98.79
68.44
89.35
94.67
98.11
98.95
91.70
99.10
99.50
99.80
99.90
93.20
99.00
99.50
✓
✗
70.92
89.51
93.90
97.27
98.64
68.00
89.12
94.17
97.84
98.93
91.30
99.00
99.30
99.80
99.90
93.00
99.10
99.50
✗
✓
78.43
92.52
95.27
97.21
98.47
76.02
91.78
94.76
97.31
98.29
90.60
98.30
99.40
99.60
99.80
90.20
98.30
99.40
✓
✓†
78.79
92.32
95.12
97.19
98.27
76.18
91.99
95.15
97.26
98.40
91.00
98.50
99.50
99.70
99.90
90.50
98.50
99.30
Table A6: γ sweep, full resolution (all four benchmarks, all ranks). Best per column in bold.
DOCCI
DCI
γ
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
γ=0
86.45
97.98
99.43
99.86
99.98
84.29
97.90
98.94
99.78
99.94
79.24
91.55
94.25
97.00
98.00
77.54
92.95
95.35
97.55
γ=0.25
87.82
98.63
99.45
99.88
99.98
86.63
98.04
99.16
99.84
99.96
80.74
92.40
95.35
97.55
98.30
79.29
92.85
95.40
97.70
γ=0.5 (default)
88.25
98.45
99.43
99.86
99.98
86.24
98.12
99.22
99.84
99.94
80.69
92.40
95.10
97.35
98.15
78.84
91.90
94.65
97.40
γ=0.75
88.20
98.67
99.49
99.82
99.96
85.73
98.00
99.25
99.78
99.98
81.04
92.20
94.70
97.20
98.25
78.59
91.05
93.90
96.90
γ=1.0
87.98
98.47
99.51
99.82
99.96
85.02
97.69
99.02
99.76
99.96
80.34
92.00
94.40
97.10
98.30
78.19
90.80
93.70
96.70
Long-DCI
Urban-1K
γ
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
γ=0
70.92
89.51
93.90
97.27
98.64
68.00
89.12
94.17
97.84
98.93
91.30
99.00
99.30
99.80
99.90
93.00
99.10
99.50
99.60
γ=0.25
76.38
92.49
95.34
97.73
98.68
74.80
91.94
95.51
97.94
98.68
91.70
98.60
99.50
99.70
99.90
93.00
98.90
99.60
99.70
γ=0.5 (default)†
78.79
92.32
95.12
97.19
98.27
76.18
91.99
95.15
97.26
98.40
91.00
98.50
99.50
99.70
99.90
90.50
98.50
99.30
99.60
γ=0.75
79.96
92.45
95.06
97.06
98.19
76.72
91.35
94.43
96.76
97.97
89.10
98.50
99.10
99.60
99.60
89.50
97.90
99.00
99.80
γ=1.0
79.94
92.17
94.79
96.76
97.89
76.59
90.78
93.91
96.49
97.70
87.90
98.00
99.10
99.50
99.60
88.40
97.50
98.70
99.80
Table A7: Cross-domain transfer DOCCI→Long-DCI. Best per column in bold.
DOCCI → Long-DCI
Method
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
Long-CLIP (zero-shot)
54.61
72.80
78.33
85.29
89.15
47.35
73.04
80.10
86.66
90.60
FineLIP†
54.65
73.04
79.46
85.87
89.90
46.70
73.01
80.51
87.10
90.64
GOAL†
55.71
74.93
80.66
86.93
90.66
56.61
75.46
81.19
87.55
91.28
StructXLIP
57.70
76.03
81.68
88.06
91.60
57.90
76.36
82.09
87.71
91.03
HN-CLIP
62.60
79.16
84.19
89.88
93.05
63.77
79.89
84.22
89.05
92.14
Table A8: Seed replication on Long-DCI. max|Δ| is the largest absolute difference across the six metrics.
Method
Seed
R@1
R@5
R@10
R@1
R@5
R@10
max|Δ|
HN-CLIP (ours)
seed 42
79.03
93.76
96.34
76.44
92.79
95.91
seed 43
78.79
92.32
95.12
76.18
91.99
95.15
1.44
GOAL
seed 42
74.29
92.77
95.58
73.31
92.14
95.77
seed 43
74.34
92.89
95.77
72.03
92.37
95.69
1.28
StructXLIP
seed 42
75.34
93.15
95.84
72.30
92.79
95.78
seed 43
75.58
93.12
96.20
72.39
93.03
95.80
0.36
Table A9: Sample efficiency, full resolution. Best per pair in bold; Δ is the average-R@1 margin of HN-CLIP over GOAL at that fraction.
Data
Method
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
Δ
DOCCI
5%
GOAL
77.88
95.35
97.88
99.59
99.84
75.45
94.65
97.76
99.35
99.84
HN-CLIP
82.73
96.94
98.69
99.69
99.84
80.08
96.22
98.35
99.57
99.98
↑ 4.74
20%
GOAL
78.22
95.65
98.22
99.55
99.86
76.47
94.88
97.96
99.39
99.84
HN-CLIP
85.43
97.86
99.20
99.80
99.96
83.75
97.25
98.80
99.76
99.98
↑ 7.25
50%
GOAL
80.20
96.47
98.33
99.63
99.88
78.18
95.37
97.96
99.43
99.92
HN-CLIP
87.16
98.39
99.43
99.84
99.94
85.37
97.69
99.18
99.80
99.94
↑ 7.07
100%
GOAL†
81.96
96.94
98.78
99.78
99.92
80.84
96.33
98.61
99.67
99.92
HN-CLIP
88.25
98.45
99.43
99.86
99.98
86.24
98.12
99.22
99.84
99.94
↑ 5.85
DCI
5%
GOAL
70.94
87.19
90.95
94.95
96.95
69.03
86.89
90.80
94.75
96.85
HN-CLIP
73.59
87.39
92.30
95.10
96.95
74.59
89.19
93.20
96.00
97.15
↑ 4.11
20%
GOAL
73.69
87.99
91.40
95.40
97.10
71.39
87.24
91.55
95.45
97.30
HN-CLIP
78.09
90.60
93.70
96.30
97.65
76.79
90.40
93.55
96.55
97.70
↑ 4.90
50%
GOAL
75.79
89.19
93.00
95.55
96.80
73.79
88.54
92.65
95.90
97.30
HN-CLIP
79.24
91.60
94.55
97.05
97.95
78.44
91.35
94.25
96.70
98.00
↑ 4.05
100%
GOAL
77.29
90.25
93.30
96.10
97.40
74.84
89.94
93.25
96.50
98.15
HN-CLIP
80.69
92.40
95.10
97.35
98.15
78.84
91.90
94.65
97.40
98.30
↑ 3.70
Table A10: Cross-domain generalization between DCI and DOCCI. Train on one dataset and test on another. Values are Recall@K (%), using Text→Image and Image→Text retrieval. In-domain best in italic bold, cross-domain best in bold.
Figure 1: Dense-caption benchmarks are dominated by hard negatives. Distributions of (a) all pairwise caption-caption cosine similarities and (b) each caption’s hardest-negative similarity, measured with the pre-trained Long-CLIP-L text encoder on the test sets. Dotted lines mark the means. The pairs in (b) receive the strongest boosts.
Figure 2: Overview of HN-CLIP. For a batch of B image–text pairs, HN-CLIP encodes images and captions with the dual encoder being fine-tuned, and forms the image–text similarity matrix S=vt⊤ and a detached, diagonal-masked caption-similarity matrix G¯ from G=tt⊤. Their combination yields boosted logits s(S+γG¯), so near-twin negatives receive larger margins. The token-level term of Eq. 3 acts alongside this objective, improving supervision on hard negatives during training while keeping inference unchanged.
Figure 3: Illustration of HN-CLIP. A real DOCCI query and its hardest in-batch negative (cos 0.89). The strongest baseline ranks the ground truth 52nd (its top-1 (pink) is the near-twin’s own image); HN-CLIP ranks it first.
Figure 4: Empirical gradient analysis (Long-CLIP-L, DOCCI, 10 epochs; losses and full-parameter gradients measured every 5 optimizer steps, 148 measurements). Standard InfoNCE declares 80% of batches solved within the first epoch and its gradient is exactly zero in fp32 in 47% of measurements (median zero at five epochs; plotted clamped to 10−10). The boosted loss keeps at least 26% of batches active in every epoch and, where the standard gradient is nonzero, exceeds it by 102–106× in per-epoch median, with interquartile bands disjoint at nine epochs, and the retrieval error it buys keeps falling.
Figure 5: Convergence comparison. Average R@1 (T→I, I→T) per fine-tuning epoch. HN-CLIP’s first epoch matches or exceeds most baselines’ final accuracy on all four benchmarks.
Figure 6: Sample efficiency (avg R@1, identical subsets). HN-CLIP at 20% of the data already clears the strongest baseline trained on 100% (dotted line, from Table 1); full numbers in Table A9.
Figure A1: Near-duplicate caption pairs on DOCCI. Each row: a test image, its caption, the three nearest other captions’ images (cosine printed underneath), and the caption texts with shared content words highlighted.
Figure A2: Near-duplicate caption pairs on DCI.
Figure A3: Near-duplicate caption pairs on Long-DCI (full-length captions).
Figure A4: Near-duplicate caption pairs on Urban-1K.
Figure A5: One real training batch per benchmark. Text–text similarity matrices G¯ (pre-trained encoder; diagonal masked, max off-diagonal entry annotated). Uniformly dark = every negative is hard; this matrix, detached and scaled by γ, is the entire mechanism of HN-CLIP.
Figure A6: Statistical analysis of the caption geometry across the four benchmarks (pre-trained Long-CLIP-L text encoder, full test splits). (A) Mean pairwise caption similarity. (B) Mean similarity of each caption’s hardest companion. (C) Composition of in-batch negatives by hardness bucket.
Figure A8: Gradient dynamics across data scales. Rows: fraction of the training set; columns: dataset. Per cell: (a) raw loss traces (thin) with running medians (thick), (b) gradient norms and their ratio (dashed), (c) gradient cosine. Log axes in (a,b); values clamped at 10−10 (exact zeros at fp32).
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^{-3} on 80% of batches within the first epoch, while its gradient reaches exact zero in fp32 in 47% of measurements. We find that this behavior is closely related to the large number of near-duplicate captions in dense-caption benchmarks, where a few highly similar negatives remain unresolved after the easy majority has already been separated. As a remedy, we introduce HN-CLIP, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins. Specifically, a detached caption-similarity matrix is added to the negative logits, assigning larger margins to more similar captions without mining, synthesizing, or resampling negatives. The resulting objective requires only one caption-similarity matrix and a masked logit addition during training, with no auxiliary data, additional parameters, offline preprocessing, or inference-time overhead. Extensive experiments on four dense-caption retrieval benchmarks show that HN-CLIP improves over the strongest competitors by +2.4--+4.3 R@1 while training 2.4x faster than GOAL and 5.4x faster than StructXLIP. Moreover, the proposed objective improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.
作者 · Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang