이미지-긴문장 검색 AI가 '이미 다 맞혔다'고 착각해서 정작 헷갈리는 문제를 못 배우던 버릇을 고쳤다
arXiv:2608.185212026-08-20
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
이미지-긴문장 검색 AI가 '이미 다 맞혔다'고 착각해서 정작 헷갈리는 문제를 못 배우던 버릇을 고쳤다
이미지와 긴 설명 문장을 짝짓는 AI(CLIP 계열)를 학습시킬 때 쓰는 기존 방식은, 학습 초반부터 손실값이 거의 0으로 떨어져 정작 헷갈리는 어려운 오답들을 더 배우지 못하는 문제가 있었다. 연구팀은 문장들끼리 얼마나 비슷한지를 미리 계산해 헷갈리는 오답에는 더 큰 페널티(마진)를 자동으로 부여하는 HN-CLIP을 제안했다. 별도 데이터나 모델 구조 추가 없이 네 개의 벤치마크에서 기존 최강 방법보다 순위1위 정확도를 2.4~4.3포인트 높였고, 학습 속도도 2.4~5.4배 빨랐다.
METAL MEDIA 해설 도표
이미지-긴문장 검색 AI가 '이미 다 맞혔다'고 착각해서 정작 헷갈리는 문제를 못 배우던 버릇을 고쳤다
01세그멘테이션, 외곽선 지도, LLM 필터링 문장 등 복잡한 장치를 추가해온 기존 연구들과 달리, 이 논문은 학습에 쓰이는 손실 함수(모델을 얼마나 틀렸는지 계산해 학습 신호를 주는 공식) 자체의 결함을 짚었다.
02긴 설명 문장이 많은 데이터셋에서는 서로 거의 똑같은 내용의 문장(니어 듀플리케이트)이 많아, 학습 1에포크(전체 데이터를 한 바퀴 도는 것) 안에 배치의 80%에서 손실이 거의 0이 되고, 47%의 측정에서는 그래디언트(모델을 어느 방향으로 고칠지 알려주는 값)가 정확히 0이 되어버렸다.
03해결책으로 문장 임베딩을 이용해 문장끼리 얼마나 비슷한지 계산한 행렬(텍스트-텍스트 유사도 행렬)을 만들고, 이를 학습에 영향은 주지 않는 상태로 떼어내(디태치) 정답과 오답 점수 차이(로짓)에 더해 비슷한 문장일수록 더 큰 마진(구별을 위해 넘어야 할 점수 차)을 요구하도록 만들었다.
04이 방법은 부가 데이터, 추가 모델 구조, 추론 시점 추가 연산 없이 단 하나의 행렬 곱셈과 덧셈만 추가되며, DOCCI·DCI·Long-DCI·Urban-1K 네 벤치마크의 8개 검색 방향 모두에서 최고 R@1(1위 정답률)을 기록했고, 전체 학습 데이터의 20%만으로도 기존 최강 기법을 100% 데이터로 학습한 것보다 앞섰다.
05기존 6개의 파인튜닝 프레임워크(Long-CLIP, FineLIP, GOAL, StructXLIP, LoRA, DoRA)에 이 손실 함수만 갈아 끼워도 모두 인도메인 성능이 향상되어, 특정 구조에 종속되지 않는 범용적인 개선임을 보였다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
세그멘테이션, 외곽선 지도, LLM 필터링 문장 등 복잡한 장치를 추가해온 기존 연구들과 달리, 이 논문은 학습에 쓰이는 손실 함수(모델을 얼마나 틀렸는지 계산해 학습 신호를 주는 공식) 자체의 결함을 짚었다.
긴 설명 문장이 많은 데이터셋에서는 서로 거의 똑같은 내용의 문장(니어 듀플리케이트)이 많아, 학습 1에포크(전체 데이터를 한 바퀴 도는 것) 안에 배치의 80%에서 손실이 거의 0이 되고, 47%의 측정에서는 그래디언트(모델을 어느 방향으로 고칠지 알려주는 값)가 정확히 0이 되어버렸다.
해결책으로 문장 임베딩을 이용해 문장끼리 얼마나 비슷한지 계산한 행렬(텍스트-텍스트 유사도 행렬)을 만들고, 이를 학습에 영향은 주지 않는 상태로 떼어내(디태치) 정답과 오답 점수 차이(로짓)에 더해 비슷한 문장일수록 더 큰 마진(구별을 위해 넘어야 할 점수 차)을 요구하도록 만들었다.
이 방법은 부가 데이터, 추가 모델 구조, 추론 시점 추가 연산 없이 단 하나의 행렬 곱셈과 덧셈만 추가되며, DOCCI·DCI·Long-DCI·Urban-1K 네 벤치마크의 8개 검색 방향 모두에서 최고 R@1(1위 정답률)을 기록했고, 전체 학습 데이터의 20%만으로도 기존 최강 기법을 100% 데이터로 학습한 것보다 앞섰다.
기존 6개의 파인튜닝 프레임워크(Long-CLIP, FineLIP, GOAL, StructXLIP, LoRA, DoRA)에 이 손실 함수만 갈아 끼워도 모두 인도메인 성능이 향상되어, 특정 구조에 종속되지 않는 범용적인 개선임을 보였다.
Table 1: Cross-modal retrieval performance of CLIP-based fine-tuning methods on four dense-caption benchmarks. We report Recall@K (%) on both Text→Image and Image→Text settings. All fine-tuned methods start from the same Long-CLIP-L backbone with an identical training budget; Long-CLIP denotes the released checkpoint. Best results in bold; second best underlined. Δ denotes the margin over the best competitor per column, with gain in ↑ green.
DOCCI
DCI
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP[ECCV’24]
78.78
95.24
98.02
66.75
91.92
96.31
67.83
83.19
87.69
64.13
84.84
89.74
FineLIP[CVPR’25]
77.51
96.02
98.41
69.90
93.43
97.45
72.69
87.14
90.65
65.48
86.84
91.00
GOAL[CVPR’25]
81.53
97.02
98.80
80.86
96.24
98.63
77.29
90.25
93.30
74.84
89.94
93.25
StructXLIP[CVPR’26]
84.73
97.69
99.00
82.61
97.08
98.71
75.84
89.94
93.65
74.49
90.05
93.40
HN-CLIP
88.25
98.45
99.43
86.24
98.12
99.22
80.69
92.40
95.10
78.84
91.90
94.65
Δ
↑ 3.52
↑ 0.76
↑ 0.43
↑ 3.63
↑ 1.04
↑ 0.51
↑ 3.40
↑ 2.15
↑ 1.45
↑ 4.00
↑ 1.85
↑ 1.25
Table 2: Plug-and-play enhancement of our ℒHN on CLIP-based fine-tuning. Results on DOCCI and Long-DCI for Text→Image and Image→Text retrieval. Upper: full-parameter fine-tuning; lower: parameter-efficient tuning. Our method consistently boosts diverse CLIP variants, and the gains grow with caption length. Best in bold, with gain in ↑ green.
DOCCI
Long-DCI
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP[ECCV’24]
86.45
98.00
99.31
84.10
97.84
99.04
70.24
89.62
94.00
68.44
89.35
94.67
+ our ℒHN
88.31
98.75
99.45
86.39
98.08
99.20
78.43
92.52
95.27
76.02
91.78
94.76
Δ
↑ 1.86
↑ 0.75
↑ 0.14
↑ 2.29
↑ 0.24
↑ 0.16
↑ 8.19
↑ 2.90
↑ 1.27
↑ 7.58
↑ 2.43
↑ 0.09
FineLIP[CVPR’25]
77.51
96.02
98.41
69.90
93.43
97.45
59.24
77.86
83.19
49.52
75.08
82.39
+ our ℒHN
85.47
97.71
99.04
83.88
97.41
98.82
74.92
90.52
93.62
73.39
89.87
93.23
Δ
↑ 7.96
↑ 1.69
↑ 0.63
↑ 13.98
↑ 3.98
↑ 1.37
↑ 15.68
↑ 12.66
↑ 10.43
↑ 23.87
↑ 14.79
↑ 10.84
GOAL[CVPR’25]
81.53
97.02
98.80
80.86
96.24
98.63
74.29
92.77
95.58
73.31
92.14
95.77
+ our ℒHN
86.24
98.14
99.35
84.86
97.69
99.20
84.27
94.99
96.57
82.28
94.26
96.18
Δ
↑ 4.71
↑ 1.12
↑ 0.55
↑ 4.00
↑ 1.45
↑ 0.57
↑ 9.98
↑ 2.22
↑ 0.99
↑ 8.97
↑ 2.12
↑ 0.41
StructXLIP[CVPR’26]
84.73
97.69
99.00
82.61
97.08
98.71
75.34
93.15
95.84
72.30
92.79
95.78
+ our ℒHN
85.45
97.92
99.12
83.31
97.45
98.92
82.04
94.18
96.23
78.98
93.22
95.65
Δ
↑ 0.72
↑ 0.23
↑ 0.12
↑ 0.70
↑ 0.37
↑ 0.21
↑ 6.70
↑ 1.03
↑ 0.39
↑ 6.68
↑ 0.43
↓ 0.13
LoRA[ICLR’22]
80.49
96.16
98.55
77.96
95.80
98.04
60.08
80.68
86.86
58.97
80.02
86.53
+ our ℒHN
82.82
97.14
98.86
80.84
96.49
98.63
64.87
83.09
88.35
62.79
81.77
87.31
Δ
↑ 2.33
↑ 0.98
↑ 0.31
↑ 2.88
↑ 0.69
↑ 0.59
↑ 4.79
↑ 2.41
↑ 1.49
↑ 3.82
↑ 1.75
↑ 0.78
DoRA[ICML’24]
80.76
96.25
98.61
78.20
95.90
98.10
60.76
81.19
87.48
59.65
80.31
87.13
+ our ℒHN
83.41
97.33
99.00
81.41
96.65
98.67
65.46
83.57
88.61
63.38
82.15
87.56
Δ
↑ 2.65
↑ 1.08
↑ 0.39
↑ 3.21
↑ 0.75
↑ 0.57
↑ 4.70
↑ 2.38
↑ 1.13
↑ 3.73
↑ 1.84
↑ 0.43
Table 3: Ablations. (a) Sensitivity to the boost strength γ (Recall@1): every γ∈[0.25,1] beats γ=0 in-domain, while under transfer (Urban-1K) milder boosts generalize better. (b) Loss components under an identical recipe; only the loss changes. (c) G¯ recomputed from the live encoder (default) vs. frozen at initialization; R@1/5/10 in Table A2. Best per column in bold.
DOCCI
DCI
Long-DCI
Urban-1K
γ
R@1
R@1
R@1
R@1
R@1
R@1
R@1
R@1
Avg
γ=0
86.45
84.29
79.24
77.54
70.92
68.00
91.30
93.00
81.34
γ=0.25
87.82
86.63
80.74
79.29
76.38
74.80
91.70
93.00
83.80
γ=0.5 (default)
88.25
86.24
80.69
78.84
79.03
76.44
91.10
90.30
83.86
γ=0.75
88.20
85.73
81.04
78.59
79.96
76.72
89.10
89.50
83.61
γ=1.0
87.98
85.02
80.34
78.19
79.94
76.59
87.90
88.40
83.04
Table A1: Benchmark statistics. Mean caption length is measured on the evaluation split in words; similarities use the pre-trained Long-CLIP-L text encoder.
Benchmark
#train
#test
words
pairwise sim
hardest sim
DOCCI
9450
5100
123
0.84
0.93
DCI
5445
1999
133
0.85
0.92
Long-DCI
5445 (=DCI)
7444
134
0.85
0.93
Urban-1K
14579 (VG)
1000
107
0.88
0.94
Table A2: Recomputed vs. frozen G¯ (Recall@K, Text→Image and Image→Text). Building the boost from a frozen copy of the pre-trained text encoder makes the margins exactly stationary but removes their implicit annealing: accuracy peaks at epoch 1 on DCI and Urban-1K and then declines, and the default wins 22 of 24 columns. Best per column in bold. Upper band: DOCCI and DCI; lower band: Long-DCI and the transfer setting Urban-1K.
DOCCI
DCI
G¯ source
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Recomputed G¯ (default)
88.25
98.45
99.43
86.24
98.12
99.22
80.69
92.40
95.10
78.84
91.90
94.65
Frozen G¯
85.86
97.61
99.18
82.43
96.82
98.73
77.79
90.80
93.35
75.84
89.49
92.95
Δ (default − frozen)
↑ 2.39
↑ 0.84
↑ 0.25
↑ 3.81
↑ 1.30
↑ 0.49
↑ 2.90
↑ 1.60
↑ 1.75
↑ 3.00
↑ 2.41
↑ 1.70
best epoch (default / frozen)
6 / 5
5 / 1
Long-DCI
Urban-1K
G¯ source
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Recomputed G¯ (default)
79.03
93.76
96.34
76.44
92.79
95.91
91.10
98.10
99.40
90.30
98.20
99.30
Frozen G¯
79.93
90.84
93.82
76.75
89.54
92.62
87.30
97.70
98.90
83.50
95.90
97.90
Δ (default − frozen)
↓ 0.90
↑ 2.92
↑ 2.52
↓ 0.31
↑ 3.25
↑ 3.29
↑ 3.80
↑ 0.40
↑ 0.50
↑ 6.80
↑ 2.30
↑ 1.40
best epoch (default / frozen)
10 / 10
4 / 1
Table A3: Training-efficiency comparison on DCI fine-tuning (10 epochs, 5.4k images, single Ascend 910B, effective batch 128). Wall-clock for GOAL/StructXLIP is measured from same-device sequential runs minus evaluation overhead; FineLIP ran in parallel and is not attributable. Offline preprocessing time (segmentation, edge extraction, LLM filtering) is not included in the wall-clock.
Method
Aux. training inputs
Offline prep.
Wall-clock
Throughput
Long-CLIP (plain FT)
none
none
≈16 min
55.5 img/s
FineLIP
none
none
—
—
GOAL
SAM segments
required
≈41 min
≈22 img/s
StructXLIP
edges + LLM lexicon
required
≈92 min
≈10 img/s
HN-CLIP (ours)
none
none
17 min
53 img/s
Table A4: Plug-and-play enhancement of ℒHN on CLIP-based fine-tuning, all four benchmarks. Results on DOCCI, DCI, Long-DCI, and Urban-1K for Text→Image and Image→Text retrieval; per framework we report the official baseline, the same recipe with ℒHN replacing its global InfoNCE term, and the per-column difference Δ (gain in ↑ green, drop in ↓ gray; differences within ±0.15 are shown as ≈0). Bold marks the +ℒHN value where it is the better of the pair. Upper band: DOCCI and DCI; lower band: Long-DCI and the transfer setting Urban-1K.
DOCCI
DCI
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP
86.45
98.00
99.31
84.10
97.84
99.04
79.39
91.60
94.65
78.34
92.85
95.25
+our ℒHN
88.31
98.75
99.45
86.39
98.08
99.20
81.04
92.20
95.00
78.54
91.95
94.60
Δ
↑ 1.86
↑ 0.75
≈0
↑ 2.29
↑ 0.24
↑ 0.16
↑ 1.65
↑ 0.60
↑ 0.35
↑ 0.20
↓ 0.90
↓ 0.65
FineLIP†
77.12
95.94
98.29
70.16
93.14
97.29
72.69
87.14
90.65
65.48
86.84
91.00
+our ℒHN
85.47
97.71
99.04
83.88
97.41
98.82
80.74
91.90
94.55
79.99
91.80
94.35
Δ
↑ 8.35
↑ 1.77
↑ 0.75
↑ 13.72
↑ 4.27
↑ 1.53
↑ 8.05
↑ 4.76
↑ 3.90
↑ 14.51
↑ 4.96
↑ 3.35
GOAL†
81.96
96.94
98.78
80.84
96.33
98.61
77.29
90.25
93.30
74.84
89.94
93.25
+our ℒHN
86.24
98.14
99.35
84.86
97.69
99.20
79.59
90.80
93.30
77.09
90.10
93.00
Δ
↑ 4.28
↑ 1.20
↑ 0.57
↑ 4.02
↑ 1.36
↑ 0.59
↑ 2.30
↑ 0.55
≈0
↑ 2.25
↑ 0.16
↓ 0.25
StructXLIP
84.73
97.69
99.00
82.61
97.08
98.71
75.84
89.94
93.65
74.49
90.05
93.40
+our ℒHN
85.45
97.92
99.12
83.31
97.45
98.92
77.34
90.75
93.10
74.14
89.89
92.30
Δ
↑ 0.72
↑ 0.23
≈0
↑ 0.70
↑ 0.37
↑ 0.21
↑ 1.50
↑ 0.81
↓ 0.55
↓ 0.35
↓ 0.16
↓ 1.10
LoRA
80.49
96.16
98.55
77.96
95.80
98.04
73.84
89.44
92.85
71.89
88.74
92.85
+our ℒHN
82.82
97.14
98.86
80.84
96.49
98.63
75.84
90.20
93.00
73.44
88.54
92.25
Δ
↑ 2.33
↑ 0.98
↑ 0.31
↑ 2.88
↑ 0.69
↑ 0.59
↑ 2.00
↑ 0.76
↑ 0.15
↑ 1.55
↓ 0.20
↓ 0.60
DoRA
80.76
96.25
98.61
78.20
95.90
98.10
74.14
89.74
93.00
72.44
89.24
92.95
+our ℒHN
83.41
97.33
99.00
81.41
96.65
98.67
76.24
90.45
93.15
73.89
88.59
92.30
Δ
↑ 2.65
↑ 1.08
↑ 0.39
↑ 3.21
↑ 0.75
↑ 0.57
↑ 2.10
↑ 0.71
↑ 0.15
↑ 1.45
↓ 0.65
↓ 0.65
Long-DCI
Urban-1K
Method
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Long-CLIP
70.24
89.62
94.00
68.44
89.35
94.67
91.70
99.10
99.50
93.20
99.00
99.50
+our ℒHN
78.43
92.52
95.27
76.02
91.78
94.76
90.60
98.30
99.40
90.20
98.30
99.40
Δ
↑ 8.19
↑ 2.90
↑ 1.27
↑ 7.58
↑ 2.43
≈0
↓ 1.10
↓ 0.80
≈0
↓ 3.00
↓ 0.70
≈0
FineLIP†
59.24
77.86
83.19
49.52
75.08
82.39
81.10
94.50
97.40
77.50
94.80
97.50
+our ℒHN
74.92
90.52
93.62
73.39
89.87
93.23
91.80
98.10
99.30
93.20
98.80
99.30
Δ
↑ 15.68
↑ 12.66
↑ 10.43
↑ 23.87
↑ 14.79
↑ 10.84
↑ 10.70
↑ 3.60
↑ 1.90
↑ 15.70
↑ 4.00
↑ 1.80
GOAL†
74.34
92.89
95.77
72.03
92.37
95.69
86.30
97.00
98.90
86.10
97.40
98.90
+our ℒHN
84.27
94.99
96.57
82.28
94.26
96.18
85.70
96.70
98.20
86.80
96.60
98.00
Δ
↑ 9.93
↑ 2.10
↑ 0.80
↑ 10.25
↑ 1.89
↑ 0.49
↓ 0.60
↓ 0.30
↓ 0.70
↑ 0.70
↓ 0.80
↓ 0.90
Table A5: Loss ablation, full resolution (all four benchmarks, all ranks). Best per column in bold.
DOCCI
DCI
ℒtok
ℒHN
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
✗
✗
86.45
98.00
99.31
99.86
99.98
84.10
97.84
99.04
99.78
99.98
79.39
91.60
94.65
97.25
97.95
78.34
92.85
95.25
✓
✗
86.45
97.98
99.43
99.86
99.98
84.29
97.90
98.94
99.78
99.94
79.24
91.55
94.25
97.00
98.00
77.54
92.95
95.35
✗
✓
88.31
98.75
99.45
99.84
99.98
86.39
98.08
99.20
99.82
99.96
81.04
92.20
95.00
97.40
98.10
78.54
91.95
94.60
✓
✓
88.25
98.45
99.43
99.86
99.98
86.24
98.12
99.22
99.84
99.94
80.69
92.40
95.10
97.35
98.15
78.84
91.90
94.65
Long-DCI
Urban-1K
ℒtok
ℒHN
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
✗
✗
70.24
89.62
94.00
97.66
98.79
68.44
89.35
94.67
98.11
98.95
91.70
99.10
99.50
99.80
99.90
93.20
99.00
99.50
✓
✗
70.92
89.51
93.90
97.27
98.64
68.00
89.12
94.17
97.84
98.93
91.30
99.00
99.30
99.80
99.90
93.00
99.10
99.50
✗
✓
78.43
92.52
95.27
97.21
98.47
76.02
91.78
94.76
97.31
98.29
90.60
98.30
99.40
99.60
99.80
90.20
98.30
99.40
✓
✓†
78.79
92.32
95.12
97.19
98.27
76.18
91.99
95.15
97.26
98.40
91.00
98.50
99.50
99.70
99.90
90.50
98.50
99.30
Table A6: γ sweep, full resolution (all four benchmarks, all ranks). Best per column in bold.
DOCCI
DCI
γ
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
γ=0
86.45
97.98
99.43
99.86
99.98
84.29
97.90
98.94
99.78
99.94
79.24
91.55
94.25
97.00
98.00
77.54
92.95
95.35
97.55
γ=0.25
87.82
98.63
99.45
99.88
99.98
86.63
98.04
99.16
99.84
99.96
80.74
92.40
95.35
97.55
98.30
79.29
92.85
95.40
97.70
γ=0.5 (default)
88.25
98.45
99.43
99.86
99.98
86.24
98.12
99.22
99.84
99.94
80.69
92.40
95.10
97.35
98.15
78.84
91.90
94.65
97.40
γ=0.75
88.20
98.67
99.49
99.82
99.96
85.73
98.00
99.25
99.78
99.98
81.04
92.20
94.70
97.20
98.25
78.59
91.05
93.90
96.90
γ=1.0
87.98
98.47
99.51
99.82
99.96
85.02
97.69
99.02
99.76
99.96
80.34
92.00
94.40
97.10
98.30
78.19
90.80
93.70
96.70
Long-DCI
Urban-1K
γ
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
γ=0
70.92
89.51
93.90
97.27
98.64
68.00
89.12
94.17
97.84
98.93
91.30
99.00
99.30
99.80
99.90
93.00
99.10
99.50
99.60
γ=0.25
76.38
92.49
95.34
97.73
98.68
74.80
91.94
95.51
97.94
98.68
91.70
98.60
99.50
99.70
99.90
93.00
98.90
99.60
99.70
γ=0.5 (default)†
78.79
92.32
95.12
97.19
98.27
76.18
91.99
95.15
97.26
98.40
91.00
98.50
99.50
99.70
99.90
90.50
98.50
99.30
99.60
γ=0.75
79.96
92.45
95.06
97.06
98.19
76.72
91.35
94.43
96.76
97.97
89.10
98.50
99.10
99.60
99.60
89.50
97.90
99.00
99.80
γ=1.0
79.94
92.17
94.79
96.76
97.89
76.59
90.78
93.91
96.49
97.70
87.90
98.00
99.10
99.50
99.60
88.40
97.50
98.70
99.80
Table A7: Cross-domain transfer DOCCI→Long-DCI. Best per column in bold.
DOCCI → Long-DCI
Method
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
Long-CLIP (zero-shot)
54.61
72.80
78.33
85.29
89.15
47.35
73.04
80.10
86.66
90.60
FineLIP†
54.65
73.04
79.46
85.87
89.90
46.70
73.01
80.51
87.10
90.64
GOAL†
55.71
74.93
80.66
86.93
90.66
56.61
75.46
81.19
87.55
91.28
StructXLIP
57.70
76.03
81.68
88.06
91.60
57.90
76.36
82.09
87.71
91.03
HN-CLIP
62.60
79.16
84.19
89.88
93.05
63.77
79.89
84.22
89.05
92.14
Table A8: Seed replication on Long-DCI. max|Δ| is the largest absolute difference across the six metrics.
Method
Seed
R@1
R@5
R@10
R@1
R@5
R@10
max|Δ|
HN-CLIP (ours)
seed 42
79.03
93.76
96.34
76.44
92.79
95.91
seed 43
78.79
92.32
95.12
76.18
91.99
95.15
1.44
GOAL
seed 42
74.29
92.77
95.58
73.31
92.14
95.77
seed 43
74.34
92.89
95.77
72.03
92.37
95.69
1.28
StructXLIP
seed 42
75.34
93.15
95.84
72.30
92.79
95.78
seed 43
75.58
93.12
96.20
72.39
93.03
95.80
0.36
Table A9: Sample efficiency, full resolution. Best per pair in bold; Δ is the average-R@1 margin of HN-CLIP over GOAL at that fraction.
Data
Method
R@1
R@5
R@10
R@25
R@50
R@1
R@5
R@10
R@25
R@50
Δ
DOCCI
5%
GOAL
77.88
95.35
97.88
99.59
99.84
75.45
94.65
97.76
99.35
99.84
HN-CLIP
82.73
96.94
98.69
99.69
99.84
80.08
96.22
98.35
99.57
99.98
↑ 4.74
20%
GOAL
78.22
95.65
98.22
99.55
99.86
76.47
94.88
97.96
99.39
99.84
HN-CLIP
85.43
97.86
99.20
99.80
99.96
83.75
97.25
98.80
99.76
99.98
↑ 7.25
50%
GOAL
80.20
96.47
98.33
99.63
99.88
78.18
95.37
97.96
99.43
99.92
HN-CLIP
87.16
98.39
99.43
99.84
99.94
85.37
97.69
99.18
99.80
99.94
↑ 7.07
100%
GOAL†
81.96
96.94
98.78
99.78
99.92
80.84
96.33
98.61
99.67
99.92
HN-CLIP
88.25
98.45
99.43
99.86
99.98
86.24
98.12
99.22
99.84
99.94
↑ 5.85
DCI
5%
GOAL
70.94
87.19
90.95
94.95
96.95
69.03
86.89
90.80
94.75
96.85
HN-CLIP
73.59
87.39
92.30
95.10
96.95
74.59
89.19
93.20
96.00
97.15
↑ 4.11
20%
GOAL
73.69
87.99
91.40
95.40
97.10
71.39
87.24
91.55
95.45
97.30
HN-CLIP
78.09
90.60
93.70
96.30
97.65
76.79
90.40
93.55
96.55
97.70
↑ 4.90
50%
GOAL
75.79
89.19
93.00
95.55
96.80
73.79
88.54
92.65
95.90
97.30
HN-CLIP
79.24
91.60
94.55
97.05
97.95
78.44
91.35
94.25
96.70
98.00
↑ 4.05
100%
GOAL
77.29
90.25
93.30
96.10
97.40
74.84
89.94
93.25
96.50
98.15
HN-CLIP
80.69
92.40
95.10
97.35
98.15
78.84
91.90
94.65
97.40
98.30
↑ 3.70
Table A10: Cross-domain generalization between DCI and DOCCI. Train on one dataset and test on another. Values are Recall@K (%), using Text→Image and Image→Text retrieval. In-domain best in italic bold, cross-domain best in bold.
Setting
R@1
R@5
R@10
R@1
R@5
R@10
Train on DCI → Test on DCI vs. DOCCI
Long-CLIP (DCI→DCI)
67.83
83.19
87.69
64.13
84.84
89.74
Long-CLIP (DCI→DOCCI)
78.78
95.24
98.02
66.75
91.92
96.31
FineLIP (DCI→DCI)
72.69
87.14
90.65
65.48
86.84
91.00
FineLIP (DCI→DOCCI)
80.69
96.47
98.53
65.63
91.51
96.31
GOAL (DCI→DCI)
77.29
90.25
93.30
74.84
89.94
93.25
GOAL (DCI→DOCCI)
80.14
96.25
98.39
77.24
95.22
97.90
StructXLIP (DCI→DCI)
75.84
89.94
93.65
74.49
90.05
93.40
StructXLIP (DCI→DOCCI)
76.94
95.06
97.76
74.24
94.12
97.51
HN-CLIP (DCI→DCI)
80.69
92.40
95.10
78.84
91.90
94.65
HN-CLIP (DCI→DOCCI)
85.00
97.90
99.25
83.14
96.98
98.75
Train on DOCCI → Test on DOCCI vs. DCI
Long-CLIP (DOCCI→DOCCI)
78.78
95.24
98.02
66.75
91.92
96.31
Long-CLIP (DOCCI→DCI)
67.83
83.19
87.69
64.13
84.84
89.74
FineLIP (DOCCI→DOCCI)
77.51
96.02
98.41
69.90
93.43
97.45
FineLIP (DOCCI→DCI)
68.23
83.84
88.74
63.48
85.79
89.84
GOAL (DOCCI→DOCCI)
81.53
97.02
98.80
80.86
96.24
98.63
GOAL (DOCCI→DCI)
69.58
85.24
89.39
69.88
85.34
89.79
StructXLIP (DOCCI→DOCCI)
84.73
97.69
99.00
82.61
97.08
98.71
StructXLIP (DOCCI→DCI)
69.93
86.74
91.05
71.39
86.99
90.60
HN-CLIP (DOCCI→DOCCI)
88.25
98.45
99.43
86.24
98.12
99.22
HN-CLIP (DOCCI→DCI)
74.84
88.54
92.40
75.64
88.04
91.35
왜 중요한가
복잡한 전처리나 추가 모델 구조 없이 손실 함수만 바꿔도 이미지-텍스트 검색 성능을 크게 끌어올릴 수 있다는 점은, 실무에서 검색 시스템을 개선할 때 비용과 속도 면에서 큰 이점이 된다. 또한 '학습이 잘 되고 있는 것처럼 손실값이 낮아 보여도 실제로는 중요한 부분을 못 배우고 있을 수 있다'는 진단은 다른 대조학습 기반 AI 개발에도 참고할 만한 통찰이다.
니어 듀플리케이트(near-duplicate) 캡션 · 내용이 거의 비슷해서 구별하기 어려운 설명 문장들
그래디언트(gradient) · 모델 파라미터를 어느 방향으로 얼마나 수정할지 알려주는 값
디태치(detach)/스톱그래디언트 · 특정 계산 결과를 학습(역전파) 대상에서 제외해 값만 참고하고 그 자체는 업데이트되지 않게 하는 기법
R@1(Recall@1) · 검색했을 때 정답이 1위로 나온 비율
본문에 싣지 못한 그림
Figure 1: Dense-caption benchmarks are dominated by hard negatives. Distributions of (a) all pairwise caption-caption cosine similarities and (b) each caption’s hardest-negative similarity, measured with the pre-trained Long-CLIP-L text encoder on the test sets. Dotted lines mark the means. The pairs in (b) receive the strongest boosts.
Figure 2: Overview of HN-CLIP. For a batch of B image–text pairs, HN-CLIP encodes images and captions with the dual encoder being fine-tuned, and forms the image–text similarity matrix S=vt⊤ and a detached, diagonal-masked caption-similarity matrix G¯ from G=tt⊤. Their combination yields boosted logits s(S+γG¯), so near-twin negatives receive larger margins. The token-level term of Eq. 3 acts alongside this objective, improving supervision on hard negatives during training while keeping inference unchanged.
Figure 3: Illustration of HN-CLIP. A real DOCCI query and its hardest in-batch negative (cos 0.89). The strongest baseline ranks the ground truth 52nd (its top-1 (pink) is the near-twin’s own image); HN-CLIP ranks it first.
Figure 4: Empirical gradient analysis (Long-CLIP-L, DOCCI, 10 epochs; losses and full-parameter gradients measured every 5 optimizer steps, 148 measurements). Standard InfoNCE declares 80% of batches solved within the first epoch and its gradient is exactly zero in fp32 in 47% of measurements (median zero at five epochs; plotted clamped to 10−10). The boosted loss keeps at least 26% of batches active in every epoch and, where the standard gradient is nonzero, exceeds it by 102–106× in per-epoch median, with interquartile bands disjoint at nine epochs, and the retrieval error it buys keeps falling.
Figure 5: Convergence comparison. Average R@1 (T→I, I→T) per fine-tuning epoch. HN-CLIP’s first epoch matches or exceeds most baselines’ final accuracy on all four benchmarks.
Figure 6: Sample efficiency (avg R@1, identical subsets). HN-CLIP at 20% of the data already clears the strongest baseline trained on 100% (dotted line, from Table 1); full numbers in Table A9.
Figure A1: Near-duplicate caption pairs on DOCCI. Each row: a test image, its caption, the three nearest other captions’ images (cosine printed underneath), and the caption texts with shared content words highlighted.
Figure A2: Near-duplicate caption pairs on DCI.
Figure A3: Near-duplicate caption pairs on Long-DCI (full-length captions).
Figure A4: Near-duplicate caption pairs on Urban-1K.
Figure A5: One real training batch per benchmark. Text–text similarity matrices G¯ (pre-trained encoder; diagonal masked, max off-diagonal entry annotated). Uniformly dark = every negative is hard; this matrix, detached and scaled by γ, is the entire mechanism of HN-CLIP.
Figure A6: Statistical analysis of the caption geometry across the four benchmarks (pre-trained Long-CLIP-L text encoder, full test splits). (A) Mean pairwise caption similarity. (B) Mean similarity of each caption’s hardest companion. (C) Composition of in-batch negatives by hardness bucket.
Figure A8: Gradient dynamics across data scales. Rows: fraction of the training set; columns: dataset. Per cell: (a) raw loss traces (thin) with running medians (thick), (b) gradient norms and their ratio (dashed), (c) gradient cosine. Log axes in (a,b); values clamped at 10−10 (exact zeros at fp32).