이미지·영상·문서를 아우르는 검색 AI가 후보를 비교하며 '왜 맞는지' 스스로 따져보게 만들었다
arXiv:2608.185042026-08-20
UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
이미지·영상·문서를 아우르는 검색 AI가 후보를 비교하며 '왜 맞는지' 스스로 따져보게 만들었다
멀티모달 검색 AI는 보통 질문과 후보를 각각 따로 살펴본 뒤 비슷한 정도만 계산해서 순위를 매긴다. UMER은 질문과 후보 쌍을 나란히 놓고 어디가 맞고 어디가 다른지 비교 추론을 시킨 뒤, 빠른 임베딩 검색과 정밀한 순위 판정을 한 모델 안에서 함께 학습시켜 서로의 약점을 보완하게 했다. 그 결과 MMEB-V2라는 78개 과제 벤치마크에서 기존 방법들보다 높은 점수를 내면서도, 필요에 따라 속도와 정확도를 조절할 수 있었다.
METAL MEDIA 해설 도표
이미지·영상·문서를 아우르는 검색 AI가 후보를 비교하며 '왜 맞는지' 스스로 따져보게 만들었다
01문제의식: 기존 방식은 질문과 후보 이미지를 각각 따로 설명하는 추론(아이템별 CoT)만 하기 때문에, 비슷하게 생긴 오답(하드 네거티브)과 정답을 구별할 근거를 만들지 못한다는 한계를 지적했다.
02해결책: 정답 후보와 헷갈리는 오답 후보를 한 쌍으로 묶어 모델에게 같이 보여주고, 무엇이 일치하고 무엇이 어긋나는지 비교하며 설명하는 '쌍 기반 추론(Pair-Aware Discriminative Reasoning)'을 학습시켰다.
03구조: 하나의 멀티모달 대형언어모델(MLLM) 안에서 빠른 벡터 검색용 임베딩 학습과, 쌍을 직접 비교해 관련도를 점수로 매기는 순위 학습을 동시에 진행했고, 서로의 판단이 믿을 만할 때만 지식을 주고받는 '상호 증류(CMD)'로 두 기능을 맞물려 개선했다.
04데이터: Qwen3.5-9B로 정답·오답 쌍마다 '질문 의도'와 '후보 관찰 내용'을 담은 추론 문장을 만들고, 작은 검증 모델(Qwen3.5-0.8B)이 그 설명만 보고 정답/오답을 맞힐 수 있는지 걸러내 품질을 확보했다.
05성능: 78개 과제로 구성된 MMEB-V2 벤치마크에서 전체 평균 65.5점으로 최고 수준을 기록했고, 순수 임베딩 검색만 쓸 때도 기존 추론 기반 방법(UME-R1) 대비 최대 118.6배 빠른 속도를 보였다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
문제의식: 기존 방식은 질문과 후보 이미지를 각각 따로 설명하는 추론(아이템별 CoT)만 하기 때문에, 비슷하게 생긴 오답(하드 네거티브)과 정답을 구별할 근거를 만들지 못한다는 한계를 지적했다.
해결책: 정답 후보와 헷갈리는 오답 후보를 한 쌍으로 묶어 모델에게 같이 보여주고, 무엇이 일치하고 무엇이 어긋나는지 비교하며 설명하는 '쌍 기반 추론(Pair-Aware Discriminative Reasoning)'을 학습시켰다.
구조: 하나의 멀티모달 대형언어모델(MLLM) 안에서 빠른 벡터 검색용 임베딩 학습과, 쌍을 직접 비교해 관련도를 점수로 매기는 순위 학습을 동시에 진행했고, 서로의 판단이 믿을 만할 때만 지식을 주고받는 '상호 증류(CMD)'로 두 기능을 맞물려 개선했다.
데이터: Qwen3.5-9B로 정답·오답 쌍마다 '질문 의도'와 '후보 관찰 내용'을 담은 추론 문장을 만들고, 작은 검증 모델(Qwen3.5-0.8B)이 그 설명만 보고 정답/오답을 맞힐 수 있는지 걸러내 품질을 확보했다.
성능: 78개 과제로 구성된 MMEB-V2 벤치마크에서 전체 평균 65.5점으로 최고 수준을 기록했고, 순수 임베딩 검색만 쓸 때도 기존 추론 기반 방법(UME-R1) 대비 최대 118.6배 빠른 속도를 보였다.
Table 1: Main results on MMEB-V2. The best and second-best scores in each column are in bold and underlined, respectively.
Model
Image
Video
VisDoc
All
CLS
QA
RET
GD
Overall
CLS
QA
RET
MRET
Overall
VDRv1
VDRv2
VR
OOD
Overall
# of Datasets
10
10
12
4
36
5
5
5
3
18
10
4
6
4
24
78
Baseline Models
GME
54.4
29.9
66.9
55.5
51.9
34.9
42.0
25.6
32.4
33.9
86.1
54.0
82.5
43.1
72.7
54.1
VLM2Vec
58.7
49.3
65.0
72.9
59.7
33.4
30.5
20.6
33.0
29.0
49.8
13.5
51.8
33.5
41.6
47.0
VLM2Vec-V2
62.9
56.3
69.5
77.3
64.9
39.3
34.3
28.8
38.5
34.9
75.5
44.9
79.4
39.4
65.4
58.0
DUME
59.3
55.0
66.3
78.0
62.5
37.7
46.6
17.1
30.0
33.2
67.6
43.3
47.1
33.8
52.8
52.7
BToks
64.3
59.8
68.8
77.4
66.0
43.7
47.0
33.0
33.6
39.9
71.1
38.6
81.3
38.1
62.7
59.0
UME-R1
64.8
62.8
67.6
77.2
66.6
44.3
51.2
32.9
39.7
42.2
72.4
46.2
79.2
37.2
63.9
60.1
PLUME
66.5
59.2
67.6
79.7
66.3
45.0
52.3
33.5
46.7
44.1
72.1
49.8
78.1
57.4
67.5
61.6
RIME
67.9
64.4
69.8
82.1
69.1
48.0
52.1
33.6
39.2
43.7
76.4
51.4
81.7
63.9
71.4
64.1
Ours
UMER-E
64.4
64.9
70.2
78.4
68.0
43.0
50.6
34.7
32.8
41.1
76.1
50.1
82.8
68.3
72.2
63.1
UMER-R
64.9
68.4
69.7
74.1
68.5
51.4
51.6
35.3
33.3
44.0
75.1
48.3
84.6
67.9
71.8
63.9
UMER-H
66.1
68.9
71.7
80.8
70.4
50.6
52.8
37.7
35.0
45.0
78.0
50.5
85.0
68.8
73.6
65.5
Table 2: Controlled ablations of pair-aware CoT, ranking supervision and complementary mutual distillation (CMD). The best and second-best scores in each column are in bold and underlined, respectively.
Configuration
Embedding mode (UMER-E)
Ranking mode (UMER-R)
Hybrid mode (UMER-H)
Image
Video
VisDoc
All
Image
Video
VisDoc
All
Image
Video
VisDoc
All
Embedding only
66.5
37.8
69.2
60.7
–
–
–
–
–
–
–
–
+ pair-aware CoT, w/o ranking losses
68.2
41.1
71.2
62.9
44.4
31.8
57.1
45.4
68.6
42.2
71.3
63.3
+ ranking losses, w/o CoT
68.0
40.2
69.3
62.0
68.6
43.1
69.4
63.0
70.8
45.1
71.2
65.0
+ pair-aware CoT and ranking losses
67.9
39.6
71.1
62.4
67.2
44.2
72.1
63.4
70.3
44.9
73.4
65.4
Pair-aware model, w/o CMD
67.9
39.6
71.1
62.4
67.2
44.2
72.1
63.4
70.3
44.9
73.4
65.4
+ ranking → embedding only
67.9
40.4
71.5
62.7
68.5
43.3
70.9
63.4
70.6
44.7
72.7
65.3
+ embedding → ranking only
67.8
40.4
70.1
62.2
69.2
43.9
71.1
63.9
70.8
45.1
72.6
65.5
+ always-on bidirectional CMD
68.0
40.2
71.1
62.6
68.5
43.6
71.0
63.6
70.7
45.3
73.0
65.5
+ selective bidirectional CMD (full)
68.0
41.1
72.2
63.1
68.5
44.0
71.8
63.9
70.4
45.0
73.6
65.5
Table 3: Accuracy–efficiency on MMEB-V2 using one NVIDIA A100 GPU (batch size=1). UMER-H reuses the UMER-E index; 0. denotes no extra indexing cost.
Model
K
Score ↑
Reasoning Tokens
Query Latency
Indexing Time
(/query)
(s/query) ↓
(s/candidate) ↓
UME-R1
–
60.1
352
9.963
11.755
PLUME
–
61.6
8
0.329
0.366
UMER-E
–
63.1
0
0.084
0.118
UMER-H
3
64.8
410
13.566
0.
5
65.5
665
19.880
10
66.0
1305
42.764
20
66.1
2508
81.726
Table 1: Implementation and evaluation settings for UMER.
Setting
Value
Model and optimization
Initialization
Qwen2-VL-2B-Instruct.
Embedding tokens
Four learnable tokens per input (M=4).
Embedding temperature
0.02.
Objective weights
λcot=0.2, λbce=0.1, λmargin=0.2, and λcmd=0.2.
Optimizer and schedule
AdamW; per-device batch size 64; one accumulation step; linear schedule with peak rate 5×10−5, 100 warmup steps and 10,000 maximum steps.
Adaptation
LoRA rank 16, scaling 64 and dropout 0.1 on attention/MLP projections and the language-model head; visual encoder frozen.
Precision and preprocessing
BF16 with FlashAttention-2; maximum image pixels 2,359,296.
Inference
Candidate selection
Retrieve with the embedding branch and rerank the top K=5 candidates.
Ranking generation
At most 256 newly generated tokens per candidate.
Hybrid fusion
Separately z-score normalize embedding and ranking scores over the five candidates.
Ranking-score weight
α=2.0.
Evaluation and reporting
Benchmark
MMEB-V2: 78 datasets across image, video and visual-document modalities.
Benchmark versions
Corrected ViDoSeek-page and MMLongBench-page versions.
Metrics
Hit@1 for image and video tasks; NDCG@5 for visual-document tasks.
Aggregation
Unweighted macro averages over datasets, including all 78 tasks for All.
Table 4: Ranking-score-weight sensitivity of UMER-H at K=5, with candidates, decoding configuration, and normalization fixed.
Ranking-score weight
0.5
0.75
1
1.5
2
3
5
UMER-H
64.7
65.0
65.2
65.5
65.5
65.4
65.3
Table 5: Decision overlap of embedding retrieval (E) and pair-aware ranking (R) over 83,530 MMEB-V2 queries. Both, E only, R only, and Neither indicate which branch produces the correct top decision. Each cell is a within-family percentage.
Family
Both
E only
R only
Neither
Reasoning / semantic
48.8
9.0
12.1
30.1
Content matching
45.4
8.9
8.3
37.4
All
47.2
9.0
10.3
33.5
왜 중요한가
검색 서비스나 추천 시스템처럼 대량의 이미지·영상·문서 중에서 빠르게 후보를 걸러내면서도 헷갈리는 오답을 정확히 걸러내야 하는 실무 환경에 바로 적용할 수 있는 방법을 제시했다. 속도와 정확도 중 상황에 맞게 조절할 수 있다는 점도 실제 서비스 설계에 유용하다.
이 논문의 용어
임베딩(Embedding) · 텍스트나 이미지를 숫자 벡터로 바꿔 비슷한 것끼리 가깝게 배치하는 표현 방식
하드 네거티브(Hard Negative) · 정답과 겉모습이나 주제가 비슷해서 헷갈리기 쉬운 오답 후보
CoT(Chain-of-Thought) · 결론을 내기 전에 중간 추론 과정을 글로 풀어써서 모델이 단계적으로 생각하게 하는 기법
리랭킹(Reranking) · 1차로 빠르게 뽑은 후보들을 더 정밀한 모델로 다시 순위를 매기는 2단계 검색 방식
지식 증류(Distillation) · 한 모델(또는 기능)의 판단을 다른 모델(또는 기능)이 학습해서 닮아가게 하는 방법
본문에 싣지 못한 그림
Figure 1: Two overlooked issues in universal multimodal retrieval. (a) Item-wise self-reflective CoT lacks pair-aware evidence for hard-negative discrimination. (b) Different meta- tasks require different capabilities, with embedding and ranking offering complementary strengths.
Figure 2: Overview of UMER, a unified multimodal embedding and ranking framework. (a) Query and candidate inputs are first encoded independently through masked branch attention to extract their embeddings, and are then jointly used for pair-aware autoregressive reasoning followed by a ranking token. (b) Multi-task heads support metric embedding learning, generative CoT learning and discriminative ranking learning. (c) Unified co-optimization combines multi-task supervision and complementary mutual distillation to jointly improve embedding and ranking.
Figure 3: Embedding separation under item-wise and pair-aware supervision on 3,600 queries from the 36 image tasks of MMEB-V2. (a) Bootstrap distribution of the macro-averaged positive–hard-negative margin. (b) Fraction of queries exceeding margin thresholds.
Figure 4: Capability specialization and transfer via CMD. (a) Embedding is stronger for content matching, while ranking is stronger for reasoning-intensive relevance judgment. (b) CMD transfers these complementary strengths between the two functions.