SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection
객체 탐지기가 이미 알고 있던 것을 5개의 숫자로 뽑아내 엉뚱한 물체 오탐을 잡아낸다
자율주행이나 로봇 비전에 쓰이는 객체 탐지 AI는 학습 때 본 적 없는 물체를 마주치면 엉뚱하게 확신에 찬 오탐(할루시네이션)을 낸다. 이 논문은 탐지기 내부에 이미 숨겨져 있던 부분별 의미 정보, 크기 정보, 배경 맥락 정보를 끄집어내 5차원짜리 압축된 표현으로 정리하는 SPK라는 방법을 제안한다. 이 5차원 표현만으로 기존 방법보다 더 정확하게, 그리고 왜 오탐인지 설명 가능하게 판별해낸다.
METAL MEDIA 해설 도표
객체 탐지기가 이미 알고 있던 것을 5개의 숫자로 뽑아내 엉뚱한 물체 오탐을 잡아낸다
01기존 방법들은 탐지기가 뽑아낸 복잡한 고차원 특징 위에 판별 규칙을 얹거나 탐지기 자체를 다시 학습시켜 오탐을 줄이려 했는데, 이 논문은 탐지기 안에 원래 숨어 있던 지식을 있는 그대로 끄집어내는 쪽을 택했다.
02정상 범주와 비슷해 보이는 가짜 후보(Proximal OoD)와 배경만 있는 이미지를 진단용 자료로 써서, GPT-5로 만든 부위 이름(날개, 부리 등)을 OWLv2와 SAM 2로 자동 라벨링한 뒤 이를 근거로 부위별 의미 반응을 학습시켰다.
03여기서 얻은 3가지 의미 반응(정상 범주 유사도, 가짜 후보 유사도, 배경 유사도)에 물체 크기 비율(기하 정보), 이미지 전체 맥락 유사도(맥락 정보)를 더해 5차원의 SPK 표현을 만들었다.
04PASCAL-VOC와 BDD-100K 데이터셋, YOLO·Faster R-CNN·RT-DETR 세 종류 탐지기 구조에서 실험한 결과, 같은 판별 알고리즘(KNN, Isolation Forest 등)을 SPK 표현에 적용했을 때 원래 고차원 특징을 쓸 때보다 FPR95(정상 물체 95%를 통과시켰을 때 오탐률)가 일관되게 낮아졌고, 기존 최고 성능 방법이었던 Proximal-OoD보다도 오탐 제거 개수가 더 많았다.
05탐지기 자체는 전혀 손대지 않고도 이런 성능을 냈으며, YOLO 기준 원래 추론 10.15ms에 SPK 전체 처리를 더해도 12.65ms로 실시간성을 유지했다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
기존 방법들은 탐지기가 뽑아낸 복잡한 고차원 특징 위에 판별 규칙을 얹거나 탐지기 자체를 다시 학습시켜 오탐을 줄이려 했는데, 이 논문은 탐지기 안에 원래 숨어 있던 지식을 있는 그대로 끄집어내는 쪽을 택했다.
정상 범주와 비슷해 보이는 가짜 후보(Proximal OoD)와 배경만 있는 이미지를 진단용 자료로 써서, GPT-5로 만든 부위 이름(날개, 부리 등)을 OWLv2와 SAM 2로 자동 라벨링한 뒤 이를 근거로 부위별 의미 반응을 학습시켰다.
여기서 얻은 3가지 의미 반응(정상 범주 유사도, 가짜 후보 유사도, 배경 유사도)에 물체 크기 비율(기하 정보), 이미지 전체 맥락 유사도(맥락 정보)를 더해 5차원의 SPK 표현을 만들었다.
PASCAL-VOC와 BDD-100K 데이터셋, YOLO·Faster R-CNN·RT-DETR 세 종류 탐지기 구조에서 실험한 결과, 같은 판별 알고리즘(KNN, Isolation Forest 등)을 SPK 표현에 적용했을 때 원래 고차원 특징을 쓸 때보다 FPR95(정상 물체 95%를 통과시켰을 때 오탐률)가 일관되게 낮아졌고, 기존 최고 성능 방법이었던 Proximal-OoD보다도 오탐 제거 개수가 더 많았다.
탐지기 자체는 전혀 손대지 않고도 이런 성능을 냈으며, YOLO 기준 원래 추론 10.15ms에 SPK 전체 처리를 더해도 12.65ms로 실시간성을 유지했다.
Figure 1: The proposed SPK framework, a proactive OoD hallucination mitigation framework, further reduces OoD-induced hallucinations beyond previous state-of-the-art methods (48), achieving additional improvements in challenging high-performance regimes.
Table 1: Comparison of OoD detection performance using FPR95 across different detector architectures trained on PASCAL-VOC and BDD-100K. Lower is better.
Method
YOLO
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
MSP
67.48
67.18
72.93
74.12
68.36
78.69
77.46
74.41
70.44
67.81
77.46
74.41
EBO
90.49
90.84
87.22
87.06
60.62
56.21
94.83
94.37
98.03
96.69
94.83
94.37
MLS
89.88
90.08
86.47
87.06
59.62
57.89
92.55
91.08
92.84
89.75
92.55
91.08
SCALE
80.67
80.92
77.44
70.59
92.34
80.70
86.35
84.98
81.12
76.92
86.35
84.98
MDS
57.67
69.47
68.42
82.35
49.96
56.38
78.90
79.34
48.65
52.90
78.90
79.34
BAM
45.36
43.72
49.63
52.18
65.44
42.16
65.73
61.34
75.61
65.27
75.48
68.44
KNN
48.20
39.50
41.95
45.24
61.95
37.53
50.54
49.76
77.10
63.00
62.10
58.00
iForest
70.27
67.82
60.42
65.23
75.43
52.38
63.25
62.78
79.52
65.92
68.23
62.30
SPK-MDS
14.99
17.28
23.35
6.46
21.03
23.50
42.42
42.73
18.43
17.83
41.77
16.48
SPK-BAM
21.96
21.73
16.75
3.07
18.98
18.17
9.07
6.18
23.12
25.19
18.76
16.84
SPK-KNN
19.64
17.99
13.47
1.17
15.50
13.69
4.70
3.09
18.26
20.43
16.97
13.74
SPK-iForest
14.25
11.84
9.86
0.70
13.92
10.52
2.31
1.52
15.48
17.32
11.42
9.25
Figure 2: Overview of the proposed SPK framework. SPK elicits semantic, geometric, and contextual priors from a pretrained object detector and organizes them into a compact five-dimensional representation for OoD hallucination detection.
Table 2: OoD detection counts (Near-OoD/Far-OoD) across different detector architectures. Lower is better.
Model
Method
VOC (N/F)
BDD (N/F)
YOLO
Original
946 / 440
701 / 666
Proximal-OoD
134 / 60
80 / 47
SPK
135 / 52
69 / 5
Faster R-CNN
Original
2150 / 1335
2576 / 1634
Proximal-OoD
710 / 253
207 / 167
SPK
299 / 140
60 / 25
RT-DETR
Original
2311 / 1589
3145 / 1220
Proximal-OoD
386 / 470
525 / 240
SPK
358 / 275
359 / 113
Figure 3: Automated part-level annotation pipeline. Given an RoI crop and its object category, OWLv2 grounds the corresponding GPT-5-generated concept vocabulary into part bounding boxes. Each box prompts SAM 2 to produce a refined pixel-level part mask. Masks satisfying the object-mask coverage threshold are projected into detector-RoI coordinates and rasterized as binary 7×7 targets. Concepts without a retained mask receive an all-zero target.
Table 3: Ablation study of the SPK loss components on YOLO trained on PASCAL-VOC and BDD-100K. Results are reported as Near-OoD / Far-OoD FPR95. The average is computed over all four results. Lower is better.
ℒdice
ℒsuppress
ℒgroup
VOC
BDD
Average
Near / Far
Near / Far
FPR95 ↓
✗
✓
✓
25.80 / 21.30
17.85 / 18.26
20.80
✓
✓
✗
22.10 / 16.40
15.29 / 15.97
17.44
✓
✗
✓
20.50 / 15.80
14.18 / 12.93
15.85
✓
✓
✓
14.25 / 11.84
9.86 / 0.70
9.16
Figure 4: Part-Level Concept Annotation Example. We use a bird RoI crop to illustrate the step-by-step annotation process for the concepts wing, head, beak, and torso. Step 1: OWLv2 grounds each concept within the detector RoI and generates a bounding-box proposal (a). Step 2: SAM 2 refines each proposal into a pixel-level part mask, which is retained only if at least 70% of its pixels overlap with the corresponding SAM 2 object mask (b). Step 3: Each retained mask is projected into detector-RoI coordinates and converted into a binary 7×7 supervision target (c). Step 4: The binary target is overlaid on the RoI crop for visualization (d).
Table 4: Ablation study of different prior components on YOLO trained on PASCAL-VOC and BDD-100K. Each dataset column reports Near-OoD / Far-OoD FPR95. The average is computed over all four results. Lower is better.
Prior components
VOC
BDD
Average
Near / Far
Near / Far
FPR95 ↓
Semantic
15.43 / 13.28
30.37 / 25.58
21.17
Semantic + Geometric
13.23 / 11.50
20.88 / 3.84
12.36
Semantic + Geometric + Contextual
14.25 / 11.84
9.86 / 0.70
9.16
Figure 5: Qualitative visualization of part-level semantic responses. The learned concept maps are well aligned with the corresponding regions, demonstrating that the semantic elicitation head successfully decodes spatially grounded semantic evidence from detector RoI features.
Table 5: Quality of the automated part-level annotations. We report the overall concept-mIoU, Recall@0.5, number of covered concepts, and per-class concept-mIoU. All values except concept coverage are percentages. Best results are shown in bold.
Method
Overall
Per-class concept-mIoU
concept-mIoU
Recall@0.5
# Concepts
bird
bus
car
cat
cow
dog
horse
Ours
40.8
38.4
25
46.5
37.9
43.9
40.4
37.4
41.8
38.2
VLPart (37)
37.0
36.8
25
35.6
32.8
32.5
38.1
36.8
44.1
36.5
Grounded SAM (31)
32.3
30.7
25
32.1
31.7
31.3
33.0
29.5
38.3
29.4
(b) Prediction: wing
Table 6: RoI feature sources and dimensions for each detector. YOLO and RT-DETR concatenate RoI-aligned features from three feature scales, whereas Faster R-CNN uses the Detectron2 box_pooler.
Detector
Feature source
Channels
RoI feature 𝐅i
YOLO
Detect neck
128+256+512
896×7×7
RT-DETR
Hybrid-encoder neck
256+256+256
768×7×7
Faster R-CNN
ResNet-FPN
256
256×7×7
(c) Prediction: torso
Table 7: Hyperparameters of the Semantic Elicitation Head.
Semantic Elicitation Head
Hyperparameter
Value
Architecture
Hidden channels
256
Dropout
0.1
Normalization
GroupNorm
Activation
GELU
Residual blocks
2
Output head
1×1 conv → num_concepts
Inference activation
Sigmoid
RoI spatial size
7×7
Inference pooling
LogSumExp, τ=0.5
Training
Training epochs
80
Batch size
2000
Optimizer
AdamW
Learning rate
2×10−4
Weight decay
5×10−4
Random seed
42
Validation split
10%
Training sampler
WeightedRandomSampler
Suppress loss weight
0.25
Group loss weight
0.75
Early-stopping patience
10 epochs
(d) Prediction: foot
Table 8: Image-level embedding sources and dimensions for each detector. We globally pool each selected feature map by its spatial mean and standard deviation, then concatenate the resulting statistics. YOLO and RT-DETR use their deepest selected backbone stage, whereas Faster R-CNN aggregates all four ResNet-FPN levels.
Detector
Feature source
Channels
Embedding dimension
YOLO
Backbone L6 (stride 16)
256
mean+std:512
RT-DETR
HGBlock L9 backbone (stride 32)
2048
mean+std:4096
Faster R-CNN
ResNet-FPN P2–P5
4×256
mean+std:2048
Figure 6: Distributions of the learned semantic group responses. Samples from different data sources predominantly activate their corresponding semantic groups, validating the effectiveness of the proposed group objective.
Table 9: Semantic-head group-classification accuracy. Computed by applying argmax to concept activations for YOLO on PASCAL-VOC.
Data split
Accuracy (%)
ID training set
96.5
Proximal OoD
81.0
Background
77.0
Figure 7: UMAP visualization of representation spaces for detections from the dog, sheep, and cat categories. The visualizations are obtained using YOLO trained on PASCAL-VOC. We compare detector classification logits (left) with the proposed SPK representations (right), using the same ID-validation, Near-OoD, and Far-OoD samples. SPK yields a more structured representation space with clearer distributional differences between ID and OoD samples.
Table 10: AUROC comparison across detector architectures on PASCAL-VOC and BDD-100K. Higher is better.
Method
YOLO
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
MSP
81.24
79.47
77.63
75.12
78.71
73.84
72.48
76.39
79.16
78.72
74.31
75.06
EBO
60.73
57.92
65.41
62.86
82.58
86.14
52.07
50.83
40.12
43.31
49.68
52.74
MLS
59.41
61.26
63.78
65.02
84.76
83.58
54.82
59.46
56.37
59.74
54.91
57.43
SCALE
69.84
71.73
72.46
79.31
57.49
69.72
64.18
67.82
69.37
75.16
66.42
65.73
MDS
85.62
77.81
80.34
68.46
88.39
84.41
73.68
70.94
87.43
85.71
73.27
71.18
BAM
90.14
89.38
89.57
86.94
80.96
91.56
80.75
83.65
75.75
81.65
76.50
80.45
KNN
89.52
90.31
91.26
88.47
81.73
92.58
88.43
86.72
75.12
81.46
83.68
85.29
iForest
77.38
80.71
82.46
80.19
76.31
85.64
81.37
83.29
70.82
81.94
78.46
83.57
SPK-MDS
96.12
93.55
95.80
98.65
96.10
94.20
92.10
91.80
94.35
95.54
89.10
99.21
SPK-BAM
94.55
92.50
96.90
99.35
96.45
95.40
98.30
98.85
93.75
93.90
93.45
99.18
SPK-KNN
94.91
93.31
97.40
99.71
97.07
96.46
99.06
99.36
94.40
94.87
93.94
99.36
SPK-iForest
96.31
95.43
98.10
99.81
97.38
97.25
99.50
99.65
95.23
95.67
95.78
99.60
(b) sheep
Table 11: Ablation of SPK loss components. Evaluated across Faster R-CNN and RT-DETR on PASCAL-VOC and BDD-100K. Lower FPR95 is better.
ℒdice
ℒsuppress
ℒgroup
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
✗
✓
✓
27.87
19.98
18.36
18.74
32.31
27.08
25.35
29.14
✓
✓
✗
24.42
15.42
14.15
16.11
27.86
22.53
20.71
24.36
✓
✗
✓
22.19
14.73
11.34
13.27
24.86
21.64
18.15
20.81
✓
✓
✓
13.92
10.52
2.31
1.52
15.48
17.32
11.42
9.25
(c) cat
Table 12: Ablation of SPK prior components. Evaluated across Faster R-CNN and RT-DETR on PASCAL-VOC and BDD-100K. Lower FPR95 is better.
Prior components
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Semantic
18.17
18.91
23.92
25.21
18.34
29.17
33.36
34.75
Semantic + Geometric
15.78
16.43
14.52
15.71
15.85
25.35
26.51
26.82
Semantic + Geometric + Contextual
13.92
10.52
2.31
1.52
15.48
17.32
11.42
9.25
Table 13: Comparison with competitive OoD detection methods for Deformable-DETR. Results are reported on PASCAL-VOC and BDD-100K as ID datasets, with MS-COCO and OpenImages as OoD datasets. Higher AUROC and lower FPR95 indicate better OoD detection performance. Methods marked with † employ an external DINO ViT encoder to extract additional visual representations for OoD detection, rather than relying solely on Deformable-DETR features. SPK (DINO ViT) is included to provide a fair comparison with UNO-Adapter, as both methods operate under this setting. The best results are highlighted in bold.
Method
ID: PASCAL-VOC
ID: BDD-100K
OoD: MS-COCO
OoD: OpenImages
OoD: MS-COCO
OoD: OpenImages
FPR95↓
AUROC↑
FPR95↓
AUROC↑
FPR95↓
AUROC↑
FPR95↓
AUROC↑
MDS (18)
97.39
50.28
97.88
49.08
70.86
76.83
71.43
77.98
Gram matrices (32)
94.16
43.97
95.29
38.81
73.81
60.13
71.56
57.14
KNN (38)
91.80
62.15
91.36
59.64
64.75
80.90
61.13
79.64
CSI (39)
84.00
55.07
79.16
51.37
70.27
77.93
71.30
76.42
VOS (8)
97.46
54.40
97.07
52.77
76.44
77.33
72.58
76.62
OW-DETR (11)
93.09
55.70
93.82
57.80
80.78
70.29
77.37
73.78
DisMax (25)
82.05
75.21
76.37
70.66
77.62
72.14
81.23
67.18
SIREN-vMF (7)
75.49
76.10
78.36
71.05
67.54
80.06
66.31
79.77
SIREN-KNN (7)
64.77
78.23
65.99
74.93
53.97
86.56
47.28
89.00
SAFE (42)
48.88
78.88
8.99
96.73
39.18
85.95
21.10
94.31
InfoBound (52)
44.88
89.76
43.89
88.00
44.88
89.76
43.89
88.00
UNO-Adapter† (28)
32.61
91.68
19.90
95.40
9.88
97.61
3.80
99.04
SPK
52.32
75.84
24.38
90.20
1.68
99.42
0.37
99.93
SPK (DINO ViT)†
28.55
92.38
14.85
96.25
0.00
99.80
0.00
99.97
Table 14: Comparison with competitive OoD detection methods for Faster R-CNN. Results are reported on PASCAL-VOC as the ID dataset, with MS COCO and OpenImages as OoD datasets. Higher AUROC and lower FPR95 indicate better OoD detection performance. Methods marked with † employ an external DINO ViT encoder to extract additional visual representations for OoD detection, rather than relying solely on Faster R-CNN features. SPK (DINO ViT) is included to provide a fair comparison with UNO-Adapter, as both methods operate under this setting. The best results are highlighted in bold.
Method
MS-COCO
OpenImages
AUROC↑
FPR95↓
AUROC↑
FPR95↓
CSI (39)
82.95
57.41
81.83
59.91
GAN-Synthesis (17)
82.67
59.97
83.67
60.93
VOS (8)
85.23
51.33
88.70
47.53
SIREN (7)
85.36
64.68
82.78
68.53
TIB (44)
90.36
41.55
88.09
47.19
DFDD (43)
90.79
41.34
88.65
44.52
WFS (45)
89.01
40.05
90.35
39.17
UNO-Adapter† (28)
91.25
38.73
92.40
35.74
SPK
91.58
41.20
95.50
25.61
SPK (DINO ViT)†
95.48
26.28
98.16
10.97
Table 15: Per-image runtime of the complete SPK inference pipeline. Reported for a YOLO model pretrained on PASCAL-VOC. The complete SPK pipeline introduces an additional 2.72 ms latency per image, corresponding to a 26.8% runtime overhead over the original detector inference. *SPK inference obtains detection outputs, image-level contextual embeddings, and RoI features within the same forward pass.
Component
Cost (ms)
Original inference
10.15
SPK inference*
12.65
Semantic prior elicitation
0.17
Isolation Forest
0.05
왜 중요한가
자율주행차나 로봇이 학습 때 못 본 물체를 보고도 엉뚱하게 확신에 찬 판단을 내리면 사고로 이어질 수 있는데, 이 방법은 탐지기를 재학습시키지 않고도 가볍게 붙여서 오탐을 걸러내고 왜 걸렀는지도 설명해준다는 점에서 실무 적용 가치가 크다. 또한 복잡한 판별 알고리즘을 개발하는 것보다 좋은 표현 공간을 만드는 것이 더 중요하다는 시사점도 준다.
이 논문의 용어
OoD 할루시네이션 · 탐지기가 학습 때 배운 적 없는 물체나 배경을 마치 아는 물체인 것처럼 확신을 가지고 잘못 인식하는 현상
FPR95 · 정상 데이터의 95%를 정확히 통과시켰을 때, 이상치가 잘못 통과되는 비율을 나타내는 지표로 낮을수록 좋음
RoI 특징 · 탐지기가 예측한 바운딩 박스 영역에서 뽑아낸 세부 특징 벡터
Isolation Forest · 데이터를 무작위로 나누는 방식을 반복해 이상치를 빠르게 골라내는 가벼운 이상 탐지 알고리즘
Proximal OoD · 정상 범주와 시각적으로 비슷해 보여서 탐지기를 혼동시키는 학습 범주 밖 물체