Figure 1: Comparison of conventional, refusal, and our MMOOC benchmarks. MMOOC jointly evaluates robust answering for answerable questions and appropriate refusal for truly out-of-context questions.
Table 1: Comparison of existing out-of-context evaluation benchmarks. Data Scale denotes the total number of question-answer pairs. QA Format indicates the supported question formats. Shift Types denotes the number of defined shift types in each benchmark. Visual Scenarios indicates the range of visual understanding and reasoning settings covered by each benchmark. QA Construction indicates how the question-answer pairs are constructed. Distractor Robustness indicates whether the benchmark evaluates correct answering under distracting contexts.
Figure 2: Examples of the eight MMOOC scenarios. The top row presents five Out-of-Context categories requiring refusal: Multimodal Ambiguity (MA), Visual False Premises (VFP), Uncertain Spatial & Physical Context (USPC), Unclear Logical & Symbolic (ULS), and Missing Knowledge & Background (MKB). The bottom row presents three answerable Shifted In-Context scenarios: Misleading Premise, Partial Answerability, and Image–Question Mismatch.
Table 3: Average performance on Out-of-Context tasks across various models, computed from Refusal Rate and Refusal Rationality. The OOC category abbreviations are: MA: Multimodal Ambiguity; VFP: Visual False Premises; USPC: Uncertain Spatial & Physical Context; ULS: Unclear Logical & Symbolic; and MKB: Missing Knowledge & Background. Complete detailed results are presented in the Appendix.
YesNo
MCQ
VQA
Model
MA
VFP
USPC
ULS
MKB
MA
VFP
USPC
ULS
MKB
MA
VFP
USPC
ULS
MKB
Open-source LMMs
Qwen3-VL-2B
13.75
15.75
5.75
10.50
19.00
40.00
29.25
45.50
22.75
21.50
26.00
29.50
8.25
12.00
21.75
Qwen3-VL-8B
28.00
58.75
15.75
35.50
30.25
36.50
22.00
71.50
39.50
34.75
55.25
86.00
28.50
35.00
49.00
Qwen3-VL-30B
44.75
58.00
22.50
26.75
35.50
37.25
36.25
76.75
32.75
38.50
42.25
54.75
20.00
20.00
19.25
Qwen3.5-27B
36.25
63.25
22.50
28.00
43.25
32.50
27.50
69.25
16.75
35.50
46.00
84.25
27.75
40.25
39.00
Qwen3.5-122B-A10B
30.50
65.25
26.25
28.25
32.75
28.25
32.25
67.75
35.25
23.00
46.00
85.25
26.50
42.00
41.50
LLaVA-1.5-7B
1.75
8.75
2.25
6.25
3.00
1.75
5.25
1.75
3.00
2.50
3.25
5.75
5.75
14.75
8.00
InternVL3-2B
21.25
39.50
3.75
17.00
28.50
4.50
14.75
4.25
11.25
5.00
28.25
26.50
11.50
15.50
33.75
InternVL3-8B
48.25
55.50
12.00
39.25
47.50
29.75
5.00
46.25
21.25
18.50
21.75
43.00
7.00
15.50
32.00
Gemma-4-26B
39.50
66.25
27.75
49.00
76.25
70.25
28.50
85.00
57.25
61.25
79.25
86.50
39.50
58.00
80.00
Gemma-4-31B
54.50
54.25
35.25
46.75
72.50
52.25
34.25
88.75
47.50
43.50
71.50
83.00
40.50
54.00
70.50
Llama-4-Maverick
40.00
51.25
19.00
27.75
34.00
32.25
26.25
47.25
27.00
36.25
52.25
78.25
20.75
31.75
58.00
Ministral-3-8B
46.75
60.50
26.50
31.00
61.75
41.50
31.25
54.25
27.75
44.50
54.50
59.25
33.00
37.25
58.00
Ministral-3-14B
45.25
52.75
28.50
45.75
61.50
50.25
24.50
75.50
42.00
40.50
52.50
50.75
15.00
36.50
36.50
Closed-source LMMs
Gemini-3.1-Pro
4.75
9.00
21.75
27.75
20.75
4.50
14.25
24.25
12.50
15.00
7.75
6.50
16.50
9.50
20.25
GPT-4o
43.00
59.25
30.00
33.75
53.75
51.25
24.75
64.50
34.00
34.25
43.50
58.25
21.00
54.00
57.75
o1
57.50
71.75
44.75
54.75
76.25
45.50
22.75
68.25
30.00
45.75
75.00
43.50
40.75
47.50
53.75
o3
29.75
35.50
18.00
34.25
43.25
37.75
22.25
37.50
16.00
19.75
47.75
61.00
24.75
34.50
41.25
Claude-Opus-4.6
34.25
67.75
15.25
30.50
30.75
14.25
5.25
31.75
6.75
4.75
14.75
35.00
15.50
23.50
46.50
Figure 4: Performance comparison of different models.
Table 4: Average performance on in-context (IC) shift tasks. MP: Misleading Premise, PA: Partial Answerability, and IQM: Image–Question Mismatch. Detailed results for all IC categories are provided in the Appendix.
YN
MCQ
VQA
Model
MP
PA
IQM
MP
PA
IQM
MP
PA
IQM
Open-source LMMs
Qwen3-VL-2B
86.00
70.25
71.50
78.00
67.50
77.25
63.00
36.25
82.75
Qwen3-VL-8B
90.00
72.25
75.25
84.75
75.25
90.25
83.25
58.00
87.75
Qwen3-VL-30B
88.50
79.50
78.50
85.50
82.00
95.00
76.50
55.00
82.75
Qwen3.5-27B
88.25
75.75
79.25
93.75
85.50
94.50
90.50
63.00
89.00
Qwen3.5-122B-A10B
80.25
78.25
75.75
87.25
90.50
94.50
90.25
72.25
86.75
LLaVA-1.5-7B
55.50
55.00
58.00
31.00
34.00
38.00
28.00
11.00
55.25
InternVL3-2B
76.25
63.75
59.75
64.00
68.00
70.50
52.00
41.50
73.25
InternVL3-8B
71.75
66.75
62.00
70.00
71.50
69.25
62.50
53.00
70.75
Gemma-4-26B
81.75
79.00
74.00
89.25
85.25
83.25
74.25
67.50
78.00
Gemma-4-31B
75.75
77.25
67.75
84.25
84.75
93.25
75.75
70.75
80.50
Llama-4-Maverick
76.00
74.25
71.75
84.50
86.50
85.75
77.50
58.00
80.00
Ministral-3-8B
73.25
72.75
71.00
80.25
71.75
90.00
76.75
64.25
80.50
Ministral-3-14B
75.25
81.00
72.50
83.00
78.50
79.25
65.25
55.50
73.50
Closed-source LMMs
Gemini-3.1-Pro
70.75
62.75
59.50
59.50
58.75
75.75
61.50
22.75
52.00
GPT-4o
82.50
78.25
72.00
70.50
76.25
82.00
76.75
62.50
80.25
o1
79.25
81.25
71.00
79.50
73.25
85.25
74.75
56.50
81.00
o3
78.25
77.25
63.75
75.00
75.25
79.00
71.00
44.25
85.25
Claude-Opus-4.6
88.25
82.75
72.50
58.00
45.25
40.00
21.75
32.25
35.75
Figure 6: Performance under different prompts.
Table A1: Question-only refusal performance, where models receive only the question without the associated image. Ref. denotes the refusal score. Higher values indicate a stronger tendency to identify the question as unanswerable in the absence of visual evidence.
Model
Ref.
Open-source LMMs
Qwen3-VL-2B
42.00
Qwen3-VL-8B
30.00
Qwen3-VL-30B
58.00
Qwen3.5-27B
32.00
Qwen3.5-122B-A10B
20.00
LLaVA-1.5-7B
10.00
InternVL3-2B
8.00
InternVL3-8B
24.00
Gemma-4-26B
92.00
Gemma-4-31B
76.00
Llama-4-Maverick
52.00
Ministral-3-8B
78.00
Ministral-3-14B
84.00
Closed-source LMMs
Gemini-3.1-Pro
4.00
GPT-4o
26.00
o1
78.00
o3
46.00
Claude-Opus-4.6
62.00
Figure 7: Robustness to misleading prompts, evaluated by our core metrics Rref (Refusal Rate) and Rrat (Reasoning Rationality).
Table A2: Detailed performance on the OOC YesNo tasks. Ref., Rat., and Mean denote Refusal Rate, Refusal Rationality, and their average, respectively.
MA
VFP
USPC
ULS
MKB
Model
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Open-source LMMs
Qwen3-VL-2B
10.00
17.50
13.75
8.00
23.50
15.75
2.00
9.50
5.75
6.00
15.00
10.50
14.00
24.00
19.00
Qwen3-VL-8B
16.00
40.00
28.00
48.00
69.50
58.75
8.00
23.50
15.75
26.00
45.00
35.50
16.00
44.50
30.25
Qwen3-VL-30B
36.00
53.50
44.75
46.00
70.00
58.00
12.00
33.00
22.50
18.00
35.50
26.75
28.00
43.00
35.50
Qwen3.5-27B
22.00
50.50
36.25
44.00
82.50
63.25
10.00
35.00
22.50
14.00
42.00
28.00
28.00
58.50
43.25
Qwen3.5-122B-A10B
16.00
45.00
30.50
50.00
80.50
65.25
12.00
40.50
26.25
14.00
42.50
28.25
18.00
47.50
32.75
LLaVA-1.5-7B
0.00
3.50
1.75
2.00
15.50
8.75
0.00
4.50
2.25
2.00
10.50
6.25
0.00
6.00
3.00
InternVL3-2B
18.00
24.50
21.25
30.00
49.00
39.50
0.00
7.50
3.75
12.00
22.00
17.00
22.00
35.00
28.50
InternVL3-8B
46.00
50.50
48.25
46.00
65.00
55.50
6.00
18.00
12.00
34.00
44.50
39.25
42.00
53.00
47.50
Gemma-4-26B
32.00
47.00
39.50
58.00
74.50
66.25
22.00
33.50
27.75
42.00
56.00
49.00
74.00
78.50
76.25
Gemma-4-31B
52.00
57.00
54.50
46.00
62.50
54.25
32.00
38.50
35.25
42.00
51.50
46.75
70.00
75.00
72.50
Llama-4-Maverick
32.00
48.00
40.00
42.00
60.50
51.25
12.00
26.00
19.00
18.00
37.50
27.75
24.00
44.00
34.00
Ministral-3-8B
38.00
55.50
46.75
50.00
71.00
60.50
18.00
35.00
26.50
24.00
38.00
31.00
56.00
67.50
61.75
Ministral-3-14B
34.00
56.50
45.25
42.00
63.50
52.75
22.00
35.00
28.50
38.00
53.50
45.75
58.00
65.00
61.50
Closed-source LMMs
Gemini-3.1-Pro
4.00
5.50
4.75
12.00
6.00
9.00
28.00
15.50
21.75
18.00
37.50
27.75
32.00
9.50
20.75
GPT-4o
38.00
48.00
43.00
46.00
72.50
59.25
26.00
34.00
30.00
26.00
41.50
33.75
44.00
63.50
53.75
o1
52.00
63.00
57.50
60.00
83.50
71.75
38.00
51.50
44.75
48.00
61.50
54.75
72.00
80.50
76.25
o3
26.00
33.50
29.75
27.00
43.75
35.50
12.00
24.00
18.00
26.00
42.50
34.25
36.00
50.50
43.25
Claude-Opus-4.6
26.00
42.50
34.25
60.00
75.50
67.75
4.00
26.50
15.25
22.00
39.00
30.50
20.00
41.50
30.75
Figure 8: Robustness to Gaussian noise, evaluated by accuracy (Acc) and reasoning rationality (Accrat).
Table A3: Detailed performance on the OOC MCQ tasks. Ref., Rat., and Mean denote Refusal Rate, Refusal Rationality, and their average, respectively.
MA
VFP
USPC
ULS
MKB
Model
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Open-source LMMs
Qwen3-VL-2B
36.00
44.00
40.00
22.00
36.50
29.25
44.00
47.00
45.50
16.00
29.50
22.75
16.00
27.00
21.50
Qwen3-VL-8B
32.00
41.00
36.50
14.00
30.00
22.00
70.00
73.00
71.50
34.00
45.00
39.50
28.00
41.50
34.75
Qwen3-VL-30B
32.00
42.50
37.25
28.00
44.50
36.25
76.00
77.50
76.75
26.00
39.50
32.75
26.00
51.00
38.50
Qwen3.5-27B
26.00
39.00
32.50
16.00
39.00
27.50
68.00
70.50
69.25
10.00
23.50
16.75
26.00
45.00
35.50
Qwen3.5-122B-A10B
22.00
34.50
28.25
24.00
40.50
32.25
64.00
71.50
67.75
28.00
42.50
35.25
12.00
34.00
23.00
LLaVA-1.5-7B
2.00
1.50
1.75
4.00
6.50
5.25
2.00
1.50
1.75
2.00
4.00
3.00
2.00
3.00
2.50
InternVL3-2B
4.00
5.00
4.50
12.00
17.50
14.75
4.00
4.50
4.25
10.00
12.50
11.25
4.00
6.00
5.00
InternVL3-8B
28.00
31.50
29.75
2.00
8.00
5.00
48.00
44.50
46.25
18.00
24.50
21.25
16.00
21.00
18.50
Gemma-4-26B
66.00
74.50
70.25
20.00
37.00
28.50
86.00
84.00
85.00
50.00
64.50
57.25
56.00
66.50
61.25
Gemma-4-31B
48.00
56.50
52.25
26.00
42.50
34.25
88.00
89.50
88.75
46.00
49.00
47.50
40.00
47.00
43.50
Llama-4-Maverick
26.00
38.50
32.25
16.00
36.50
26.25
38.00
57.50
47.25
20.00
34.00
27.00
32.00
40.50
36.25
Ministral-3-8B
34.00
49.00
41.50
26.00
36.50
31.25
52.00
56.50
54.25
18.00
37.50
27.75
38.00
51.00
44.50
Ministral-3-14B
44.00
56.50
50.25
14.00
35.00
24.50
76.00
75.00
75.50
34.00
50.00
42.00
30.00
51.00
40.50
Closed-source LMMs
Gemini-3.1-Pro
4.00
5.00
4.50
20.00
8.50
14.25
36.00
12.50
24.25
20.00
5.00
12.50
20.00
10.00
15.00
GPT-4o
50.00
52.50
51.25
18.00
31.50
24.75
64.00
65.00
64.50
30.00
38.00
34.00
30.00
38.50
34.25
o1
44.00
47.00
45.50
18.00
27.50
22.75
70.00
66.50
68.25
26.00
34.00
30.00
40.00
51.50
45.75
o3
38.00
37.50
37.75
20.00
24.50
22.25
38.00
37.00
37.50
14.00
18.00
16.00
18.00
21.50
19.75
Claude-Opus-4.6
10.00
18.50
14.25
2.00
8.50
5.25
20.00
43.50
31.75
4.00
9.50
6.75
0.00
9.50
4.75
Table A4: Detailed performance on the OOC VQA tasks. Ref., Rat., and Mean denote Refusal Rate, Refusal Rationality, and their average, respectively.
MA
VFP
USPC
ULS
MKB
Model
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Open-source LMMs
Qwen3-VL-2B
26.00
26.00
26.00
30.00
29.00
29.50
6.00
10.50
8.25
12.00
12.00
12.00
20.00
23.50
21.75
Qwen3-VL-8B
52.00
58.50
55.25
86.00
86.00
86.00
22.00
35.00
28.50
30.00
40.00
35.00
46.00
52.00
49.00
Qwen3-VL-30B
40.00
44.50
42.25
54.00
55.50
54.75
14.00
26.00
20.00
16.00
24.00
20.00
18.00
20.50
19.25
Qwen3.5-27B
38.00
54.00
46.00
84.00
84.50
84.25
14.00
41.50
27.75
38.00
42.50
40.25
34.00
44.00
39.00
Qwen3.5-122B-A10B
40.00
52.00
46.00
84.00
86.50
85.25
14.00
39.00
26.50
36.00
48.00
42.00
36.00
47.00
41.50
LLaVA-1.5-7B
2.00
4.50
3.25
6.00
5.50
5.75
4.00
7.50
5.75
14.00
15.50
14.75
6.00
10.00
8.00
InternVL3-2B
28.00
28.50
28.25
30.00
23.00
26.50
8.00
15.00
11.50
12.00
19.00
15.50
32.00
35.50
33.75
InternVL3-8B
18.00
25.50
21.75
44.00
42.00
43.00
2.00
12.00
7.00
12.00
19.00
15.50
30.00
34.00
32.00
Gemma-4-26B
78.00
80.50
79.25
88.00
85.00
86.50
34.00
45.00
39.50
54.00
62.00
58.00
78.00
82.00
80.00
Gemma-4-31B
72.00
71.00
71.50
82.00
84.00
83.00
34.00
47.00
40.50
52.00
56.00
54.00
70.00
71.00
70.50
Llama-4-Maverick
48.00
56.50
52.25
78.00
78.50
78.25
12.00
29.50
20.75
26.00
37.50
31.75
54.00
62.00
58.00
Ministral-3-8B
46.00
63.00
54.50
56.00
62.50
59.25
30.00
36.00
33.00
28.00
46.50
37.25
54.00
62.00
58.00
Ministral-3-14B
49.25
46.00
52.50
50.00
51.50
50.75
10.00
20.00
15.00
32.00
41.00
36.50
30.00
43.00
36.50
Closed-source LMMs
Gemini-3.1-Pro
8.00
7.50
7.75
10.00
3.00
6.50
22.00
11.00
16.50
14.00
5.00
9.50
30.00
10.50
20.25
GPT-4o
42.00
45.00
43.50
60.00
56.50
58.25
16.00
26.00
21.00
52.00
56.00
54.00
54.00
61.50
57.75
o1
78.00
72.00
75.00
46.00
41.00
43.50
36.00
45.50
40.75
44.00
51.00
47.50
52.00
55.50
53.75
o3
48.00
47.50
47.75
63.00
59.00
61.00
22.00
27.50
24.75
32.00
37.00
34.50
42.00
40.50
41.25
Claude-Opus-4.6
12.00
17.50
14.75
34.00
36.00
35.00
10.00
21.00
15.50
20.00
27.00
23.50
44.00
49.00
46.50
Table A5: Detailed performance on the shifted in-context YesNo tasks. ACC, Acc.rat, and Mean denote answer accuracy, answer rationality, and their average, respectively.
MP
PA
IQM
Model
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
Open-source LMMs
Qwen3-VL-2B
88.00
84.00
86.00
72.00
68.50
70.25
62.00
81.00
71.50
Qwen3-VL-8B
92.00
88.00
90.00
74.00
70.50
72.25
68.00
82.50
75.25
Qwen3-VL-30B
86.00
91.00
88.50
80.00
79.00
79.50
70.00
87.00
78.50
Qwen3.5-27B
84.00
92.50
88.25
76.00
75.50
75.75
68.00
90.50
79.25
Qwen3.5-122B-A10B
70.00
90.50
80.25
74.00
82.50
78.25
66.00
85.50
75.75
LLaVA-1.5-7B
56.00
55.00
55.50
54.00
56.00
55.00
50.00
66.00
58.00
InternVL3-2B
82.00
70.50
76.25
62.00
65.50
63.75
48.00
71.50
59.75
InternVL3-8B
64.00
79.50
71.75
60.00
73.50
66.75
48.00
76.00
62.00
Gemma-4-26B
86.00
77.50
81.75
82.00
76.00
79.00
72.00
76.00
74.00
Gemma-4-31B
72.00
79.50
75.75
80.00
74.50
77.25
60.00
75.50
67.75
Llama-4-Maverick
66.00
86.00
76.00
66.00
82.50
74.25
68.00
75.50
71.75
Ministral-3-8B
64.00
82.50
73.25
68.00
77.50
72.75
60.00
82.00
71.00
Ministral-3-14B
74.00
76.50
75.25
82.00
80.00
81.00
64.00
81.00
72.50
Closed-source LMMs
Gemini-3.1-Pro
74.00
67.50
70.75
62.00
63.50
62.75
54.00
65.00
59.50
GPT-4o
78.00
87.00
82.50
74.00
82.50
78.25
58.00
86.00
72.00
o1
70.00
88.50
79.25
76.00
86.50
81.25
56.00
86.00
71.00
o3
76.00
80.50
78.25
74.00
80.50
77.25
50.00
77.50
63.75
Claude-Opus-4.6
82.00
94.50
88.25
84.00
81.50
82.75
58.00
87.00
72.50
Table A6: Detailed performance on the shifted in-context MCQ tasks. ACC, Acc.rat, and Mean denote answer accuracy, answer rationality, and their average, respectively.
MP
PA
IQM
Model
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
Open-source LMMs
Qwen3-VL-2B
90.00
66.00
78.00
66.00
69.00
67.50
84.00
70.50
77.25
Qwen3-VL-8B
86.00
83.50
84.75
74.00
76.50
75.25
92.00
88.50
90.25
Qwen3-VL-30B
92.00
79.00
85.50
84.00
80.00
82.00
94.00
96.00
95.00
Qwen3.5-27B
98.00
89.50
93.75
88.00
83.00
85.50
96.00
93.00
94.50
Qwen3.5-122B-A10B
88.00
86.50
87.25
92.00
89.00
90.50
96.00
93.00
94.50
LLaVA-1.5-7B
36.00
26.00
31.00
30.00
38.00
34.00
42.00
34.00
38.00
InternVL3-2B
78.00
50.00
64.00
70.00
66.00
68.00
84.00
57.00
70.50
InternVL3-8B
84.00
56.00
70.00
74.00
69.00
71.50
78.00
60.50
69.25
Gemma-4-26B
98.00
80.50
89.25
86.00
84.50
85.25
86.00
80.50
83.25
Gemma-4-31B
90.00
78.50
84.25
86.00
83.50
84.75
94.00
92.50
93.25
Llama-4-Maverick
94.00
75.00
84.50
90.00
83.00
86.50
92.00
79.50
85.75
Ministral-3-8B
88.00
72.50
80.25
68.00
75.50
71.75
94.00
86.00
90.00
Ministral-3-14B
90.00
76.00
83.00
76.00
81.00
78.50
80.00
78.50
79.25
Closed-source LMMs
Gemini-3.1-Pro
58.00
61.00
59.50
56.00
61.50
58.75
78.00
73.50
75.75
GPT-4o
76.00
65.00
70.50
80.00
72.50
76.25
90.00
74.00
82.00
o1
90.00
69.00
79.50
74.00
72.50
73.25
90.00
80.50
85.25
o3
88.00
62.00
75.00
78.00
72.50
75.25
88.00
70.00
79.00
Claude-Opus-4.6
62.00
54.00
58.00
48.00
42.50
45.25
42.00
38.00
40.00
Table A7: Detailed performance on the shifted in-context VQA tasks. ACC, Acc.rat, and Mean denote answer accuracy, answer rationality, and their average, respectively.
MP
PA
IQM
Model
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
Open-source LMMs
Qwen3-VL-2B
68.00
58.00
63.00
22.00
50.50
36.25
88.00
77.50
82.75
Qwen3-VL-8B
86.00
80.50
83.25
62.00
54.00
58.00
90.00
85.50
87.75
Qwen3-VL-30B
78.00
75.00
76.50
44.00
66.00
55.00
82.00
83.50
82.75
Qwen3.5-27B
92.00
89.00
90.50
56.00
70.00
63.00
88.00
90.00
89.00
Qwen3.5-122B-A10B
90.00
90.50
90.25
68.00
76.50
72.25
86.00
87.50
86.75
LLaVA-1.5-7B
28.00
28.00
28.00
1.00
21.00
11.00
58.00
52.50
55.25
InternVL3-2B
56.00
48.00
52.00
30.00
53.00
41.50
78.00
68.50
73.25
InternVL3-8B
70.00
55.00
62.50
44.00
62.00
53.00
76.00
65.50
70.75
Gemma-4-26B
74.00
74.50
74.25
58.00
77.00
67.50
78.00
78.00
78.00
Gemma-4-31B
78.00
73.50
75.75
66.00
75.50
70.75
80.00
81.00
80.50
Llama-4-Maverick
80.00
75.00
77.50
50.00
66.00
58.00
80.00
80.00
80.00
Ministral-3-8B
80.00
73.50
76.75
58.00
70.50
64.25
80.00
81.00
80.50
Ministral-3-14B
60.00
70.50
65.25
46.00
65.00
55.50
72.00
75.00
73.50
Closed-source LMMs
Gemini-3.1-Pro
64.00
59.00
61.50
18.00
27.50
22.75
50.00
54.00
52.00
GPT-4o
80.00
73.50
76.75
52.00
73.00
62.50
82.00
78.50
80.25
o1
80.00
69.50
74.75
52.00
61.00
56.50
84.00
78.00
81.00
o3
80.00
62.00
71.00
30.00
58.50
44.25
92.00
78.50
85.25
Claude-Opus-4.6
22.00
21.50
21.75
26.00
38.50
32.25
34.00
37.50
35.75
Table A8: Refusal performance on image–question mismatch samples derived from MME, MMStar, and OK-VQA. The three subsets correspond to YesNo, multiple-choice, and open-ended VQA formats, respectively. Higher scores indicate stronger capability to identify and appropriately refuse image–question mismatches.
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.