Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection
文字和图片表面上看起来很搭,实则暗藏矛盾,这就是讽刺,这套AI能识别出来
识别图文帖子中讽刺意味的AI系统常常会被表面一致但实际含义相反的文字图片组合迷惑。这篇论文提出一个框架,能针对每条帖子动态判断文字和图片哪个更关键,并引入一种对比学习方法,把讽刺样本里的表面一致性当作陷阱而非真实证据来处理。在MMSD和MMSD2.0两个基准数据集上,该方法的表现持续超过现有的强基线模型。
METAL MEDIA 解读图
文字和图片表面上看起来很搭,实则暗藏矛盾,这就是讽刺,这套AI能识别出来
- 01由于有些讽刺帖子主要靠文字线索,有些则靠图片矛盾,作者设计了一个动态门控融合模块,对文字和图片这两种模态的信息进行双向过滤,并针对每个样本单独调整两者的权重
- 02讽刺类图文对常常字面上看起来一致,实际含义却相反,为此论文提出了SaCR(讽刺感知对比正则化),对非讽刺样本提高图文相似度、对讽刺样本则压低相似度,防止模型被表面一致性误导
- 03使用预训练的视觉-语言模型CLIP提取文字和图像特征,再通过双向交叉注意力机制在数值层面加门控,过滤掉无用或有误导性的跨模态信号
- 04整个模型通过多目标损失进行端到端训练,同时优化最终分类损失、文本/图像单模态辅助分类损失以及SaCR对比损失
- 05在MMSD和MMSD2.0两个数据集上,该方法的F1分数均优于对比的基线模型;消融实验显示,去掉跨模态交互模块或动态融合门会导致性能下降最明显
他们做了什么
- 由于有些讽刺帖子主要靠文字线索,有些则靠图片矛盾,作者设计了一个动态门控融合模块,对文字和图片这两种模态的信息进行双向过滤,并针对每个样本单独调整两者的权重
- 讽刺类图文对常常字面上看起来一致,实际含义却相反,为此论文提出了SaCR(讽刺感知对比正则化),对非讽刺样本提高图文相似度、对讽刺样本则压低相似度,防止模型被表面一致性误导
- 使用预训练的视觉-语言模型CLIP提取文字和图像特征,再通过双向交叉注意力机制在数值层面加门控,过滤掉无用或有误导性的跨模态信号
- 整个模型通过多目标损失进行端到端训练,同时优化最终分类损失、文本/图像单模态辅助分类损失以及SaCR对比损失
- 在MMSD和MMSD2.0两个数据集上,该方法的F1分数均优于对比的基线模型;消融实验显示,去掉跨模态交互模块或动态融合门会导致性能下降最明显

| Dataset | Split | Total | Sarcastic | Non-sarcastic |
|---|---|---|---|---|
| MMSD | Train | 19,816 | 8,642 | 11,174 |
| Val | 2,410 | 959 | 1,451 | |
| Test | 2,409 | 959 | 1,450 | |
| MMSD2.0 | Train | 19,816 | 9,576 | 10,240 |
| Val | 2,410 | 1,042 | 1,368 | |
| Test | 2,409 | 1,037 | 1,372 |

| Modality | Method | MMSD | MMSD2.0 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Acc.% | P% | R% | F1% | Acc.% | P% | R% | F1% | ||
| Text | TextCNN | 80.03 | 74.29 | 76.39 | 75.32 | 71.61 | 64.62 | 75.22 | 69.52 |
| SMSD | 80.90 | 76.46 | 75.18 | 75.82 | 73.56 | 68.45 | 71.55 | 69.97 | |
| BERT | 83.60 | 78.50 | 82.51 | 80.45 | 76.52 | 74.48 | 73.09 | 73.91 | |
| Image | ResNet | 64.76 | 54.41 | 70.80 | 61.53 | 65.50 | 61.17 | 54.39 | 57.58 |
| ViT | 67.83 | 57.93 | 70.07 | 63.40 | 72.02 | 65.26 | 74.83 | 69.72 | |
| Multimodal | DIP | 89.59 | 87.76 | 86.58 | 87.17 | 80.96 | 78.02 | 77.56 | 77.79 |
| Multi-view CLIP | 88.33 | 82.66 | 88.65 | 85.55 | 85.64 | 80.33 | 88.24 | 84.10 | |
| MoBA | 88.96 | 82.84 | 88.12 | 85.40 | 85.83 | 80.42 | 88.67 | 84.34 | |
| G2SAM | 90.48 | 87.95 | 89.02 | 88.48 | 79.43 | 72.04 | 78.07 | 78.07 | |
| TFCD | 89.57 | 84.83 | 89.43 | 88.13 | 86.54 | 82.46 | 87.95 | 84.31 | |
| DGLF | 89.43 | 85.81 | 89.27 | 87.51 | 86.82 | 81.90 | 89.85 | 85.69 | |
| LLaVA+RAG | 89.97 | 89.26 | 89.58 | 89.42 | 86.43 | 87.00 | 86.30 | 86.34 | |
| ESAM | 90.11 | 86.87 | 89.54 | 88.19 | 85.87 | 83.12 | 86.05 | 84.56 | |
| GPT-5.4 (zeroshot) | 71.05 | 76.51 | 75.50 | 71.01 | 72.85 | 78.79 | 75.78 | 72.55 | |
| Ours | 92.62 | 91.96 | 92.82 | 92.33 | 89.66 | 89.36 | 89.74 | 89.51 |

| Variant | MMSD | MMSD2.0 | ||
|---|---|---|---|---|
| Acc.% | F1% | Acc.% | F1% | |
| Full | 92.62 | 92.33 | 89.66 | 89.36 |
| w/o CMI | 87.91 | 87.42 | 85.10 | 84.97 |
| w/o BiXAtt (only v→t) | 87.48 | 87.03 | 81.86 | 81.70 |
| w/o BiXAtt (only t→v) | 81.84 | 81.07 | 80.53 | 80.37 |
| w/o VGate | 92.08 | 91.77 | 88.71 | 88.59 |
| w/o DFGate | 90.35 | 89.80 | 88.34 | 88.28 |
| w/o SaCR | 91.28 | 91.00 | 88.54 | 88.49 |
| w/o UniAux | 91.31 | 90.92 | 89.16 | 89.03 |
为什么重要
任何需要从社交媒体帖子中读取语气或情感的系统(如内容审核、舆情分析、聊天机器人应答)如果漏掉被表面一致文字图片掩盖的讽刺,就可能完全误判说话人的真实意图,因此专门针对这一失效模式的方法具有直接的实用价值。文字图片表面一致却可能掩盖矛盾意图这一发现,也可能对讽刺检测之外的其他多模态理解任务有参考意义。
本文术语
- 多模态讽刺检测(MSD) · 通过同时分析文字、图片等不同形式的内容来判断说话者是否在讽刺
- CLIP · 一种预训练的视觉-语言模型,能把文字和图像映射到同一个可比较的空间
- 门控(gate) · 一种可学习的机制,用0到1之间的数值控制信号通过的多少
- 对比正则化 · 一种辅助训练方法,让相似的表示彼此靠近,不相似的表示彼此远离
- Grad-CAM · 一种可视化技术,能显示模型做判断时主要依据图像的哪些区域
论文原文摘要(英文)
Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challenging due to instance-dependent modality contributions and misleading semantic consistency, where surface-level alignment masks underlying contradictory intent. Existing methods often rely on fixed fusion strategies and treat sarcasm as generic cross-modal mismatch, limiting their ability to capture subtle sarcasm cues and instance-specific modality interactions. To address these challenges, we propose a novel MSD framework that integrates Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization (SaCR). Specifically, a bidirectional gated interaction module performs cross-modal feature filtering and adaptively calibrates textual and visual contributions at the instance level. A dynamic fusion gate further balances modality importance to generate more robust multimodal representations. Furthermore, SaCR is introduced as a label-aware contrastive regularization objective that encourages semantic consistency for non-sarcastic samples while suppressing misleading consistency in sarcastic cases. The proposed framework is trained end-to-end with a multi-objective learning strategy that jointly optimizes multimodal classification and auxiliary unimodal supervision. Extensive experiments on MMSD and MMSD2.0 demonstrate that the proposed method consistently outperforms strong baselines.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Hao Guo et al., arXiv:2608.19942, arxiv-nonexclusive