GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection

Published in 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026), 2026

TL;DR This paper proposes GDCNet for multimodal sarcasm detection, using LLM-generated image-grounded captions to anchor cross-modal semantics and capture image–text incongruity. It fuses discrepancy and multimodal features via gated fusion, achieving state-of-the-art results on MMSD2.0.

Abstract Multimodal sarcasm detection (MSD) aims to identify sarcasm within image–text pairs by modeling semantic incongruities across modalities. Existing methods often exploit cross-modal embedding misalignment to detect inconsistency, but struggle when visual and textual content are loosely related or semantically indirect. Recent approaches leverage large language models (LLMs) to generate sarcastic cues or explanations, but the diverse perspectives of the generated sarcastic text from LLMs may amplify noise. To remedy these issues, we propose the Generative Discrepancy Comparison Network (GDCNet), a simple yet effective framework that captures cross-modal conflicts by introducing descriptive and factually grounded image captions generated by LLMs as stable semantic anchors. Specifically, GDCNet first employs a multi-modal large language model to generate factual, image-consistent textual descriptions. Then, it computes semantic and sentiment-level discrepancies between the generated description and the original text, while also measuring the visual-textual consistency. These discrepancies are fused with visual and textual features via a gated module to adaptively balance modality contributions. Extensive experiments on MSD benchmarks validate GDCNet’s substantial gains in accuracy and robustness, setting a new state-of-the-art on the MMSD2.0 benchmark. Our framework implementation and code are available at https://anonymous.4open.science/r/GDCNet.

Recommended citation: Zhang, Shuguang, et al. "GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection." Proceedings of the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026). 2026.

Show citation
@inproceedings{zhang2026gdcnet,
  title = {GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection},
  author = {Shuguang Zhang and Junhong Lian and Guoxin Yu and Baoxun Xu and Xiang Ao},
  year = {2026},
  booktitle = {ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  publisher = {IEEE},
  pages = {13192--13196},
  doi = {10.1109/icassp55912.2026.11461651},
  url = { https://doi.org/10.1109/icassp55912.2026.11461651 }
}

Leave a Comment