TY - GEN
T1 - An Unsupervised Method for Sarcasm Detection with Prompts
AU - Lin, Qihui
AU - Lou, Chenwei
AU - Liang, Bin
AU - Wang, Qianlong
AU - Wen, Zhiyuan
AU - Mao, Ruibin
AU - Xu, Ruifeng
N1 - Publisher Copyright:
© 2024, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2024
Y1 - 2024
N2 - Sarcasm detection is challenging in natural language processing since its peculiar linguistic expression. Thanks in part to the availability of considerable annotated resources for some datasets, current supervised learning-based approaches can achieve promising performance in sarcasm detection. In real-world scenarios, annotating data for the peculiar language expression of sarcasm proves challenging. Consequently, recent studies have delved into unsupervised learning approaches for sarcasm detection, seeking to mitigate the labor-intensive process of annotation. In this paper, we present a novel unsupervised sarcasm detection method leveraging abundant unlabeled social media data. Our approach revolves around employing prompts as a cornerstone. Initially, we gathered approximately 3 million texts from Twitter through targeted hashtag-based searches, segregating them into sarcasm and non-sarcasm categories based on associated hashtags. Subsequently, these collected texts undergo training using a pre-trained BERT model, customized for masked language modeling and coined as SarcasmBERT. This step aims to enhance the model’s grasp of sarcastic cues within the text. Finally, we devise prompts tailored for the unlabeled data to execute unsupervised sarcasm detection effectively. Our experimental findings across six benchmark datasets highlight the superiority of our method over state-of-the-art unsupervised baselines. Additionally, the integration of our SarcasmBERT into established BERT-based sarcasm detection methods showcases a direct avenue for enhancing performance, thereby illustrating its potential for immediate and substantial improvements.
AB - Sarcasm detection is challenging in natural language processing since its peculiar linguistic expression. Thanks in part to the availability of considerable annotated resources for some datasets, current supervised learning-based approaches can achieve promising performance in sarcasm detection. In real-world scenarios, annotating data for the peculiar language expression of sarcasm proves challenging. Consequently, recent studies have delved into unsupervised learning approaches for sarcasm detection, seeking to mitigate the labor-intensive process of annotation. In this paper, we present a novel unsupervised sarcasm detection method leveraging abundant unlabeled social media data. Our approach revolves around employing prompts as a cornerstone. Initially, we gathered approximately 3 million texts from Twitter through targeted hashtag-based searches, segregating them into sarcasm and non-sarcasm categories based on associated hashtags. Subsequently, these collected texts undergo training using a pre-trained BERT model, customized for masked language modeling and coined as SarcasmBERT. This step aims to enhance the model’s grasp of sarcastic cues within the text. Finally, we devise prompts tailored for the unlabeled data to execute unsupervised sarcasm detection effectively. Our experimental findings across six benchmark datasets highlight the superiority of our method over state-of-the-art unsupervised baselines. Additionally, the integration of our SarcasmBERT into established BERT-based sarcasm detection methods showcases a direct avenue for enhancing performance, thereby illustrating its potential for immediate and substantial improvements.
KW - pre-trained language model
KW - prompt
KW - sentiment analysis
KW - unsupervised sarcasm detection
UR - https://www.scopus.com/pages/publications/85181977734
U2 - 10.1007/978-3-031-51671-9_3
DO - 10.1007/978-3-031-51671-9_3
M3 - 会议稿件
AN - SCOPUS:85181977734
SN - 9783031516702
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 34
EP - 46
BT - Cognitive Computing – ICCC 2023 - 7th International Conference Held as Part of the Services Conference Federation, SCF 2023, Proceedings
A2 - Pan, Xiuqin
A2 - Jin, Ting
A2 - Zhang, Liang-Jie
PB - Springer Science and Business Media Deutschland GmbH
T2 - 7th International Conference on Cognitive Computing, ICCC 2023, Held as Part of the Services Conference Federation, SCF 2023
Y2 - 17 December 2023 through 18 December 2023
ER -