TY - GEN
T1 - Sparse Activation Editing for Reliable Instruction Following in Narratives
AU - Zhao, Runcong
AU - Cao, Chengyu
AU - Zhu, Qinglin
AU - Lv, Xiucheng
AU - Shao, Shun
AU - Gui, Lin
AU - Xu, Ruifeng
AU - He, Yulan
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Complex narrative contexts often challenge language models' ability to follow instructions, and existing benchmarks fail to capture these difficulties. To address this, we propose Concise-SAE, a training-free framework that improves instruction following by identifying and editing instruction-relevant neurons using only natural language instructions, without requiring labelled data. To thoroughly evaluate our method, we introduce FREEINSTRUCT, a diverse and realistic benchmark of 1,212 examples that highlights the challenges of instruction following in narrative-rich settings. While initially motivated by complex narratives, Concise-SAE demonstrates state-of-the-art instruction adherence across varied tasks without compromising generation quality. The data and code are available at https://github.com/Chacioc/Concise-SAE.
AB - Complex narrative contexts often challenge language models' ability to follow instructions, and existing benchmarks fail to capture these difficulties. To address this, we propose Concise-SAE, a training-free framework that improves instruction following by identifying and editing instruction-relevant neurons using only natural language instructions, without requiring labelled data. To thoroughly evaluate our method, we introduce FREEINSTRUCT, a diverse and realistic benchmark of 1,212 examples that highlights the challenges of instruction following in narrative-rich settings. While initially motivated by complex narratives, Concise-SAE demonstrates state-of-the-art instruction adherence across varied tasks without compromising generation quality. The data and code are available at https://github.com/Chacioc/Concise-SAE.
UR - https://www.scopus.com/pages/publications/105040218798
U2 - 10.18653/v1/2025.emnlp-main.1311
DO - 10.18653/v1/2025.emnlp-main.1311
M3 - 会议稿件
AN - SCOPUS:105040218798
T3 - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
SP - 25817
EP - 25832
BT - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
A2 - Christodoulopoulos, Christos
A2 - Chakraborty, Tanmoy
A2 - Rose, Carolyn
A2 - Peng, Violet
PB - Association for Computational Linguistics (ACL)
T2 - 30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Y2 - 4 November 2025 through 9 November 2025
ER -