Skip to main navigation Skip to search Skip to main content

TIANWEN: A Comprehensive Benchmark for Evaluating LLMs in Chinese Classical Poetry Understanding and Reasoning

  • Zhenwu Pei
  • , Rongbo Chen
  • , Xuefeng Bai*
  • , Kehai Chen
  • , Yingjie Zhu
  • , Andong Chen
  • , Min Zhang
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Chinese classical poetry, with its rich cultural heritage and intricate linguistic features, presents significant challenges for natural language processing systems. Despite recent advancements in Large Language Models (LLMs), their abilities in the domain of Chinese classical poetry remain largely underexplored. To bridge this gap, we introduce TianWen, a comprehensive benchmark designed to evaluate LLMs’ understanding and reasoning capabilities with respect to Chinese classical poetry. TianWen consists of 4k test instances across four understanding tasks and four reasoning tasks, incorporating varying levels of granularity and diverse task formats to enable thorough and multifaceted evaluation. We extensively evaluated 16 representative LLMs, including open-source and closed-source models on our benchmark, and the results indicate substantial room for improvement. Further analysis confirms the validity of scaling laws and highlights the challenges of Chinese classical poetry processing. We release our dataset and evaluation script at https://github.com/pzwstudy/TianWen-benchmark.

Original languageEnglish
Title of host publicationNatural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
EditorsXian-Ling Mao, Zhaochun Ren, Muyun Yang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages516-528
Number of pages13
ISBN (Print)9789819533428
DOIs
StatePublished - 2026
Externally publishedYes
Event14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025 - Urumqi, China
Duration: 7 Aug 20259 Aug 2025

Publication series

NameLecture Notes in Computer Science
Volume16102 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
Country/TerritoryChina
CityUrumqi
Period7/08/259/08/25

Keywords

  • Chinese Classical Poetry
  • Evaluation
  • LLM
  • Understanding and Reasoning

Fingerprint

Dive into the research topics of 'TIANWEN: A Comprehensive Benchmark for Evaluating LLMs in Chinese Classical Poetry Understanding and Reasoning'. Together they form a unique fingerprint.

Cite this