Skip to main navigation Skip to search Skip to main content

Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation

  • Tianhao Niu
  • , Yiming Cui
  • , Baoxin Wang
  • , Xiao Xu
  • , Xin Yao
  • , Qingfu Zhu*
  • , Dayong Wu
  • , Shijin Wang
  • , Wanxiang Che*
  • *Corresponding author for this work
  • Research Center for Social Computing and Interactive Robotics
  • Harbin Institute of Technology
  • IFLYTEK Co., Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Chart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts. However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and (3) inadequate complexity. To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code. (2) Synthesize based on the chart images in the academic paper. We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline. Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-of-the-art performance on multiple Chart2Code benchmarks within open-source models.

Original languageEnglish
Title of host publicationEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
EditorsChristos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
PublisherAssociation for Computational Linguistics (ACL)
Pages15828-15844
Number of pages17
ISBN (Electronic)9798891763326
DOIs
StatePublished - 2025
Event30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - Suzhou, China
Duration: 4 Nov 20259 Nov 2025

Publication series

NameEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference

Conference

Conference30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Country/TerritoryChina
CitySuzhou
Period4/11/259/11/25

Fingerprint

Dive into the research topics of 'Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation'. Together they form a unique fingerprint.

Cite this