TY - GEN
T1 - Chart2Code53
T2 - 30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
AU - Niu, Tianhao
AU - Cui, Yiming
AU - Wang, Baoxin
AU - Xu, Xiao
AU - Yao, Xin
AU - Zhu, Qingfu
AU - Wu, Dayong
AU - Wang, Shijin
AU - Che, Wanxiang
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Chart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts. However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and (3) inadequate complexity. To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code. (2) Synthesize based on the chart images in the academic paper. We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline. Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-of-the-art performance on multiple Chart2Code benchmarks within open-source models.
AB - Chart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts. However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and (3) inadequate complexity. To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code. (2) Synthesize based on the chart images in the academic paper. We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline. Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-of-the-art performance on multiple Chart2Code benchmarks within open-source models.
UR - https://www.scopus.com/pages/publications/105040240488
U2 - 10.18653/v1/2025.emnlp-main.799
DO - 10.18653/v1/2025.emnlp-main.799
M3 - 会议稿件
AN - SCOPUS:105040240488
T3 - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
SP - 15828
EP - 15844
BT - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
A2 - Christodoulopoulos, Christos
A2 - Chakraborty, Tanmoy
A2 - Rose, Carolyn
A2 - Peng, Violet
PB - Association for Computational Linguistics (ACL)
Y2 - 4 November 2025 through 9 November 2025
ER -