Abstract
Since the release of ChatGPT, pre-trained large language models (LLMs) have garnered worldwide attention for their extensive “general capabilities“across multiple industries and tasks. LLMs have also shed light on the construction of AI models in the civil engineering (CE) industry. To evaluate the capability boundaries of pre-trained LLMs, researchers have proposed a series of evaluation datasets in general domains to test the fundamental abilities of these models. However, these datasets do not include knowledge specific to CE. Moreover, the performance of large models in specialized domains still lags significantly behind that of human experts. In particular, within the CE domain, where domain-specific LLMs research is still in its infancy, universally recognized domain-specific LLM models have not been proposed, even for basic tasks, such as understanding and reasoning with single-modal textual data. Therefore, this study focused on evaluating the fundamental capabilities of large industry-specific models in professional domains. A Chinese evaluation dataset for large models in the CE domain, Civil- Eval, was constructed based on the national-level professional certification examinations of the CE industry. This dataset includes "26 single-choice questions and 91 multiple-choice questions across eight participants, which were organized and verified through manual screening and review. Ten representative LLMs, along with the reasoning model OpenAI-o1/DeepSeek-R1 and CE knowledge model CivllGPT, were evaluated using this dataset. The results indicate that general-purpose LLMs, without training and fine-tuning of CE-specific corpora, still exhibit significant gaps in industry applications compared with human experts. In terms of subject performance, LLMs demonstrate stronger language comprehension for common-sense questions, but perform poorly on tasks requiring mathematical reasoning and multiple-choice questions. The reasoning model OpenA-o1 outperforms in both easy and complex questions through built-in chain-of-thought techniques, and CivilGPT only shows improvements for simple questions. The evaluation concludes that the future development of domain-specific LLMs should be pre-trained and fine-tuned based on high-quality foundational general-purpose LLMs with strong reasoning and multi-modal capabilities. Additionally, high-quality chain-of-thought annotations and expert feedback datasets are crucial for domain-specific LLMs.
| Translated title of the contribution | 基于 Civil-Eval 的土木和交通工程 大模型中文测评方法 |
|---|---|
| Original language | English |
| Pages (from-to) | 148-161 |
| Number of pages | 14 |
| Journal | Zhongguo Gonglu Xuebao/China Journal of Highway and Transport |
| Volume | 39 |
| Issue number | 1 |
| DOIs | |
| State | Published - 17 Jun 2025 |
Keywords
- artificial intelligence
- evaluation method
- instruction fine-tuning
- large language model for the civil and transportation engineering
- pre-trained large language model
- transportation engineering
- 交通工程
- 人工智能
- 土木和交通工程行业大模型
- 指令微调
- 评测方法
- 预训练大语言模型
Fingerprint
Dive into the research topics of 'Evaluation Method for LLMs of Civil and Transportation Engineering Based on the Civil-Eval Benchmark Dataset'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver