TY - GEN
T1 - Overview of Medical NLP Code Generation with FHIR for Clinical Trial Screening
AU - Tao, Liang
AU - Mi, Lan
AU - Wu, Chunxiao
AU - Wen, Dong
AU - Tang, Buzhou
AU - Zhang, Xiaoyan
AU - Zong, Hui
AU - Li, Zuofeng
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - This study presents a comprehensive evaluation of large language model based medical NLP code generation for clinical trial eligibility screening. Using 51 criteria across 19 categories and 51 case reports, we assess systems that transform natural-language rules into FHIR-compliant structured representations and executable patient-retrieval logic. The benchmark requires participants to generate FHIR Bundles rather than only end-to-end code to promote transparency, standards alignment, and interoperability. Top-performing teams adopted strategies such as intermediate-schema decoupling, dynamic few-shot prompting, iterative refinement, structured prompt engineering, and agent-based synthesis. Results demonstrate that LLMs can generate clinically meaningful, standards-compliant code, while also highlighting the need for explicit semantic layers to ensure safety and interpretability. This task provides the first large-scale evaluation of medical NLP code generation grounded in FHIR Profiles and offers a foundation for future development of verifiable, trustworthy clinical AI systems. Additional details, datasets, and evaluation materials are available at the CHIP 2025 website: http://cips-chip.org.cn/2025/eval3.
AB - This study presents a comprehensive evaluation of large language model based medical NLP code generation for clinical trial eligibility screening. Using 51 criteria across 19 categories and 51 case reports, we assess systems that transform natural-language rules into FHIR-compliant structured representations and executable patient-retrieval logic. The benchmark requires participants to generate FHIR Bundles rather than only end-to-end code to promote transparency, standards alignment, and interoperability. Top-performing teams adopted strategies such as intermediate-schema decoupling, dynamic few-shot prompting, iterative refinement, structured prompt engineering, and agent-based synthesis. Results demonstrate that LLMs can generate clinically meaningful, standards-compliant code, while also highlighting the need for explicit semantic layers to ensure safety and interpretability. This task provides the first large-scale evaluation of medical NLP code generation grounded in FHIR Profiles and offers a foundation for future development of verifiable, trustworthy clinical AI systems. Additional details, datasets, and evaluation materials are available at the CHIP 2025 website: http://cips-chip.org.cn/2025/eval3.
KW - CHIP
KW - Clinical trial recruitment
KW - Code generation
KW - FHIR
KW - Large language models
UR - https://www.scopus.com/pages/publications/105041240996
U2 - 10.1007/978-981-95-7299-1_34
DO - 10.1007/978-981-95-7299-1_34
M3 - 会议稿件
AN - SCOPUS:105041240996
SN - 9789819572984
T3 - Communications in Computer and Information Science
SP - 549
EP - 558
BT - Health Information Processing - 11th China Health Information Processing Conference, CHIP 2025, Proceedings
A2 - Zhang, Yanchun
A2 - Chen, Qingcai
A2 - Tang, Buzhou
A2 - Lin, Hongfei
A2 - Jin, Bo
A2 - Liu, Lei
A2 - Hao, Tianyong
A2 - Huang, Zhengxing
PB - Springer Science and Business Media Deutschland GmbH
T2 - 11th China Health Information Processing Conference, CHIP 2025
Y2 - 22 November 2025 through 24 November 2025
ER -