Skip to main navigation Skip to search Skip to main content

Understanding Large Language Model Performance in Software Engineering: A Large-scale Question Answering Benchmark

  • Ruida Hu
  • , Chao Peng
  • , Jingyi Ren
  • , Bo Jiang
  • , Xiangxin Meng
  • , Qinyun Wu
  • , Pengfei Gao
  • , Xinchen Wang
  • , Cuiyun Gao*
  • *Corresponding author for this work
  • Haribin Institute of Technology
  • ByteDance Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this work, we introduce CodeRepoQA, a large-scale benchmark specifically designed for evaluating repository-level question-answering capabilities in the field of software engineering. CodeRepoQA encompasses five programming languages and covers a wide range of scenarios, enabling comprehensive evaluation of language models. To construct this dataset, we crawl data from 30 well-known repositories in GitHub, the largest platform for hosting and collaborating on code, and carefully filter the raw data. In total, CodeRepoQA is a multi-turn question-answering benchmark with 585,687 entries. It covers a diverse array of software engineering scenarios, with an average of 6.62 dialogue turns per entry. We evaluate ten popular large language models on our dataset and provide in-depth analysis. We find that LLMs still have limitations in question-answering capabilities in the field of software engineering, and medium-length contexts are more conducive to their performance.

Original languageEnglish
Title of host publicationSIGIR 2025 - Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval
PublisherAssociation for Computing Machinery, Inc
Pages3025-3029
Number of pages5
ISBN (Electronic)9798400715921
DOIs
StatePublished - 13 Jul 2025
Externally publishedYes
Event48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025 - Padua, Italy
Duration: 13 Jul 202518 Jul 2025

Publication series

NameSIGIR 2025 - Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval

Conference

Conference48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025
Country/TerritoryItaly
CityPadua
Period13/07/2518/07/25

Keywords

  • Language Model
  • Mining Software Repository
  • Question Answering

Fingerprint

Dive into the research topics of 'Understanding Large Language Model Performance in Software Engineering: A Large-scale Question Answering Benchmark'. Together they form a unique fingerprint.

Cite this