Skip to main navigation Skip to search Skip to main content

Data Foundations of Long-Context Language Models: A Survey

  • Zechen Sun
  • , Yuyang Sun
  • , Zhaochen Su
  • , Zecheng Tang
  • , Juntao Li*
  • , Ao Zhou
  • , Wenliang Chen
  • , Min Zhang
  • *Corresponding author for this work
  • Institute of Computer Science and Technology
  • Soochow University
  • Hong Kong University of Science and Technology
  • Nanjing University

Research output: Contribution to journalArticlepeer-review

Abstract

As the context window of Large Language Models (LLMs) continues to expand, the data required to effectively train and evaluate these capabilities remains underexplored. With existing research primarily focuses on architectural optimization, there is a need for a systematic, data-centric review. This survey bridges this gap by investigating the data foundations of Long-Context Language Models (LCMs). We begin by examining current data strategies alongside their strengths and limitations, mapping the required data to desired model capabilities. Building on this, we explore how targeted training data designs drive core, often interconnected skills such as retrieval, reasoning, and aggregation. Furthermore, we analyze the evaluation landscape, illustrating how selecting appropriate benchmarks is crucial for probing capability boundaries and guiding effective model selection. Finally, we synthesize actionable guidelines for data construction and outline critical future directions to propel the advancement of long-context language models, including quantifying data quality, establishing scaling laws for length distributions, and developing dynamic evaluation frameworks.

Original languageEnglish
Pages (from-to)1803-1825
Number of pages23
JournalTransactions of the Association for Computational Linguistics
Volume14
DOIs
StatePublished - 2026
Externally publishedYes

Fingerprint

Dive into the research topics of 'Data Foundations of Long-Context Language Models: A Survey'. Together they form a unique fingerprint.

Cite this