Context Window

A context window is the maximum number of tokens a language model can process in a single request, counting both the input and the output it generates. Everything the model uses for an answer, including the system prompt, tool definitions, retrieved documents, and conversation history, has to fit inside this limit, and text beyond it is truncated or never sent. The context window differs from memory in the everyday sense, because nothing persists between requests unless the application sends it again, and from training data, which shapes the model's weights but is not visible as text at inference time. In agentic systems, the window fills with tool results and intermediate steps over a long task, so its size sets how much work an agent can do before history must be compacted or offloaded.

Window size marks an upper bound on what the model can read, and models do not use all of it equally well. Studies have shown that models retrieve information less reliably from the middle of a long context, and that accuracy tends to decline as input length grows, an effect described as context rot. Longer inputs also raise latency and cost per request, so retrieval, summarization, and careful selection of what enters the context usually give better results than filling the window to its limit.

Work with us

Ready to put agentic AI to work?

Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.