Everything the model considers has to fit in this one budget: the system prompt, any documents you retrieved, the whole conversation so far and the response it is about to write. When a long chat starts forgetting what was said at the beginning, the usual reason is that the beginning was dropped to make room.
Bigger windows have made a class of workarounds unnecessary, but not the discipline. Filling a large window with everything available costs money on every call, adds latency, and measurably degrades accuracy — models attend less reliably to material buried in the middle of a very long context. Sending the right five pages still beats sending all two hundred.
Related terms
Large language model (LLM)
A model trained on enormous amounts of text to predict what comes next, which turns out to be enough to write, summarise, translate and reason through many tasks.
RAG (retrieval-augmented generation)
Looking up relevant documents first and putting them in the prompt, so the model answers from your material instead of from memory.
System prompt
The standing instructions sent ahead of every conversation, setting what the model is supposed to do and how it should behave.
The bench this belongs to
AI automationThe repetitive half of your week, handed to software that does not get bored. Inbox triage, follow-ups, reporting, data entry between tools that were never meant to talk.
