Models do not see text, they see tokens: common words become one token, rare words split into several, and a hundred tokens make roughly seventy-five English words. Every model call is metered in tokens in and tokens out, which is why a chatbot that drags a long document into every message gets expensive fast.
The practical consequence is budgeting: prompt length, conversation history and the answer itself all share the same context window, and pricing is per million tokens, not per request. Techniques like RAG exist largely to spend tokens only on the fragments that matter.
Related terms
Context window
The maximum amount of text a model can take into account at once, counting the instructions, the conversation and the answer together.
RAG (retrieval-augmented generation)
Looking up relevant documents first and putting them in the prompt, so the model answers from your material instead of from memory.
Large language model (LLM)
A model trained on enormous amounts of text to predict what comes next, which turns out to be enough to write, summarise, translate and reason through many tasks.
The bench this belongs to
AI automationThe repetitive half of your week, handed to software that does not get bored. Inbox triage, follow-ups, reporting, data entry between tools that were never meant to talk.
