Introduced in the 2017 paper 'Attention Is All You Need', the transformer replaced sequential reading with attention: every token looks at every earlier token at once and decides which are relevant. That parallelism is what made training on internet-scale text feasible.
Almost everything the industry calls an LLM — the GPT series, Claude, Gemini, Llama — is a transformer trained to predict the next token. The architecture's limits are the field's limits: a fixed context window, no built-in memory, and cost that grows with length.
Related terms
Large language model (LLM)
A model trained on enormous amounts of text to predict what comes next, which turns out to be enough to write, summarise, translate and reason through many tasks.
Context window
The maximum amount of text a model can take into account at once, counting the instructions, the conversation and the answer together.
Token
The unit a language model reads and writes — a word fragment, word or punctuation mark — and the unit its cost and context limits are counted in.
