# Transformer

> The neural network architecture behind modern language models — attention layers that weigh which earlier words matter when predicting the next one.

Introduced in the 2017 paper 'Attention Is All You Need', the transformer replaced sequential reading with attention: every token looks at every earlier token at once and decides which are relevant. That parallelism is what made training on internet-scale text feasible.

Almost everything the industry calls an LLM — the GPT series, Claude, Gemini, Llama — is a transformer trained to predict the next token. The architecture's limits are the field's limits: a fixed context window, no built-in memory, and cost that grows with length.

## Related terms

- https://dfieldsolutions.com/en/glossary/llm.md
- https://dfieldsolutions.com/en/glossary/context-window.md
- https://dfieldsolutions.com/en/glossary/token.md

---

Source: https://dfieldsolutions.com/en/glossary/transformer
DField Solutions — Dunakeszi, Hungary — dezso@dfieldsolutions.com
Booking: see https://dfieldsolutions.com/en/contact
