DFIELDSOLUTIONS

AI and language models

GlossaryTransformer

The neural network architecture behind modern language models — attention layers that weigh which earlier words matter when predicting the next one.

Introduced in the 2017 paper 'Attention Is All You Need', the transformer replaced sequential reading with attention: every token looks at every earlier token at once and decides which are relevant. That parallelism is what made training on internet-scale text feasible.

Almost everything the industry calls an LLM — the GPT series, Claude, Gemini, Llama — is a transformer trained to predict the next token. The architecture's limits are the field's limits: a fixed context window, no built-in memory, and cost that grows with length.

Related terms

All termsStart a conversationMarkdown version

DField Bt. · Dunakeszi · dezso@dfieldsolutions.com
5.0
“From LinkedIn DM to live site. Two tiny tweaks, then shipped.”Michael J Ringer · Vilya ProtectionFounder · Spain