The training objective is unglamorous: given a stretch of text, predict the next piece. Done at sufficient scale, the model has to internalise grammar, facts, argument structure and a great deal about how people write in order to predict well, and those capabilities are what you use when you ask it to do something.
Two consequences follow from that and explain most surprises. The model has no separate store of facts it can check, so a plausible-sounding wrong answer is produced by exactly the same machinery as a right one. And it has no memory between calls: everything it knows about your conversation was sent to it again, as text, on every single request.
Related terms
Context window
The maximum amount of text a model can take into account at once, counting the instructions, the conversation and the answer together.
Hallucination
A confident, fluent, entirely invented answer — a citation, a figure or an API that does not exist.
Fine-tuning
Continuing to train an existing model on your own examples, so it adopts a behaviour or a style it did not have.
Embedding
A list of numbers representing a piece of text, arranged so that texts about similar things end up close together.
The bench this belongs to
AI automationThe repetitive half of your week, handed to software that does not get bored. Inbox triage, follow-ups, reporting, data entry between tools that were never meant to talk.
