A model does not emit one answer, it emits a probability for every possible next token. Temperature reshapes those probabilities before one is sampled: near zero the model almost always takes the front-runner, near one it explores. Same model, same prompt, different personality.
The setting should follow the task. Support answers, extraction and anything a user will fact-check want low temperature; brainstorms and first drafts can take more. What temperature never does is create knowledge — a high setting makes a wrong answer more creative, not more right.
Related terms
Large language model (LLM)
A model trained on enormous amounts of text to predict what comes next, which turns out to be enough to write, summarise, translate and reason through many tasks.
Hallucination
A confident, fluent, entirely invented answer — a citation, a figure or an API that does not exist.
System prompt
The standing instructions sent ahead of every conversation, setting what the model is supposed to do and how it should behave.
The bench this belongs to
Generative AIStable Diffusion and ComfyUI wired into the place where your content actually gets made, with a model fine-tuned so everything comes out looking like you.
