Early chatbots were text in, text out; multimodal models add eyes and ears. The same architecture that predicts the next token handles image patches or audio frames, so a user can photograph a broken part, a receipt or a screenshot and get an answer grounded in what is actually there.
In practice this collapses pipelines that used to need OCR plus rules plus a person: document intake, damage reports, accessibility descriptions. The caveats scale with the input — images cost more tokens than they look, and a model that can read your invoice can also read what it should not have seen.
Related terms
Large language model (LLM)
A model trained on enormous amounts of text to predict what comes next, which turns out to be enough to write, summarise, translate and reason through many tasks.
OCR
Optical character recognition — turning images of text into real text, so scanned invoices, receipts and forms become data instead of pictures.
Speech to text
Transcribing spoken audio into written text with a model — the front end of voice notes, call summaries and 'talk to the computer' interfaces.
The bench this belongs to
Generative AIStable Diffusion and ComfyUI wired into the place where your content actually gets made, with a model fine-tuned so everything comes out looking like you.
