# Multimodal model

> A model that takes and produces more than text — images, audio, documents — so 'what is wrong in this photo' or 'read this invoice' is one call.

Early chatbots were text in, text out; multimodal models add eyes and ears. The same architecture that predicts the next token handles image patches or audio frames, so a user can photograph a broken part, a receipt or a screenshot and get an answer grounded in what is actually there.

In practice this collapses pipelines that used to need OCR plus rules plus a person: document intake, damage reports, accessibility descriptions. The caveats scale with the input — images cost more tokens than they look, and a model that can read your invoice can also read what it should not have seen.

## Related terms

- https://dfieldsolutions.com/en/glossary/llm.md
- https://dfieldsolutions.com/en/glossary/ocr.md
- https://dfieldsolutions.com/en/glossary/speech-to-text.md

---

Source: https://dfieldsolutions.com/en/glossary/multimodal
DField Solutions — Dunakeszi, Hungary — dezso@dfieldsolutions.com
Booking: see https://dfieldsolutions.com/en/contact
