# Speech to text

> Transcribing spoken audio into written text with a model — the front end of voice notes, call summaries and 'talk to the computer' interfaces.

Modern speech recognition is a solved-enough problem that the interesting questions moved upstream: open models like Whisper transcribe accurately across languages, accents and noisy rooms, cheaply enough to run on every call recording a business makes.

The value appears when transcription meets the rest of the pipeline: a WhatsApp voice order becomes structured intake, a sales call becomes a searchable CRM note, a dictated field report becomes the document nobody had to type. Transcription is rarely the product; it is the door the product walks through.

## Related terms

- https://dfieldsolutions.com/en/glossary/multimodal.md
- https://dfieldsolutions.com/en/glossary/ocr.md
- https://dfieldsolutions.com/en/glossary/workflow-automation.md

---

Source: https://dfieldsolutions.com/en/glossary/speech-to-text
DField Solutions — Dunakeszi, Hungary — dezso@dfieldsolutions.com
Booking: see https://dfieldsolutions.com/en/contact
