
Build notes3 sections · 2 min readBy Dezső Mező · Published October 3, 2026.mdRAGLLMEmbeddingsData
The demo works because the demo's data is clean. Production fails on the boring parts: chunking that splits a price from its product, retrieval that returns last year's catalogue, and a prompt that lets the model fill gaps from memory.
01Chunk around the fact, not the token count
If a product's name, price and stock land in different chunks, retrieval can hand the model the price of a different product and nothing in the prompt will notice. Chunk by document structure — one item, one chunk — and let the size vary.
02Freshness is a field, not a hope
Every chunk carries the timestamp of the source it came from, and retrieval prefers recent sources when facts conflict. A catalogue update that does not trigger a re-index is a hallucination scheduled for next week.
03The prompt's job is refusal, not eloquence
Instruct the model to answer only from retrieved context and to say so when the context does not cover the question. A visible 'I don't have that in the current catalogue' is a feature; a confidently wrong price is a liability.
What to take away
- Chunk by document structure, not by token count.
- Timestamp every chunk; re-index on every source update.
- Prompt for refusal when context is missing, not for fluency.
- Cite the chunk in the answer so wrong answers are debuggable.
More from the lab
Browse all entries
Prompts are code: eval them like code
Self-hosted n8n or Zapier — where the bill actually differs
MCP servers in production, not in the demo
The invoice pipeline that ended the Friday admin block
Want this looked at on your own system?Start a conversation
