Retrieval before generation: practical RAG for operations teams
Most of the quality in a retrieval-augmented system is decided before the model writes a word. A field guide to chunking, retrieval and evaluation for internal knowledge assistants.

Retrieval-augmented generation is often described as “give the model your documents.” In practice, the model is the easy part. The hard part is making sure the right passage reaches it, every time, for questions your users actually ask.
Start from the questions
Before choosing an embedding model or a vector database, collect fifty to a hundred real questions from the people who will use the system. Support tickets, chat logs and email threads are good sources. These questions become your evaluation set, and they reveal what kind of retrieval you need: policy lookups, procedure steps, or answers that combine several documents.
Chunk by meaning, not by size
Fixed-length chunks are simple, but they split procedures in half and separate a rule from its exceptions. Chunk along the document’s own structure — headings, sections, list items — and attach context to each chunk: the document title, section path and last-updated date. A chunk that says “Refunds are processed within 7 days” is far more useful when it also carries “Supplier policy › Airline X › Refunds.”
Use hybrid retrieval
Vector search is good at meaning; keyword search is good at exact terms like booking references, product codes and error messages. Operations questions contain both. Combining the two, then re-ranking the merged candidates, is usually a better starting point than tuning either alone.
Ground the answer and show the source
Instruct the model to answer only from retrieved passages and to say when they do not contain the answer. Return citations with every response so users can verify the source. In internal tools, a confident wrong answer costs more than “I couldn’t find this in the knowledge base.”
Evaluate retrieval separately from generation
When an answer is wrong, find out where it went wrong:
- Retrieval miss — the right passage was never retrieved.
- Ranking miss — it was retrieved but ranked too low to be used.
- Generation miss — it was present, but the model ignored or misread it.
Each failure has a different fix. Measuring retrieval recall on your question set, independent of the model, tells you whether to work on chunking and search or on prompting.
Keep the index fresh
Operational knowledge changes. Re-index on document updates, store versions, and prefer the most recent document when two conflict. A RAG system that answers from last year’s policy is worse than no system at all.
The takeaway
Treat RAG as a search problem with a language model at the end. Invest in questions, structure and evaluation first, and model choice becomes a much smaller decision.