A retrieval-augmented generation demo takes an afternoon. Point a model at a folder of documents, ask a question, get a plausible answer. That speed is exactly why so many enterprise RAG projects stall: the demo hides every decision that matters in production.
When we onboard AI for a client, these are the five decisions we settle before a pilot is allowed near real users.
1. Where the data lives, and where it never goes
The first question is not which model to use. It is which boundary your data must stay inside. For most enterprises that means the vector store, the embedding process and the orchestration layer run within your own cloud tenancy, under your existing identity and network controls. Decide this first, because it constrains every other choice.
2. Permissions travel with the content
Your document systems already know who can read what. A RAG system that ignores those permissions will happily summarize a board memo for an intern. Access controls must be captured at ingestion and enforced at retrieval time, so a user's question only searches the content they are entitled to see.
If retrieval does not respect permissions, the system is a data leak with a chat interface.
3. Ingestion is a pipeline, not a one-time load
Knowledge changes daily. Policies are revised, tickets are closed, products are renamed. Treat ingestion as a production pipeline with scheduling, change detection, deletion handling and monitoring. Stale or orphaned content is the most common source of confident wrong answers.
4. Measure quality before you ship
Without an evaluation set, every prompt change is a guess. Build a representative set of real questions with known good answers, then score retrieval relevance and answer accuracy on every release. Add guardrails for sensitive data, off-topic requests and unsupported claims, and log enough to trace any answer back to its sources.
5. Plan for cost and latency on day one
A pilot with ten users hides costs that a rollout to ten thousand will not. Track tokens, latency and cache rates per use case from the first deployment. Right-sizing models, trimming context and caching frequent retrievals often cut spend substantially without hurting quality, but only if you are measuring.
From pilot to production
None of these decisions are exotic. They are the same disciplines that make any enterprise system trustworthy: clear boundaries, enforced access, reliable pipelines, measurable quality and managed cost. Settling them early is what lets a working prototype in week five become a production system your security team is comfortable signing off.