Production RAG architecture¶
Intermediate · 20 min read · diagrams + worked examples
The RAG chatbot template is about 60 lines of Python. A RAG system serving thousands of people at a company needs more parts — and this page explains every part, why it exists, and how to choose between the options.
Before you read
You'll get the most out of this if you've built the RAG chatbot template or know what chunks, retrieval and prompts are. Every other term is explained as it appears.
What you'll learn
- What changes between a demo and a production RAG system
- The two paths every RAG system has — and each component on them
- How hybrid search and re-ranking work (with a worked example and code)
- How to size a system: chunks, storage and tokens
- How to keep it secure, measurable and affordable
1. From demo to production — what changes?¶
| Demo (the template) | Production | |
|---|---|---|
| Documents | A few files in a folder | Thousands to millions, from Drive, Confluence, websites, databases |
| Freshness | Re-read everything at start-up | Documents change daily — only re-process what changed |
| Users | Just you | Many people, each allowed to see different documents |
| Search | Keyword search | Keyword and meaning-based search, then a re-ranker |
| Quality | "Looks right" | Measured on a fixed test set before every release |
| Cost & speed | Doesn't matter | Budgeted per question; answers stream in under a few seconds |
| Failures | Crash and restart | Retries, fallbacks, alerts — and logs to explain bad answers |
Everything on the rest of this page exists to solve one of the rows in this table.
2. The idea in plain English¶
Think of a company library with a research assistant:
- Librarians keep the shelves up to date — new books arrive, old editions are removed. (the offline ingestion path)
- The catalogue lets you find books by exact title and by topic. (keyword index + vector index)
- When someone asks a question, the assistant pulls a pile of likely books, picks the best few pages, and writes an answer quoting those pages. (retrieval → re-ranking → generation with citations)
- Some shelves are restricted — the assistant only uses books you are allowed to read. (permissions)
3. The big picture¶
Every RAG system has two paths that meet at the indexes.
Offline path — runs whenever documents change
flowchart LR
S[Sources<br/>docs · web · DB] --> P[Parse & clean] --> C[Chunk + metadata]
C --> E[Embed] --> V[(Vector index)]
C --> K[(Keyword index<br/>BM25)]
Online path — runs for every question
flowchart LR
U([User]) --> G[Gateway<br/>auth · rate limit] --> Q[Query rewrite] --> R[Hybrid retrieve]
I[(Vector + BM25<br/>indexes)] -.-> R
R --> RR[Re-rank] --> PB[Prompt builder] --> L[LLM] --> GR[Guardrails<br/>+ citations] --> A([Answer])
The next two sections walk through each path one step at a time.
4. Walkthrough: follow one document in¶
Imagine HR updates leave-policy.pdf in Google Drive. Here's what happens to it.
Step 1 — Detect the change. A connector checks the source on a schedule (or receives a webhook) and notices the file's modified time or content hash changed. Only this file is queued — nothing else is re-processed.
Step 2 — Parse and clean. The PDF becomes text. Headers, footers, page numbers and cookie banners are removed; headings and tables are kept, because they carry meaning.
Step 3 — Chunk. The text is split into pieces of roughly 200–500 words, following the document's headings where possible. Each chunk gets metadata:
{
"chunk_id": "leave-policy.pdf#4",
"text": "Employees get 24 days of paid leave per year...",
"source": "drive://HR/leave-policy.pdf",
"section": "Annual leave",
"allowed_groups": ["all-employees"],
"updated_at": "2026-09-30T10:15:00Z"
}
Step 4 — Embed. Each chunk's text is turned into an embedding — a list of numbers (for example 1,536 of them) that captures its meaning. Texts about similar topics get similar numbers.
Step 5 — Index. The embedding goes into the vector index; the text goes into the
keyword (BM25) index. The old chunks of leave-policy.pdf are deleted from both — otherwise
the bot would quote the outdated policy.
Run ingestion as a queue
Put one message per changed document on a queue (e.g. a database table, Redis, SQS) and let workers process them. A broken file then retries on its own instead of blocking everything else.
5. Walkthrough: follow one question out¶
An employee asks: "How many leave days can I carry over to next year?"
| # | Step | What happens | Typical time |
|---|---|---|---|
| 1 | Gateway | Checks who the user is, applies rate limits, attaches their groups (e.g. all-employees, engineering) |
~10 ms |
| 2 | Query rewrite (optional) | If it's a follow-up ("what about interns?"), an LLM rewrites it into a full question using the chat history | 200–500 ms |
| 3 | Hybrid retrieve | Runs keyword and vector search in parallel, filtered to documents the user may see, then merges the two result lists | 50–150 ms |
| 4 | Re-rank | A re-ranking model reads the top ~30 candidates next to the question and keeps the best 4–8 | 100–300 ms |
| 5 | Prompt builder | Puts instructions + numbered chunks + question together, within a token budget | ~1 ms |
| 6 | LLM | Generates the answer, streaming it word by word | first words in 300–800 ms |
| 7 | Guardrails | Checks citations point to real chunks, redacts personal data, refuses if the context didn't contain the answer | ~10 ms |
The user sees text appearing in under a second — streaming hides most of the remaining time. (Timings are typical ranges; measure your own.)
6. The components, one by one¶
For each part: what it does → why it's needed → your options → how to choose.
6.1 Connectors & sync¶
- What: pull documents from where they live (Google Drive, SharePoint, Confluence, Notion, websites, databases).
- Why: documents change constantly; the index must follow without a full rebuild.
- Options: scheduled polling · webhooks from the source · nightly full sync as a safety net.
- Choose: webhooks where available for speed, plus a periodic full comparison to catch anything missed — including deletions.
6.2 Parsing & cleaning¶
- What: turn PDFs, Word files and HTML into clean text, keeping headings and tables.
- Why: this caps your quality — if a table is scrambled here, no model can answer from it later.
- Options: simple text extraction (
pypdf,python-docx) · layout-aware parsers for complex PDFs · OCR for scans. - Choose: start simple; switch to layout-aware parsing for documents with tables and multi-column pages.
6.3 Chunking¶
- What: split text into retrievable pieces, each with metadata (source, section, permissions, date).
- Why: search works on chunks — too big and they mix topics; too small and they lose context.
- Options: fixed size with overlap · split by headings (structure-aware) · split where the topic changes (semantic).
- Choose: split by headings and prefix each chunk with its heading path (
Leave policy > Carry-over), ~200–500 words with 10–20 % overlap. Then tune on real questions.
6.4 Embeddings & the vector index¶
- What: store each chunk's meaning as numbers so you can find chunks by meaning, not exact words.
- Why: "Can I take my unused holidays into January?" should find a chunk that says "carry-over of annual leave".
- Options: hosted embedding APIs or open-source models · vector stores such as pgvector (Postgres), Pinecone, Qdrant, Weaviate, Elasticsearch/OpenSearch.
- Choose: if you already run Postgres, pgvector is enough for up to a few million chunks. Use a managed vector database when you need more scale or less operations work.
6.5 The keyword index (BM25)¶
- What: classic search on exact words — the same idea as in the template.
- Why: embeddings are weak on exact terms: product codes (
SKU-4471), names, error messages, acronyms. - Options: Elasticsearch/OpenSearch, Postgres full-text search, or an in-memory BM25 for small sets.
6.6 Hybrid retrieval — combining both searches¶
Keyword and vector search each return a ranked list, and their scores aren't comparable. The standard
way to merge them is Reciprocal Rank Fusion (RRF): each document earns 1 / (60 + rank) from every
list it appears in, and the totals are sorted. Being near the top of both lists wins.
Worked example
| Document | Keyword rank | Vector rank | RRF score |
|---|---|---|---|
| A | 1 | 2 | 1/61 + 1/62 = 0.03252 |
| C | 3 | 1 | 1/63 + 1/61 = 0.03227 |
| B | 2 | — | 1/62 = 0.01613 |
| D | — | 3 | 1/63 = 0.01587 |
Final order: A, C, B, D — A and C appear in both lists, so they beat documents found by only one search.
def reciprocal_rank_fusion(*ranked_lists, k=60):
"""Merge several ranked lists of IDs into one, rewarding IDs that rank well in many lists."""
scores = {}
for ranking in ranked_lists:
for rank, doc_id in enumerate(ranking, start=1): # ranks start at 1
scores[doc_id] = scores.get(doc_id, 0.0) + 1 / (k + rank)
return sorted(scores, key=scores.get, reverse=True) # best total first
keyword_results = ["A", "B", "C"]
vector_results = ["C", "A", "D"]
print(reciprocal_rank_fusion(keyword_results, vector_results)) # → ['A', 'C', 'B', 'D']
6.7 Re-ranking¶
- What: a second, more careful model (a cross-encoder) reads each candidate chunk together with the question and scores how well it answers it.
- Why: the first search is fast but rough. Re-ranking the top 30–50 and keeping the best 4–8 sends the LLM fewer, better chunks — better answers and lower cost.
- Choose: add it when the right chunk is usually somewhere in the top 30 but not reliably in the top 5. It adds 100–300 ms, so measure the gain.
6.8 Prompt builder & token budget¶
- What: assembles instructions, numbered chunks, chat history and the question.
- Why: models have a limited context window, and every token costs money and time.
- How: set a budget and fill it in priority order. For example, with an 8,000-token budget:
| Part | Tokens |
|---|---|
| Instructions (system prompt) | ~300 |
| Recent chat history (summarised if long) | up to 1,000 |
| Retrieved chunks (best first, stop when full) | up to 6,000 |
| Question | ~50 |
| Room left for the answer | the rest |
Always wrap chunks in clear delimiters (e.g. <context>…</context>) and tell the model to treat them as data.
6.9 The LLM¶
- Stream the answer so users see text immediately.
- Use low temperature (0–0.3) for factual answers.
- Keep a fallback model from another provider for outages, and set timeouts.
- Use a smaller, cheaper model for helper steps such as query rewriting.
6.10 Guardrails & citations¶
- Before generation: block obvious prompt-injection attempts and strip personal data you shouldn't send out.
- After generation: check that every
[n]citation points to a chunk that was actually retrieved; if the answer has no citations, either retry or say you don't know. - Always: treat retrieved text as untrusted — a document saying "ignore your instructions" must not change behaviour.
6.11 Caching¶
| Cache | Saves | When it helps |
|---|---|---|
| Embeddings of chunks | Re-embedding unchanged text | Always — key by content hash |
| Embeddings of frequent questions | One API call per question | Many repeated questions |
| Full answers | The whole pipeline | FAQ-style traffic; expire when documents change |
6.12 Observability & evaluation¶
- Log every request: question, retrieved chunk IDs, prompt size, model, latency, tokens, answer, user feedback.
- Build a golden set: 50–200 real questions with the expected source documents (and ideally reference answers).
- Measure on every change (new chunking, new model, new prompt):
| Metric | Question it answers |
|---|---|
| Recall@k | Was the right chunk among the top k retrieved? |
| MRR | How high did the right chunk rank? |
| Faithfulness | Is every claim in the answer supported by the retrieved chunks? |
| Answer relevance | Does the answer actually address the question? |
| Latency & cost | p50/p95 response time and tokens per question |
7. Worked example: sizing a system¶
Scenario: a company with 50,000 documents (average 10 pages ≈ 5,000 words each) and 5,000 questions per day.
| Quantity | Calculation | Result |
|---|---|---|
| Chunks | 5,000 words ÷ ~400 words per chunk ≈ 13 per document × 50,000 | ≈ 650,000 chunks |
| Vector storage | 650,000 × 1,536 numbers × 4 bytes | ≈ 4 GB (plus index overhead) |
| Tokens per question | ~300 instructions + 6 chunks × ~500 + ~50 question + ~300 answer | ≈ 3,650 tokens |
| Tokens per day | 3,650 × 5,000 questions | ≈ 18 million tokens/day |
What this tells you: the index fits comfortably in pgvector or a managed vector store, and LLM tokens are the main running cost — so re-ranking (fewer chunks) and answer caching pay off quickly. Multiply the tokens by your model's price to get the daily cost.
8. Security & permissions¶
- Filter by permission during retrieval, using the
allowed_groupsmetadata on each chunk. Never retrieve everything and ask the LLM to "hide" what the user shouldn't see — it can't be trusted to. - Separate tenants (customers) with separate indexes or namespaces, plus filters.
- Keep secrets out of prompts and redact personal data before logging.
- Propagate deletions — when a document is removed or access is revoked, remove its chunks quickly.
9. Common mistakes¶
| Mistake | What happens | Fix |
|---|---|---|
| Only vector search | Misses exact codes, names and acronyms | Add keyword search and merge with RRF |
| Chunks without headings | Retrieved chunks lack context ("it costs ₹900" — what does?) | Prefix each chunk with its heading path |
| Re-indexing everything nightly | Slow, expensive, stale during the day | Incremental sync by content hash |
| Old chunks not deleted | Bot quotes outdated policies | Delete a document's old chunks when it changes |
| No golden set | Every change is a guess | Measure recall@k and faithfulness before each release |
| Too many chunks in the prompt | Higher cost, worse answers | Re-rank and send only the best 4–8 |
10. Production checklist¶
- Incremental ingestion with deletes, running on a queue
- Structure-aware chunking with heading prefixes and metadata
- Hybrid search (keyword + vector) with permission filters
- Re-ranking measured against the golden set
- Prompt with delimiters, a token budget and an "I don't know" rule
- Streaming, timeouts and a fallback model
- Citation checks and personal-data redaction
- Logging of chunks, tokens, latency and feedback
- Golden-set evaluation in CI before each release
Interview questions¶
Walk me through a production RAG architecture.
Two paths. Offline: connectors detect changes → parse and clean → structure-aware chunking with metadata → embed → write to vector and keyword indexes, deleting old chunks. Online: gateway (auth, rate limits, user groups) → optional query rewrite → hybrid retrieval with permission filters → RRF merge → re-rank → prompt with a token budget → streaming LLM → guardrails and citations. Around it: caching, logging, and golden-set evaluation.
Why hybrid search instead of only embeddings?
Embeddings match meaning but miss exact strings — product codes, names, error messages. BM25 matches those exactly. Running both and merging with reciprocal rank fusion gives the best of each, with no need to make their scores comparable.
How do you keep users from seeing documents they shouldn't?
Store permissions (groups, tenant) as metadata on each chunk and filter at retrieval time using the authenticated user's groups. Never rely on the prompt to hide data the model has already been given.
The bot gives a wrong answer. How do you debug it?
Check the logs for that request. If the right chunk wasn't retrieved, it's a retrieval problem — parsing, chunking, search or re-ranking. If it was retrieved but the answer is wrong, it's a generation problem — prompt instructions, too many noisy chunks, or the model. Then add the question to the golden set so it can't regress.
How would you reduce cost by 50 %?
Send fewer, better chunks (re-ranking), cache embeddings and repeated answers, use a smaller model for helper steps, summarise long chat histories, and check the logs for wasted tokens such as duplicate chunks.
Next: AI agent architecture →