AI System Architecture¶
How real GenAI systems are put together — the parts, how data flows between them, and the trade-offs you'll defend in a system-design interview or a client review. Each page is a guided walkthrough: plain-English idea first, then diagrams, every component explained, worked examples, common mistakes, a checklist and interview questions.
-
Production RAG architecture
Follow a document in and a question out: ingestion, chunking, hybrid search, re-ranking, prompts, guardrails, caching, evaluation — plus a full sizing example.
Intermediate · 20 min
-
AI agent architecture
The agent loop step by step and in runnable Python, tool schemas, permissions, memory, when to use workflows vs agents, and the failure modes to design for.
Advanced · 20 min
Suggested reading order¶
- Build the basics first — the RAG chatbot template gives you the core idea in 60 lines.
- Production RAG architecture — what changes when thousands of people and documents are involved.
- AI agent architecture — what changes when the model can take actions, not just answer.
Architecture vocabulary¶
The terms used across these pages, in one place.
| Term | Meaning |
|---|---|
| Offline path | Work done ahead of time, when data changes — e.g. processing and indexing documents |
| Online path | Work done for every request, while the user waits |
| Chunk | A small piece of a document, the unit that search works on |
| Embedding | A list of numbers representing the meaning of a text, used for meaning-based search |
| Vector index | A database that finds the embeddings closest to a question's embedding |
| BM25 / keyword index | Classic search that ranks text by matching words |
| Hybrid search | Running keyword and vector search together and merging the results |
| Re-ranking | A second, more careful model that reorders the top search results |
| Context window | The maximum text (in tokens) a model can read in one request |
| Token | A piece of a word; models read, charge and limit by tokens (~¾ of an English word on average) |
| Tool calling | A model replying with a structured request to run a function, instead of text |
| Orchestrator | Your code that runs the agent loop, tools and limits |
| Guardrails | Checks before and after the model — permissions, validation, redaction, refusals |
| Golden set | A fixed list of test questions or tasks with expected results, used to measure quality |
| p50 / p95 latency | The response time that 50 % / 95 % of requests are faster than |
How to answer an AI system-design question¶
- Clarify — who the users are, scale (requests per day, number of documents), latency target, accuracy bar, budget, and how sensitive the data is.
- Draw the data flow — the offline (ingest/index) and online (request → response) paths separately.
- Pick components and justify each one — retrieval type, model size, where state and permissions live.
- Cover the "-ilities" — latency, cost, reliability, security (prompt injection, personal data) and evaluation.
- Name the trade-offs — and what you would measure to decide between options.
Coming next
Website chatbot with lead capture · Document extraction pipeline · LLM gateway & cost control · Multi-agent workflows