Skip to content

AI System Architecture

How real GenAI systems are put together — the parts, how data flows between them, and the trade-offs you'll defend in a system-design interview or a client review. Each page is a guided walkthrough: plain-English idea first, then diagrams, every component explained, worked examples, common mistakes, a checklist and interview questions.

  • Production RAG architecture


    Follow a document in and a question out: ingestion, chunking, hybrid search, re-ranking, prompts, guardrails, caching, evaluation — plus a full sizing example.

    Intermediate · 20 min

    Read the walkthrough

  • AI agent architecture


    The agent loop step by step and in runnable Python, tool schemas, permissions, memory, when to use workflows vs agents, and the failure modes to design for.

    Advanced · 20 min

    Read the walkthrough

Suggested reading order

  1. Build the basics first — the RAG chatbot template gives you the core idea in 60 lines.
  2. Production RAG architecture — what changes when thousands of people and documents are involved.
  3. AI agent architecture — what changes when the model can take actions, not just answer.

Architecture vocabulary

The terms used across these pages, in one place.

Term Meaning
Offline path Work done ahead of time, when data changes — e.g. processing and indexing documents
Online path Work done for every request, while the user waits
Chunk A small piece of a document, the unit that search works on
Embedding A list of numbers representing the meaning of a text, used for meaning-based search
Vector index A database that finds the embeddings closest to a question's embedding
BM25 / keyword index Classic search that ranks text by matching words
Hybrid search Running keyword and vector search together and merging the results
Re-ranking A second, more careful model that reorders the top search results
Context window The maximum text (in tokens) a model can read in one request
Token A piece of a word; models read, charge and limit by tokens (~¾ of an English word on average)
Tool calling A model replying with a structured request to run a function, instead of text
Orchestrator Your code that runs the agent loop, tools and limits
Guardrails Checks before and after the model — permissions, validation, redaction, refusals
Golden set A fixed list of test questions or tasks with expected results, used to measure quality
p50 / p95 latency The response time that 50 % / 95 % of requests are faster than

How to answer an AI system-design question

  1. Clarify — who the users are, scale (requests per day, number of documents), latency target, accuracy bar, budget, and how sensitive the data is.
  2. Draw the data flow — the offline (ingest/index) and online (request → response) paths separately.
  3. Pick components and justify each one — retrieval type, model size, where state and permissions live.
  4. Cover the "-ilities" — latency, cost, reliability, security (prompt injection, personal data) and evaluation.
  5. Name the trade-offs — and what you would measure to decide between options.

Coming next

Website chatbot with lead capture · Document extraction pipeline · LLM gateway & cost control · Multi-agent workflows