Skip to content

5. Error handling & logging

Intermediate · 8 min read

Production code fails in expected ways (timeouts, rate limits, bad input) and unexpected ones. Good error handling recovers from the first kind; good logging lets you debug the second.

5.1 Custom exceptions

class LLMError(Exception):
    """Base class for everything that can go wrong talking to the model."""

class RateLimitError(LLMError):
    def __init__(self, retry_after: float):
        super().__init__(f"rate limited, retry after {retry_after}s")
        self.retry_after = retry_after          # extra data the caller can use

try:
    raise RateLimitError(2.5)
except LLMError as err:                         # catching the base catches subclasses too
    print(type(err).__name__, err.retry_after)  # → RateLimitError 2.5

A small hierarchy lets callers choose: catch RateLimitError to retry, or LLMError to fail gracefully.

5.2 Chaining: raise ... from

import json

class BadModelOutput(Exception):
    pass

def parse_reply(text):
    try:
        return json.loads(text)
    except json.JSONDecodeError as err:
        # Wrap the low-level error in a meaningful one, keeping the original as the cause.
        raise BadModelOutput(f"model did not return JSON: {text[:20]!r}") from err

try:
    parse_reply("Sure! Here you go")
except BadModelOutput as err:
    print(err)                         # → model did not return JSON: 'Sure! Here you go'
    print(type(err.__cause__).__name__)   # → JSONDecodeError

5.3 Logging instead of print

import logging

logging.basicConfig(
    level=logging.INFO,                                    # show INFO and above
    format="%(asctime)s %(levelname)s %(name)s: %(message)s",
)
log = logging.getLogger("rag.pipeline")                    # one logger per module

log.debug("chunk details…")         # hidden: below the INFO level
log.info("retrieved %d chunks in %.0f ms", 4, 82.3)       # use %-style args, not f-strings
log.warning("context is close to the token limit")

try:
    1 / 0
except ZeroDivisionError:
    log.exception("scoring failed")   # logs the message AND the full traceback

Why not print? Logs have levels you can turn up or down, timestamps, the module name, and can be sent to files or monitoring tools without changing your code.

5.4 What to log in an LLM app

Log Why
Model, latency, prompt & completion tokens Cost and speed tracking
Retrieved chunk IDs / sources Debugging wrong answers
Retries and error types Spotting rate limits and outages
A request ID on every line Following one request through the pipeline

Don't log API keys or full personal data.

Why it matters for GenAI

When a user reports "the bot gave a wrong answer", the only way to find out why is a log of what was retrieved, what was sent and what came back.

Practice

  • Create a ToolError exception and a function that raises it from a KeyError when a tool name isn't in a registry dict.