Skip to content

10. Performance & caching

Advanced · 8 min read

The golden rule: measure first. In LLM apps the slow part is usually the network call — so caching often beats any code tweak.

10.1 Measuring

import timeit

items = list(range(10_000))
as_set = set(items)

# Time 1,000 membership checks in a list vs a set.
list_time = timeit.timeit(lambda: 9_999 in items, number=1_000)
set_time = timeit.timeit(lambda: 9_999 in as_set, number=1_000)
print(set_time < list_time)          # → True   a set lookup doesn't scan every item

To find where a whole program spends its time, run it under the profiler:

python -m cProfile -s cumulative my_script.py

10.2 Choosing the right structure

Need Slow Fast
"Is X in this collection?" x in big_list x in big_set
Look up by key search a list of dicts dict[key]
Build a long string s += piece in a loop "".join(pieces)
Count things manual dict updates collections.Counter

10.3 In-memory caching with lru_cache

import functools

calls = 0

@functools.lru_cache(maxsize=1024)       # keep up to 1024 recent results
def embed(text: str) -> tuple[float, ...]:
    """Pretend embedding call — the expensive part we don't want to repeat."""
    global calls
    calls += 1
    return tuple(float(ord(c)) for c in text[:3])

for q in ["refund", "shipping", "refund", "refund"]:
    embed(q)

print(calls)                             # → 2   two unique texts, two real calls
print(embed.cache_info().hits)           # → 2

Cached arguments must be hashable (strings, numbers, tuples) — not lists or dicts. Return tuples rather than lists so callers can't modify the cached value.

10.4 Caching across runs (on disk)

import hashlib
import json
from pathlib import Path

CACHE = Path(".llm_cache")
CACHE.mkdir(exist_ok=True)

def cached_completion(prompt: str, model: str = "gpt-4o-mini") -> str:
    """Return a saved answer if this exact (model, prompt) was asked before."""
    key = hashlib.sha256(f"{model}\n{prompt}".encode()).hexdigest()   # stable file name
    path = CACHE / f"{key}.json"
    if path.exists():
        return json.loads(path.read_text(encoding="utf-8"))["answer"]
    answer = f"(model answer to: {prompt})"          # ← the real API call goes here
    path.write_text(json.dumps({"answer": answer}), encoding="utf-8")
    return answer

print(cached_completion("What is RAG?") == cached_completion("What is RAG?"))   # → True

During development this makes re-running a pipeline free and instant. In production, a shared cache (Redis, a database) does the same across servers.

10.5 Batching

APIs often accept many inputs per request — embedding 100 texts in one call instead of 100 calls cuts latency and overhead dramatically. Combine with the batched() generator from Iterators & generators.

Why it matters for GenAI

Cache embeddings and repeated answers, batch requests, run independent calls concurrently (async) — these three cut LLM-app latency and cost far more than micro-optimising Python.

Practice

  • Profile a script that chunks a large text file. Is most of the time in reading, splitting or joining?