Skip to content

8. Testing with pytest

Intermediate · 9 min read

Tests prove your code works — and keep proving it after every change. pytest (pip install pytest) is the standard tool.

8.1 Your first test

text_utils.py
def chunk_words(text: str, size: int) -> list[str]:
    """Split text into chunks of `size` words."""
    if size <= 0:
        raise ValueError("size must be positive")
    words = text.split()
    return [" ".join(words[i:i + size]) for i in range(0, len(words), size)]
test_text_utils.py
from text_utils import chunk_words

def test_splits_into_chunks():                 # any function named test_* is a test
    assert chunk_words("a b c d e", 2) == ["a b", "c d", "e"]

def test_empty_text_gives_no_chunks():
    assert chunk_words("", 3) == []

Run pytest in the folder — it finds files named test_*.py and reports each pass or failure, showing exactly which values differed.

8.2 Testing errors

test_text_utils.py
import pytest
from text_utils import chunk_words

def test_rejects_zero_size():
    with pytest.raises(ValueError, match="positive"):   # the block must raise this error
        chunk_words("a b", 0)

8.3 Many cases with parametrize

test_text_utils.py
import pytest
from text_utils import chunk_words

@pytest.mark.parametrize("text, size, expected", [
    ("a b c", 1, 3),          # each tuple becomes its own test
    ("a b c", 2, 2),
    ("a b c", 10, 1),
])
def test_chunk_count(text, size, expected):
    assert len(chunk_words(text, size)) == expected

8.4 Fixtures: shared setup

test_retrieval.py
import pytest

@pytest.fixture
def docs():                                    # runs before each test that asks for it
    return ["refunds within 30 days", "free shipping over 500"]

def test_finds_refund_doc(docs):               # ask for a fixture by naming it as a parameter
    assert any("refund" in d for d in docs)

def test_temp_files(tmp_path):                 # built-in fixture: a fresh temporary folder
    f = tmp_path / "note.txt"
    f.write_text("hello", encoding="utf-8")
    assert f.read_text(encoding="utf-8") == "hello"

8.5 Faking the LLM with monkeypatch

Tests must not call real APIs — that's slow, costs money and gives different answers each time. Replace the call with a fake:

app.py
def call_llm(prompt: str) -> str:
    raise RuntimeError("would call the real API")    # the real network call lives here

def summarise(text: str) -> str:
    return call_llm(f"Summarise: {text}").strip()
test_app.py
import app

def test_summarise_uses_llm_and_strips(monkeypatch):
    seen = []

    def fake_llm(prompt):
        seen.append(prompt)                    # record what we were asked
        return "  short summary  "             # canned reply

    monkeypatch.setattr(app, "call_llm", fake_llm)   # swap it in for this test only

    assert app.summarise("long text") == "short summary"
    assert seen == ["Summarise: long text"]

Why it matters for GenAI

Test the deterministic parts — chunking, prompt building, parsing model output, retries — with fakes. Evaluate the model's quality separately, with a fixed question set.

Practice

  • Write a parametrized test for a clean(text) function that collapses whitespace, with three input/output pairs.