Skip to content

How to write clean code

Beginner-friendly · 15 min · before/after examples

Clean code is code that someone else — or you in three months — can read, change and trust without fear. It isn't about being clever; it's about being clear. Below are twelve habits, each with a before/after, then a full refactor of a messy GenAI script.


1. Use names that explain themselves

def f(d, n):
    r = []
    for x in d:
        if x[1] > n:
            r.append(x[0])
    return r
def sources_above(scored_chunks, min_score):
    """Return the source names of chunks scoring above `min_score`."""
    return [source for source, score in scored_chunks if score > min_score]

print(sources_above([("faq.md", 0.9), ("old.md", 0.2)], min_score=0.5))   # → ['faq.md']

Name things by what they are (chunks, api_key, retry_count), functions by what they do (load_documents, build_prompt). Booleans read best as questions: is_valid, has_sources.

2. Keep functions small — one job each

If you need "and" to describe a function ("loads the files and chunks them and calls the LLM"), split it. Small functions are easier to name, test and reuse.

def load_documents(folder): ...
def chunk(text, size=200): ...
def retrieve(question, chunks, k=4): ...
def build_prompt(question, chunks): ...
def ask_llm(prompt): ...

3. Replace magic numbers with named constants

def trim(text):
    return text[:12000]          # why 12000?
MAX_CONTEXT_CHARS = 12_000       # ~3,000 tokens — leaves room for the answer

def trim(text):
    return text[:MAX_CONTEXT_CHARS]

4. Return early instead of nesting

def answer(question, chunks):
    if question:
        if chunks:
            return f"Answering {question!r} from {len(chunks)} chunks"
        else:
            return "No documents found."
    else:
        return "Please ask a question."
def answer(question, chunks):
    if not question:
        return "Please ask a question."
    if not chunks:
        return "No documents found."
    return f"Answering {question!r} from {len(chunks)} chunks"

print(answer("What is RAG?", ["c1", "c2"]))   # → Answering 'What is RAG?' from 2 chunks

Handle the special cases first and get them out of the way — the main path stays unindented.

5. Add type hints and a docstring

def build_prompt(question: str, chunks: list[str], max_chunks: int = 4) -> str:
    """Combine the question with up to `max_chunks` numbered context chunks."""
    context = "\n".join(f"[{i}] {c}" for i, c in enumerate(chunks[:max_chunks], start=1))
    return f"<context>\n{context}\n</context>\n\nQuestion: {question}"

print(build_prompt("What is RAG?", ["Retrieval + generation."]).splitlines()[1])
# → [1] Retrieval + generation.

The hints tell readers (and your editor) what goes in and comes out; the docstring says why the function exists. More in Type hints.

6. Don't repeat yourself (DRY)

summary = client_call("Summarise: " + text, temperature=0.2, max_tokens=300)
title = client_call("Write a title for: " + text, temperature=0.2, max_tokens=300)
tags = client_call("List tags for: " + text, temperature=0.2, max_tokens=300)
def ask(instruction: str, text: str) -> str:
    """One place for model settings — change them once, not three times."""
    return client_call(f"{instruction}: {text}", temperature=0.2, max_tokens=300)

summary = ask("Summarise", text)
title = ask("Write a title for", text)
tags = ask("List tags for", text)

But don't over-do it: two similar lines are fine. Extract when you see the third copy.

7. Comments should explain why, not what

i = i + 1            # add one to i
# The API counts pages from 1, our list from 0.
page_number = index + 1

If a comment just repeats the code, delete it — or rename things so the code says it itself.

8. Keep settings out of the code

import os

MODEL = os.getenv("LLM_MODEL", "gpt-4o-mini")             # change without editing code
TEMPERATURE = float(os.getenv("LLM_TEMPERATURE", "0.2"))  # environment values are strings

Keys, model names, URLs and limits belong in environment variables or a settings file — see Project structure.

9. Separate logic from input/output

Functions that only take inputs and return outputs (no printing, no files, no network) are easy to test and reuse. Keep the "talking to the outside world" at the edges.

def score_and_print(text):
    words = text.split()
    print(f"{len(words)} words")     # logic and output tangled together
def word_count(text: str) -> int:    # pure: easy to test
    return len(text.split())

print(f"{word_count('retrieval augmented generation')} words")   # → 3 words

10. Handle errors on purpose

Catch the specific errors you expect, at the place you can do something useful about them — and let unexpected ones surface. See Common mistakes #8.

11. Let tools enforce the style

Don't argue about spaces and quotes — automate them:

uvx ruff format .        # formats every file consistently
uvx ruff check . --fix   # finds bugs and style problems, fixes what it safely can

Add them to your editor (format on save) and CI. See Manage a project with uv.

12. Write a test for anything important

def test_build_prompt_numbers_chunks():
    prompt = build_prompt("Q?", ["first", "second"])
    assert "[1] first" in prompt and "[2] second" in prompt

A test documents what the code should do and catches the day it stops. See Testing with pytest.


Before & after: a full refactor

A small script that summarises the .txt files in a folder. Both versions do the same job — but only one is easy to change, test and trust.

import os

def go(p):
    r = []
    for f in os.listdir(p):
        if f.endswith(".txt"):
            t = open(os.path.join(p, f)).read()
            if len(t) > 0:
                t = t[:2000]
                try:
                    s = llm("Summarise in one line: " + t)
                except:
                    s = "ERR"
                r.append(f + ": " + s)
    return r
import logging
from pathlib import Path

MAX_CHARS_PER_DOC = 2_000          # keep each prompt small and cheap
SUMMARY_INSTRUCTION = "Summarise in one line"

log = logging.getLogger(__name__)


def read_text_files(folder: Path) -> dict[str, str]:
    """Return {file name: text} for every non-empty .txt file in `folder`."""
    texts = {}
    for path in sorted(folder.glob("*.txt")):
        text = path.read_text(encoding="utf-8").strip()
        if text:
            texts[path.name] = text
    return texts


def summarise(text: str, llm) -> str:
    """One-line summary of `text`; returns a placeholder if the model call fails."""
    prompt = f"{SUMMARY_INSTRUCTION}: {text[:MAX_CHARS_PER_DOC]}"
    try:
        return llm(prompt)
    except ConnectionError as err:  # the failure we expect from a network call
        log.warning("Summary failed: %s", err)
        return "(summary unavailable)"


def summarise_folder(folder: Path, llm) -> list[str]:
    """'file: summary' lines for every text file in `folder`."""
    return [f"{name}: {summarise(text, llm)}" for name, text in read_text_files(folder).items()]

What changed — and why it matters

Change Why
go(p) → three named functions Each does one job and can be tested on its own
2000 → MAX_CHARS_PER_DOC The limit is explained and changed in one place
os.listdir + open() → pathlib with encoding="utf-8" Cross-platform, files always closed, no encoding surprises
sorted(...) Same order on every computer — results are reproducible
bare except → except ConnectionError + log Real bugs aren't hidden; failures are visible in logs
llm passed in as a parameter Tests can pass a fake model — no API key, no cost
Type hints + docstrings Readers know what goes in, what comes out and why

Clean-code checklist

Before you commit, ask:

  • Could a teammate understand every name without asking me?
  • Does each function do one thing, in under ~30 lines?
  • Are there unexplained numbers or strings? Make them named constants.
  • Are secrets and settings coming from the environment, not the code?
  • Am I catching only the errors I can actually handle?
  • Did ruff format and ruff check pass?
  • Is there a test for the part most likely to break?

Previous: Avoid common coding mistakes ←