Skip to content

1. Iterators & generators

Intermediate · 8 min read

Generators produce values one at a time, on demand — perfect for huge files and for streaming LLM tokens as they arrive.

1.1 How for really works

models = ["a", "b"]
it = iter(models)          # every iterable can give you an iterator
print(next(it))            # → a     next() asks for the next value
print(next(it))            # → b
# One more next(it) would raise StopIteration — that's how a for-loop knows to stop.

1.2 Generators with yield

def count_up(limit):
    n = 1
    while n <= limit:
        yield n            # hand back one value, pause here until asked again
        n += 1

gen = count_up(3)
print(next(gen))           # → 1
print(list(gen))           # → [2, 3]   the rest of the values

A function with yield returns a generator; nothing runs until you start iterating.

1.3 Processing big files lazily

from pathlib import Path

Path("big.txt").write_text("\n".join(f"line {i}" for i in range(1, 6)), encoding="utf-8")

def read_chunks(path, lines_per_chunk=2):
    """Yield the file a few lines at a time instead of loading it all."""
    batch = []
    with open(path, encoding="utf-8") as f:
        for line in f:
            batch.append(line.strip())
            if len(batch) == lines_per_chunk:
                yield batch
                batch = []
    if batch:                          # leftover lines at the end
        yield batch

for chunk in read_chunks("big.txt"):
    print(chunk)                       # first → ['line 1', 'line 2']

Memory stays flat whether the file has 10 lines or 10 million.

1.4 Generator expressions

squares = (n * n for n in range(1_000_000))   # parentheses → generator, not a list
print(next(squares), next(squares))           # → 0 1

total = sum(len(w) for w in ["rag", "agents"])   # no extra brackets needed inside a call
print(total)                                  # → 9

1.5 Streaming tokens

import time

def fake_llm_stream(text):
    """Pretend to be a streaming LLM: yield one word at a time."""
    for word in text.split():
        time.sleep(0.01)               # a real API would be waiting on the network here
        yield word + " "

for token in fake_llm_stream("Streaming feels much faster"):
    print(token, end="", flush=True)   # print as it arrives, on one line
print()

With the OpenAI SDK, client.chat.completions.create(..., stream=True) returns exactly this kind of iterator — you loop over it and print each piece.

Why it matters for GenAI

Streaming responses, reading large document sets, and batching texts for embedding are all generator patterns.

Practice

  • Write a generator batched(items, size) that yields lists of size items (the last one may be shorter).
Answer
def batched(items, size):
    for i in range(0, len(items), size):
        yield items[i:i + size]

print(list(batched([1, 2, 3, 4, 5], 2)))   # → [[1, 2], [3, 4], [5]]