7. Concurrency: threads vs processes¶
Advanced · 8 min read
Python offers three ways to do several things at once. Picking the right one depends on what you're waiting for.
| Your task is… | Use | Why |
|---|---|---|
| Waiting on network / disk (API calls, downloads) | asyncio or threads |
Waiting releases the CPU, so many can overlap |
| Heavy computation (parsing PDFs, number crunching) | processes | Each process has its own interpreter and runs on its own CPU core |
| A synchronous library you can't change | threads | Works without rewriting code as async |
7.1 The GIL in one paragraph¶
In standard Python (CPython), the Global Interpreter Lock lets only one thread run Python bytecode at a time. Threads still help for I/O — a thread waiting on the network releases the lock — but they don't speed up pure-Python CPU work. For that, use processes. (Free-threaded "no-GIL" builds exist from Python 3.13 but aren't the default yet.)
7.2 Threads with ThreadPoolExecutor¶
from concurrent.futures import ThreadPoolExecutor
import time
def fetch(url: str) -> str:
time.sleep(0.2) # stands in for a blocking HTTP request
return f"fetched {url}"
urls = ["a.com", "b.com", "c.com", "d.com"]
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as pool:
results = list(pool.map(fetch, urls)) # run fetch() on each url, results in order
print(results[0]) # → fetched a.com
print(time.perf_counter() - start < 0.5) # → True ~0.2 s instead of 0.8 s
7.3 Processes with ProcessPoolExecutor¶
from concurrent.futures import ProcessPoolExecutor
def count_primes(limit: int) -> int:
"""Deliberately CPU-heavy work."""
return sum(all(n % d for d in range(2, int(n ** 0.5) + 1)) for n in range(2, limit))
# The guard is required for processes on Windows and macOS: each worker re-imports this file.
if __name__ == "__main__":
with ProcessPoolExecutor() as pool: # one worker per CPU core by default
print(list(pool.map(count_primes, [20_000, 30_000]))) # → [2262, 3245]
Work sent to processes must be picklable (plain functions defined at the top of a module, simple data) — lambdas and open connections can't be sent.
7.4 Handling results as they finish¶
from concurrent.futures import ThreadPoolExecutor, as_completed
import time
def summarise(doc_id: int) -> str:
time.sleep(0.05 * (3 - doc_id)) # later docs finish sooner
return f"summary {doc_id}"
with ThreadPoolExecutor() as pool:
futures = {pool.submit(summarise, i): i for i in range(3)}
for future in as_completed(futures): # yields each one as soon as it's done
print(future.result()) # first → summary 2
Why it matters for GenAI
Calling a synchronous SDK for 100 documents? A ThreadPoolExecutor with a sensible
max_workers is the quickest speed-up. Parsing thousands of PDFs locally? Use processes.
Practice¶
- Use a
ThreadPoolExecutorto "fetch" 10 fake URLs (each sleeping 0.1 s) withmax_workers=5. How long should it take?
Answer
About 0.2 s: two rounds of five requests running at the same time.