Skip to content

7. Concurrency: threads vs processes

Advanced · 8 min read

Python offers three ways to do several things at once. Picking the right one depends on what you're waiting for.

Your task is… Use Why
Waiting on network / disk (API calls, downloads) asyncio or threads Waiting releases the CPU, so many can overlap
Heavy computation (parsing PDFs, number crunching) processes Each process has its own interpreter and runs on its own CPU core
A synchronous library you can't change threads Works without rewriting code as async

7.1 The GIL in one paragraph

In standard Python (CPython), the Global Interpreter Lock lets only one thread run Python bytecode at a time. Threads still help for I/O — a thread waiting on the network releases the lock — but they don't speed up pure-Python CPU work. For that, use processes. (Free-threaded "no-GIL" builds exist from Python 3.13 but aren't the default yet.)

7.2 Threads with ThreadPoolExecutor

from concurrent.futures import ThreadPoolExecutor
import time

def fetch(url: str) -> str:
    time.sleep(0.2)                     # stands in for a blocking HTTP request
    return f"fetched {url}"

urls = ["a.com", "b.com", "c.com", "d.com"]
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as pool:
    results = list(pool.map(fetch, urls))   # run fetch() on each url, results in order

print(results[0])                       # → fetched a.com
print(time.perf_counter() - start < 0.5)   # → True   ~0.2 s instead of 0.8 s

7.3 Processes with ProcessPoolExecutor

parse_parallel.py
from concurrent.futures import ProcessPoolExecutor

def count_primes(limit: int) -> int:
    """Deliberately CPU-heavy work."""
    return sum(all(n % d for d in range(2, int(n ** 0.5) + 1)) for n in range(2, limit))

# The guard is required for processes on Windows and macOS: each worker re-imports this file.
if __name__ == "__main__":
    with ProcessPoolExecutor() as pool:            # one worker per CPU core by default
        print(list(pool.map(count_primes, [20_000, 30_000])))   # → [2262, 3245]

Work sent to processes must be picklable (plain functions defined at the top of a module, simple data) — lambdas and open connections can't be sent.

7.4 Handling results as they finish

from concurrent.futures import ThreadPoolExecutor, as_completed
import time

def summarise(doc_id: int) -> str:
    time.sleep(0.05 * (3 - doc_id))     # later docs finish sooner
    return f"summary {doc_id}"

with ThreadPoolExecutor() as pool:
    futures = {pool.submit(summarise, i): i for i in range(3)}
    for future in as_completed(futures):          # yields each one as soon as it's done
        print(future.result())                    # first → summary 2

Why it matters for GenAI

Calling a synchronous SDK for 100 documents? A ThreadPoolExecutor with a sensible max_workers is the quickest speed-up. Parsing thousands of PDFs locally? Use processes.

Practice

  • Use a ThreadPoolExecutor to "fetch" 10 fake URLs (each sleeping 0.1 s) with max_workers=5. How long should it take?
Answer

About 0.2 s: two rounds of five requests running at the same time.