Skip to content

6. Async programming (asyncio)

Advanced · 10 min read

LLM and API calls spend most of their time waiting on the network. async code lets one program wait on many requests at once instead of one after another.

6.1 async and await

import asyncio

async def call_model(prompt: str, seconds: float) -> str:   # "async def" = a coroutine
    await asyncio.sleep(seconds)    # `await` = "pause me here; run something else meanwhile"
    return f"reply to {prompt!r}"

async def main():
    reply = await call_model("hi", 0.1)
    print(reply)                    # → reply to 'hi'

asyncio.run(main())                 # start the event loop and run main()

In Jupyter, the loop is already running — use await main() directly instead of asyncio.run.

6.2 Running calls concurrently with gather

import asyncio
import time

async def call_model(prompt: str) -> str:
    await asyncio.sleep(0.2)                    # pretend each call takes 200 ms
    return prompt.upper()

async def main():
    prompts = ["summarise", "translate", "classify"]
    start = time.perf_counter()

    # Start all three at once and wait for every result (returned in the same order).
    results = await asyncio.gather(*(call_model(p) for p in prompts))

    elapsed = time.perf_counter() - start
    print(results)                              # → ['SUMMARISE', 'TRANSLATE', 'CLASSIFY']
    print(elapsed < 0.4)                        # → True   ~0.2 s total, not 0.6 s

asyncio.run(main())

6.3 Limiting concurrency with a semaphore

Firing 500 requests at once will hit rate limits. A semaphore caps how many run together.

import asyncio

async def embed(text: str, limit: asyncio.Semaphore) -> int:
    async with limit:                           # at most N tasks inside this block at a time
        await asyncio.sleep(0.05)
        return len(text)

async def main():
    limit = asyncio.Semaphore(5)                # 5 requests in flight, max
    texts = [f"doc {i}" for i in range(20)]
    sizes = await asyncio.gather(*(embed(t, limit) for t in texts))
    print(len(sizes), sum(sizes))               # → 20 110

asyncio.run(main())

6.4 Timeouts and errors

import asyncio

async def slow_call():
    await asyncio.sleep(5)
    return "late"

async def main():
    try:
        await asyncio.wait_for(slow_call(), timeout=0.1)   # give up after 100 ms
    except asyncio.TimeoutError:
        print("timed out")                                # → timed out

    # return_exceptions=True: one failure doesn't cancel the rest
    async def ok():
        return "ok"
    async def boom():
        raise ValueError("bad input")
    results = await asyncio.gather(ok(), boom(), return_exceptions=True)
    print([type(r).__name__ for r in results])           # → ['str', 'ValueError']

asyncio.run(main())

6.5 Real SDKs

Most LLM SDKs ship an async client — e.g. from openai import AsyncOpenAI, then await client.chat.completions.create(...). Web frameworks like FastAPI run async def endpoints natively, so one server can handle many slow LLM calls at once.

Don't block the loop

Inside async def, use async libraries (httpx.AsyncClient, asyncio.sleep). A normal time.sleep() or requests.get() freezes every other task until it finishes.

Why it matters for GenAI

Embedding 10,000 chunks, evaluating a model on 200 questions, or calling three tools in parallel goes from minutes to seconds with gather + a semaphore.

Practice

  • Use asyncio.gather to run five fake calls that each sleep 0.1 s, and confirm the total time is about 0.1 s.