Skip to content

Call an LLM from Python

Beginner · 5 min

Most providers speak the OpenAI-compatible chat API, so one SDK covers several of them.

1. Install and set your key

pip install openai python-dotenv

Create a .env file (and add .env to .gitignore — never commit keys):

.env
OPENAI_API_KEY=sk-...

2. Make a call

from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()      # copy the values from your .env file into environment variables
client = OpenAI()  # creates the API client; it reads OPENAI_API_KEY automatically

# Send a conversation and get the model's next message back.
reply = client.chat.completions.create(
    model="gpt-4o-mini",  # which model answers (cheaper/faster or smarter/slower)
    messages=[
        # "system" sets the assistant's behaviour for the whole chat
        {"role": "system", "content": "You are a concise assistant."},
        # "user" is the actual question
        {"role": "user", "content": "Explain RAG in one sentence."},
    ],
    temperature=0.3,      # low = focused and consistent, high = more creative
)

# The reply can contain several choices; we asked for one, so take the first.
print(reply.choices[0].message.content)
import os
from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()  # loads GROQ_API_KEY from your .env file

# Same SDK, different server: point base_url at Groq and pass Groq's key.
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])

reply = client.chat.completions.create(
    model="llama-3.1-8b-instant",  # a fast, low-cost open model hosted by Groq
    messages=[{"role": "user", "content": "Explain RAG in one sentence."}],
)
print(reply.choices[0].message.content)  # the model's answer text

Model names change over time — check your provider's model list if one is retired.

3. What the parameters mean

messages
The conversation so far. system sets behaviour, user is the question, assistant holds earlier replies.
temperature
Randomness. Use 0–0.3 for factual or extraction tasks, higher for brainstorming.
max_tokens
Upper limit on the reply length (and therefore cost).

Keep keys out of code

Load keys from environment variables, add .env to .gitignore, and rotate a key immediately if it ever lands in a public repo.

Next: use this in a real project — Build a RAG chatbot.