5. Data structures¶
Beginner · 10 min read
Four built-in containers cover almost everything. LLM APIs speak in lists of dictionaries, so these are essential.
| Type | Example | Ordered | Changeable | Duplicates |
|---|---|---|---|---|
list |
[1, 2, 2] |
✅ | ✅ | ✅ |
tuple |
(1, 2) |
✅ | ❌ | ✅ |
set |
{1, 2} |
❌ | ✅ | ❌ |
dict |
{"k": 1} |
✅ (insertion order) | ✅ | keys unique |
5.1 Lists¶
chunks = ["intro", "pricing", "faq"]
chunks.append("contact") # add to the end
chunks.insert(0, "title") # add at a position
print(chunks) # → ['title', 'intro', 'pricing', 'faq', 'contact']
print(chunks[1], chunks[-1]) # → intro contact
print(chunks[1:3]) # → ['intro', 'pricing'] slicing works like strings
chunks.remove("faq") # remove by value
last = chunks.pop() # remove and return the last item
print(last, len(chunks)) # → contact 3
scores = [0.4, 0.9, 0.7]
print(sorted(scores, reverse=True)) # → [0.9, 0.7, 0.4] new sorted list
print(max(scores), sum(scores)) # → 0.9 2.0
5.2 Tuples¶
point = (12.97, 77.59) # fixed group of values — can't be changed
lat, lon = point # "unpacking" into variables
print(lat) # → 12.97
def min_max(values):
return min(values), max(values) # functions often return tuples
low, high = min_max([3, 9, 1])
print(low, high) # → 1 9
5.3 Sets¶
tags = {"rag", "agents", "rag"} # duplicates are dropped automatically
print(len(tags)) # → 2
print("rag" in tags) # → True very fast membership check
a, b = {"python", "sql"}, {"python", "docker"}
print(a & b) # → {'python'} in both (intersection)
print(sorted(a | b)) # → ['docker', 'python', 'sql'] in either (union)
5.4 Dictionaries¶
message = {"role": "user", "content": "What is RAG?"} # key → value pairs
print(message["role"]) # → user
print(message.get("name", "anonymous")) # → anonymous .get() avoids errors for missing keys
message["content"] = "Explain RAG simply" # update a value
message["name"] = "Priya" # add a new key
for key, value in message.items(): # loop over pairs
print(key, "=", value)
print(list(message.keys())) # → ['role', 'content', 'name']
5.5 Lists of dictionaries — the LLM message format¶
messages = [
{"role": "system", "content": "You are concise."},
{"role": "user", "content": "Hi!"},
]
messages.append({"role": "assistant", "content": "Hello!"}) # keep the conversation going
user_turns = [m["content"] for m in messages if m["role"] == "user"]
print(user_turns) # → ['Hi!']
5.6 Comprehensions¶
A comprehension builds a new collection in one readable line.
words = ["RAG", "agents", "LLM", "eval"]
lengths = [len(w) for w in words] # list: transform each item
print(lengths) # → [3, 6, 3, 4]
short = [w for w in words if len(w) <= 3] # list: filter
print(short) # → ['RAG', 'LLM']
word_len = {w: len(w) for w in words} # dict comprehension
print(word_len["agents"]) # → 6
unique_lengths = {len(w) for w in words} # set comprehension
print(sorted(unique_lengths)) # → [3, 4, 6]
Why it matters for GenAI
Chat history is a list of dicts, retrieved chunks are a list of dicts with text and
source, and API responses are nested dicts. Comprehensions are how you reshape them.
Practice¶
- From
docs = [{"source": "a.md", "score": 0.9}, {"source": "b.md", "score": 0.3}], build a list of sources with a score above 0.5.