The previous article ended with a loop of about thirty lines: call the model, append its message, run any tools it asked for, append the results, call again. That loop is correct, but "append the results" hides the part that actually decides whether your agent works. This article opens one turn of the loop and stays inside it: what you append, in what order, with which fields, and why the API refuses to run if you get it slightly wrong.
The loop has a grammar
A chatbot sends a list of messages and prints what comes back. An agent sends a list of messages and keeps everything, including its own tool requests, in that same list. So the transcript stops being a conversation and becomes a log with roles that the API validates:
user— the goal, plus anything you inject.assistant— what the model produced. This is also where a tool request lives. It is not a separate event type; it is an assistant message whosetool_callsfield is non-empty.tool— one result per requested call. This message is written by your code, not the model.
That third role is the one that did not exist in your chatbot. It is the observation in ReAct machinery: the only channel through which anything from the outside world enters the model's context 3.
One turn, message by message
Take "Is it warmer in Paris or Dublin right now?" and a single get_weather(city) tool. The transcript after the first model call looks like this.
Read the arrows as "then appended". The model's response in the middle is one object containing an array of two calls, each with an id, a function name, and arguments. In OpenAI's Chat Completions API those are call.id, call.function.name, and call.function.arguments; arguments is a JSON string, not a parsed object, which is the first thing most people trip over. On Anthropic's API the same thing appears as a tool_use content block with an input object that is already parsed 12.
The two tool messages are yours. Each carries tool_call_id (Anthropic: tool_use_id) pointing back at the request it answers, and a string content holding whatever your function returned. Anthropic wraps them as tool_result blocks inside the next user message; the shape differs, the pairing does not.
Note that the assistant message can carry prose and a tool request at once. Anthropic's documentation shows exactly this: a text block ("I'll help you check the current weather…") followed by a tool_use block in the same message 2. That prose is where the model's plan is visible, and it is worth keeping rather than discarding.
Why the ids exist
The pairing is enforced, not decorative. On OpenAI, if an assistant message has tool_calls, the next messages must be tool messages covering every tool_call_id; violating it gives a 400: An assistant message with 'tool_calls' must be followed by tool messages responding to each 'tool_call_id' 1. If the model requested two calls, you owe two results before the transcript can continue.
The reason is that the model has no other way to match answers to questions. "19 C" and "12 C" are unlabeled; the ids are what let it know the 19 belongs to Paris. With parallel calls to the same function — which the API explicitly tells you to expect, since a response may contain zero, one, or several calls — that matching is the difference between a correct answer and a coin flip 1.
Two practical consequences. Append the assistant message exactly as returned; do not re-render it as text, and if you use a reasoning model, pass back any reasoning items that came with the tool calls, because they are part of the same turn 1. And never "summarize" tool output into a user message to save a step — you lose the ids and break the contract.
The observation is your whole interface to the world
Everything the model will ever learn about the outside world arrives as the content of a tool message. Three design choices follow.
You decide what the world says back. content is a string; OpenAI's guide says the format is up to you — JSON, plain text, an error code — and the model interprets it 1. A weather API that returns 40 fields will spend tokens on 38 of them you do not need. Returning a small, deliberately chosen projection is normal good practice, not cheating.
Errors are observations too. If the tool throws, do not let the exception kill the loop. Catch it and return something like {"error": "city not found: Dublin"}. Anthropic has an explicit is_error: true flag on tool_result for this; on OpenAI you just put the error text in content 2. This is the practical payoff of the ReAct framing: the model sees a failed action and can reshape its plan — retry with a different argument, or pick another tool — instead of the process dying. This is why the paper lists "handle exceptions and adjust action plans" among the things thoughts do with observations 3.
Treat tool output as data, not instructions. The content you append is written partly by third parties (a web page, a file, a database row). It lands in the model's context with the same authority as your own prompt. Keeping tool results clearly structured and instructing the model to treat fetched text as information only is the cheapest mitigation you have.
Where the plan lives
In the original ReAct prompt, a trajectory is one flat text stream: Thought → Action → Observation, repeated, and your code parses the action out with string matching 3. The paper's point was that interleaving the two beats reasoning alone or acting alone — reasoning sets up the next action, and the observation grounds the next reasoning step.
Modern tool-calling APIs keep the interleaving but move the parsing into the API: the action comes back as a validated object instead of text you must regex. The thought survives as the assistant message's text content, which you are not obliged to prompt for — the model often emits it anyway. Nothing else changed. Planning is not a separate module you build for a small agent; it is the model reasoning over messages and the observations you keep feeding it. When people ask where the planner is, the answer is: in the same list you are already managing.
A minimal implementation
Here is the whole turn written out. It is deliberately explicit about the two steps the previous article compressed.
1import json
2from openai import OpenAI
3
4client = OpenAI()
5
6TOOLS = {"get_weather": {
7 "type": "function",
8 "function": {
9 "name": "get_weather",
10 "description": "Get current weather for a city. Use for any question about "
11 "current or today's conditions; do not use for forecasts.",
12 "parameters": {
13 "type": "object",
14 "properties": {"city": {"type": "string", "description": "City name, e.g. 'Paris'"}},
15 "required": ["city"],
16 },
17 },
18}}
19
20def execute(name, args): # your code, your credentials
21 if name == "get_weather":
22 return {"city": args["city"], "temp_c": 19, "conditions": "clear"}
23 return {"error": f"unknown tool: {name}"}
24
25def run_agent(user_message, max_steps=8):
26 messages = [{"role": "user", "content": user_message}]
27 for _ in range(max_steps):
28 resp = client.chat.completions.create(
29 model="gpt-4o-mini", messages=messages, tools=list(TOOLS.values())
30 )
31 msg = resp.choices[0].message
32 messages.append(msg) # step 2: keep the assistant turn verbatim
33
34 if not msg.tool_calls: # step 3: no action requested -> done
35 return msg.content
36
37 for call in msg.tool_calls: # step 4: one observation per call, by id
38 try:
39 args = json.loads(call.function.arguments)
40 content = json.dumps(execute(call.function.name, args))
41 except Exception as exc:
42 content = json.dumps({"error": str(exc)})
43 messages.append({
44 "role": "tool",
45 "tool_call_id": call.id, # the pairing the API validates
46 "content": content,
47 })
48 return "Stopped: step budget exhausted"Two details are load-bearing. messages.append(msg) appends the SDK's response object, which serializes back into a valid assistant message with its tool_calls intact — reconstructing it by hand is how the id mismatch bugs happen. And the for loop appends all results before the next model call: the request is only valid once every id has been answered.
Trace it once with the Paris/Dublin question. Iteration 1 sends [user] and gets an assistant message with two calls. The loop appends that message, then two tool messages. Iteration 2 sends all four messages and gets plain text, so the function returns. If the model instead calls get_weather("Dublin") again because the first result was an error, that is iteration 2 asking a new question — the loop does not care, it just keeps answering.
What the loop still assumes
The model will stop on its own once it has an answer, so the normal exit is not msg.tool_calls. Everything else is failure handling: max_steps for a model that keeps gathering data, an unknown-tool branch for a hallucinated name, and json.loads wrapped because a malformed arguments string is a legitimate thing for a model to emit. Schema-strict tool modes reduce the last one by validating arguments before they reach you, at the cost of tighter schema rules.
Two more boundaries. Every iteration resends the entire transcript plus the tool schemas, so cost and latency grow with turn count, and step budgets are also token budgets. And observations accumulate forever — the point at which you must compress, summarize, or drop old tool results is the point where the loop stops being thirty lines. That is the next problem, and it is a different one.
Exercise
Without running it, predict what happens when you (a) append only the first of two tool results, and (b) append the two tool messages in the reverse order. Then decide which of the two the API catches.
Answer
(a) The next request is rejected with a 400 naming the unanswered tool_call_id; the loop never reaches the model. (b) The request is accepted — order among sibling tool messages is not the constraint — but nothing guarantees the model will attribute the right result to the right request, and with two calls to the same function that is exactly the confusion the ids were meant to prevent. The API catches missing; it does not catch mislabeled.