You already know how to call an LLM API: send a list of messages, get back an assistant message, print it. That is a chatbot. An agent uses the same API call, the same model, and often the same prompt — the difference is a loop you write around that call. This article is about that loop: what exactly goes around in it, why the result is qualitatively different from a chatbot, and how to write it in about 30 lines.
Start from what a single call can do
One chat completion takes messages and returns one assistant message. That is the entire contract. Two consequences matter here:
- The call has no side effects. It cannot read a file, call your database, or check a price. It produces text, and nothing else changes in the world.
- The call has no memory of its own. Everything the model "knows" about this conversation is sitting in the
messagesarray you just sent. If a fact is not in there, the model does not have it.
That second point is the lever. If we want the model's next answer to depend on something real — today's temperature, a row from your table, the contents of a file — that thing has to end up in messages before the answer is generated. A chatbot never puts it there. An agent does.
Tool calling does not run anything
This is the step most people get wrong on their first agent.
When you pass a tools array, you are sending the model a list of functions with JSON schemas and descriptions. In return, the model may stop producing prose and instead return a structured request: a tool name plus arguments. On OpenAI's API that shows up as message.tool_calls with call.id, call.function.name, and call.function.arguments; on Anthropic's it shows up as a tool_use block and stop_reason: "tool_use".
Here is what has not happened: nothing has been executed. The arguments are model-generated text, and if you print them and go to lunch, the weather API is never called. The model decides; your code executes. That split is the whole architecture, and it is also where your safety controls live, because your code is the only thing with credentials and side effects.
So a single tool round trip is still not an agent. It is one request, one execution, one result — a slightly fancier chatbot, if you stop there.
The loop
The loop is small enough to state in four sentences:
- Send the current messages (plus the tool schemas) to the model.
- Append the returned assistant message to the messages, exactly as it came back.
- If it contains no tool calls, you are done — return its text.
- If it does contain tool calls, execute them, append each result as a new message, and go back to step 1.
Step 4 sending you back to step 1 is the only thing in this article that a chatbot does not have. Everything else you have already used.
In code, against the OpenAI Chat Completions API:
1import json
2from openai import OpenAI
3
4client = OpenAI()
5
6TOOLS = [{
7 "type": "function",
8 "function": {
9 "name": "get_weather",
10 "description": "Return the current temperature in Celsius for a city.",
11 "parameters": {
12 "type": "object",
13 "properties": {"city": {"type": "string"}},
14 "required": ["city"],
15 },
16 },
17}]
18
19def execute_tool(name: str, args: dict) -> str:
20 if name == "get_weather":
21 return "15" # pretend this hit a real API
22 return f"Error: unknown tool {name}"
23
24def run_agent(user_message: str, max_steps: int = 8) -> str:
25 messages = [
26 {"role": "system", "content": "Use the tools when you need facts you don't have."},
27 {"role": "user", "content": user_message},
28 ]
29
30 for _ in range(max_steps):
31 # 1. Ask the model what to do next, given everything so far.
32 response = client.chat.completions.create(
33 model="gpt-4o-mini",
34 messages=messages,
35 tools=TOOLS,
36 temperature=0,
37 )
38 message = response.choices[0].message
39 messages.append(message) # 2. Record the model's turn verbatim
40
41 if not message.tool_calls: # 3. Nothing requested -> we are done
42 return message.content
43
44 for call in message.tool_calls: # 4. Execute, then feed results back
45 args = json.loads(call.function.arguments)
46 result = execute_tool(call.function.name, args)
47 messages.append({
48 "role": "tool",
49 "tool_call_id": call.id,
50 "content": result,
51 })
52
53 return "Stopped: hit the step limit without a final answer."Two details that cause real bugs if you skip them. First, append the assistant message before the tool results — the API rejects a tool result that does not answer a preceding tool call, and it matches them by call.id, so the id has to survive untouched. Second, a tool result is a message with a role of its own (tool here; on Anthropic's API a user message containing only tool_result blocks), not a system note you improvise. Some SDK versions also prefer you re-serialize the assistant message (message.model_dump()) rather than appending the object; either way, the rule is to put the turn back in exactly as the API returned it.
What the loop actually buys you
Run this agent on "It's 12 °C in Berlin today. What's the temperature in Rome right now? Is Rome warmer?" The trace looks like this:
1turn 1 assistant: tool_calls=[get_weather(city="Rome")]
2 (your code runs it)
3 tool: "15"
4turn 2 assistant: "Rome is 15 °C, so it's 3 degrees warmer than Berlin right now."Two model calls, one tool call. Note what turn 2 is doing: it reads a number that entered the context window one message ago and reasons over it. A chatbot given the same question has no such number, so it either refuses or invents one. The loop is the mechanism by which the model's own output changes the world and the world's reaction comes back as input. That is the entire qualitative jump.
Three things follow directly from it:
- Recovery. If a search tool returns "no results for this exact phrase," that string is now in context, and the next model call can choose a broader query instead. There is no separate error-handling logic to write beyond making your tools return errors as strings rather than raise, so the model can see them.
- Chaining. The output of one tool can be the input of the next: look up an order id, then look up the customer for that id. The plan is not stored anywhere as a plan — it exists in the growing message history, and each step is chosen after the previous result is known. Asssistant messages often mix a short line of prose with the tool call ("Let me check the current temperature first"); keeping that text preserves the reasoning trail for later turns, which makes failures much easier to debug.
- Multi-step tasks. Any task where step 2 depends on what step 1 returns needs this shape. Anything where all the steps are known up front does not — you could just call two functions yourself.
How it stops, and why it needs guards
The normal exit is step 3: the model answers without asking for a tool. On OpenAI's API that is finish_reason: "stop"; on Anthropic's it is stop_reason: "end_turn". Your loop ends there because there is nothing left to feed back.
The abnormal exits need your help, because nothing about the API stops a confused model from calling the same tool with the same arguments forever. Three cheap guards cover almost everything: a maximum number of iterations (the max_steps above), a check for repeated tool-plus-argument pairs, and a token budget, since every iteration appends to messages and each new call resends the whole history — cost and latency grow with the length of the loop, not with the difficulty of the question. When a guard trips, return the best partial answer plus a clear status, rather than raising into the void.
What the loop does not fix
The loop is the architecture. It is not the quality. Most first agents that misbehave do so because of the tools, not the loop:
- The description is the model's only documentation. The model picks a tool from its name, its description, and its schema, and it has never read your function body. A vague description ("handles users") gets a tool called in the wrong places; an over-strict schema makes it never call the tool at all.
- Tools are real code with real effects. A read-only search is safe to retry; sending an email or charging a card is not. For side-effecting tools, make the operation idempotent or put a confirmation step in front of it. The loop will happily execute the same write twice if the model decides to.
- The model can request a tool that doesn't exist. Handle the unknown-name case by returning a string the model can read and react to, as
execute_toolabove does.
The mental model to keep
An agent is a state machine whose decision engine is an LLM: your code holds the state (the message list), the model chooses the next transition, and the loop runs until the model stops requesting actions. Everything else — ReAct-style prompts, frameworks, planners, memory stores, MCP tool servers — is elaboration on that one loop. Build it once by hand so you know what the frameworks are doing for you.
Practice: predict the trace
Add a tool celsius_to_fahrenheit(celsius: float) and run the agent on: "What's the temperature in Rome right now, in Fahrenheit?"
Expected trace:
1turn 1 assistant: tool_calls=[get_weather(city="Rome")]
2 tool: "15"
3turn 2 assistant: tool_calls=[celsius_to_fahrenheit(celsius=15)]
4 tool: "59"
5turn 3 assistant: "It's 59 °F in Rome right now."Three model calls, two tool rounds. Notice that the loop count is not something you decide in advance — turn 2 only exists because turn 1's result had to be converted, and a one-shot design could not have known that. Also notice that the same for loop ran three times without changing; if you had to add a branch per extra tool call, you would have written a workflow, not an agent.