Skip to main content

Build Your First AI Agent in Code

A working AI agent in about sixty lines of Python, no framework: tools, the loop, and limits enforced in code. Then what to add before production.

IntermediateVerdeshell Team · 8 min read · Last reviewed

An agent is a model, some tools, and a loop: send the goal and tool definitions, run whichever tool the model asks for, send back the result, and repeat until it answers. Write that loop yourself once before you use a framework.

Key takeaways

  • An AI agent in code is a model, a set of tool definitions, and a loop that runs the tools the model asks for.
  • The model never executes anything itself — it returns a structured request, and your code decides whether and how to run it.
  • Put limits such as refund caps and permissions in the tool code, never only in the prompt.
  • A turn limit is the backstop stop condition every loop needs.
  • Before production, add input validation, logging of every step, timeouts, an evaluation set and approval for consequential actions.
How a tool call travels between your code, the model and a toolYour code sends the goal and the tool definitions to the model. The model replies with a tool use request, for example get_order for order A1001. Your code runs the tool, applying its own limits, and receives the result. Your code sends the result back to the model as a tool result with the same ID. This repeats until the model replies with a final answer and no more tool calls.Your codeModelToolgoal + tool definitionstool_use: get_order(A1001)run it — your limits applyresulttool_result (same ID)final answer — no more tool callsThe middle four steps repeat for every tool the model asks for. The model never runs anything itself.
The model never runs anything. It asks; your code runs the tool and sends back the result — until the model stops asking.

Hover or tap the diagram to replay the animation.

What you are building

The smallest useful agent has three parts: a model, a set of tools it can ask to use, and a loop. Anthropic’s guidance on building agents calls a model with tools, retrieval and memory an “augmented” model; put it in a loop where it chooses its own next step, and you have an agent.

The example below is a support agent with two tools — look up an order, and issue a refund — and a refund limit. It uses Anthropic’s Python SDK because the code has to use some provider; the same structure works with any model that supports tool calling, and the last section shows the equivalent names for OpenAI.

Step 1 · Describe the tools

Each tool has a name, a description and an input schema — a JSON Schema saying which arguments it takes. The model sees only these descriptions, never your code, so the description is where you tell it when to use the tool and what it returns. Anthropic’s documentation calls detailed descriptions by far the most important factor in tool performance.

Keep tools narrow. “issue_refund(order_id, amount)” is easier for a model to use correctly, and for you to limit, than “run_sql(query)”.

Step 2 · Write the loop

import json
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY
MODEL = "claude-sonnet-5-5"     # any current model id

ORDERS = {"A1001": {"status": "delivered", "total": 40.0}}
REFUND_LIMIT = 50.0  # enforced here, in code, not in the prompt

def get_order(order_id):
    return ORDERS.get(order_id, {"error": "No such order"})

def issue_refund(order_id, amount):
    if amount > REFUND_LIMIT:
        return {"error": "Over the limit. Needs a person."}
    return {"refunded": amount, "order_id": order_id}

TOOLS = [
    {"name": "get_order",
     "description": "Look up an order by ID. "
                    "Returns its status and total.",
     "input_schema": {
         "type": "object",
         "properties": {"order_id": {"type": "string"}},
         "required": ["order_id"]}},
    {"name": "issue_refund",
     "description": "Refund a delivered order, "
                    "up to the order total.",
     "input_schema": {
         "type": "object",
         "properties": {"order_id": {"type": "string"},
                        "amount": {"type": "number"}},
         "required": ["order_id", "amount"]}},
]
HANDLERS = {"get_order": get_order,
            "issue_refund": issue_refund}

def run_agent(goal, max_turns=10):
    messages = [{"role": "user", "content": goal}]
    for _ in range(max_turns):
        response = client.messages.create(
            model=MODEL, max_tokens=1024,
            tools=TOOLS, messages=messages,
            system="You are a support agent. Use the tools. "
                   "Never guess order details.")
        messages.append({"role": "assistant",
                         "content": response.content})
        if response.stop_reason != "tool_use":  # finished
            return "".join(b.text for b in response.content
                           if b.type == "text")
        results = []
        for block in response.content:
            if block.type == "tool_use":
                output = HANDLERS[block.name](**block.input)
                results.append({"type": "tool_result",
                                "tool_use_id": block.id,
                                "content": json.dumps(output)})
        messages.append({"role": "user", "content": results})
    return "Stopped: turn limit reached."

print(run_agent("Order A1001 arrived damaged. Refund it in full."))

The loop sends the conversation and the tool definitions to the model. If the response’s stop reason is “tool_use”, the model is asking for one or more tools: the code runs each one and sends the results back as “tool_result” blocks, matched to each request by its ID. When the model replies without asking for a tool, that reply is the answer.

Two details trip people up. The model’s full response, including its tool requests, must be added to the conversation before the results; and all results for one turn go back together in a single message.

Step 3 · Put the limits in code

Notice where the refund limit lives: in “issue_refund”, not in the instructions. A prompt saying “never refund more than 50” is a request the model usually follows. Code that refuses is a guarantee.

This matters because tool results are untrusted input. An order note, an email or a web page the agent reads can contain text written to manipulate it — a prompt injection. If the only thing between that text and a large refund is a sentence in the prompt, the limit is not real.

The same principle covers permissions. Give each tool the narrowest access it needs: read-only where possible, scoped to one customer or one folder, and with consequential actions returning “needs approval” rather than acting.

Step 4 · Stop conditions

This agent stops when the model answers without asking for a tool — the natural end. The “max_turns” limit is the backstop: without it, a confused model can call tools indefinitely and run up cost.

For real tasks, the stronger stop condition is one the code can test — the refund recorded, the tests passing, the ticket closed. See anatomy of a good agent loop for why “testable” is most of the engineering.

The same loop with other providers

With OpenAI’s Responses API the shape is identical and the names differ: tools are declared as functions with a JSON Schema for their parameters, the model returns “function_call” items with a call ID, and you send back “function_call_output” items carrying that ID.

SDKs and frameworks will run this loop for you — Anthropic’s SDK has a tool runner helper, and every agent framework includes one. Use them once you know what they are doing; when something goes wrong, you will need to read the turn-by-turn messages, and having written the loop yourself is what makes them readable.

Before production

Validate tool inputs. The model can return an unknown tool name or malformed arguments; the example would crash. Check every call against its schema — strict schema modes help — and return a clear error the model can recover from.

Log every step: each request, tool call, result and answer. This is your debugging record and your audit trail.

Add timeouts and retries around tools, and a budget per run.

Build an evaluation set of real cases with known correct outcomes, and run it on every change to the prompt, tools or model.

Keep a person approving anything consequential until the evaluation results justify removing them. Then add memory and standard tool connections as the task requires.

Next in the pathTool Use, Function Calling and MCP, Explained

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.