← BlogAI Agents

smolagents: planning, tools, and delegation

Isaac Kargar4 min read

  • AI Agents
  • smolagents
  • Tools
  • HuggingFace

smolagents documentation provides a small framework for agents that use tools. The common base is MultiStepAgent; CodeAgent expresses tool calls as Python code, while ToolCallingAgent expresses them as structured tool calls. This walkthrough targets the documented v1.26.0 interface.

The agent loop is simple: the model chooses an action, a tool or executor produces an observation, and the next step uses the updated memory. A run ends when the agent returns an answer or reaches its step limit. The library’s planning option inserts a planning step at a configured cadence.

CodeAgent writes a Python action for the executor. ToolCallingAgent emits a structured function call that the framework dispatches. The comparison is about the action interface; both still use the same step and memory lifecycle.

Install one supported version

Pin the version used by the application and test the model adapter with it.

python -m pip install "smolagents==1.26.0"

The first example uses InferenceClientModel with a public instruct model identifier. A local or self-hosted adapter can replace it without changing the agent code; credentials and the endpoint belong in the runtime configuration.

A tool and a CodeAgent

Tools are ordinary typed Python functions with a description that tells the model when to use them. The decorator builds the tool schema.

from smolagents import CodeAgent, InferenceClientModel, tool


model = InferenceClientModel(model_id="meta-llama/Meta-Llama-3.1-8B-Instruct")


@tool
def word_count(text: str) -> int:
    """Return the number of whitespace-separated words in text.

    Args:
        text: The passage whose words should be counted.
    """
    return len(text.split())


agent = CodeAgent(
    tools=[word_count],
    model=model,
    max_steps=6,
    additional_authorized_imports=["math"],
)

answer = agent.run("Count the words in 'a small tool call'.")
print(answer)

CodeAgent lets the model write a short code action that calls the tool. The executor runs that action, records the result, and gives the observation to the next step. Keep the authorized import list explicit. Allowing an import tells the executor that the module may be used; it is not a security sandbox. Use an isolated executor such as a container or a hosted sandbox when generated code is not trusted.

ToolCallingAgent

ToolCallingAgent uses the model’s structured tool-call interface instead of generated Python. It is useful when the provider has reliable tool-call support or when the application wants to avoid a code action for every step.

from smolagents import ToolCallingAgent


structured_agent = ToolCallingAgent(
    tools=[word_count],
    model=model,
    max_steps=6,
)

answer = structured_agent.run("Count the words in 'structured tool calls'.")
print(answer)

The two agents share the MultiStepAgent run lifecycle and memory model. Their action representation differs, so a callback or message consumer should handle the corresponding message and tool-call types rather than parsing a text representation by hand.

The MultiStepAgent lifecycle

The public run method accepts a task and options such as stream, reset, images, additional_args, max_steps, and return_full_result. With reset=True, a new run starts with a fresh memory. With streaming enabled, the method yields intermediate step results; otherwise it returns the final result.

result = agent.run(
    "Count the words in 'one two three'.",
    reset=True,
    max_steps=4,
)

Each step has a task, an action, an observation, and the next action or final answer. The step limit is a termination bound, not a promise that the task will be solved. A caller should handle the result and any execution error explicitly.

Planning cadence

Set planning_interval on an agent when a task benefits from a plan that is refreshed during execution.

planned_agent = CodeAgent(
    tools=[word_count],
    model=model,
    planning_interval=3,
    max_steps=8,
)

In the current implementation, planning runs on step 1 and then when (step_number - 1) % planning_interval == 0. With an interval of 3, the planned steps are 1, 4, and 7. This cadence is different from checking step_number % planning_interval == 0.

The planning prompt has separate templates for an initial plan and a later update. When customizing the later message, use the post-planning key:

post_planning_template = agent.prompt_templates["planning"][
    "update_plan_post_messages"
]

The key is a message template, not the pre-planning template. Keep the task, known facts, observations, and remaining steps in the update prompt so the plan reflects current state.

Memory and callbacks

The agent stores the system prompt and task steps in its memory. A callback can inspect a step after it has finished and record a bounded application metric.

from smolagents import ActionStep, CodeAgent


def record_step(memory_step: ActionStep, agent: CodeAgent) -> None:
    timing = memory_step.timing
    if timing is not None and timing.end_time is not None:
        duration = timing.end_time - timing.start_time
        print(f"step {memory_step.step_number} duration: {duration:.3f}s")


observed_agent = CodeAgent(
    tools=[word_count],
    model=model,
    step_callbacks={ActionStep: record_step},
)

The callback receives the completed memory step and the agent. The current step’s timing object supplies both timestamps, so the final step is measured from its own start time. Persist only the memory needed by the application and treat model-generated memory as untrusted input.

Delegating to a managed agent

An agent can expose another agent as a managed capability. The manager receives the managed agent’s name and description and can delegate a task to it.

from smolagents import CodeAgent


researcher = CodeAgent(
    tools=[word_count],
    model=model,
    max_steps=4,
    name="researcher",
    description="Counts words in a supplied passage and reports the result.",
)

manager = CodeAgent(
    tools=[],
    model=model,
    managed_agents=[researcher],
    max_steps=6,
)

The manager’s tool description should state the task the managed agent can perform and the shape of the result it returns. Give the managed agent only the tools and imports needed for that task. Delegation adds another model call and another failure boundary, so direct tool use is simpler when no separate context or role is needed.

Choosing an agent type

CodeAgent is a good fit when a short sequence of calculations or several tool calls are easier to express as code and the executor is appropriately isolated. ToolCallingAgent fits providers with dependable structured tool calls and tasks that do not need generated code. Both are built on MultiStepAgent, so both support step limits, memory, streaming, callbacks, and optional planning.

The library’s import authorization is a capability list, not a complete sandbox. Use an explicit list of modules, avoid wildcard imports, and run untrusted generated code in a separate execution boundary. Tool descriptions, argument validation, and resource limits should be enforced by the application as well as described in the prompt.

References

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →