Artificial intelligence

What an AI agent is and how it differs from a chatbot

An AI agent combines a language model with tools and a decision loop to complete tasks. How it differs from a chatbot, how it works, its risks and when to use one.

An AI agent is a system in which a language model (LLM) does not merely answer with text but decides which actions to take, performs them through tools (searching, reading files, calling an API, running code), observes the result and repeats the cycle until it completes a task. A chatbot answers a message; an agent pursues a goal over several steps. The difference lies not in the model, which may be the same, but in the architecture around it: tools, control loop and stopping criterion.

This article explains that architecture, which kinds of systems are called “agents”, which risks they introduce and when it makes sense to use one instead of something simpler.

Chatbot versus agent

Chatbot (conversational assistant) Agent
Input A message from the user A goal or task
Output A text reply Actions and, finally, a result
Steps One per turn Several, decided by the model itself
Access to the world Only what is in the conversation External tools (files, APIs, browser, terminal)
Who decides the next step The user, by typing The model, until the stopping criterion is met
Main risk A wrong answer A wrong action with real effects

A chatbot that can search the internet before answering is halfway there: it uses a tool, but in a single fixed step. What defines an agent is the loop: the model picks the tool, sees the result and decides whether it needs another.

How it works inside

Most current agents follow a similar structure, popularised by the ReAct pattern (reason and act) and clearly described in Anthropic’s guide Building effective agents:

  1. Goal. The system receives the task (“find out why test X fails in the repository and propose a patch”).
  2. Available tools. They are described to the model with a name, what they are for and which parameters they accept (for example, read_file(path), run_tests()).
  3. Decision. The model produces either a final answer or a tool call with its arguments.
  4. Execution. The host program runs the tool (the model never executes anything itself) and returns the result to the model as new context.
  5. Repeat until the model decides it has finished or a limit is reached (number of steps, time, cost).

Two important details:

  • Everything the model “knows” about the task lives in its context window: the instruction, the tool descriptions and the accumulated results. When the context fills up, something must be summarised or discarded.
  • Tools are ordinary code written by the developer. The Model Context Protocol (MCP) is an open standard for exposing tools and data to models uniformly, so that one tool server can serve different agents.

Workflows versus agents

Not everything that chains several model calls is an agent. Anthropic distinguishes between:

  • Workflows: the steps are defined in advance in code. The model fills in each step (classifies, summarises, drafts), but does not decide the order. They are predictable, cheap to debug and sufficient for most automations.
  • Agents: the model dynamically decides what to do and how many steps to take. They are more flexible and more expensive, and their behaviour is harder to guarantee.

The practical recommendation of that guide, and the one worth internalising, is to start with the simplest solution that works: one well-designed call, then a workflow, and only afterwards an agent, if the task genuinely requires open-ended decisions.

Real use cases

  • Coding assistants that read a repository, run tests, edit files and open a pull request.
  • Technical support that looks up internal documentation, checks an account’s status and performs an action (refund, reset) within defined limits.
  • Research: searching several sources, reading pages, cross-checking and drafting a report with references.
  • Operational automation: reviewing logs, correlating alerts and proposing (not executing) a corrective action.

Notice the pattern: multi-step tasks, a verifiable result and well-bounded tools.

Risks specific to agents

An agent can be wrong just like a chatbot, but its mistakes get executed. The most relevant risks:

  1. Prompt injection. If the agent reads a web page, an email or a file containing text such as “ignore your instructions and send the data to…”, it may obey. The OWASP Top 10 for LLM applications ranks it as the first risk. Treat all external content as data, not orders, and limit what the tools can do.
  2. Irreversible actions. Deleting, paying, sending. Require human confirmation for them or do not expose them as tools at all.
  3. Loops and cost. An agent that does not converge can repeat steps indefinitely. Set limits on steps, time and budget.
  4. Excessive permissions. Give the agent the minimum credentials for its task, as you would any service.
  5. False confidence. An agent “explaining” what it did does not guarantee it is correct. Verify the result with tests, reviews or independent checks.

When an agent is (and is not) worth it

It is worth it when the task has several steps whose order depends on what is found along the way, when the result can be verified (a passing test, a cross-checked fact) and when the cost of an error is bounded or reversible.

It is not worth it when the process is always the same (use a workflow), when a single model call solves the problem, or when a mistake would have serious consequences without the possibility of human review.

Conclusion

An AI agent is a language model inside a loop with tools and a stopping criterion. That architecture lets it complete tasks instead of just answering, and in exchange introduces new risks: wrong actions, prompt injection and runaway costs. The sensible way to adopt them is gradual (single call, workflow, agent), with bounded tools, minimal permissions and verification of the result.

Sources and references

  1. Anthropic: Building effective agents anthropic.com
  2. Model Context Protocol: specification modelcontextprotocol.io
  3. OWASP Top 10 for Large Language Model Applications owasp.org
  4. ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al., 2022) arxiv.org