Artificial intelligence

How to count the tokens in a prompt and estimate what a call to a language model costs

Why models charge per token, how to count tokens exactly for GPT and approximately for Claude, and how to cut the cost of each call without losing quality.

Language models charge and set limits in tokens, not words or characters. A token is the fragment of text the model works with: in English it averages about 4 characters, and in ordinary prose in other European languages the figure is similar, although short sentences, punctuation and rare words split into more pieces. To know what a call costs you count the tokens of the input (your prompt plus the context) and of the output (the reply), which usually have different prices, and multiply them by the model’s rate per million tokens.

This article explains how to count tokens exactly or approximately depending on the model, what figures to expect, where the tokens you do not see hide, and how to reduce the bill.

Why tokens and not words

As explained in how a language model works, the model does not see letters or words but tokens from a fixed vocabulary built during training. The computational cost of processing a text depends on the number of tokens, so prices and limits (context window, maximum reply length) are expressed in that unit.

Practical consequence: the same text costs differently depending on the language and on the model, because each family has its own tokenizer.

How to count tokens

OpenAI models: exact counts

OpenAI publishes its tokenizers. o200k_base is the one for GPT-4o and later models; cl100k_base the one for GPT-4 and GPT-3.5. You can count locally with tiktoken (Python) or gpt-tokenizer (JavaScript):

import { encode } from 'gpt-tokenizer'; // o200k_base by default
const tokens = encode('La indexación no está garantizada.');
console.log(tokens.length); // 8
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
len(enc.encode("La indexación no está garantizada."))  # 8

The token counter on AIMRAN Tools uses those same tokenizers in the browser and shows the count for o200k and cl100k, the characters per token of the text you paste and the estimated cost if you enter the model’s rate per million tokens. The text never leaves your machine.

Claude and other models: estimate or API

Anthropic does not publish Claude’s tokenizer; for an exact count you call its token counting endpoint, which does not consume generation tokens. To estimate without calling the API, the usual practice is to start from the cl100k count and apply a correction factor; that is what the AIMRAN Tools counter does, marking that figure as approximate. Good enough for budgeting; for billing, use the API.

How many tokens a text has

As a guide, measured with o200k_base while writing this article:

Text Characters Tokens Characters per token
“Hola, ¿qué tal?” (Spanish) 15 6 2.5
“Hello, how are you?” 19 6 3.2
The first paragraph of the Spanish version of this article 590 127 4.6

Short sentences are proportionally expensive (punctuation and opening marks count as tokens of their own); in long prose the average approaches 4 to 4.5 characters per token in English and Spanish alike. That Spanish paragraph has 105 words and 127 tokens, that is, about 1.2 tokens per word. JSON code and URLs are considerably more expensive because every punctuation symbol tends to be a token.

The tokens you do not see

When estimating the cost of a call it is easy to forget that the input is not just your question:

  • The system prompt and the application’s fixed instructions, which travel with every request.
  • The conversation history, resent in full (or summarised) on every turn. A long conversation can cost more for the history than for the new question.
  • Tool descriptions in agents and the results they return.
  • Attached documents (RAG): a twenty-page PDF is thousands of tokens per call if you include it whole.
  • Formatting tokens the provider adds per message (a few per turn).
  • The output, which almost always has a higher price per token than the input and which you can cap with the maximum-tokens parameter.

API responses return the real usage (usage with input and output tokens): log those values so you do not depend on estimates.

How to calculate the cost

cost = (input_tokens / 1,000,000) × input_price + (output_tokens / 1,000,000) × output_price

Prices vary between models by more than an order of magnitude and change every few months, so always take them from the provider’s pricing page at the time of calculating. Many providers also offer prompt caching: the repeated part (system prompt, documents) is billed far cheaper when reused in nearby calls, which changes the sums considerably for applications with a large fixed context.

How to reduce tokens without losing quality

  1. Trim the system prompt: remove repetitions and examples the model no longer needs.
  2. Summarise the history after a certain number of turns instead of resending it whole.
  3. Chunk documents and send only the relevant fragments (retrieval first) instead of the whole document.
  4. Ask for concise outputs and set a reasonable maximum of reply tokens.
  5. Use the right model: classifying or extracting data does not need the most expensive model.
  6. Exploit prompt caching by placing the stable part first.
  7. Measure before optimising: paste the real prompt into a counter and see where the tokens go; it is often a long example or an unminified JSON.

Conclusion

Counting tokens is the only way to know what a call costs and what fits in it. For OpenAI models the count is exact with their public tokenizers; for Claude, approximate unless you use its counting API. Expect roughly 4 characters per token in prose, remember that the history and the system prompt count too, and log the real usage each response returns: that is the number that ends up on the invoice.

Sources and references

  1. OpenAI Help: What are tokens and how to count them? help.openai.com
  2. OpenAI: Tokenizer platform.openai.com
  3. tiktoken (OpenAI, GitHub) github.com
  4. Anthropic: Token counting docs.anthropic.com
  5. gpt-tokenizer (npm) npmjs.com

Related tools

Herramientas gratuitas de AIMRAN Tools que funcionan en tu navegador, sin registro.

Ver todas las herramientas
  • Token counter

    Counts a text’s tokens with the GPT tokenizers, estimates Claude’s and calculates the cost per million tokens.

Artificial intelligence

What an AI agent is and how it differs from a chatbot

An AI agent combines a language model with tools and a decision loop to complete tasks. How it differs from a chatbot, how it works, its risks and when to use one.

4 min read