What are tokens, and what do they cost?

AI language models do not read letters or words. They read tokens: short pieces of text that the model treats as single units. Tokens matter for two practical reasons. Providers charge by the token, and every model has a limit on how many tokens it can read at once.

What a token looks like

A token can be a whole word, part of a word, a number group or a punctuation mark. In English, short common words are usually one token each. Longer or rarer words are split into several pieces. "Hello, world!" is four tokens: "Hello", the comma, " world" and the exclamation mark.

The piece of software that does the splitting is called a tokenizer. Each model family has its own, built from the text it was trained on, so the same sentence can be 20 tokens on one model and 24 on another.

Rules of thumb

These are estimates. The token counter gives an estimate for any text you paste; for exact, billing-grade numbers use the provider's own counter for the model you intend to use.

How pricing works

Prices are quoted per million tokens, with two rates:

Output is priced several times higher than input, often around five times, because writing text takes more computation than reading it. A short question with a long answer can therefore cost more than a long prompt with a one-line answer.

Working out a cost

The sum is the same for every provider:

cost per request = (input tokens × input price + output tokens × output price) ÷ 1,000,000

Suppose a prompt of 1,000 tokens, an answer of 300 tokens, and a model priced at $1 per million input tokens and $5 per million output tokens. The input costs 1,000 × $1 ÷ 1,000,000 = $0.001. The output costs 300 × $5 ÷ 1,000,000 = $0.0015. One request costs $0.0025. At 1,000 requests a day that is $2.50 a day, or about $75 a month.

Notice that the 300-token answer cost more than the 1,000-token prompt.

Tokens in RAG

In a RAG system the prompt is mostly retrieved passages, so two settings control most of the input cost: the passage size and the number of passages retrieved. Five passages of 200 words is about 1,300 tokens of context on every question; three passages of 60 words is about 240.

There is also a one-time cost to embed the documents, charged on the total tokens of all passages. Overlap increases it, because repeated words are embedded twice. The chunking visualizer shows the estimated total for your settings.

Context windows

Every model has a context window: the most tokens it can handle in one request, input and output together. Current models have large windows, but filling them is rarely a good idea. Long prompts cost more, respond more slowly, and models tend to make less use of material buried in the middle of a very long context.

Ways to spend less

Prices change often. Compare current rates across models in the token counter and cost estimator.

More guides

What is RAG? · How to choose a chunk size for RAG · What are embeddings? · Hybrid search: words plus meaning · How to test RAG retrieval · Seven common RAG mistakes