What are tokens, and what do they cost?
AI language models do not read letters or words. They read tokens: short pieces of text that the model treats as single units. Tokens matter for two practical reasons. Providers charge by the token, and every model has a limit on how many tokens it can read at once.
What a token looks like
A token can be a whole word, part of a word, a number group or a punctuation mark. In English, short common words are usually one token each. Longer or rarer words are split into several pieces. "Hello, world!" is four tokens: "Hello", the comma, " world" and the exclamation mark.
The piece of software that does the splitting is called a tokenizer. Each model family has its own, built from the text it was trained on, so the same sentence can be 20 tokens on one model and 24 on another.
Rules of thumb
- In English, 100 tokens is roughly 75 words, or one token is about four characters.
- A page of plain prose, around 500 words, is about 650 to 700 tokens.
- Code, tables, numbers and unusual names use more tokens per word than ordinary sentences.
- Languages other than English generally use more tokens for the same meaning.
These are estimates. The token counter gives an estimate for any text you paste; for exact, billing-grade numbers use the provider's own counter for the model you intend to use.
How pricing works
Prices are quoted per million tokens, with two rates:
- Input tokens: everything you send, including instructions, retrieved passages and the question.
- Output tokens: everything the model writes back.
Output is priced several times higher than input, often around five times, because writing text takes more computation than reading it. A short question with a long answer can therefore cost more than a long prompt with a one-line answer.
Working out a cost
The sum is the same for every provider:
cost per request = (input tokens × input price + output tokens × output price) ÷ 1,000,000
Suppose a prompt of 1,000 tokens, an answer of 300 tokens, and a model priced at $1 per million input tokens and $5 per million output tokens. The input costs 1,000 × $1 ÷ 1,000,000 = $0.001. The output costs 300 × $5 ÷ 1,000,000 = $0.0015. One request costs $0.0025. At 1,000 requests a day that is $2.50 a day, or about $75 a month.
Notice that the 300-token answer cost more than the 1,000-token prompt.
Tokens in RAG
In a RAG system the prompt is mostly retrieved passages, so two settings control most of the input cost: the passage size and the number of passages retrieved. Five passages of 200 words is about 1,300 tokens of context on every question; three passages of 60 words is about 240.
There is also a one-time cost to embed the documents, charged on the total tokens of all passages. Overlap increases it, because repeated words are embedded twice. The chunking visualizer shows the estimated total for your settings.
Context windows
Every model has a context window: the most tokens it can handle in one request, input and output together. Current models have large windows, but filling them is rarely a good idea. Long prompts cost more, respond more slowly, and models tend to make less use of material buried in the middle of a very long context.
Ways to spend less
- Retrieve fewer or smaller passages when testing shows the answer is still found.
- Ask for shorter answers. Output is the expensive side.
- Use a smaller model for simple questions. The price gap between a provider's largest and smallest model is often ten times or more.
- Use provider discounts such as caching of repeated prompt text and batch processing, where your use allows them.
Prices change often. Compare current rates across models in the token counter and cost estimator.
More guides
What is RAG? · How to choose a chunk size for RAG · What are embeddings? · Hybrid search: words plus meaning · How to test RAG retrieval · Seven common RAG mistakes