Tokens and Context Windows, Explained
What a token is, why cost and limits are counted in tokens rather than words, what a context window holds, and why filling it is not the same as using it well.
BeginnerVerdeshell Team · 8 min read · Last reviewed
Models read and write tokens — pieces of words — not words. Price, speed and limits are all counted in them. The context window is everything the model can consider for one response, including its own answer, and more of it is not automatically better.
Key takeaways
- A token is a piece of text — a word, part of a word or punctuation; in English one token is roughly four characters, or three-quarters of a word.
- Cost, speed and limits are counted in tokens, and the same text can produce different counts on different models and in different languages.
- The context window is everything a model can consider for one response — instructions, documents, history and its own answer.
- A bigger context window is not a better one: accuracy can degrade as it fills, so curate what goes in.
What a token is
A language model does not read words or letters. Before any text reaches it, a tokenizer splits the text into tokens: common words become one token, rarer words are split into pieces, and punctuation, numbers and even leading spaces have tokens of their own. The model reads a sequence of tokens and writes one token at a time — the mechanism described in how large language models work.
OpenAI’s rule of thumb for English is that one token is roughly four characters, or about three-quarters of a word — so 100 tokens is around 75 words. It is an estimate, not a conversion rate.
Why the count matters
Almost everything practical about a model is measured in tokens: the price of a request, how fast it responds, how much it can read at once and how long its answer can be.
Counts are not portable. The same text produces different token counts on different models, because each model family has its own tokenizer, and many languages other than English produce more tokens for the same meaning. If you are budgeting for content in Hindi, Tamil or another language, measure with your actual text rather than applying the English rule of thumb.
Some tokens are invisible. Reasoning models generate internal reasoning tokens before their answer; you do not see them as text, but they count as output and are billed. A short answer can cost much more than its length suggests.
Tokenization also explains some odd failures. A model that sees a word as one or two pieces does not see its individual letters, which is why tasks like counting characters or reversing words can trip up an otherwise capable model.
What a context window is
The context window is the maximum amount of text a model can consider when producing one response. Anthropic’s documentation describes it as the model’s working memory — distinct from everything the model learned in training.
It holds everything in the request: the system prompt, the conversation so far, any documents or tool results you supply, tool definitions — and the answer the model is about to write. If the input fills the window, there is no room left for the output.
Anything outside the window does not exist for that response. A model has no memory of earlier conversations unless the application puts them back into the context.
Bigger is not automatically better
Context windows have grown very large, and it is tempting to put everything in. Two things argue against it.
Quality can degrade as the window fills. Anthropic’s documentation calls this context rot: as the token count grows, accuracy and recall degrade. Research published in 2023 (Liu et al., “Lost in the Middle”) found that models use information at the start and end of a long context more reliably than information in the middle.
Cost and speed scale with it. Every token in the window is processed on every request. Sending a whole manual to answer a question about one page is slower and more expensive than sending the right page.
The better approach is to select what goes in — the relevant passages rather than the whole library. That is what retrieval does, and why practitioners now talk about context engineering.
Managing tokens in practice
Measure, do not guess. Model providers offer tokenizers and token-counting tools; use them on your real prompts and documents before estimating costs.
Keep the stable part of the prompt stable. Several providers offer prompt caching, which makes a repeated prefix — long instructions, a fixed document — cheaper on later requests. It changes what you pay for those tokens, not whether they count towards the window.
Trim long conversations. In chat and agent applications the history grows with every turn; summarise or drop what is no longer needed rather than letting old turns crowd out the current task.
Compare models on total cost per task, not price per token. OpenAI’s own guidance makes the point: a lower price per million tokens does not necessarily mean a lower total cost, because models tokenize differently and produce different amounts of output and reasoning. Test with representative tasks — more on this in choosing a language model.
More in Large Language Models
Want this built properly?
We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.