Chain-of-Thought Prompting, Explained
How asking a model to reason step by step improves multi-step answers, the variants built on it, and why reasoning models change the advice.
IntermediateVerdeshell Team · 9 min read · Last reviewed
Chain-of-thought prompting asks a model to work through intermediate steps before giving its answer. It helps on problems that genuinely have steps — arithmetic, logic, planning — and does little for simple lookups. Reasoning models now do it on their own.
Key takeaways
- Chain-of-thought (CoT) prompting asks the model to reason through intermediate steps before answering.
- It helps most on multi-step problems such as arithmetic, logic and planning, and little on simple factual questions.
- Self-consistency samples several reasoning paths and takes the most common answer.
- Reasoning models reason by default — do not tell them to “think step by step”.
What chain-of-thought prompting is
Chain-of-thought (CoT) prompting asks a model to produce intermediate reasoning before its final answer. The technique was described in a 2022 Google paper by Jason Wei and colleagues, which showed that giving large models a few examples with the working written out improved their results on arithmetic, commonsense and symbolic reasoning tasks.
The intuition follows from how language models generate text: one token at a time. Asked for the answer immediately, the model has to produce it in a single step. Asked to write out the steps first, each step becomes context for the next — the model gets room to work.
Few-shot and zero-shot chain-of-thought
The original method was few-shot: show worked examples that include the reasoning, then pose the new question.
Q: A shop had 23 boxes. It sold 9 and received 2 deliveries of 12. How many now?
A: Start with 23. After selling 9: 23 − 9 = 14. Two deliveries of 12 add 24: 14 + 24 = 38. The answer is 38.
Q: A warehouse holds 120 crates. 45 ship out and 3 trucks deliver 18 each. How many now?
A:A follow-up paper the same year (Kojima et al.) found a shortcut: simply adding “Let’s think step by step” to the prompt, with no examples, produced much of the same improvement. That is zero-shot chain-of-thought, and it is why the phrase became so widely used.
Variants that build on it
Self-consistency (Wang et al., 2022) samples several independent reasoning paths for the same question and takes the most common final answer. Different paths make different mistakes; the majority is more often right. The cost is several generations per answer.
Least-to-most prompting (Zhou et al., 2022) first asks the model to break a hard problem into simpler subproblems, then solves them in order, each building on the last — useful when a problem is harder than the examples it was shown.
Tree of Thoughts (Yao et al., 2023) goes further, letting the model explore and evaluate several branches of reasoning and back up when one fails. It is powerful for search-like problems and considerably more expensive; most business tasks do not need it.
When it helps, and when it does not
Chain-of-thought helps where a problem really has steps: calculations, multi-condition rules, planning, debugging, comparing options against criteria. It does little for single-step tasks — a factual lookup or a simple classification has no chain to write out.
It also costs more. Every reasoning token is generated and billed, and the answer takes longer. For a high-volume task, test whether the improvement is worth it on your own inputs.
Do not mistake the reasoning for an explanation
The written steps are not guaranteed to be how the model actually reached its answer. Research published in 2023 (Turpin et al.) showed that models’ chain-of-thought explanations can be systematically unfaithful — a plausible-sounding rationale for an answer that was in fact driven by something else in the prompt.
Treat the reasoning as working, useful for spotting errors, not as an audit trail. In a product, ask for the reasoning and the final answer in separate, clearly tagged parts so your code can use the answer without parsing the working.
Reasoning models change the advice
Reasoning models are trained to produce their own chain of thought before answering. The providers’ guidance is consistent: do not tell them to “think step by step”, and do not prescribe the steps — describe the goal and the constraints, and let the model plan. Anthropic’s guidance notes that a general instruction to think thoroughly often produces better reasoning than a hand-written plan.
Chain-of-thought prompting is still the right tool for standard models, and for the many production tasks where a smaller, faster model is the better choice.
Want this built properly?
We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.