Why AI Hallucinates, and How to Reduce It
Why language models give fluent, confident answers that are wrong, the two kinds of hallucination, and the techniques that actually reduce them.
BeginnerVerdeshell Team · 9 min read · Last reviewed
A language model produces the most plausible continuation, and plausible is not the same as true. Hallucinations cannot be switched off — but grounding the model in your documents, letting it say “I don’t know” and checking its claims reduce them substantially.
Key takeaways
- A hallucination is fluent, confident output that is false or unsupported — a consequence of how models generate text, not a bug that gets patched.
- Models guess partly because training and evaluation reward a confident answer over “I don’t know”.
- Factuality errors contradict the world; faithfulness errors contradict the documents or instructions you gave the model.
- The most effective fixes: supply the facts (retrieval), allow “I don’t know”, ask for supporting quotes, and verify anything that matters.
What a hallucination is
In AI, a hallucination is output that sounds right and is not: an invented citation, a confident wrong date, a policy clause that does not exist, a function call to an API that was never written. The defining feature is the confidence. A model that is wrong usually sounds exactly as sure as a model that is right.
It is worth separating two kinds, a distinction used in research surveys of the problem. A factuality hallucination contradicts the world — the wrong capital city, a made-up statistic. A faithfulness hallucination contradicts what you gave the model — a summary that adds a point the document never made, or an answer that ignores your instruction. They have different fixes, so it helps to know which one you are looking at.
Why models hallucinate
A model writes one token at a time, each time choosing a plausible continuation of everything before it — the mechanism in how large language models work. When the right answer is well represented in what it learned, the plausible continuation is usually correct. When it is not — a rare fact, a recent event, something only your company knows — the model still produces a plausible continuation, and plausible is not the same as true.
Research published by OpenAI in 2025 (Kalai et al.) argued that two things keep the problem alive. Pre-training makes some errors statistically inevitable for facts that appear rarely. And the way models are evaluated rewards guessing: most benchmarks give credit for a right answer and nothing for “I don’t know”, so a model that always guesses scores better than one that admits uncertainty — like a student guessing on a multiple-choice exam.
Two common beliefs do not hold up. Lowering the temperature makes output more repeatable, not more accurate; a model can be consistently wrong. And a longer context does not guarantee faithfulness: a model can be given the right document and still answer from its general knowledge, or miss the relevant passage in a long input.
Ground the model in your sources
The most effective single fix for factuality errors is to stop relying on what the model remembers. Put the relevant documents in front of it at the time of the question — retrieval-augmented generation — and tell it to answer only from them.
Anthropic’s guidance on reducing hallucinations adds a step for long documents: ask the model to extract the word-for-word quotes relevant to the question first, then answer using only those quotes. The answer is then anchored in text you can check.
Answer the question using only the policy below.
First, copy the exact sentences from the policy that are relevant, in <quotes> tags.
Then answer using only those quotes. If no quote answers the question,
reply "The policy does not cover this."
<policy>
{policy text}
</policy>
Question: {question}Let it say “I don’t know”
If the instructions imply that an answer is always expected, the model will produce one. Explicitly permitting uncertainty — “if the documents do not contain the answer, say so” — is one of the simplest and most effective techniques in Anthropic’s guidance.
Make the escape route specific and easy to detect. A fixed phrase such as “I don’t have enough information to answer this” lets your application route the question to a person instead of showing a guess.
Check before you trust
Ask for citations and verify them. Have the model cite a supporting quote for each claim, then check that each quote really appears in the source — automatically where you can. Anthropic suggests a second pass in which the model must find a supporting quote for every claim and remove any it cannot support.
Compare several answers. Running the same question more than once and comparing the results can expose hallucinations: where the answers disagree, the model is not sure. This is the same idea as self-consistency in chain-of-thought prompting.
Ask for reasoning on multi-step questions. Seeing the steps can reveal a wrong assumption — remembering that the written reasoning is not a guaranteed account of how the answer was reached.
Keep a person in the loop where errors are expensive. For legal, medical, financial or customer-facing commitments, the model drafts and a person approves.
What this means when you are building
None of these techniques eliminates hallucination — Anthropic’s own guidance says as much — so design for it. Decide which errors would be costly, and put grounding, verification or human review exactly there rather than everywhere.
Measure it. Keep a set of real questions with known answers, including ones the documents do not cover, and track how often the system answers wrongly or answers when it should have declined. That number is the one to improve, and the one to check every time you change the model or the prompt.
More in Large Language Models
Want this built properly?
We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.