Large Language Models, Explained
What is actually happening between your prompt and the answer — tokens, attention, pre-training and post-training — why models make things up and how to reduce it, and how to choose the right model for a task.
4 articles · start with the first and read in order, or jump to what you need
- 1 · BeginnerHow Large Language Models Work, Without the MathsTokens, attention, pre-training, instruction tuning and next-token prediction — what happens between your prompt and the answer, and why models fail.9 min read
- 2 · BeginnerTokens and Context Windows, ExplainedWhat a token is, why cost and limits are counted in tokens rather than words, what a context window holds, and why filling it is not the same as using it well.8 min read
- 3 · BeginnerWhy AI Hallucinates, and How to Reduce ItWhy language models give fluent, confident answers that are wrong, the two kinds of hallucination, and the techniques that actually reduce them.9 min read
- 4 · IntermediateHow to Choose a Language ModelHosted or open-weight, large or small, reasoning or standard — the trade-offs that matter, and how to decide with your own tasks, not a leaderboard.9 min read
Common questions: Large Language Models
How does a large language model work?
It splits text into tokens and predicts the next token, one at a time, based on everything before it. Pre-training on a very large body of text teaches it to continue text; further training on instructions and on feedback about which answers are better turns it into a usable assistant.
Why do large language models hallucinate?
A model produces the most plausible continuation, and plausible is not the same as true. When the right answer is not well represented in what it learned, the plausible answer is often fluent and wrong — and most evaluations reward guessing over saying “I don’t know”.
What is a token in AI?
A token is the unit a language model reads and writes — a word, part of a word, a number or punctuation. In English one token is roughly four characters, or about three-quarters of a word. Cost, speed and context limits are all counted in tokens.
How do I choose the right language model?
Start from the task, not a leaderboard. Write down what good output looks like, collect real examples with known answers, shortlist two or three models that meet your constraints on data location, cost and speed, and test each on your own examples. Pick the cheapest one that meets the bar.
What does temperature do in a language model?
Temperature controls how varied the output is. Low values make answers more repeatable; higher values make them more varied. It does not make answers more or less accurate.