Skip to main content

Large Language Models, Explained

What is actually happening between your prompt and the answer — tokens, attention, pre-training and post-training — why models make things up and how to reduce it, and how to choose the right model for a task.

4 articles · start with the first and read in order, or jump to what you need

Common questions: Large Language Models

How does a large language model work?

It splits text into tokens and predicts the next token, one at a time, based on everything before it. Pre-training on a very large body of text teaches it to continue text; further training on instructions and on feedback about which answers are better turns it into a usable assistant.

Why do large language models hallucinate?

A model produces the most plausible continuation, and plausible is not the same as true. When the right answer is not well represented in what it learned, the plausible answer is often fluent and wrong — and most evaluations reward guessing over saying “I don’t know”.

What is a token in AI?

A token is the unit a language model reads and writes — a word, part of a word, a number or punctuation. In English one token is roughly four characters, or about three-quarters of a word. Cost, speed and context limits are all counted in tokens.

How do I choose the right language model?

Start from the task, not a leaderboard. Write down what good output looks like, collect real examples with known answers, shortlist two or three models that meet your constraints on data location, cost and speed, and test each on your own examples. Pick the cheapest one that meets the bar.

What does temperature do in a language model?

Temperature controls how varied the output is. Low values make answers more repeatable; higher values make them more varied. It does not make answers more or less accurate.