How Large Language Models Work, Without the Maths
Tokens, attention, pre-training, instruction tuning and next-token prediction — what happens between your prompt and the answer, and why models fail.
BeginnerVerdeshell Team · 9 min read · Last reviewed
A language model predicts the next piece of text, one piece at a time, based on everything before it. Training on a huge amount of text makes those predictions remarkably capable; further training on instructions and feedback makes them useful. Neither makes them reliably true.
Key takeaways
- A language model predicts the next token, one at a time, based on everything before it.
- Pre-training produces a base model that continues text; instruction and preference tuning turn it into an assistant.
- Temperature changes how varied the output is, not how accurate it is.
- Models make things up because they produce plausible text, and most evaluations reward guessing over saying “I don’t know”.
Tokens: what the model actually reads
A model does not read words. Text is first split into tokens — common words, pieces of words and punctuation — using a method such as byte-pair encoding, which builds a vocabulary of frequent fragments so that even rare or invented words can be represented as a few pieces.
This detail has practical consequences. Cost and context limits are counted in tokens, not words. Languages and formats that split into more tokens cost more to process. And the fact that a model sees fragments rather than letters is why tasks like counting characters in a word can trip up otherwise capable models.
Attention: how the pieces relate to each other
Each token is turned into a list of numbers — an embedding — that places it in a space where related meanings sit close together. The model then passes these through many layers of a Transformer, the architecture introduced in 2017.
The key mechanism is attention. At every layer, each token looks at every other token in the input and decides how much each one matters to it. That is how the model works out that “it” in a sentence refers to a noun twenty words earlier, or that a clause at the end of a paragraph changes the meaning of the start.
Because attention looks at everything at once rather than reading left to right, Transformers can be trained in parallel on very large amounts of text. That property, more than any other, is why they scaled when earlier designs did not.
Pre-training: learning to predict the next token
A model starts with random settings — billions of adjustable numbers called parameters. Pre-training shows it an enormous body of text, a large slice of the public web, books and code, and asks one question over and over: given everything so far, what comes next? Each wrong guess nudges the parameters slightly towards a better one.
This simple objective turns out to force a great deal of learning. To predict the next word of a legal contract, a physics explanation or a Python function well, a model has to pick up grammar, facts, styles, reasoning patterns and code structure along the way.
The result is a base model. It is very good at continuing text, and not at all good at being helpful: ask it a question and it may continue with three more questions, because that is a plausible continuation too.
Post-training: from text predictor to assistant
What people actually use has been through further training. Instruction tuning trains the base model on many examples of instructions paired with good responses, so that following an instruction becomes its default behaviour.
Preference tuning then refines it using judgements of which of two answers is better. The best-known version, reinforcement learning from human feedback (RLHF), trains a separate model to predict human preferences and uses it to steer the language model. Some developers use AI-generated feedback guided by a written set of principles instead of, or as well as, human ratings.
This stage is where much of a model’s personality lives — its helpfulness, its tone, what it refuses to do. It is also why two models built on similar pre-training can feel very different to use.
Generation: one token at a time
When you send a prompt, it is tokenised and passed through the model, which produces a probability for every possible next token. One is chosen, appended to the input, and the whole process runs again for the next token — until the answer is complete. A long answer is hundreds of these steps.
Which token gets chosen is controlled by sampling settings. Temperature adjusts how sharply the model favours its most likely options: low values make output more repeatable, higher values more varied. Research on text generation has found that always choosing the single most likely token produces bland, repetitive text, which is why some randomness is normal. Temperature changes variety, not accuracy.
Everything the model can take into account for one response — your prompt, any documents supplied, the conversation so far and its own answer — has to fit in its context window. That is its working memory. Anything outside it simply does not exist for that response, and research has shown that models tend to use information at the start and end of a long context more reliably than information buried in the middle.
Why models make things up
A model produces the most plausible continuation, and plausible is not the same as true. When the right answer is not well represented in what it learned, the most plausible continuation is often a confident, fluent, wrong one.
Research published by OpenAI in 2025 argued that the problem is reinforced by how models are evaluated: most benchmarks give credit for a right answer and nothing for saying “I don’t know”, which rewards guessing over admitting uncertainty — much like a multiple-choice exam.
The practical implications follow directly — covered in depth in why AI hallucinates. Give the model the facts it needs rather than relying on what it remembers (that is what retrieval is for); tell it that saying it does not know is acceptable; and check outputs where being wrong matters.
Reasoning models
Reasoning models are trained, largely with reinforcement learning, to work through a problem in intermediate steps before giving their final answer, spending more computation on harder questions. They are noticeably better at maths, code and multi-step logic, and slower and more expensive per answer.
They also change how you should prompt, which chain-of-thought prompting covers: with a reasoning model you describe the goal and let it plan, rather than dictating each step.
What this means in practice
Most of the behaviour people find puzzling makes sense once you hold the mechanism in mind. Same prompt, different answer: sampling. Forgot something from earlier in a long chat: the context window. Confidently wrong: plausible, not verified. Better when you show an example: you have made the pattern you want the most plausible continuation.
And the design rule that runs through the rest of this hub follows from it: a language model is a component that generates plausible text. Reliability comes from what you build around it — the context you give it, the checks you run on its output, and the decisions you do or do not delegate to it.
More in Large Language Models
Want this built properly?
We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.