Prompt Engineering: Techniques That Hold Up
The techniques every major model provider recommends, how prompting changes for reasoning models, and an honest look at CO-STAR, RISEN and other prompt frameworks.
Verdeshell Team · 10 min read · Last reviewed
Prompt engineering is mostly clear writing for a reader with no context: say what you want, give the material, show an example, specify the format. The tricks matter far less than the clarity.
What prompt engineering is, and is not
Prompt engineering is the practice of writing the input to a language model so that it reliably produces the output you need. For one-off use in a chat window it is just good communication. In a product, where the same prompt runs thousands of times on inputs you have not seen, it becomes engineering: versioned, tested and measured.
It is not a collection of magic phrases. Many early tricks — threatening the model, promising it a tip, insisting it is an expert — were artefacts of particular models and fade as models improve. What has held up is what would help a capable new colleague who knows nothing about your situation.
The techniques every major provider recommends
OpenAI, Anthropic and Google each publish prompting guidance for their models. Read side by side, they agree on a short list.
Be clear and specific. State the task directly, including what you do not want. Vague instructions get average answers, because the model fills every gap with the most typical choice.
Give it a role and a purpose. A system instruction that sets who the model is acting as, and who the answer is for, shapes vocabulary, depth and tone more efficiently than piling on adjectives.
Provide the context. Include the material the answer should be based on rather than hoping the model knows it, and say so: answer only from the document below.
Show examples. One or two examples of the output you want — few-shot prompting — communicate format and style faster than any description. Vary them, or the model will copy them too literally.
Use structure. Separate instructions from data with delimiters, headings or XML-style tags, so the model can tell which text is the instruction and which is the material. This also makes long prompts far easier to maintain.
Specify the output format. Length, structure, fields, a JSON schema. If code will read the answer, say exactly what shape it must be — and validate it, because stated formats are usually but not always followed.
Break complex tasks down. One prompt that summarises, classifies and drafts a reply will do all three worse than three focused prompts chained together, each of which you can test on its own.
Let the model think, where it does not already. For models without built-in reasoning, asking it to work through the problem before answering — chain-of-thought prompting — improves multi-step tasks. Anthropic’s guidance adds one more that deserves wider use: explicitly allow the model to say it does not know. It is one of the simplest ways to reduce made-up answers.
Prompting reasoning models is different
Reasoning models plan and check their own work before answering, and the providers’ guidance for them runs against some of the advice above. OpenAI’s comparison is useful: a standard model is like a junior colleague who needs explicit steps; a reasoning model is like a senior one you give a goal.
In practice: keep the prompt simple and direct, describe the goal and the constraints rather than a step-by-step procedure, and do not ask it to “think step by step” — it already does, and prescribing the steps can make its reasoning worse. Try without examples first, and add them only if the output format needs pinning down.
Clarity, context and a defined output format still matter exactly as much.
Prompt frameworks: CO-STAR, RISEN and the rest
A family of acronyms circulates for structuring prompts. CO-STAR — Context, Objective, Style, Tone, Audience, Response — is the best documented: it comes from GovTech Singapore’s prompt engineering playbook (2023) and was popularised by the winner of Singapore’s GPT-4 prompt engineering competition.
The others have thinner origins. CRISPE (Capacity and role, Insight, Statement, Personality, Experiment) first appeared in a community prompt list on GitHub in early 2023. RACE (Role, Action, Context, Execute) comes from a marketing analytics consultancy. RISEN (Role, Instructions, Steps, End goal, Narrowing) traces to a social-media educator’s post. RTF (Role, Task, Format) has no traceable origin at all.
The honest assessment: these are checklists, not methods. Every one of them is a subset of the provider guidance above — role, context, task, format — and none has been shown to beat simply writing that guidance down clearly. They are useful for training a team to stop writing one-line prompts. They are not worth adopting as a standard, and there is no reason to pick one over another.
Treat prompts like code
A production prompt is a piece of software. Keep it in version control, not pasted into a dashboard. Give it a small set of real test inputs with known-good outputs, and run them whenever the prompt or the model changes — a prompt tuned for one model version can quietly degrade on the next.
Change one thing at a time. When a prompt edit improves one case, check that it has not broken three others. That discipline is the difference between prompt engineering and prompt guessing.
And read the failures. The fastest improvements come from looking at the worst outputs on real inputs and asking what the prompt failed to say.
From prompt engineering to context engineering
As systems grew from single prompts into agents that run for many steps, the wording of the instruction stopped being the main lever. What matters more is everything that ends up in the model’s context window at each step: the instructions, the retrieved documents, the tool results, the conversation history — and what is left out.
Practitioners began calling this context engineering in mid-2025. Anthropic’s engineering team describes it as the natural progression of prompt engineering: curating the smallest set of information that gives the model what it needs for the next step, because an overfull context degrades answers as surely as a missing fact does.
The techniques above are still the foundation. Context engineering is what they become when the prompt is assembled by a system rather than written by a person.
Want this built properly?
We design and run these systems for clients. Tell us the problem and we will tell you whether an agent is the right shape for it.