Skip to main content

How to Choose a Language Model

Hosted or open-weight, large or small, reasoning or standard — the trade-offs that matter, and how to decide with your own tasks, not a leaderboard.

IntermediateVerdeshell Team · 9 min read · Last reviewed

There is no best model, only the best model for a task, a budget and a set of constraints. Decide what you need, shortlist two or three candidates, and test them on your own examples — leaderboards measure someone else’s task.

Key takeaways

  • Choose a model for a specific task, measured on your own examples — not from a leaderboard.
  • Hosted models are the quickest start; open-weight models give you control over where data goes, at the cost of running them yourself.
  • Smaller and standard models are faster and cheaper; reasoning models earn their cost on genuinely multi-step problems.
  • Compare total cost per completed task, and keep prompts and evaluations portable so that switching later is easy.

Start from the task, not the model

Model announcements arrive every few weeks, each claiming to top some leaderboard. A benchmark measures performance on someone else’s task, under someone else’s conditions; what matters is how a model performs on yours.

So start by writing down what the system has to do, what a good output looks like, and what it must never do. Collect twenty to fifty real examples with known-good answers. That small evaluation set does more for a model decision than any published ranking — and you will need it anyway to check every later change.

Hosted or open-weight

Hosted models are used through a provider’s API. They are the fastest way to start, include the latest capabilities, and need no infrastructure. You send your data to the provider, so read their data-handling and retention terms, and check where your data is processed if you have obligations under laws such as India’s DPDP Act or the GDPR.

Open-weight models publish their trained parameters, so you can run them on your own infrastructure or a private cloud. That gives you control over where data goes and freedom from a provider’s pricing and deprecation decisions — at the cost of hosting, scaling and updating them yourself. “Open-weight” is not always “open source”: licences vary, and some restrict commercial use. Read the licence.

Many organisations use both: hosted models for general work, an open-weight model for data that must not leave their environment.

Size, speed and cost

Larger models generally handle harder, more open-ended tasks better. Smaller models are faster and much cheaper per request, and on narrow, well-defined tasks — classification, extraction, routing — a small model with a good prompt is often enough. A smaller model fine-tuned for one task can sometimes match a much larger general one.

Many production systems use more than one: a small, fast model for high-volume steps and a stronger one for the step that needs judgement. A prompt chain makes that easy, because each step can use a different model.

Compare total cost per completed task, not price per token. Models tokenize differently and produce different amounts of output, so the cheaper rate is not always the cheaper bill — a point covered in tokens and context windows.

Reasoning or standard

Reasoning models work through a problem before answering. They are noticeably better at maths, code, planning and multi-step logic, and slower and more expensive per answer, because they generate reasoning tokens you pay for but do not see.

Use them where the problem genuinely has steps. For summarising, drafting, extraction and most customer-facing chat, a standard model is usually faster, cheaper and just as good. And prompt them differently — goals rather than steps, as covered in chain-of-thought prompting.

The other things that decide it

Context window: enough room for your largest realistic input plus the answer — and remember that filling it is not the same as using it well.

Modalities: whether you need images, PDFs, audio or video as input, or images as output.

Tool use and structured output: if the model will call functions or return JSON that code depends on, check how reliably it does so, and whether the API can enforce a schema.

Languages: test quality in every language you serve — performance outside English varies more between models than headline results suggest.

Latency: measure time to the first token and to a complete answer, under realistic load, if users are waiting.

Stability: providers retire model versions and APIs on their own schedule. Prefer pinned model versions, and expect to migrate.

A simple selection process

Write down the task, the constraints (data location, budget, latency, languages) and what good looks like. Shortlist two or three models that meet the hard constraints. Run your evaluation set on each with the same prompt, then adjust the prompt for each and run it again — a model judged on a prompt written for another often looks worse than it is.

Pick the cheapest model that meets the bar, not the one that wins by the widest margin. Then keep the parts that let you switch: prompts in version control, tools behind a standard interface such as MCP, and the evaluation set ready to run against the next candidate. The best model for your task will change; your ability to measure it is what lasts.

The main trade-offs at a glance
DecisionChoose the first when…Choose the second when…
Hosted vs open-weightYou want the fastest start and the newest capabilitiesData must stay in your environment, or you need control over versions and cost
Large vs smallTasks are open-ended or need broad knowledge and judgementTasks are narrow and high-volume, and speed or cost matter most
Reasoning vs standardProblems have genuine steps — maths, code, planning, multi-step logicTasks are summarising, drafting, extraction or chat
General vs fine-tunedGood prompts already meet the barA consistent behaviour must hold at volume and prompts cannot hold it
Next in the path · Prompt EngineeringWhat Is Prompt Engineering? Techniques That Work

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.