Every domain-specific assistant begins with the same quiet decision: which base model to build on. It is easy to skip, because the interesting work, data and prompts and evaluation, all arrives later. But the base you choose fixes your ceiling for context length, language coverage, licensing and cost per answer. Reaching for the largest model available is rarely the right move, and going smaller than you need usually costs more time than it saves in hosting.
Start with the job, not the leaderboard
Write down what the assistant must actually do before you open a single benchmark table. A support assistant answering from a fixed documentation set has very different needs from a coding helper or a multilingual triage bot. Answer these first:
- What is the longest input it must read in a single request?
- Which languages must it handle well, and how good is good enough?
- Does it only produce text, or does it need to call tools and follow strict output formats?
- How fast must a reply arrive before users notice a delay?
- Where will it run, and who is allowed to see the data?
Five filters that decide most of the shortlist
- Size and active parameters. Architectures that activate only a fraction of their parameters per token can deliver large-model quality at small-model speed. For most assistants, a mid-size model that fits comfortably on your target hardware beats a larger one that does not.
- Context window. Long context is genuinely useful, but it is not free. Attention cost grows with input length, and quality can drift when a model is stuffed near its limit. Choose a window that fits your worst realistic document plus your instructions, not the biggest number on the card.
- Licence. Check commercial use, redistribution terms, and whether derivatives must carry a specific name. Some permissive-sounding licences still add conditions that matter to a product.
- Tooling and formats. Does the model ship in quantised formats that run on your hardware? Are fine-tuning and serving stacks well supported? An excellent model with poor tooling is a slow project.
- Provenance and maintenance. A base that is actively maintained and documented is easier to debug months later than an abandoned one, even if the abandoned one scores higher today.
Match the context window to your real documents
Long-context models are marketed as a way to skip retrieval. In practice, a focused retrieval step plus a moderate context window is often cheaper and more accurate than feeding entire manuals into every request. Measure your own inputs: if most questions resolve against two or three short passages, a modest window is enough, and you can reserve the remainder for examples and instructions.
Language coverage and domain vocabulary
Benchmarks are usually dominated by English, so a strong average score can hide weak performance in the language your users actually speak. Test the candidates on real sentences from your domain, including jargon, product names and abbreviations. A model that is merely average in English but strong in your target language is the better choice for a regional assistant.
Run a fair bake-off
Do not pick on vibes. Assemble a test set of fifty to one hundred real questions with known good answers. Send each candidate the same prompts, then score factuality, format compliance and latency side by side. Keep the benchmark honest by holding prompts constant and refusing to hand-tune one model more than another. The result is usually clear after a few hours, and far more reliable than a public leaderboard.
Plan for the second model
Most production assistants end up with at least two models. A small, fast model handles the common, easy majority of traffic, while a larger one is reserved for hard or ambiguous requests. Design the routing rule early, even if you launch with a single model, so you can add the second without rewriting your pipeline.
A short checklist before you commit
- The model clears your hardest realistic questions, not just easy ones.
- It fits your hardware and latency budget in a quantised format.
- The licence permits your intended commercial use.
- It performs acceptably in your real languages and domain vocabulary.
- You have a documented way to fine-tune, serve and update it.
Choose deliberately, write the decision down with the evidence behind it, and revisit it once a quarter. The base model is not a permanent commitment, but a well-reasoned first choice saves months of rework later.



