Blog · AI Engineering

Choosing a Base Model for a Domain-Specific Assistant

2026-09-28·4 min read·MD ABU SAYEED
In short: Pick the smallest base model that clears your hardest real questions, then spend your budget on data and evaluation rather than raw parameter count.
Close-up of programming code
Close-up of programming code

Every domain-specific assistant begins with the same quiet decision: which base model to build on. It is easy to skip, because the interesting work, data and prompts and evaluation, all arrives later. But the base you choose fixes your ceiling for context length, language coverage, licensing and cost per answer. Reaching for the largest model available is rarely the right move, and going smaller than you need usually costs more time than it saves in hosting.

Start with the job, not the leaderboard

Write down what the assistant must actually do before you open a single benchmark table. A support assistant answering from a fixed documentation set has very different needs from a coding helper or a multilingual triage bot. Answer these first:

Five filters that decide most of the shortlist

Match the context window to your real documents

Long-context models are marketed as a way to skip retrieval. In practice, a focused retrieval step plus a moderate context window is often cheaper and more accurate than feeding entire manuals into every request. Measure your own inputs: if most questions resolve against two or three short passages, a modest window is enough, and you can reserve the remainder for examples and instructions.

Language coverage and domain vocabulary

Benchmarks are usually dominated by English, so a strong average score can hide weak performance in the language your users actually speak. Test the candidates on real sentences from your domain, including jargon, product names and abbreviations. A model that is merely average in English but strong in your target language is the better choice for a regional assistant.

Run a fair bake-off

Do not pick on vibes. Assemble a test set of fifty to one hundred real questions with known good answers. Send each candidate the same prompts, then score factuality, format compliance and latency side by side. Keep the benchmark honest by holding prompts constant and refusing to hand-tune one model more than another. The result is usually clear after a few hours, and far more reliable than a public leaderboard.

Plan for the second model

Most production assistants end up with at least two models. A small, fast model handles the common, easy majority of traffic, while a larger one is reserved for hard or ambiguous requests. Design the routing rule early, even if you launch with a single model, so you can add the second without rewriting your pipeline.

A short checklist before you commit

Choose deliberately, write the decision down with the evidence behind it, and revisit it once a quarter. The base model is not a permanent commitment, but a well-reasoned first choice saves months of rework later.

Frequently asked questions

Should I always choose the largest base model I can afford?
No. Larger models cost more to run and respond more slowly. Choose the smallest model that passes your own evaluation set, then invest the remaining budget in better data and testing.
How important is the licence when choosing a base model?
Very. Licence terms decide whether you may use the model commercially, redistribute it, or how derivatives must be named. Read them before you build, not after launch.
Do I need a long context window to avoid building retrieval?
Usually not. A focused retrieval step plus a moderate context window is often cheaper and more accurate than sending entire documents with every request.
How many test questions do I need to compare two base models?
Fifty to one hundred realistic questions with known good answers are enough to expose clear differences when you hold the prompts constant across candidates.

Want a model trained on your own data?

Tell us what you want to build and we will send a scoped plan within 1–2 business days.

Request a fit assessment

More from the blog

How to Keep a Trained Model From Going Stale

How to Keep a Trained Model From Going Stale

A trained model starts drifting the day you ship it. A practical routine for spotting staleness, refreshing data and retraining without breaking what works.
2026-09-28
Choosing Between Hosted and Self-Hosted AI: A Practical Guide for Teams in 2026

Choosing Between Hosted and Self-Hosted AI: A Practical Guide for Teams in 2026

Hosted vs self-hosted AI compared: cost, control, latency, and compliance. A practical decision framework for teams evaluating where to run their models.
2026-09-25
Building a Support Assistant on Your Own Documentation

Building a Support Assistant on Your Own Documentation

Build a support AI assistant on your own documentation: audit content, chunk it well, retrieve before generating, add guardrails, and evaluate.
2026-09-24