Ask three teams that fine-tune language models what improved their results most, and most give the same answer: not the base model, not the hyperparameters — the instruction data. Instruction tuning teaches a model how to behave, and it learns those habits entirely from the examples you hand it. Write them carelessly and even a strong model turns vague, repetitive, or wrong. Write them well and a modest model can outperform expectations. Here is how to write instruction data that actually teaches.
Why quality beats quantity
A fine-tuned model reproduces patterns, including the bad ones. If your training examples are inconsistent, the model learns inconsistency; if they are padded with filler, it learns to pad. Roughly 500 to 5,000 well-written examples is usually enough to teach a specific behaviour, and stacking thousands of mediocre ones on top tends to dilute the signal rather than strengthen it. Treat every example like production copy, not like test data.
Match the format the model will see
Write every example in the same conversational shape your assistant will use in production: a system message defining its role, a user message, and an assistant response. If your product answers in two tight paragraphs followed by a bullet list, your examples should look exactly like that. If it should politely refuse out-of-scope requests, include refusal examples too. The training file is where behaviour is defined — the model cannot read your product spec.
Cover the range of real usage
Diversity is the biggest driver of generalisation. Vary phrasing, length, tone, and difficulty: short questions and long ones, typos and clean text, happy paths and edge cases. Include the way real users actually ask things, not the way your team talks internally. A useful rule of thumb is to write several examples per behaviour that differ in surface detail but share the same underlying intent — that teaches the model the intent rather than one fixed wording.
Balance the difficulty curve
A common mistake is training only on easy, well-formed requests. Real traffic includes multi-part questions, follow-ups that reference earlier turns, ambiguous wording, and requests the assistant should decline. If those situations matter to your product, they deserve explicit examples. Deliberately include a slice of hard cases: questions that need reasoning across several points, prompts with missing context where the right move is to ask a clarifying question, and hostile or off-topic inputs where the correct response is a polite redirect. A model that has never seen a hard example will improvise one when it matters most.
Keep the voice consistent
Models absorb tone quickly, so every response should sound like one writer produced it. Decide up front how formal the voice is, which spelling convention it follows, how answers are structured, whether the assistant says "I", and how it expresses uncertainty. Then apply those decisions everywhere. Inconsistent voice is one of the most common causes of a model that feels different every day. One effective technique: have a single editor rewrite every response at the end, no matter who drafted it. It feels slower, but it removes the drift that makes a model unpredictable.
Write outputs you would ship
Every response in your dataset should be one you would happily publish verbatim. That means accurate facts, complete answers, no references to things the user never said, and no invented details. If you source responses from documentation or from a stronger model, review each one manually — anything copied into the training set becomes a habit. Cut anything you would not stand behind.
A checklist before you train
- Every example matches your production chat format exactly
- Responses demonstrate the style, length, and structure you want
- Coverage spans your main intents plus refusals and edge cases
- A human reviewed every response for accuracy and tone
- Duplicates removed and near-duplicates paraphrased
- A held-out validation set of 50 to 100 examples kept out of training
Split off a validation set
Before training, hold back 50 to 100 examples that the model never sees. These become your yardstick: after each training run, check responses against them to see whether quality is genuinely improving or the model is simply memorising the training file. Skipping this step leaves you tuning blind, judging progress by vibes rather than evidence.
Instruction data is unglamorous work, but it is the highest-leverage few days you will spend on a fine-tuning project. Write examples the way a careful product writer would, keep them consistent, and the model will meet you much more than halfway.



