Blog · AI Training & Data

How to Write Good Instruction Data for Fine-Tuning

2026-10-04·6 min read·MD ABU SAYEED
In short: Instruction data quality, not quantity, determines how well a fine-tuned model behaves, so write a small set of diverse, consistent, correctly formatted examples that mirror real usage.
Monitor showing security data and code
Monitor showing security data and code

Ask three teams that fine-tune language models what improved their results most, and most give the same answer: not the base model, not the hyperparameters — the instruction data. Instruction tuning teaches a model how to behave, and it learns those habits entirely from the examples you hand it. Write them carelessly and even a strong model turns vague, repetitive, or wrong. Write them well and a modest model can outperform expectations. Here is how to write instruction data that actually teaches.

Why quality beats quantity

A fine-tuned model reproduces patterns, including the bad ones. If your training examples are inconsistent, the model learns inconsistency; if they are padded with filler, it learns to pad. Roughly 500 to 5,000 well-written examples is usually enough to teach a specific behaviour, and stacking thousands of mediocre ones on top tends to dilute the signal rather than strengthen it. Treat every example like production copy, not like test data.

Match the format the model will see

Write every example in the same conversational shape your assistant will use in production: a system message defining its role, a user message, and an assistant response. If your product answers in two tight paragraphs followed by a bullet list, your examples should look exactly like that. If it should politely refuse out-of-scope requests, include refusal examples too. The training file is where behaviour is defined — the model cannot read your product spec.

Cover the range of real usage

Diversity is the biggest driver of generalisation. Vary phrasing, length, tone, and difficulty: short questions and long ones, typos and clean text, happy paths and edge cases. Include the way real users actually ask things, not the way your team talks internally. A useful rule of thumb is to write several examples per behaviour that differ in surface detail but share the same underlying intent — that teaches the model the intent rather than one fixed wording.

Balance the difficulty curve

A common mistake is training only on easy, well-formed requests. Real traffic includes multi-part questions, follow-ups that reference earlier turns, ambiguous wording, and requests the assistant should decline. If those situations matter to your product, they deserve explicit examples. Deliberately include a slice of hard cases: questions that need reasoning across several points, prompts with missing context where the right move is to ask a clarifying question, and hostile or off-topic inputs where the correct response is a polite redirect. A model that has never seen a hard example will improvise one when it matters most.

Keep the voice consistent

Models absorb tone quickly, so every response should sound like one writer produced it. Decide up front how formal the voice is, which spelling convention it follows, how answers are structured, whether the assistant says "I", and how it expresses uncertainty. Then apply those decisions everywhere. Inconsistent voice is one of the most common causes of a model that feels different every day. One effective technique: have a single editor rewrite every response at the end, no matter who drafted it. It feels slower, but it removes the drift that makes a model unpredictable.

Write outputs you would ship

Every response in your dataset should be one you would happily publish verbatim. That means accurate facts, complete answers, no references to things the user never said, and no invented details. If you source responses from documentation or from a stronger model, review each one manually — anything copied into the training set becomes a habit. Cut anything you would not stand behind.

A checklist before you train

Split off a validation set

Before training, hold back 50 to 100 examples that the model never sees. These become your yardstick: after each training run, check responses against them to see whether quality is genuinely improving or the model is simply memorising the training file. Skipping this step leaves you tuning blind, judging progress by vibes rather than evidence.

Instruction data is unglamorous work, but it is the highest-leverage few days you will spend on a fine-tuning project. Write examples the way a careful product writer would, keep them consistent, and the model will meet you much more than halfway.

Frequently asked questions

How many instruction examples do I need for fine-tuning?
For a focused assistant, 500 to 5,000 high-quality examples is typical. Quality and diversity matter far more than raw volume; a small clean dataset usually beats a large noisy one.
Should instruction data be written by humans or AI?
Either can produce a first draft, but a human should review every response for accuracy, tone, and consistency. Unreviewed AI-generated data tends to pass its own mistakes on to your model.
What format should instruction data be in?
Most teams use JSONL with system, user, and assistant messages per example, matching the chat template of the base model. The exact structure matters less than consistency across every example.
Can I turn existing documentation into instruction data?
Documentation is a strong source of grounded answers, but it is not instruction data on its own. Wrap each passage in a realistic user question and a response written in your assistant's voice.

Want a model trained on your own data?

Tell us what you want to build and we will send a scoped plan within 1–2 business days.

Request a fit assessment

More from the blog

Running a Large Model Locally: What Hardware You Actually Need

Running a Large Model Locally: What Hardware You Actually Need

A practical guide to running large language models on your own hardware: how VRAM, RAM and storage decide what you can run, and how fast it responds.
2026-10-03
LoRA vs Full Fine-Tuning: How to Choose

LoRA vs Full Fine-Tuning: How to Choose

LoRA or full fine-tuning? Compare cost, quality, hardware and the signals that tell you which method fits when adapting an open model to your own domain.
2026-09-29
How to Keep a Trained Model From Going Stale

How to Keep a Trained Model From Going Stale

A trained model starts drifting the day you ship it. A practical routine for spotting staleness, refreshing data and retraining without breaking what works.
2026-09-28