Training a model on real customer interactions — support tickets, chat transcripts, transaction records — can produce an assistant that genuinely understands your business. It can also create one of the most serious privacy exposures a company can face. The good news is that the two goals are not in conflict. A few deliberate choices let you train on customer data without betraying the trust of the people who provided it. Most of the work happens before any training code runs, in how you choose, clean and store the data itself.
Start with the legal basis
Before a single row is exported, answer one question: on what basis are you allowed to use this data for training? For most teams that means consent, a legitimate interest or an explicit contract clause — and the answer varies by jurisdiction. GDPR-style rules also impose purpose limitation: data collected to resolve a support ticket is not automatically fair game for model training. Write the basis down and keep it with the dataset. If you cannot state it in one sentence, you are not ready to train.
Collect less, as a training habit
Data minimisation is the cheapest privacy control you will ever implement. Ask what the model actually needs to learn: a tone of voice, a catalogue of policy answers, a routing decision. Much of that can be captured from a small, curated sample instead of the full archive. Exporting ten thousand records “just in case” increases exposure, storage cost and legal risk for no measurable gain. For fine-tuning, a well-chosen few hundred examples often outperform a sprawling, noisy archive — so the smallest dataset is frequently the better one anyway. Sample deliberately, and delete the working copies once the run finishes.
De-identification is a process, not a checkbox
Removing names and email addresses is the first step, not the last. Free text is full of quasi-identifiers: order numbers, rare job titles, precise timestamps, a postcode combined with a birth date. Individually they look harmless; combined, they re-identify people surprisingly often. A workable pipeline removes direct identifiers, generalises the risky fields, and then measures the result — a re-identification test on a sample, repeated whenever the underlying data changes. Automated scrubbing helps, but it fails in predictable ways: nickname lists go stale and pattern rules miss context. Human review of a random sample still catches the cases automation confidently gets wrong.
Keep the data under lock and key
Training data deserves the same protection as production customer records: encrypted at rest, access limited to the people running the job, and never quietly copied into shared drives, notebooks or vendor demos. If you use a hosted training service, check where the data goes, how long it is retained, and whether it may be used to improve another company’s models. A vendor’s contractual promises matter less than the configuration you can actually see. Finally, decide which employees and contractors can touch the working data, and give each of them the least access that still lets them work — privacy failures are far more often accidental than adversarial.
Give people a way out
Customers should be able to learn that their data is used for training and, where the law requires it, to opt out or ask for deletion. This gets genuinely hard after training: removing one person’s influence from a finished model is an open research problem, not a button. The pragmatic answer is to plan for it in advance — keep datasets small and refreshable, log which data trained which model version, and be honest about retention and the limits of deletion in your privacy notice.
Make the choices provable
If a regulator or a customer asks how their data was used, “carefully” is not an answer. Keep an audit trail that links every training run to its dataset version, its purpose and the approvals that cleared it. Log when data was deleted and when a model was retrained without it. That trail turns a privacy policy from a promise into evidence, and it makes incident response dramatically faster if a leak ever happens.
A seven-point checklist
- Write down the legal basis and purpose before exporting anything.
- Sample the smallest useful dataset instead of the full archive.
- Strip direct identifiers and generalise quasi-identifiers.
- Test for re-identification on a sample, and repeat the test when the data changes.
- Encrypt working data, restrict access, and delete it when training ends.
- Record which dataset version trained which model version.
- Update the privacy notice and honour opt-outs in future runs.
None of this is exotic. It is the same discipline you already apply to backups, keys and access control, pointed at a new kind of asset. Teams that get it right end up with the same benefit — a model that speaks your business fluently — plus a story about how they got there that customers can actually trust.



