Skip to content
NewEraAI

Governance

UK GDPR and AI: what you can and cannot put into a model

A practical guide for UK businesses, written by engineers rather than lawyers. What changes when personal data meets a third-party model, what to put in your DPA, and the four rules that keep most SMEs out of trouble.

18 June 2026 · 6 min read · New Era AI

This is a practical article, not legal advice. We are an AI consultancy, not a law firm, and anything that carries real regulatory risk for your business should go past someone qualified. What follows is the operational version (the questions that actually come up when a UK business starts feeding real data into real tools, and the answers that keep most of them out of trouble.

The single most common mistake we see is not a data breach. It is a member of staff pasting a customer list into a free consumer chatbot to "tidy it up", with nobody having told them not to. Almost everything below is downstream of that.

Start here: nothing about UK GDPR changed because it is AI

This is the point that clears up most of the confusion. UK GDPR does not have an AI chapter that switches on when you use a model. The same principles apply that always applied:

  • You need a lawful basis to process personal data.
  • You should collect and use the minimum necessary.
  • You must be transparent about what you do with it.
  • You are responsible for your processors, including model providers.
  • People retain their rights) access, erasure, objection.

What AI changes is not the law. It changes how easy it has become for data to leave your control by accident, and how many new processors appear in your stack without anyone noticing.

The four rules that cover most SMEs

If you do nothing else, do these.

1. Use business-tier services, never consumer ones, for customer data

There is a real and material difference between a consumer chatbot account and the business or enterprise tier of the same provider. The business tiers generally offer contractual commitments that your inputs are not used to train their models, plus a data processing agreement you can actually put in a file.

Consumer tiers frequently do not. Free tiers almost never do.

This is the cheapest control available and it removes the majority of the risk. Buy the business tier. It is not expensive relative to the problem it prevents.

2. Write down what staff may and may not paste

A one-page acceptable-use note, written in plain English, distributed once and mentioned again in six months. It needs to say:

  • Which tools are approved, by name.
  • That customer records, employee records, health information, financial details and anything from a supplier under NDA do not go into an unapproved tool.
  • What to do instead when someone genuinely needs help with a real document (usually "ask, and we will get you the approved way to do it".

Most breaches at this scale are not malicious. They are a helpful person under time pressure who was never told.

3. Minimise before you send

You very rarely need to send a full record. If you want a model to draft a follow-up email, it needs the job type and the tone, not the customer's full address and payment history.

Stripping identifiers before sending) sometimes called pseudonymisation (is a small amount of engineering work with a disproportionate benefit. It also makes the "what happens if this leaks" conversation much shorter.

Build this into the integration rather than trusting people to remember it. Rules that depend on human discipline at 4:50pm on a Friday are not controls.

4. Keep a list of your processors

Every AI feature you switch on potentially adds a company that now touches your data. Transcription, enrichment, a summarisation feature buried inside a CRM add-on, a note-taker that joined your calls last Tuesday.

Keep a simple list: tool, what data it sees, where it processes it, whether there is a DPA. Review it twice a year. This is not exciting and it is the thing that makes a subject access request or a client security questionnaire take an hour instead of a fortnight.

The questions worth asking any AI supplier

When we evaluate a tool on a client's behalf, these are the ones that actually discriminate between vendors:

Is our data used to train your models, or anyone else's? The answer must be no, in the contract, not in a marketing FAQ.

Where is data processed and stored? UK, EEA, or elsewhere. If elsewhere, what transfer mechanism applies. A vendor who cannot answer this quickly has not thought about it.

How long do you retain inputs and outputs? Including logs. "Indefinitely for quality purposes" is a real answer some vendors give and it is a problem.

Can we delete on request, and how fast? This is what makes an erasure request answerable.

Who are your sub-processors? The model provider behind the tool is often a different company from the one selling you the tool.

What happens on your side if we churn? Data export and deletion terms.

A supplier who answers all six without friction is usually fine. One who deflects on two or more is a risk you are taking knowingly.

Where the genuinely hard cases are

Most SME work sits comfortably inside the rules above. A few situations need more care:

Automated decisions with legal or similarly significant effects. If a model is deciding whether someone gets credit, a job, or a service) not assisting a person, but deciding (additional obligations apply and you need proper advice. In practice, keeping a human genuinely in the loop (not a rubber stamp) avoids most of this.

Special category data. Health, biometrics, ethnicity, religion, sexual orientation, trade union membership. Higher bar, different lawful bases. If your business handles it, this is the conversation to have with a lawyer before anything is built.

Call recording and transcription. Consent and notification requirements around recording are their own topic. If you are deploying a voice agent, the greeting is doing compliance work as well as customer service work, and it should be written accordingly.

Children's data. Different standards apply. If it is in scope, do not improvise.

What this looks like in a roadmap

Governance is not a separate project that happens after the fun part. In a well-ordered 12-month roadmap it appears as a small item in quarter one) approved tool list, acceptable use note, processor register (and then as a standing check on every build item after that.

The check is one line: what personal data does this touch, and is that necessary? Asked at design time it costs nothing. Asked after launch it costs a rebuild.

Our own position, for what it is worth, is stated plainly and applies to every engagement: we follow UK GDPR, we use enterprise-grade providers, we agree data-handling boundaries in writing before anything is built, and client data is never used to train public models.

The short version

Buy the business tier. Tell your staff what they may paste. Send the minimum. Keep a list of who touches your data.

Four things, none of them expensive, and between them they cover the overwhelming majority of what actually goes wrong. The rest is worth a proper conversation) with us about the architecture, and with a lawyer about anything that carries regulatory weight.

If you want the AI part of that reviewed against how your business actually works, that is what consulting is for, and the first assessment is free.

Next step

Start with the free AI Opportunity Assessment.

A short, no-obligation conversation about where enquiries, hours and revenue leak today. You do not have to pick a tier to have it, and what comes out of it feeds Discover, so the first paid day starts from evidence rather than a blank sheet.