GDPR and AI — Personal Data in Prompts, DPIA and LLM Vendor Agreements (Practically)
RETURN_TO_BLOG
AI & Security 15 min

GDPR and AI — Personal Data in Prompts, DPIA and LLM Vendor Agreements (Practically)

Paweł Wiszniewski
Paweł Wiszniewski
SEO & GEO Specialist · AI Engineer

A consultant pastes a client's email into ChatGPT to get a draft reply. An HR person uploads a candidate's CV into an AI tool to summarize their experience. A salesperson has a chatbot judge whether a lead is “hot” based on call notes. In none of these cases did anyone think “I'm processing personal data” — and each of them just did exactly that, under GDPR, with the full set of obligations that follow. Six weeks before this post was published, on August 6, 2026, the head of Poland's data protection authority (UODO) published the first official preliminary question lists for checking AI tools' GDPR compliance — a signal that the Polish regulator has stopped theorizing and started handing companies concrete self-assessment tools. It's a good moment to do the same inside your own company, before an inspection does it for you.

Pasted a customer's email into ChatGPT to speed up your reply? That's already personal data processing under GDPR — with the full weight of obligations most teams have never heard of. Six weeks ago Poland's data protection authority (UODO) published the first official self-assessment checklists for AI/GDPR compliance — proof the regulator is already watching, not just theorizing. When does a prompt trigger a DPIA, how does the “meaningful human involvement” test from Article 22 hold up against a lead-scoring chatbot, and how do OpenAI's, Anthropic's and Google Cloud's DPA agreements actually differ — a practical guide without the legal jargon.

This post opens a compliance triangle I complete across several articles: technical data security answers “will the data leak,” the AI Act answers “is the AI system allowed and properly labeled,” and GDPR — the subject of this post — answers “are you even allowed to process this data this way at all.” Three different questions, three different responsibilities, one and the same deployed chatbot.

The compliance triangle — why GDPR is a separate game

/// THE COMPLIANCE TRIANGLE — THREE DIFFERENT QUESTIONS

Meeting one condition doesn't exempt you from the other two

01
TECHNICAL SECURITY
"Will the data leak?" — encryption, access, whether the vendor trains on your data. That's engineering, not law
02
THE AI ACT
"Is the AI system allowed and labeled?" — risk classification, documentation duties, marking AI-generated content
03
GDPR
"Are you even allowed to process this data this way?" — legal basis, DPIA, vendor agreement. Independent of how secure the system is

Companies deploying AI most often confuse these three layers, or handle only one of them. Technical security asks whether data is encrypted and whether the vendor won't use it for training — that's engineering. The AI Act asks whether the system is properly classified, labeled and documented — that's product regulation compliance. GDPR asks something more fundamental: do you even have a legal basis to process this specific personal data in this specific way — regardless of how secure the system is or whether the AI Act permits it. You can have a system that is fully secure technically and compliant with the AI Act, and still be breaking GDPR, because nobody checked the legal basis or ran the required impact assessment.

When a prompt becomes personal data processing

The answer is simpler than most teams want to hear: if a prompt, an attachment, or the context of a conversation with an AI model contains any information that identifies a specific person — a full name, an email address, a phone number, the content of a customer message, data from a CV — that is already personal data processing under GDPR, the moment the query is sent to the model. It doesn't matter that nobody “consciously saves” that data — simply transmitting it to a system (let alone the vendor retaining it for logging or service-improvement purposes) is itself a processing operation.

Two questions follow that must be answered before any deployment: who is the controller and who is the processor (usually your company is the controller and the model vendor — OpenAI, Anthropic, Google — is a processor acting on your instructions, which requires a data processing agreement), and what is the legal basis for this specific processing (consent, contractual necessity, legitimate interest — each with different requirements and different risks). A company that cannot answer both questions for every AI deployment touching personal data is not ready for an inspection.

DPIA for an AI deployment — when it's mandatory and how to do it

A Data Protection Impact Assessment (DPIA) is required when processing may pose a high risk to people's rights and freedoms — and EDPB guidelines list nine criteria, of which meeting two or more automatically triggers the obligation: evaluation or scoring, automated decision-making with a significant effect, systematic monitoring, sensitive data, large-scale processing, combining datasets, data of vulnerable subjects (e.g. children, patients), innovative use of technology, and blocking access to a service. Systematic, extensive profiling with a significant legal effect triggers the DPIA obligation on its own, without needing a second criterion. Most customer-service chatbot or lead-scoring deployments meet at least two of these criteria at once — evaluation plus processing at the scale of the entire customer base — so a DPIA is the rule, not the exception.

/// THE 9 EDPB CRITERIA — TWO OR MORE = DPIA MANDATORY

Systematic, extensive profiling with a significant effect triggers it on its own

01Evaluation or scoring
02Automated decisions with a significant effect
03Systematic monitoring
04Sensitive (special-category) data
05Large-scale processing
06Combining datasets
07Data of vulnerable subjects
08Innovative use of technology
09Blocking access to a service

In 2026 the EDPB adopted its first harmonized DPIA template along with explanatory guidance — a signal of how structured and well-documented regulators expect the assessment to be. A practical process for an AI deployment:

  1. 1.Describe the processing — what data, for what purpose, who has access, how long it's retained (including retention at the model vendor).
  2. 2.Assess necessity and proportionality — could the goal be achieved with less data or a less invasive tool.
  3. 3.Identify risks to individuals — a wrong AI decision, a leak, a discriminatory model error, no way to appeal.
  4. 4.Plan mitigating measures — anonymization before the prompt, limited retention, a human-review threshold, a DPA with the vendor.
  5. 5.Consult your DPO (Data Protection Officer), if you're required to appoint one, before running the system in production.

If the AI system is also a high-risk system under the AI Act, the DPIA doesn't replace a fundamental rights impact assessment (FRIA) — the two documents need to line up with each other, which I cover in the post on the AI Act.

Article 22 in practice — chatbots, scoring, and the “meaningful human involvement” test

Article 22 GDPR prohibits decisions based solely on automated processing that produce legal effects or similarly significantly affect a person — loans, insurance, hiring, credit scoring are the classic examples, but exactly the same mechanism applies to a chatbot that automatically rejects a complaint, or a system that scores a lead and automatically drops it from the sales pipeline without any human involved.

The key word is “solely” — and this is where most companies fall into a trap. Adding a formal “human review” step isn't enough if that review is a fiction. The EDPB, in its 2024 opinion on AI models, clarified what genuine human involvement requires:

/// ARTICLE 22 — THE "MEANINGFUL HUMAN INVOLVEMENT" TEST

Missing even one element = the decision still counts as fully automated

01
REAL AUTHORITY
The reviewer can change or reject the decision — not a single-click "approve" button
02
ACCESS TO THE DATA
They see all the data the model used to reach the decision, not just the output
03
UNDERSTANDING THE LOGIC
They understand the model's criteria well enough to genuinely evaluate the decision
04
ADDITIONAL INFORMATION
They can factor in information the model didn't have or didn't process
  • The person reviewing the decision has real authority to change or reject it — not a single-click “approve” button.
  • They have access to all the data the model used to reach the decision, not just the output.
  • They understand the logic and criteria behind the model's decision well enough to genuinely evaluate it.
  • They can factor in additional information the model didn't have or didn't process.

A mere “rubber stamp” without these four elements does not take the processing outside the scope of Article 22 — meaning the person affected by the decision still has the right to obtain human intervention, express their point of view, and contest the decision. When designing a chatbot or scoring system, build this mechanism in from the start, not as a patch after a customer complaint.

DPAs and data residency — how OpenAI, Anthropic and Google differ

Signing a data processing agreement (DPA) with a model vendor is a necessary condition for legally using AI with personal data — without it, the vendor is processing data with no legal basis on your side, no matter how good its safeguards are. The three major vendors differ in their defaults, though, and the differences matter in practice:

VendorDPAEU data residencyPractical note
OpenAISigned through the account dashboard, not on by defaultAvailable on higher-tier plans, requires explicit configurationZero Data Retention for select API endpoints mitigates retention risk
AnthropicBuilt into the Commercial Terms of Service — accepted along with the termsNo native EU region for most tiers — API processing happens in the USTransferring data to the US requires your own transfer assessment (SCCs) in your records of processing
Google (Vertex AI / Gemini Enterprise)Cloud Data Processing Addendum as the standard agreementConfigurable data residency in a chosen EU regionPaid API tiers don't train on your data by default — verify this explicitly for the specific product's terms

None of these agreements relieves you of controller responsibility — a DPA governs the relationship with the processor, but you are the one answerable to the regulator for whether you had a legal basis to send that data there in the first place. Check the vendor's DPA before signing the deployment contract with your client, not after.

Anonymization and pseudonymization before the prompt — the cheapest safeguard you have

The most effective way to reduce GDPR risk in AI doesn't require any agreement or consent at all — it consists of making sure personal data never reaches the model in identifiable form. A pre-processing layer before the prompt is sent:

/// ANONYMIZATION LEVELS — FROM RISKY TO SAFE

The cheapest safeguard: data that never reaches the model in identifiable form

01
NEVER SEND DIRECTLY
National ID, card data, special-category data — without a separate, documented legal basis
02
PSEUDONYMIZE BEFORE THE PROMPT
"John Smith, john.smith@company.com" → token "Customer_A482"; substitute real data back in only after the model responds
03
AGGREGATE WHERE POSSIBLE
Asking about a trend in complaints doesn't need names — it needs complaint content without identifiers
04
SELF-HOSTED MODEL
When anonymization isn't possible and the AI must know the customer's identity — the only fully safe option
  • Never send directly: national ID numbers, payment card data, special-category data (health, religion, sexual orientation) without a separate, documented legal basis.
  • Pseudonymize before the prompt whenever possible: turn “John Smith, john.smith@company.com” into a token like “Customer_A482” before sending it to the model, and substitute the real data back in only on your application side, based on the response. This requires a thin layer of code (a regex or a lightweight NER model detecting personal data), but it radically reduces exposure.
  • Aggregate wherever the purpose allows it: if you're asking the model about a trend in customer complaints, you don't need individual names — you need complaint content stripped of identifiers.
  • For the most sensitive data — a self-hosted model. If anonymization isn't possible (e.g. the AI needs to know the customer's identity to complete a task), the only fully safe option is a model running on your own infrastructure, with nothing sent externally — I describe the architecture for such deployments in AI data security.

UODO is already watching — what the new question lists cover

The lists UODO published on August 6, 2026 don't replace a full risk analysis or a DPIA, but they're a useful first filter: a separate set of questions for SMEs using off-the-shelf AI tools (without training their own model), a separate set for the public sector (accounting for the principle of legality and administrative procedure), and a separate set for organizations building or fine-tuning their own models. Until September 30, 2026, UODO is collecting feedback from practitioners on these lists at a dedicated address — a sign the document will keep being updated, so treat it as a living standard rather than a one-off publication. For a company deploying an off-the-shelf chatbot (e.g. a customer service system or a business chatbot), the first set is the relevant one — and a good starting point before your own DPIA, not a substitute for it.

Deployment checklist for SMEs

  1. 1.Map where personal data already reaches AI — both knowingly and unknowingly (see: Shadow AI, where I describe the scale of the problem of tools used without the company's knowledge).
  2. 2.Establish the controller and legal basis for every deployment touching personal data — before launch, not after.
  3. 3.Go through UODO's question list appropriate to your category as a first review.
  4. 4.Run a DPIA if you meet two or more EDPB criteria — use the new template as a starting point.
  5. 5.Sign a DPA with the model vendor and check the EU data residency option if available.
  6. 6.Implement pseudonymization before the prompt everywhere the business purpose allows it.
  7. 7.Design genuine human oversight for every AI decision with an effect on a customer — with real authority to change it, not a cosmetic review.
  8. 8.Document everything — the records of processing activities, the DPIA, the DPA, the human-verification procedure. An inspection asks for documents, not intentions.

---

I've been building GDPR-compliant AI deployments from day one: data-flow mapping, DPIAs, choosing the right DPA and data residency with a vendor, pseudonymization and self-hosting architecture for sensitive data. I do this as part of AI consulting. Reach out — I'll start with an audit of where personal data actually reaches AI tools in your company today, and what your real exposure is.

Worth reading next:

/// AUTHOR
Paweł Wiszniewski – AI & Web Engineer

Paweł Wiszniewski

SEO & GEO Specialist & AI Engineer

SEO/GEO specialist (10 years) and AI engineer (3 years). I build search visibility, AI systems and automations that reduce costs and improve operational efficiency.

Signal received?

Terminate
Silence

Initiate protocol. Establish connection. Let's build something loud.

> WAITING_FOR_INPUT...