What should never be pasted into an AI tool

Every prompt is a transfer. The interface makes it feel like a private notebook; legally and technically it is closer to sending an email to a company you have never audited.

This article sets out what actually leaves your machine when an AI assistant is used, what happens to it afterwards, and where European law makes that a problem rather than a preference. It ends with a list that can be pinned next to a screen, and with the five questions worth asking before a tool is allowed near client work.

It does not tell you which vendor is safe. Vendor terms are read clause by clause in a separate article, because a rule that depends on what a company promises this quarter is not a rule.

The short answer

Three things are true of nearly every AI assistant on the market, whatever the marketing page says.

What is typed leaves the machine. Unless a model runs locally on your own hardware, the prompt, the attachments and the file names are sent to a third party’s servers and processed there.

Deleting is a request, not a physical event. Deletion is governed by the vendor’s retention policy, by its backups, and — as 2025 demonstrated — by court orders the vendor does not control.

Something typed can, in principle, come back out. Not usually, not predictably, but the property has been demonstrated in production systems, and a regulator has built its guidance on that assumption.

None of this makes AI tools unusable. It makes one distinction load-bearing: between information that is yours to disclose, and information that belongs to someone who has not been asked.

What actually leaves your machine

More is sent than the sentence typed into the box.

The prompt itself, obviously. Any file attached to it, in full, including the parts you did not intend to be read: the tracked changes in a Word document, the hidden columns of a spreadsheet, the metadata of a PDF, the earlier draft still sitting in the revision history. A summary of a “cleaned” document is produced from the whole file, not from the visible part of it.

The conversation history, resent with each turn, which is how a client name mentioned twenty messages earlier stays in scope for the rest of the session.

Account-level information: who you are, which organisation pays, the IP address, the timestamps. That is enough to link a question to a firm even when the question itself is anonymous. “Draft a redundancy letter for a 12-person team” is not sensitive on its own; sent from a named corporate account on a Tuesday afternoon, it is.

And, for meeting tools, the recording — which contains what people said before the agenda started and after they thought it had ended.

Deleting is not always deleting

Retention policies describe intent. They do not always survive contact with litigation.

Between May and September 2025, OpenAI was placed under a court order in The New York Times Company v. Microsoft Corporation et al. directing it to preserve and segregate output log data that would otherwise have been deleted. In practice, conversations that Free, Plus, Pro and Team users had deleted, and Temporary Chats, were retained rather than removed on the usual 30-day cycle. ChatGPT Enterprise, ChatGPT Edu and Zero Data Retention API customers were outside the scope. The order ended on 26 September 2025, and OpenAI announced on 22 October 2025 that it was no longer required to retain consumer and API content indefinitely — while confirming that data captured between April and September 2025 is still held, under legal hold, for the plaintiffs’ continuing demands. That termination is narrower than its headline: data already segregated stays preserved, except for users in the EEA, Switzerland and the United Kingdom, and accounts associated with a list of 87 news domains must still be preserved going forward.

Two lessons survive the end of that order, and they are not about OpenAI.

The first is that the deletion promise shown in a consumer interface is subordinate to obligations the vendor has no ability to refuse. Any vendor, in any jurisdiction, can be placed under the same kind of order tomorrow.

The second is that scope followed the contract, not the product. The accounts excluded from the preservation order were the ones with a commercial agreement behind them. That distinction — consumer terms versus a signed data processing agreement — is the single most consequential choice a professional makes about AI tooling, and it is usually made by accident, by signing up with a work email on a Sunday evening.

What goes in can, in principle, come out

Model memorisation is not a hypothetical. In November 2023, a team of researchers from Google DeepMind, the University of Washington, Cornell, Carnegie Mellon, Berkeley and ETH Zürich published a study extracting training data from production language models, including ChatGPT. Their divergence attack pushed the model away from its assistant behaviour and caused it to emit training data “at a rate 150x higher than when behaving properly”. Their conclusion is the part that matters here: current alignment techniques do not eliminate memorisation.

Regulators have taken the same position. In its Opinion 28/2024, adopted on 17 December 2024, the European Data Protection Board held that an AI model trained on personal data cannot be assumed to be anonymous: for that claim to hold, it must be shown that it is “very unlikely” both to identify the individuals whose data was used, and to extract that personal data from the model through queries.

Two clarifications are owed to the reader, because the risk is routinely overstated.

Content typed into a chat is not automatically training data. Whether it is used for training depends on the plan and the settings, and business and API tiers commonly exclude it by contract.

And extraction is a research result obtained under adversarial conditions, not something a competitor is likely to do to your client file. The practical exposure runs the other way: a support engineer reviewing a flagged conversation, a misdirected share link, a subcontractor in a third country, a breach at the vendor. Mundane routes, and the ones that actually occur.

In Europe, pasting is processing

Under the GDPR, typing someone else’s personal data into an AI tool is processing it, and sending it to the vendor is a disclosure by transmission. The professional doing the typing does not become a bystander; in most cases the firm remains the controller, and the vendor becomes a processor. Several consequences follow, and none of them are triggered by the vendor’s conduct — they are triggered by yours.

A processor may only be used where a contract under Article 28 is in place. A consumer account accepted by clicking a checkbox is not that contract.

Purpose limitation and minimisation (Article 5) apply to the prompt. A whole client file pasted in to get a two-line summary is difficult to defend as adequate, relevant and limited to what is necessary.

If the data ends up outside the EEA, Chapter V applies, and a transfer mechanism has to exist. The Italian supervisory authority, acting as a matter of urgency on 30 January 2025, ordered the definitive limitation of DeepSeek’s processing of the data of users located in Italy, on grounds that included the storage of that data in the People’s Republic of China.

And if data is exposed through the tool, it is a personal data breach like any other: Article 33 gives the controller 72 hours from awareness to notify the supervisory authority. “It was pasted into a chatbot” is not a category of incident that regulators treat differently.

One obligation is newer, widely missed, and has just been rewritten. Article 4 of the AI Act, applicable since 2 February 2025, addressed AI literacy to providers and deployers of AI systems — and a company whose employees use a general-purpose chatbot for ordinary work is a deployer. The AI Omnibus regulation, in force since 27 July 2026, replaced that article: providers and deployers must now take measures to support the development of AI literacy among their staff, rather than guarantee a particular level of it, with more of the effort placed on the Commission and the Member States. Supervision and enforcement rules apply from 3 August 2026. The duty is lighter than it was and it has not gone away, which leaves the practical position unchanged: a firm with no written position on what may be pasted is in a weaker position than one that has written a bad policy.

For some professions, it is a criminal matter

For anyone bound by professional secrecy, the analysis stops being about data protection and becomes about criminal liability.

Article 226-13 of the French Code pénal punishes the disclosure of secret information by a person who holds it by virtue of their profession, office or a temporary assignment, with one year’s imprisonment and a fine of €15,000. It does not require harm, publication, or intent to profit. Disclosure to a single third party is disclosure. A cloud provider is a third party.

Three points are routinely got wrong.

Client consent does not always cure it. For French avocats in particular, professional secrecy is treated as a matter of public order and general interest, not as a right the client can waive on the lawyer’s behalf. Consent is worth obtaining; it should not be relied on as the whole defence.

Anonymising the name is not enough. A dispute described by sector, town, amount and date identifies the parties to anyone in the field. Re-identification is the test, not the presence of a surname.

Health, biometric, and criminal-offence data attract additional rules. Articles 9 and 10 of the GDPR apply to a therapist’s notes and to case material in the same way they apply to a hospital’s records.

Notaires, expert-comptables, médecins, avocats, and anyone handling case material under a temporary assignment are covered. So, in substance, are consultants under a confidentiality clause — with a contractual penalty rather than a criminal one, which is not always the lighter outcome.

The list

Never, into a consumer-grade account:

  • Client, patient or case material of any kind, including “anonymised” extracts that remain identifiable by context.
  • Identity documents, bank details, payroll files, medical records, contracts under negotiation.
  • Credentials, API keys, tokens, connection strings — pasting a key into a chat is the same event as publishing it, and it should be rotated afterwards, not merely deleted.
  • Anything covered by an NDA signed by your employer or your client.
  • Source code belonging to a client, or to an employer that has not authorised it.
  • Recordings of meetings whose participants have not been told a tool is being used. Consent to be recorded is not consent to be processed by a third-party vendor.

Only with a signed processing agreement, and only where necessary:

  • Internal documents containing colleagues’ personal data.
  • Commercial data whose leak would be damaging but not unlawful — pricing, pipeline, draft strategy.
  • Client material where the client has been informed in writing that an AI tool is used, and which tool.

Free of the problem entirely:

  • Public documents, published law, your own writing, invented examples.
  • Anything a client file has been rewritten into: structure, question and constraints, with the specifics removed. Most professional prompts survive this operation intact. “Draft a response to a claim of latent defects on a property sold in 2019, buyer alleging concealment” contains no personal data and produces the same answer as the version with names in it.

Five questions before a tool is allowed near client work

  1. Is the content used to train models by default, and can that be turned off contractually rather than in a settings panel? A toggle can be reset by an update; a clause cannot.
  2. How long is content retained after deletion, and where is that stated? A number in a contract, not a sentence on a marketing page.
  3. Which sub-processors receive it, and in which countries? The published sub-processor list is the answer, and it changes.
  4. Can humans read the conversations, and under what circumstances? Abuse review and quality review are legitimate; they are also human beings reading client material.
  5. Does a Data Processing Agreement exist, and has it been signed by the entity that is actually being invoiced? The commonest failure is a firm operating on personal accounts while believing it has an enterprise contract.

A tool that answers all five in writing can be used on real work. A tool that answers none of them can still be used — on the rewritten version of the problem, with the specifics kept out.

What this article does not settle

This is not legal advice, and it is not a substitute for a professional body’s own guidance, which for regulated professions takes precedence over anything written here.

It also stops deliberately short of naming which vendors satisfy which of the five questions. Those answers change with each revision of a set of terms, and stating them without a date would be worse than useless. They are established, clause by clause and with dates, in the vendor terms article.

What is claimed here is narrower: that the risk is not exotic, that it is mostly contractual rather than technical, and that the decisive step is taken long before any prompt is typed — when a professional signs up with a work email and no agreement behind it.

Sources

Consulted 13 September 2026.

Similar Posts