Sovereignty
Anonymise, then send to an AI: why that does not answer the question
Many institutions believe they have settled confidentiality by masking names before sending their data to an online AI. Here is what the CNIL says, what the European authority says about encryption, and why the requirement that “nothing leaves your network” is a question of architecture, not of text processing.
The reflex: mask the names, then send
The approach is widespread and it has genuine merits. A filter detects identifying data in a text, replaces it with tokens, sends the rest to an online model, then reinjects the real values into the answer. You keep access to the best models on the market, you deploy no infrastructure, you pay as you go.
For many use cases, this is a clear improvement on sending raw data. The problem appears when the organisation has set a strict requirement: the data must not leave its network. At that point, three pitfalls come into view.
First pitfall: pseudonymising is not anonymising
The CNIL defines pseudonymisation as processing that prevents data from being attributed to a person without additional information, and notes that it very often remains possible to recover people’s identities using third party data. The consequence: pseudonymised data remains personal data. It stays within the scope of the GDPR, with every obligation that comes with it.
True anonymisation is a far more demanding standard. It requires three conditions to be met at the same time: it must not be possible to single out an individual, nor to link together sets of data relating to that individual, nor to infer new information about them with near certainty.
On free text, these conditions are almost impossible to meet. A set of meeting minutes contains implicit identifiers that no filter masks completely: an age, a rare condition, a department, a date, a sequence of events, a family situation. The combination “47 year old, this condition, this department, this date” is often enough to re-identify someone. What the tools call anonymisation of free text is, in legal terms, almost always plain pseudonymisation.
Second pitfall: encryption does not cover the moment of processing
Encryption in transit and at rest is real and useful. But it protects the journey and the disk, not the computation. To transcribe or summarise a text, the model has to read it in the clear. At that precise moment, the data is decrypted inside the provider’s infrastructure.
The European Data Protection Board states this explicitly in its recommendations on supplementary measures: encryption is not sufficient when the provider needs access to the data in the clear in order to process it. That is exactly the case with an online AI service.
Third pitfall: the law follows the provider, not the server
A persistent idea holds that hosting “in a European region” settles the question of extraterritorial laws. It does not. The connecting factor is the provider, not the physical location of the machines: a provider subject to US law can be compelled to hand over the data it holds, including data stored in Europe.
The European authorities have further concluded that an order of this kind is not a valid legal basis for transfer under the GDPR. That does not make the provider immune: it creates a conflict of laws. And the compliance risk is carried by the data controller, which is to say by the client institution, not by the vendor.
The real question is not legal, it is architectural
This is the decisive point, and the one most often forgotten. When a leadership team lays down as a rule that its data must under no circumstances travel over the internet, it is not asking a question of legal qualification. It is asking a question of network traffic.
And sending a text, however carefully redacted, to an online interface is by definition sending data over the internet. The requirement is breached at the very first byte sent, quite apart from whether that byte is personal data or not.
Put differently, “anonymise, then send” answers a question the organisation never asked, and fails on the one it did ask. A requirement of this kind is an architectural requirement. The only answer that satisfies it literally is for the processing to run entirely inside the organisation’s network, and for nothing to be sent out.
What anonymisation is still good for
None of the above disqualifies the masking of identifying data. It is an excellent practice, provided it is put in its proper place: defence in depth inside the perimeter, and not a pass to step outside it.
Masking identifiers in a process that already runs on your own infrastructure reduces the impact of a misconfiguration or of improper access. On structured data or aggregates, true anonymisation is sometimes achievable, and there it is genuinely valuable. The difference lies in the order of operations: you anonymise because you are being careful, not to earn yourself the right to export.
Bring the model to the data
The reversal is simple to state: instead of sending the data to the model, you install the model where the data already is. The processing runs on the organisation’s own infrastructure. There is then no transfer, no foreign provider, no risk of re-identification outside the perimeter, and no conflict of laws.
This choice has a cost: it takes a machine and a solution built to work offline. In return, compliance is no longer something to be managed, it disappears by construction. For a leadership team that has set “nothing leaves your network” as a rule and not as a wish, it is the only answer that holds.
Sources
- CNIL, Anonymisation of personal data
- CNIL, Data anonymisation, a key process for open data
- EDPB, Recommendations 01/2020 on supplementary measures (post Schrems II)
- EDPB and EDPS, joint response to the LIBE Committee on the CLOUD Act
This article sets out one reading of publicly available texts and recommendations. It does not constitute legal advice: the qualification that applies to your situation must be established with your own counsel.