Your code is not the only thing leaking into AI prompts. The Excel file with your customer list, the Word contract, the HR email thread — those go too. Here's how to let the AI work on them without a single real name leaving the machine.


The request that started this

The pattern shows up in every company that adopts AI assistants, usually within weeks:

"Can you summarize this contract?"
"Clean up this workbook and add a total per region."
"Draft a reply to this email thread."

Every one of those documents is dense with exactly the data a company must not hand to a third party: customer names, email addresses, phone numbers, IBANs, salaries, negotiated prices, internal product codenames. A single invoice workbook can identify dozens of natural persons and reveal your entire pricing structure at the same time.

Under the GDPR, sending that document to a remote AI API is processing of personal data by a third party, often with a transfer outside the EU. Under any sane reading of trade-secret protection, it is also handing your commercial terms to infrastructure you don't control. The compliance answer cannot be "don't use AI" — the productivity gap is too large, and people will paste the document into a chat window anyway. Shadow usage is worse than governed usage.

So the real requirement, as stated by the people who have to sign off on this, is:

  1. The AI must still be able to reason on the document — sum the amounts, keep track of who said what in the thread, understand that the same person appears in three sheets.
  2. Nothing that identifies a person or reveals a sensitive value may leave the machine.
  3. The AI's answer — a text summary or a modified document — must come back with the real values restored, exactly.

Requirement 1 is what kills naive redaction. Replace every name with ███████ and the AI can no longer tell two parties of a contract apart; replace every amount with [REDACTED] and "add a total per region" is meaningless. The document has to stay plausible.

That is pseudonymization: each value is replaced by a same-format surrogate through a bijective mapping kept locally. Jean Dupont becomes another plausible person name, an internal email becomes another valid email address, an IBAN becomes another correctly-structured IBAN, a date becomes a date shifted by a constant offset. Same original → same surrogate everywhere, so cross-references survive. The mapping file never leaves the machine, and restoration is exact reverse substitution — longest match first, so Jean-Pierre Dupont doesn't get half-restored by the Jean Dupont entry.


Detection: patterns, NER, then a sweep

Detection runs in three layers, because each one fails differently:

Validated patterns catch the values with verifiable structure: email addresses, phone numbers, IBANs, card numbers, dates. One gotcha worth sharing: off-the-shelf recognizers often validate email domains against the public suffix list — which makes them reject addresses on internal domains ([email protected], [email protected]). Those are precisely the addresses you most need to pseudonymize, because they reveal your infrastructure. The fix was a permissive recognizer that keeps the local-part/domain shape requirement (so a @decorator in a code snippet is not matched) but drops public-suffix validation.

NER (named-entity recognition, via a local Presidio + spaCy sidecar — nothing goes to any cloud for detection either) catches the values with no structure at all: person names, organizations, locations. NER over spreadsheet fragments is noisy — a single-word column header like "Téléphone" gets tagged as an entity — so an exclusion list is a first-class feature, not an afterthought.

A dictionary sweep runs after detection: every mapped value is re-searched across all text fragments on word boundaries, catching occurrences the detector missed (the organization name inside the workbook's title property, the person mentioned mid-sentence in a cell comment). Accepted entities are expanded across case variants: detect Acme Industries once and ACME INDUSTRIES in a header is covered too, with a case-shaped surrogate, each variant restored exactly.

Numbers get their own rule system, because "hide the amounts" means different things per column: a multiplicative scale preserves order and ratios (the AI can still find the biggest customer), offset shifts, keep leaves quantities and years alone. Formulas are re-evaluated from the transformed inputs — a =SUM() over scaled cells shows the scaled total, so the workbook stays internally consistent and the AI never sees a sum that doesn't match its column.


The formats are where the real work is

The detection layer is maybe a third of the effort. The rest is format plumbing, and it is full of traps where the original value survives in a corner of the file you didn't rewrite. A few that cost real debugging time:

Format Trap
xlsx Rewriting a cell leaves the original string orphaned in sharedStrings.xml, readable by unzipping the file. Comment authors: the API call appends the new author and keeps the old one in the comments table. Hyperlink targets live in a separate .rels part.
docx / pptx Visible text is split across formatting runs — a bold surname interrupting a name breaks naive per-run matching. In pptx, a brand in a footer lives in the slide master, not the slide.
ODF (odt/ods/odp) The zip must keep the mimetype entry first and uncompressed, or applications refuse the file. Easy to break on write, invisible until a user opens the output.
PDF Input-only: text is extracted and pseudonymized as text (no layout rebuild). Scanned pages need OCR — and since OCR misreads, dictionary terms are fuzzy-matched (edit distance ≤ 2) on OCR'd content to catch a codename read with one wrong letter.
RTF Surrogates must be re-encoded through RTF escaping (non-ASCII as unicode escapes), and the metadata \info group holds author names.
HTML / XML / EPUB Text nodes only; script and style are code, not prose, and are left alone.
email (.eml/.msg) Addressing headers and body must map consistently — the sender named in the signature is the same entity as the From: header.

The lesson that generalizes: coverage must be locked by a test, not by review. The test plants a forced term in every rewritable channel of a workbook, pseudonymizes, then greps the raw output archive: the term must not survive anywhere. Anything that cannot be rewritten — defined names referenced by formulas, pivot caches, chart series, embedded images — is not silently passed through: it produces an explicit warning telling the user to check before submitting, and raster images are dropped by default (pixel-rendered text in a screenshot cannot be pseudonymized; an OCR-gated mode keeps only images whose text is clean and readable).

Fail-closed is the other non-negotiable: a document that cannot be pseudonymized is withheld and replaced by an explicit placeholder. The failure mode of a privacy tool must never be "sent in clear."


Making it work for a whole team: policy packs and suggest

Patterns and NER don't know your company. Project Nightingale, the internal product codename, is not in any model's training as a sensitive term. And you can't ask every project manager to maintain regexes.

Two configuration levels solve this:

  • A policy pack per document category (Finance, HR, pricing…), maintained centrally and distributed by URL, so every pseudonymization on every machine uses the current version. It bundles the policy (NER on/off, language, what happens to amounts and images), a [terms] dictionary and [numeric] rules. Transport is https-only, packs are cached (an unreachable server falls back to the cached copy; no cache and no server → the run fails rather than running without the company policy), and each pack is signed at publication — once the public key is installed on a machine, an unsigned, tampered or rolled-back pack is refused. The pack wins any conflict with local files: a user's exclusion cannot veto a term the company forces.
  • Local files next to a person's own documents (report-terms.txt, report-numeric.txt) complement the pack.

The cold-start problem — "what should even be in the dictionary?" — is handled by a suggest command: feed it a few representative documents of a category, and it emits a candidate terms file (NER results with occurrence counts, all-caps/acronym candidates in a commented "uncertain" section, where internal codes usually hide) and a candidate numeric-rules file (one scale line per numeric column, annotated with the header). A human prunes the false positives, flips scale to keep for quantities and years, renames the file — done. Ten minutes of review instead of a blank page.


Three ways to plug it into an AI workflow

CLI, for the one-shot case: pseudonymize report.docx writes report-pseudonymized.docx plus a local mapping file; the AI's text answer pipes through depseudonymize, and a modified document restores the same way.

An HTTP proxy, for the transparent case: the AI client's API traffic passes through a local proxy that pseudonymizes in both directions. Attachments are transformed before leaving; prompt text naming a value from an attached document reaches the AI under the same surrogate the document uses (otherwise your own question deanonymizes the attachment); tool reads of document files are swept; and the AI's streamed response is restored on the fly. The user types real names and reads real names — the provider only ever stores surrogates.

An MCP server, for clients whose network path can't be proxied (Claude Desktop): one tool reads a local document and returns its pseudonymized content into the conversation; a second restores an answer and writes it to a local file without returning the restored text to the conversation — because anything returned to the conversation is sent back to the provider on the next turn.


What this is — and is not — under the GDPR

Precision matters here, because "GDPR-compliant AI" is a marketing phrase and this is an engineering article.

What this technique implements is pseudonymization in the GDPR sense (art. 4(5)): personal data that can no longer be attributed to a person without additional information — the mapping — which is kept separately, locally, under your control. That supports data-minimization arguments (art. 5(1)(c)), data-protection-by-design (art. 25) and security-of-processing measures (art. 32): what the provider receives identifies no one without a file that never left your machine.

It is not anonymization. Context can still re-identify: if a document says "the CEO of the company headquartered in [small town]," no surrogate on the name prevents inference. Detection is best-effort by nature — that's exactly why the residual channels produce warnings instead of silence, and why documents that fail are withheld instead of forwarded. Pseudonymized data remains personal data under the GDPR; the transformation reduces risk and supports compliance, it does not exempt you from it. Any tool claiming otherwise is overclaiming.


Takeaways

  1. Plausibility is the requirement that shapes everything. The AI must reason on the document, so surrogates must keep format, cross-references and numeric coherence (formulas recalculated over transformed values). Redaction is easier to build and useless in practice.
  2. The file formats are the boss fight. Shared-strings tables, formatting runs, slide masters, .rels parts, RTF escapes — original values hide in places no visual review catches. Lock coverage with a plant-and-grep test over the raw archive, and warn loudly about every channel you can't rewrite.
  3. Team-scale privacy is a distribution problem. Central signed policy packs with fail-closed verification, local per-document overrides, and a suggest bootstrap turn "maintain a dictionary" from a chore nobody does into a ten-minute review.

The document pseudonymization described here ships in PromptCape alongside its code-obfuscation pipeline — CLI, proxy and MCP server share one engine and one session mapping. Worked examples and sample documents are in the companion repo at gitlab.com/gbreton7/promptcape-docs, and the tool is free to try for 3 months at promptcape.com.