// Project
Anonymize before inference
An LLM gateway that strips personal data from prompts before the model, puts it back into the answer and traces every call without keeping the text. With measurements.
- Published
- Reading
- 12 min
Contents
Picture a 200-person company where every team has taken up LLMs on its own. One uses an API key that gets passed from project to project. Another could not tell you how much it spends. Customer service pastes customers’ messages into its prompts as they are, with their name, their address and sometimes their IBAN.
To bring some order, I built a gateway: a single point of passage between the teams and the models. Every call goes through it. It checks who is calling and how much budget they have left, removes the personal data from the text before sending it to the model, then puts it back in place in the answer. It also keeps a trace of every call with its cost, without ever keeping the text itself.
The company is made up. The code, on the other hand, really runs, with open-source models on my own machine, and every figure in this article can be reproduced with a command from the repo.
Personal data masked
99.2%
Out of 508 names, addresses, IBANs and numbers spread over 100 test texts.
Time added
+10ms
Per request, median. The model itself takes 0.7 to 3 seconds to answer in the demo.
Faster than an LLM
~900×
9 ms per text for Presidio, over 8 seconds for a 7-billion-parameter local model, which finds less.
Personal data in the logs
0
Checked on every code change.
The problem
Sending a prompt to a model is a bit like sending a letter to a contractor. Whatever it holds ends up at their place, in their archives, and with everyone who handles the letter along the way. For an LLM, those archives are called server logs and monitoring tools.
You could ban personal data from prompts, but that does not last long. A customer service agent answers people. To refund an order, they need the customer’s name, address and account number. Without them, there is nothing left to handle.
The model, on the other hand, does not need to know that the customer is called Marie Dupont. To write its answer, it only needs to know there is a person, an address and an account, and to be able to refer to them. That is the whole idea. Before sending, each item is replaced with a label, a bit like redacting a document, except that the mapping between labels and real values is kept aside. When the answer comes back, each value goes back in its place.
Here is what it looks like end to end, with the repo’s just demo command.
The team sends its message (1). The model only receives labels such as <PERSON_1> (2). The answer goes back to the team with the real values (3). As for the recorded trace, it only holds the team, the cost, the duration and the number of items masked (4).
One gateway in front of every model
Rather than asking every team to be careful, I wanted a single way in to the models, where the rules apply by default.
So teams call a single address that speaks the same language as the OpenAI API, which means their code barely changes. Each one has its own key, with its budget and its rate limit. They never ask for a model by name, but for a use: chat-small for quick answers, chat-large for more careful ones, embed for searching documents. If a better model comes out tomorrow, the platform plugs it in behind the same name and nobody has anything to change.
Build or reuse
The question came up early. Writing my own gateway was tempting, since I would have had exactly what I needed and nothing more. But the list of what it had to do grew fast: keys per team, budgets, rate limits, switching to another model when the first one goes down, a price per token, streamed answers, tracking every call. That meant months of work for a tool that is only a means to an end. What I wanted to show is the platform built around it.
I compared five options, from a home-made gateway to SaaS offerings, and chose LiteLLM Proxy, a widely used open-source project that covers almost all of the need. The full reasoning is in the ADR, the record where I write down each architecture decision along with the options I ruled out.
Decision: LiteLLM Proxy as the LLM gateway
LiteLLM is the only option that handles keys, budgets, switching between models, masking and tracking without a paid licence, and that runs the same way locally and on Kubernetes. Envoy AI Gateway can neither give each team a key nor count a budget in euros. Kong keeps masking for its Enterprise edition. A SaaS gateway, finally, would send prompts out of the company before they could even be masked. In the end, the code I wrote around LiteLLM comes down to an 85-line class, two setup scripts and two configuration files.
Mask, process, then rebuild
On every call, three steps follow one another.
- Mask. Before sending anything, LiteLLM hands the text to Presidio, an open-source tool from Microsoft that spots personal data. Each item found is replaced with a numbered label, and LiteLLM keeps the mapping in memory.
- Process. The model works on the masked text. It sees
<PERSON_1>where there was a name, and can use it in its answer like any other word. - Rebuild. When the answer comes back, LiteLLM replaces each label with its real value. The team gets a perfectly normal answer.
Sent by the team
Bonjour, je suis Marie Dupont. Ma commande n'est jamais arrivée au 12 rue de la Paix, 75002 Paris. Pouvez-vous me rembourser sur le compte FR76 3000 6000 0112 3456 7890 189 ?
Received by the model
Bonjour, je suis <PERSON_1>. Ma commande n'est jamais arrivée au <FR_ADDRESS_2>. Pouvez-vous me rembourser sur le compte <IBAN_CODE_3> ?
Dates stay visible, on purpose. An order or delivery date is often useful to the answer, and a date alone identifies nobody.
Teaching Presidio French
Presidio recognizes a lot in English, but nothing specifically French: not the social security number, not the tax number, not the shape of French addresses. Its English language model also misses most French names and cities. So I added a French language model and wrote recognizers, small specialized detectors, one per kind of data.
When an identifier carries a check digit, like the social security number or the IBAN, the recognizer verifies it. That is what keeps any random 13-digit number from being taken for a tax number.
Names were the hardest part. A language model misses a name when nothing around it announces one, in a letter header or a signature. So the name recognizer relies on the clues a person would notice: a title (M., Docteur), a form field (Nom :), or a known first name, among the 5,114 first names given to at least 500 children in France since 1900, according to INSEE, the French statistics office.
None of this is specific to French. Presidio works language by language, and the gateway’s Analyzer already loads an English model next to the French one. Adding German or Spanish means loading the matching language model and writing recognizers for that country’s identifiers.
Two bugs along the way
While testing on real French text, I ran into two LiteLLM bugs in how labels were numbered. The first one was obvious. When an address and the city inside it were both detected, the labels got mixed up into an unreadable FR_ADDRESS_2ON_4.
The second was sneakier. Each message in a conversation was numbered from 1, so the person named in the first message and the one named in the second both became <PERSON_1>. When rebuilding, the answer put one name in place of the other, without any visible error.
Both bugs are reported to LiteLLM (#42130, #31959) and still open. Until they are fixed, I work around them in a small subclass that replaces a single method. Tests check that it stays compatible with the LiteLLM version in use (ADR-015).
In practice
For the teams, all of this stays invisible. They call the gateway the way they would call the OpenAI API, with their team’s key, and get normal answers back. Three calls are enough to cover it. They run as they are once the gateway is started with just gateway-up. TEAM_KEY holds the customer service team’s key and MASTER_KEY the platform team’s. Both are in the .env file, as TEAM_KEY_SUPPORT and LITELLM_MASTER_KEY.
The prompts are in French, since that is what the gateway is tuned for. The first one asks the model to write to Marie Dupont, in one sentence starting with “Bonjour” and her name, that her refund goes to the account FR76…
An ordinary call
The team asks for a use, chat-large, rather than a model. Its message holds a name and an IBAN, and the answer comes back with both, as if nothing had happened.
curl -s localhost:4000/v1/chat/completions \
-H "Authorization: Bearer $TEAM_KEY" \
-H 'content-type: application/json' \
-d '{"model": "chat-large", "temperature": 0, "messages": [{"role": "system", "content": "Tu es le service client. Recopie les étiquettes entre chevrons telles quelles."}, {"role": "user", "content": "Écris à Marie Dupont, en une phrase qui commence par Bonjour et son nom, que son remboursement part sur le compte FR76 3000 6000 0112 3456 7890 189."}]}' \
| jq -r '.choices[0].message.content' Bonjour Marie Dupont, votre remboursement est en cours d'approbation et sera versé sur le compte IBAN : FR76 3000 6000 0112 3456 7890 189.
import os
from openai import OpenAI
client = OpenAI(base_url="http://localhost:4000", api_key=os.environ["TEAM_KEY"])
answer = client.chat.completions.create(
model="chat-large",
temperature=0,
messages=[
{"role": "system", "content": "Tu es le service client. Recopie les étiquettes entre chevrons telles quelles."},
{"role": "user", "content": "Écris à Marie Dupont, en une phrase qui commence par Bonjour et son nom, que son remboursement part sur le compte FR76 3000 6000 0112 3456 7890 189."},
],
)
print(answer.choices[0].message.content) The system message asks the model to copy the labels without touching them. Without it, in my tests, the demo’s small model often rewrote them its own way, and a damaged label can no longer be replaced with the real value.
What the model receives
The platform team can run the masking alone, without calling a model, to see exactly what goes out. That is what just demo does to show the masked text.
curl -s localhost:4000/guardrails/apply_guardrail \
-H "Authorization: Bearer $MASTER_KEY" \
-H 'content-type: application/json' \
-d '{"guardrail_name": "pii-fr", "text": "Écris à Marie Dupont, en une phrase qui commence par Bonjour et son nom, que son remboursement part sur le compte FR76 3000 6000 0112 3456 7890 189."}' \
| jq -r .response_text Écris à <PERSON_1>, en une phrase qui commence par Bonjour et son nom, que son remboursement part sur le compte <IBAN_CODE_2>.
Rules apply by default
Each team can only use what was planned for it. If the customer service team asks for embed, the document search model it has no use for, the gateway refuses before even contacting a model. An exceeded monthly budget or too many requests per minute lead to the same kind of refusal.
curl -s localhost:4000/v1/embeddings \
-H "Authorization: Bearer $TEAM_KEY" \
-H 'content-type: application/json' \
-d '{"model": "embed", "input": "Bonjour"}' \
| jq -r .error.message team not allowed to access model. This team can only access models=['chat-small', 'chat-large']. Tried to access embed
The numbers
To know whether masking holds up, I needed texts where I knew every personal data item in advance. I generated 100 fictional French texts (letters, emails, messages, forms) holding 508 of them, all annotated.
That left a question I would surely be asked: why not just ask an LLM to find the personal data? So I ran the same test on three local models of growing size, against Presidio. Everything runs on an i7-14700KF processor, without a graphics card.
| Detector | Precision (%) | Recall (%) | Fully masked (%) | Time per text (ms) |
|---|---|---|---|---|
| Presidio + French recognizers | 88.1 | 99 | 99.2 | 9 |
| qwen2.5:0.5b | 46.9 | 9.1 | 16.9 | 538 |
| qwen2.5:1.5b | 77.2 | 36 | 40.2 | 1,592 |
| qwen2.5:7b | 91.1 | 78.9 | 82.7 | 8,239 |
A little vocabulary to read this table. Recall is the share of personal data items that were found. Precision is, among everything the tool flagged, the share that really was personal data. The Fully masked column says the most: it is the share of items with not a single character left reaching the model, even when the tool got the category wrong. An address taken for a city stays hidden, and that is all that matters here.
The smallest model finds almost nothing. Even the biggest one, with its 7 billion parameters, lets about one item in six through, and needs more than 8 seconds per text where Presidio takes 9 thousandths. The gap widens on structured identifiers. An IBAN follows strict rules and carries a check digit, which a rule verifies with certainty and an LLM can only guess: the 7b model finds only 40% of IBANs. It wins on a single point, precision on names. When it flags a name, it almost always is one (99%, against 82% for Presidio), but it misses many more. For this specific job, a specialized tool and well-written rules beat a general-purpose model, at a fraction of the cost.
Presidio’s first run was not as good. It left 31 items in clear: names without context, addresses split over two lines, and recent Mastercard cards, whose numbers start with a 2 and which Presidio did not know. Three fixes took the rate from 93.9% to 99.2%. One doubt remained: had I simply learned those 100 texts by heart? So I generated a second set with another random seed, never using it to tune the recognizers. It goes from 94.1% to 99.4%, which shows the fixes also hold on texts they had never seen.
And the time cost? Masking adds 10 milliseconds per request at the median (14.8 ms instead of 4.8 ms), measured with a mock answer to isolate its cost. Next to a model that takes one to several seconds to answer, it goes unnoticed.
The proof: a no-leak test in CI
A good benchmark score is not enough. It measures Presidio on its own, while in real conditions, data can escape elsewhere: in a technical log that says a bit too much, or in a trace that keeps the text of the prompt. That is in fact what happened during my first tracking tests. The very first trace held “Bonjour Marie Dupont” in clear, because the real values are put back into the answer before LiteLLM records it. Since then, traces keep no message at all.
To check what actually leaves the gateway, I wrote a test that replays the 100 texts through it, as a team would, then looks for every item in three places:
- in what the model receives, by restarting LiteLLM in
DEBUGmode for the duration of the test, the only level at which it logs the exact body sent to Ollama; - in the Langfuse traces of those calls;
- in the logs of LiteLLM and Ollama, at their usual level.
The test is deliberately strict, since an item counts as leaked as soon as a single one of its words gets through. That is what revealed the case of Alexandrie Toussaint. Presidio had taken the first name Alexandrie for the city and masked it as a place, but the surname stayed visible: <LOCATION_1> Toussaint. Searching for the full name, the test would have seen nothing.
| Where | With masking | Without masking |
|---|---|---|
| Text received by the model | 3 | 1,512 |
| Langfuse traces | 0 | 0 |
| LiteLLM and Ollama logs | 0 | 0 |
The three words that still get through are the three misses already known from the benchmark: Toussaint, Paris and Caen. They are listed in the test, which fails if a new miss appears, but also if one of them disappears, so that the list stays accurate. To make sure the test really can spot a leak, I watched it fail while causing each leak by hand: turning masking off, turning message recording back on in the traces, then switching LiteLLM to verbose mode. It now runs on every pull request, in a CI job of its own.
just test -m leak collected 311 items / 309 deselected / 2 selected test_no_leak.py::test_traces_and_logs_hold_no_personal_data PASSED [ 50%] test_no_leak.py::test_the_model_receives_only_the_known_misses PASSED [100%] ========== 2 passed, 309 deselected in 108.03s (0:01:48) ==========
Traces keep everything needed to run the service: the team, the model, the number of tokens, the duration and the cost of each call. The text is replaced with redacted-by-litellm. Each team’s lead only sees their own team’s calls.
The limits
- Presidio still misses some data. A city mentioned without context or a first name too rare for the INSEE list can get through. Since the 100 texts come from the same document templates, other kinds of documents will surely bring up other misses.
- The model has to copy labels exactly.
qwen2.5:1.5bsometimes drops the angle brackets and writesPERSON_1, or swaps them for square brackets. The damaged label is then not recognized and stays in the answer. An instruction in the system prompt was enough in my tests, but I have not measured how often it happens yet. - Streamed answers arrive in one piece. A label can be split between two chunks of the answer, so LiteLLM waits for the whole answer before putting the values back.
- Some LiteLLM features are paid. Turning masking on team by team is one of them. So I turned it on for everyone by default, and the teams that do not need it are opted out through a setting only the platform team can change.
What comes next
This gateway is the first block of an internal AI platform. The second deploys it on Kubernetes, in llmops-platform, with vLLM instead of Ollama, Envoy Gateway at the entrance and keys stored in a secret manager. The third will be its first real client, an agent that draws on a document base and whose quality is evaluated on every change. Each will get its own article.
Until then, two workstreams matter more than the rest: reducing Presidio’s misses and making label copying reliable, because they decide what actually reaches the model.
If you want to see it all run, the code, the architecture decisions and the benchmark are on GitHub. Two commands are enough: just gateway-up, then just demo.