// Tool
Presidio, in French
What Presidio needs to learn to find names, addresses, social security numbers and IBANs in French texts, without an LLM: recognizers, check digits and measurements on 100 texts.
- Published
- Reading
- 13 min
Contents
One morning, the customer service team receives a letter from Camille Roussel. She is writing about her mother’s medical refunds, and she gives everything needed to handle the request: her address, her mother’s social security number and tax number, an IBAN, a card number, her phone number and her email. The advisor would like a language model to help draft the answer. But before sending this letter to a model, every one of these values has to be found and replaced with a label.
In a first article, I described the gateway that does this for a whole fictional company. This one opens up the part that finds the data: Presidio, an open source tool from Microsoft. Presidio is very good at finding personal data in English text. In French, it misses a lot, and this article is about what I had to teach it, with the numbers to back it up. Everything below can be reproduced with the gateway’s repo, even without the gateway, since Presidio works perfectly well on its own.
Presidio as it comes
59.1%
Share of personal data fully masked in 100 French texts, with the official image and its English model.
With a French model
72.8%
The French language model finds the cities and most names, but still no address and no social security number.
With French recognizers
99.2%
Seven detectors written for the project, which check the check digits and read the clues around names.
Per text
9ms
Median analysis time on a CPU, no graphics card. A local LLM with 7 billion parameters takes more than 8 seconds.
What Presidio sees out of the box
The simplest way to try Presidio is to start its official image and send it a text. That is what I did with Camille Roussel’s letter. Its numbers are made up, but their check digits are right, as they would be on real numbers.
The result is puzzling. The official image only ships an English language model, so it reads the letter as English text. It does find the IBAN, the email and both full names. But it also sees a person in “joindre au 06 12” (reach me at 06 12) and a place in “Cordialement” (Kind regards). Above all, it misses the address, the social security number, the tax number and the card number.
The figure below shows the same letter, and two other texts, as seen by three versions of Presidio. These are Presidio’s real answers, recorded by the repo’s just presidio-steps script.
Camille Roussel 8 allée des Tilleuls 33000 Bordeaux Bonjour, Je vous écris au sujet du remboursement de soins de ma mère, Mme Josiane Roussel. Son numéro de sécurité sociale est le 2 52 07 33 063 112 03 et son numéro fiscal le 30 23 217 600 053. Le remboursement peut partir sur le compte FR76 3000 6000 0112 3456 7890 189, ou sur la carte 2221 0000 1234 5673. Vous pouvez me joindre au 06 12 34 56 78 ou à camille.roussel@example.fr. Cordialement, Camille
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Camille Roussel | SpacyRecognizer | 0.85 | |
| PERSON | Mme Josiane Roussel | SpacyRecognizer | 0.85 | |
| IBAN_CODE | FR76 3000 6000 0112 3456 7890 189 | IbanRecognizer | 0.5 → 1 | valid key |
| PERSON | joindre au 06 12 | SpacyRecognizer | 0.85 | false positive |
| PHONE_NUMBER | 06 12 34 56 78 | PhoneRecognizer | 0.4 | |
| EMAIL_ADDRESS | camille.roussel@example.fr | EmailRecognizer | 0.5 → 1 | valid key |
| LOCATION | Cordialement | SpacyRecognizer | 0.85 | false positive |
| PERSON | Camille | SpacyRecognizer | 0.85 |
Camille Roussel 8 allée des Tilleuls 33000 Bordeaux Bonjour, Je vous écris au sujet du remboursement de soins de ma mère, Mme Josiane Roussel. Son numéro de sécurité sociale est le 2 52 07 33 063 112 03 et son numéro fiscal le 30 23 217 600 053. Le remboursement peut partir sur le compte FR76 3000 6000 0112 3456 7890 189, ou sur la carte 2221 0000 1234 5673. Vous pouvez me joindre au 06 12 34 56 78 ou à camille.roussel@example.fr. Cordialement, Camille
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Camille Roussel | SpacyRecognizer | 0.85 | |
| PERSON | Mme Josiane Roussel | SpacyRecognizer | 0.85 | |
| IBAN_CODE | FR76 3000 6000 0112 3456 7890 189 | IbanRecognizer | 0.5 → 1 | valid key |
| LOCATION | FR76 | SpacyRecognizer | 0.85 | |
| PHONE_NUMBER | 06 12 34 56 78 | PhoneRecognizer | 0.4 | |
| EMAIL_ADDRESS | camille.roussel@example.fr | EmailRecognizer | 0.5 → 1 | valid key |
| PERSON | Camille | SpacyRecognizer | 0.85 |
Camille Roussel 8 allée des Tilleuls 33000 Bordeaux Bonjour, Je vous écris au sujet du remboursement de soins de ma mère, Mme Josiane Roussel. Son numéro de sécurité sociale est le 2 52 07 33 063 112 03 et son numéro fiscal le 30 23 217 600 053. Le remboursement peut partir sur le compte FR76 3000 6000 0112 3456 7890 189, ou sur la carte 2221 0000 1234 5673. Vous pouvez me joindre au 06 12 34 56 78 ou à camille.roussel@example.fr. Cordialement, Camille
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Camille Roussel | FrPersonRecognizer | 0.85 → 1 | context “ami” |
| FR_ADDRESS | 8 allée des Tilleuls ⏎ 33000 Bordeaux | FrAddressRecognizer | 0.6 | |
| PERSON | Mme Josiane Roussel | SpacyRecognizer | 0.85 | |
| FR_NIR | 2 52 07 33 063 112 03 | FrNirRecognizer | 0.3 → 1 | context “sécurité” · valid key |
| FR_FISCAL_NUMBER | 30 23 217 600 053 | FrFiscalNumberRecognizer | 0.3 → 1 | context “fiscal” · valid key |
| IBAN_CODE | FR76 3000 6000 0112 3456 7890 189 | IbanRecognizer | 0.5 → 1 | context “compte” · valid key |
| LOCATION | FR76 | SpacyRecognizer | 0.85 | |
| CREDIT_CARD | 2221 0000 1234 5673 | CardRecognizer | 0.3 → 1 | context “carte” · valid key |
| PHONE_NUMBER | 06 12 34 56 78 | PhoneRecognizer | 0.4 → 0.75 | context “joindre” |
| EMAIL_ADDRESS | camille.roussel@example.fr | EmailRecognizer | 0.5 → 1 | valid key |
| PERSON | Camille | FrPersonRecognizer | 0.85 → 1 | context “ami” |
Salut Anaïs, le Docteur Camus a validé l'arrêt. Tu peux le déposer chez Étienne avant vendredi ? Sinon appelle-moi au 01 46 39 90 88. Bises, Olivier
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Salut Anaïs | SpacyRecognizer | 0.85 | |
| LOCATION | Sinon appelle-moi au 01 | SpacyRecognizer | 0.85 | false positive |
| PERSON | Olivier | SpacyRecognizer | 0.85 |
Salut Anaïs, le Docteur Camus a validé l'arrêt. Tu peux le déposer chez Étienne avant vendredi ? Sinon appelle-moi au 01 46 39 90 88. Bises, Olivier
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Olivier | SpacyRecognizer | 0.85 |
Salut Anaïs, le Docteur Camus a validé l'arrêt. Tu peux le déposer chez Étienne avant vendredi ? Sinon appelle-moi au 01 46 39 90 88. Bises, Olivier
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Anaïs | FrPersonRecognizer | 0.3 → 0.65 | context “salut” |
| PERSON | Camus | FrPersonRecognizer | 0.85 → 1 | context “salut” |
| PHONE_NUMBER | 01 46 39 90 88 | PhoneRecognizer | 0.4 → 0.75 | context “appel” |
| PERSON | Olivier | SpacyRecognizer | 0.85 |
Nom : Buisson Prénom : Hortense Adresse : 21 rue Jean Moulin 74000 Annecy N° de sécurité sociale : 1 85 05 78 006 084 91 Téléphone : 04 50 12 34 56
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Hortense ⏎ Adresse | SpacyRecognizer | 0.85 | |
| PERSON | Jean Moulin | SpacyRecognizer | 0.85 | |
| LOCATION | Annecy | SpacyRecognizer | 0.85 | |
| PHONE_NUMBER | 04 50 12 34 56 | PhoneRecognizer | 0.4 → 0.75 | context “phone” |
Nom : Buisson Prénom : Hortense Adresse : 21 rue Jean Moulin 74000 Annecy N° de sécurité sociale : 1 85 05 78 006 084 91 Téléphone : 04 50 12 34 56
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Buisson ⏎ Prénom | SpacyRecognizer | 0.85 | |
| PERSON | Hortense ⏎ Adresse | SpacyRecognizer | 0.85 | |
| LOCATION | Annecy | SpacyRecognizer | 0.85 | |
| PHONE_NUMBER | 04 50 12 34 56 | PhoneRecognizer | 0.4 → 0.75 | context “phone” |
Nom : Buisson Prénom : Hortense Adresse : 21 rue Jean Moulin 74000 Annecy N° de sécurité sociale : 1 85 05 78 006 084 91 Téléphone : 04 50 12 34 56
Highlighted: found. Dotted: left in clear. Struck through: false positive.
| Type | Span | Recognizer | Score | Clue |
|---|---|---|---|---|
| PERSON | Buisson ⏎ Prénom | SpacyRecognizer | 0.85 | |
| PERSON | Hortense | FrPersonRecognizer | 0.85 → 1 | context “prénom” |
| PERSON | Hortense ⏎ Adresse | SpacyRecognizer | 0.85 | |
| FR_ADDRESS | 21 rue Jean Moulin ⏎ 74000 Annecy | FrAddressRecognizer | 0.6 → 0.95 | context “adresse” |
| PERSON | Jean Moulin | FrPersonRecognizer | 0.6 | |
| LOCATION | Annecy | SpacyRecognizer | 0.85 | |
| FR_NIR | 1 85 05 78 006 084 91 | FrNirRecognizer | 0.3 → 1 | context “sécurité” · valid key |
| PHONE_NUMBER | 04 50 12 34 56 | PhoneRecognizer | 0.4 → 0.75 | context “téléphon” |
Adding the French language model gets rid of most of these false positives, but not of the missing data. No stock detector knows the NIR (the French social security number), the tax number or the shape of a French address. And in the short message, “Anaïs”, “Étienne” and “Docteur Camus” all slip past the model. A name on its own, with no sentence around it to announce it, is not enough for it.
How Presidio decides
To see what needs adding, it helps to first see how Presidio works. Picture it as a panel of small specialists reading the text at the same time.
The first one is a language model, spaCy, trained on large amounts of annotated text to recognize the names of people and places. It judges from the sentence: “Mme Josiane Roussel” looks like a person because the words around it suggest so. The other specialists are recognizers, small detectors that each look for one thing only. The email one looks for an at sign surrounded the right way; the IBAN one looks for two letters followed by digits, and checks their check digits.
Each specialist that thinks it has found something proposes a span of the text with a score between 0 and 1, which says how sure it is. Two mechanisms then move that score:
- Check digits. When a value has them, the recognizer checks them. If they are right, the score goes to 1. If they are wrong, the span is dropped, whatever its starting score.
- Context. Each recognizer has its list of clue words. If one of them appears in the five words before the span, the score gains 0.35. “domicilié” (residing at) announces an address; “salut” (hi) announces a first name.
Finally, a threshold sorts things out. In the gateway, it is set to 0.4: anything below is ignored.
My whole strategy fits in these three rules. A pattern that looks like personal data, but could just as well be something else, starts at 0.3, just below the threshold. It only gets through if check digits or a clue word confirm it. Here is what that gives on real requests:
-
2 52 07 33 063 112 04
shape of a NIR, wrong key
0.3 → 0
dropped
-
2 52 07 33 063 112 03
shape of a NIR, right key
0.3 → 1
kept
-
« Il vit à 33000 Bordeaux » (he lives in)
no clue word
0.3
ignored
-
« Il est domicilié à 33000 Bordeaux » (he resides in)
clue “domicilié”
0.3 → 0.65
kept
-
« Tu peux le déposer chez Étienne ? » (can you drop it at Étienne’s?)
first name alone, no clue
0.3
ignored
-
« Salut Étienne, tu peux le déposer ? » (hi Étienne, can you drop it?)
clue “salut”
0.3 → 0.65
kept
-
« Le Docteur Camus a validé. » (Doctor Camus approved)
after a title
0.85
kept
Presidio can explain each of its decisions, as long as you ask with the return_decision_process option. We will come back to it at the end of the article, when it is time to run the commands yourself.
Numbers that check themselves
Many French identifiers have a very handy property: their last digits are computed from the first ones. That is the check key. It was designed to catch typos, and it also tells a real number apart from any other number of the same length.
The social security number
The NIR has 15 characters that tell a little story: sex, year and month of birth, department and town of birth, an order number, then a two-digit key. The key is 97 minus the remainder of the first 13 digits divided by 97. The Corsican departments complicate the computation a little, since 2A and 2B contain a letter. They are replaced with 19 and 18 before dividing.
The calculator below redoes this computation step by step. Try changing a single digit: the expected key changes, and the number becomes invalid.
15 characters expected: 13 digits (or 2A, 2B for Corsica) and a 2-digit key.
- Sex
- 2
- Year
- 52
- Month
- 07
- Department
- 33
- Town
- 063
- Order
- 112
- Key
- 03
- Without spaces
252073306311203 - First 13 digits, 2A → 19 and 2B → 18
2520733063112 - Remainder of the division by 97
94 - Expected key
97 − 94 = 03 - Written key
03valid
The recognizer does exactly this computation. The pattern describes the structure of the number, with or without spaces, and validate_result checks the key:
def validate_result(self, pattern_text: str) -> bool:
nir = pattern_text.replace(" ", "").upper()
digits = nir[:13].replace("2A", "19").replace("2B", "18")
return 97 - int(digits) % 97 == int(nir[13:])A random 15-digit number has only about one chance in 97 of passing this check, and it first needs the structure of a NIR, with a plausible month and department. False positives become so rare that a valid number can be masked even with no clue word around it.
The tax number
The tax number, the one printed on French tax notices, has 13 digits and starts with 0, 1, 2 or 3. Its last three digits are the first ten modulo 511. It is a little-known rule, and yet a very effective one: out of 511 random 13-digit numbers, only one follows it.
IBAN and card numbers
Presidio already checks the key of an IBAN, and that of a card number with the Luhn algorithm. For IBANs, all it took was adding French clue words such as “RIB”, “virement” (transfer) and “prélèvement” (direct debit).
Card numbers held a surprise. With the stock recognizer, the benchmark lets four card numbers out of ten through, all Mastercard numbers starting with 2. Mastercard has issued numbers between 2221 and 2720 since 2017, and Presidio’s pattern only accepts cards starting with 1, 3, 4, 5 or 6. I wrote a recognizer that inherits from Presidio’s and only changes the pattern, keeping the Luhn check. All ten cards of the test set are now found.
Addresses, without a safety net
An address has no check key. Nothing mathematically tells “12 rue de la Paix” from “12 pages de la notice” (12 pages of the manual). So the recognizer relies on shape: a number, possibly followed by bis or ter, then a street type from a list (rue, avenue, allée, impasse, chemin, lieu-dit and about twenty others), then words starting with a capital letter, possibly joined by de la, du or des. The postcode and city can follow.
These capital letters hide a trap. Presidio applies the case-insensitive option to every pattern, so a rule like “a word starting with a capital letter” would accept any word, and the address would spill over into the rest of the sentence. The fix is to turn that option off inside the pattern, for that piece only, with the (?-i:...) syntax.
The benchmark revealed another case. In the header of a letter, the address takes two lines, with the street on the first and the postcode and city on the second. The pattern only accepted a space between the two, so the postcode and city stayed in clear. It now accepts a line break.
A postcode followed by a city, on its own, can also be an address. But “33000 Bordeaux” can just as well appear in an ordinary sentence about a branch office or a trade fair. So this pattern starts at 0.3, below the threshold, and only gets through after a word such as “domicilié”, “adresse” or “habite” (lives).
Names, the hardest part
Names have neither check digits nor a fixed shape. The spaCy model does well when the sentence helps, as in “je vous écris au sujet de ma mère, Mme Josiane Roussel” (I am writing about my mother, Mrs Josiane Roussel). It misses names on their own, though. That means a name alone on the first line of a letter, a signature, a first name after “Salut”, or a surname after “Docteur”. In the 100 test texts, it let 16 of them through.
I wrote a recognizer that looks for the same clues a human reader would. It knows four situations, each with its own score:
- after a title or a form field (“M.”, “Docteur”, “Maître”, “Nom :”), the next word is very likely a name: 0.85;
- a name alone on its line that starts with a known first name, as in a letter header or a signature: 0.85;
- a known first name followed by another capitalized word, such as “Hortense Rodriguez”: 0.6;
- a known first name on its own: 0.3, below the threshold. It only gets through with a clue word such as “salut”, “bonjour”, “mon fils” (my son) or “cordialement”.
The known first names come from INSEE’s first names file, which lists the first names given in France since 1900. I kept the 5,114 spellings given to at least 500 children. Two are excluded by hand, Paris and France: they are first names too, but in a letter they nearly always mean the capital or the country.
The recognizer does not replace spaCy, it complements it. When both find the same name, Presidio keeps a single result. With this recognizer, the share of names found went from 90.5 % to 98.5 %.
Three stock recognizers to touch up
Some of Presidio’s recognizers nearly did the job and only needed a touch-up.
The phone recognizer relies on phonenumbers, the Python port of Google’s library that validates numbers from all over the world. But by default, it only looks for American, British, German, Israeli, Indian, Canadian or Brazilian numbers. A number written with +33 still gets through, since the country code is enough to place it. A “01 46 39 90 88” written the French way, however, is valid in none of these regions. The setting that adds France does exist, but Presidio’s configuration loader ignores it. Hence a small five-line subclass.
The URL recognizer also flagged the domain of every email address: in marie@example.fr, it saw a URL example.fr on top of the email. I dropped those duplicates.
And there is the Mastercard one, told above.
The numbers
For each type of personal data, the benchmark counts the share found by each of the three versions of Presidio. The difference comes almost entirely from what is French.
| Type of data | Items | Official image (%) | + French model (%) | + French recognizers (%) |
|---|---|---|---|---|
| Person names | 201 | 76.6 | 90.5 | 98.5 |
| Cities and countries | 51 | 9.8 | 96.1 | 96.1 |
| Addresses | 74 | 0 | 0 | 100 |
| Phone numbers | 61 | 88.5 | 91.8 | 100 |
| Emails | 35 | 100 | 100 | 100 |
| IBANs | 30 | 100 | 100 | 100 |
| Social security numbers | 30 | 0 | 0 | 100 |
| Card numbers | 10 | 60 | 60 | 100 |
| Tax numbers | 10 | 0 | 0 | 100 |
| IP addresses | 6 | 100 | 100 | 100 |
These 100 texts are fictional letters, emails, short messages and forms that I generated with their 508 personal data items already marked, so that every detection can be checked. The figure that really matters is the share of items with no character left reaching the model, whatever category was picked. An address taken for a city stays hidden. That is the figure that goes from 59.1 % to 99.2 %.
These recognizers were not this good from the start. Before the fixes for names on their own, two-line addresses and cards starting with 2, the rate topped out at 93.9 %. Fixing rules by looking at the errors of a test set carries a known risk: tuning them to those texts and those texts only. So I generated a second set with another random seed and never looked at it during the fixes. It went from 94.1 % to 99.4 %, which shows that the rules hold on names and numbers they had never seen.
Precision is 88.1 %. In other words, about one detection in eight is not personal data: rue Jean Moulin taken for a person, or a city that points to no one in particular. For this job, it is the right trade-off. A word masked by mistake costs the model a little context, while a missed value goes straight to it.
It all stays light. Analysis takes 9 milliseconds per text at the median, and the container uses 1.6 GB of memory with the French and English models loaded. In the first article, I compared Presidio with LLMs asked to find the same data. The best one, with 7 billion parameters, let one item in six through and took more than 8 seconds per text.
What still gets through
Four items out of 508 stay in clear in the benchmark, and other texts will surely bring out more. Here is what I know about them.
A first name on its own, with no clue. “Tu peux le déposer chez Étienne ?”: the first name is in the list, but nothing around it announces it, so its score stays at 0.3. Lowering the threshold would catch this case, but also every word that is both a first name and a common word.
A first name missing from the list. Alexandrie was not given to 500 children. In “Alexandrie Toussaint”, spaCy saw a city (Alexandria) in the first name and masked it as a place, so the surname stayed visible. The gateway’s leak test caught it, because it searches for every word separately.
A city without context. Paris and Caen, mentioned in passing, slipped past spaCy, and no recognizer looks for a city on its own. They rarely identify anyone by themselves, but they still count as leaks.
Clue words that match too much. Presidio matches clue words as parts of words, and it includes the first word of the span itself in its search. The word “ami” (friend) is inside “Camille”, so every first name containing “ami” gets the context bonus automatically. In the same way, the English clue “phone” fires inside “Téléphone”. Here, it only reinforces correct detections. But a recognizer with clue words that are too short can start confirming anything.
Your turn
You do not need the gateway to try all this. The Analyzer is a separate container with an HTTP API.
- Start the Analyzer. From the repo, copy
.env.exampleto.env, which Docker Compose reads at startup, then rundocker compose up -d presidio-analyzer. The image starts from the official one, adds the French model and the recognizers, and answers on port 5002. - Analyze a text. Send the text to
/analyzewith"language": "fr". Thereturn_decision_processoption adds Presidio’s explanation to each result. - Add your own recognizer. To try a pattern, pass it directly in the request. To keep it, add it to
config/presidio/recognizers.yaml, or write a Python class like those inpii_recognizers/when a computation is needed.
Here is the explanation for a NIR. The pattern alone gave 0.3, the word “sécurité” was recognized as a clue, and the key is right, hence the final score of 1. With a wrong key, the same request returns an empty list.
curl -s localhost:5002/analyze \
-H 'content-type: application/json' \
-d '{"text": "Mon numéro de sécurité sociale est le 2 52 07 33 063 112 03.", "language": "fr", "return_decision_process": true}' \
| jq '.[] | {entity_type, score, recognizer: .analysis_explanation.recognizer, original_score: .analysis_explanation.original_score, context: .analysis_explanation.supportive_context_word, checksum: .analysis_explanation.validation_result}' {
"entity_type": "FR_NIR",
"score": 1.0,
"recognizer": "FrNirRecognizer",
"original_score": 0.3,
"context": "sécurité",
"checksum": true
} Every company also has identifiers of its own, such as a customer number or a case number, that nobody else knows. To try one, a recognizer can be passed directly in the request, with its pattern, its starting score and its clue words. Here, the customer number starts at 0.3 and goes up to 0.65 thanks to the words “dossier” (case) and “client” (customer).
curl -s localhost:5002/analyze \
-H 'content-type: application/json' \
-d '{"text": "Votre dossier CLI-2026-00042 est clos.", "language": "fr", "score_threshold": 0.4,
"ad_hoc_recognizers": [{"name": "Numéro client", "supported_language": "fr", "supported_entity": "CUSTOMER_ID",
"patterns": [{"name": "numéro client", "regex": "\\bCLI-20\\d{2}-\\d{5}\\b", "score": 0.3}],
"context": ["dossier", "client"]}]}' \
| jq -c '.[] | {entity_type, start, end, score}' {"entity_type":"CUSTOMER_ID","start":14,"end":28,"score":0.6499999999999999} The same recognizer, written in config/presidio/recognizers.yaml with type: custom, applies to every request after a simple restart of the container. Python recognizers, on the other hand, need the image to be rebuilt.
To wrap up
Presidio is not a magic tool. It is a framework, with a good language model and a clear way of combining rules. For French, nearly all the work is writing those rules: patterns for the shape of the data, check digits where they exist, and clue words for everything else. With seven recognizers and about 300 lines of Python, the share of masked data goes from 59 % to 99 %, in 9 milliseconds per text and without a single call to a generative model.
The recognizers’ code, their tests and both benchmarks are on GitHub. The just presidio-steps command redoes this article’s measurements in about a minute.