A claims handler asks an AI assistant to summarize a policyholder's email. A recruiter pastes three CVs into a chatbot to compare candidates. An engineer pipes support tickets into an LLM API to sort them by urgency. All three prompts carry names, phone numbers and account details to a model provider outside the business.
Masking is usually the first fix teams reach for. Hide the personal details, send what is left, and the model never sees who the data is about. It works well in some places and quietly breaks AI output in others. This guide covers how masking works, the main techniques, where it fits, where it falls short and when tokenization is the better choice.
What is PII masking for LLMs and why does it matter?
PII masking for LLMs is a process of hiding personal data in a prompt with placeholders, symbols or fake values before the text reaches a language model, so the provider never sees the real details.
Take the request "Call Maria Lopez on 555-0142 about her refund." After masking, the model receives "Call [PERSON] on [PHONE] about her refund." It can still write the call script. It just never learns who Maria is or how to reach her.
Why it matters comes down to where a prompt travels once it leaves your app:
- The provider: OpenAI, for example, keeps API abuse-monitoring logs for up to 30 days by default, and retention rules differ across providers and plans
- Your own stack: Gateways, observability tools and application logs often store full copies of every prompt
- Other regions: A request may be processed outside the country where the data was collected. By 2027, more than 40% of AI-related data breaches will be caused by improper use of generative AI across borders, per a Gartner forecast
- Regulators: GDPR, HIPAA and PCI DSS all limit where personal data can go and who can see it
Masking itself is not new. Database teams have relied on it for years in two forms:
| Type | How it works | Typical use |
|---|---|---|
| Static masking | Creates a sanitized copy of a dataset | Test, staging and demo environments |
| Dynamic masking | Hides values at read time, based on who is asking | Production systems where different roles see different views |
Masking for LLMs borrows the same idea and applies it to free text, in real time, on every prompt. The catch is that LLMs are not databases. They read meaning from context, and masking takes some of that context away. How much depends on how the masking is done, which the next two sections walk through.
How does PII masking for LLMs work? A step-by-step breakdown
PII masking for LLMs works in four steps: detect personal data in the prompt, replace or hide each value, send the masked prompt to the model and return a response that keeps the masked values.
Detect PII in the prompt
Everything starts with finding the sensitive values, and they arrive from more places than the user's message. In a typical LLM app, personal data comes in through three doors: what the user types, documents pulled in by retrieval (RAG), and results returned by tools an agent calls. Masking has to cover all three, or the gap simply moves somewhere else.
Two detection methods do most of the work:
| Method | What it catches | Example |
|---|---|---|
| Regex and pattern rules | Values with a fixed format | Card numbers, SSNs, emails, phone numbers |
| Named entity recognition (NER) | Values that depend on context | "Priya in accounts", "moved from Leeds to Austin" |
Regex is fast and precise for structured data, but it has no idea that "Jordan" in one sentence is a person and in another is a country. NER models read the surrounding words to make that call. Most production tools run both.
Replace or hide the value
Once a value is found, the masking tool swaps it out. The replacement can take three forms:
- A placeholder such as [PERSON] or [EMAIL]
- Symbols such as ****1234 for a card number
- A fake but realistic value, such as "Anna Berg" in place of "Maria Lopez"
The choice decides how much the model can still understand, and it is the main difference between the techniques covered in the next section.
Send the masked prompt to the LLM
The cleaned prompt goes to OpenAI, Anthropic, Gemini or whichever provider the app uses. From here, anything the provider stores, logs or reviews holds masked text instead of real personal data.
Return the response
The model answers using whatever it received. Masked values stay masked, and because the original value was never stored anywhere, there is no way to put it back:
Prompt sent: Draft a refund email to [PERSON] at [EMAIL].
Model returns: Hi [PERSON], your refund has been approved.
User sees: Hi [PERSON], your refund has been approved.
Someone now has to fill in the name by hand, or the reply goes out with a placeholder in it. For one-way tasks such as analytics, that is fine. For anything customer-facing, it becomes a problem.
Top PII masking techniques for LLMs: how each one works
The top PII masking techniques for LLMs are redaction, character masking, generic placeholders, numbered entity replacement and synthetic data, each trading privacy against the context a model keeps.
The same sentence looks very different under each one. The examples below all start from "Maria Lopez paid with card 4417 2290 1188 5521."
Some guides also count encryption-based or vault-based tokenization as masking. Because those methods can be reversed, this guide covers them separately in the tokenization sections further down.
Redaction
Redaction deletes the value outright or swaps it for a fixed label:
Original: Maria Lopez paid with card 4417 2290 1188 5521
Redacted: [REDACTED] paid with card [REDACTED]
Nothing about the original survives, and nothing can bring it back.
- Pros: The strongest privacy of any masking method, with no way to recover the value, and simple to explain to auditors
- Cons: Data utility drops sharply, since the model loses who did what and a sentence with several blanks can become unreadable
- Best use: Compliance reports, documents shared outside the business and public records requests
Character masking
Character masking hides part of a value with symbols and leaves a few characters visible. It is the familiar card-receipt format.
| Aspect | Detail |
|---|---|
| Example | "M*** L*** paid with card **** **** **** 5521" |
| Variations | First letter only, last four digits, or the email domain kept (m***@mail.com) |
| Pros | People can still recognize the record, and the format stays intact for systems that expect it |
| Cons | Visible characters leak information, and the model gains little from a string of symbols |
| Best use | Screens, receipts and support views where staff confirm a record by its last four digits |
Generic placeholders
Every value of the same type becomes one label: "[PERSON] paid with card [CARD]." The model at least knows a person and a card are involved, which is enough for many classification and summary tasks. It is also the default output of many open-source detection libraries and AI gateways, which is why it shows up so often.
The trouble starts when a prompt mentions two people. "[PERSON] emailed [PERSON] about [PERSON]'s invoice" leaves the model guessing.
Works for: Ticket tagging, topic classification and spam filtering, where one entity per text is typical.
Breaks on: Email threads, case notes and chats, where several people interact.
Numbered entity replacement
This technique adds a counter to each placeholder so different values stay apart within one prompt:
Original: Maria Lopez emailed Daniel Kim about Maria's invoice.
Masked: [PERSON_1] emailed [PERSON_2] about [PERSON_1]'s invoice.
Now the model can follow who is who, and relationships inside a single prompt or session stay intact. Four limits remain:
- Numbering usually resets with each request, so Maria might be [PERSON_1] today and [PERSON_3] tomorrow
- The mapping is often thrown away after masking, which means real names cannot be restored
- Counters reveal how many people a document mentions, which can hint at its contents
- Analysis across large datasets suffers, because the same person gets different numbers in different files
It works best for one-off analysis of a single document, anonymized reports and shared logs where names never need to come back.
Synthetic data replacement
Synthetic replacement swaps each real value for a realistic fake one that keeps its format: "Anna Berg paid with card 4000 0012 3456 7899." The text reads naturally, so both the model and human reviewers handle it as they would real data. Test systems accept the fake card because it still passes format checks, and statistical patterns in the dataset can be kept for analytics and training.
Realism is also the weak point. Generating convincing values for every data type, language and region takes real effort. A fake name that happens to belong to a real customer can cause confusion downstream, and if fake values are derived from real ones with a simple, predictable rule, the patterns can hint at the originals.
Synthetic data shines in test environments, training datasets and review workflows where human testers need readable records and nobody needs the original values again.
Here is how the five techniques compare at a glance:
| Technique | How it works | Example |
|---|---|---|
| Redaction | Deletes the value or replaces it with a fixed label | "Maria Lopez" becomes [REDACTED] |
| Character masking | Hides part of the value with symbols and leaves a few characters visible | 4417 2290 1188 5521 becomes **** **** **** 5521 |
| Generic placeholders | Replaces every value of the same type with one label | "Maria Lopez" and "Daniel Kim" both become [PERSON] |
| Numbered entity replacement | Adds a counter to each label so values stay distinct within one prompt | "Maria Lopez" becomes [PERSON_1] and "Daniel Kim" becomes [PERSON_2] |
| Synthetic data replacement | Swaps the real value for a realistic fake one in the same format | "Maria Lopez" becomes "Anna Berg" |
Pros and cons of PII masking for large language models
PII masking for large language models is simple, fast and cheap to run, but it strips context, cannot restore real values and blurs people together across prompts, which weakens AI answers.
Where masking helps
- Open-source libraries and regex rules cover common values with no new infrastructure
- Pattern checks typically add only milliseconds to each request
- There is no vault or key management to run, which keeps costs low
- Providers and logs see less personal data, which supports data minimization
Limitations
Most of the downsides trace back to one fact: once a value is masked, the original is gone.
| Drawback | What it looks like in practice |
|---|---|
| Context lost | "[PERSON] complained about [PERSON]" gives the model nothing to reason with |
| No way back | Replies come back with placeholders that someone has to fix by hand |
| Identity resets | The same customer looks like a stranger in every new message |
| Detection gaps | A value the detector misses, such as a nickname or an unusual ID format, goes through unmasked |
| Partial leaks | Visible digits, job titles or locations can still point to one person |
For a batch job that counts complaint themes, few of these matter. For a support assistant replying to a named customer, most of them do. The first three explain why masking struggles with LLM accuracy, and the last one raises the question of re-identification. Both get a closer look below, right after the cases where masking is the right call.
Best use cases for PII masking in LLM workflows
The best use cases for PII masking in LLM workflows are training datasets, one-way analytics, log cleanup, test data and external document sharing, where the real values are never needed again.
The test is simple. If nobody downstream will ever need the original name, number or address, masking is usually enough, and five workflows pass that test more often than not.
| Use case | Why masking fits | Technique that suits it | Example |
|---|---|---|---|
| Training and fine-tuning datasets | A model that learns from real names or IDs can repeat them later, and there is no simple, reliable way to make it forget | Synthetic data or redaction | Training a claims classifier on historical claim notes |
| One-way analytics | Trends need no identities | Generic placeholders | Tagging intent across thousands of chats |
| Log and trace cleanup | Debug logs and observability tools tend to keep data for months | Redaction or character masking | An engineer reading a trace next quarter sees the error, not the customer |
| Test and staging data | Developers need records that behave like production without being production | Synthetic data | A fake card number with a valid checksum keeps test suites running |
| Documents shared outside the business | Recipients should hold nothing they could ever reverse | Redaction | Court filings, vendor handoffs and public records requests |
What these cases share is a one-way street. Data goes in, a result comes out, and the result never has to be tied back to a specific person. Live AI workflows rarely work that way, and that is where masking starts to break.
Why traditional PII masking breaks LLM context and accuracy
Traditional PII masking breaks LLM context because it erases who is who in a prompt, cannot restore real values in responses and loses track of people across multi-turn chats and agent workflows.
Picture a support thread between two customers and one agent:
Maria Lopez says the parcel Daniel Kim sent her arrived damaged. Daniel says it left his warehouse intact. Agent Sam Patel has asked both for photos.
With generic placeholders, the model gets "[PERSON] says the parcel [PERSON] sent her arrived damaged. [PERSON] says it left his warehouse intact. Agent [PERSON] has asked both for photos." Ask it who should receive the refund and it has no way to answer.
That single example shows four separate failures:
- Context disappears: Relationships between people, such as sender and receiver, are what the model reasons over. Masking hides exactly that.
- Values blur together: Three different people collapse into one label, so summaries attribute actions to the wrong person.
- Responses cannot be restored: The model's answer comes back full of placeholders, and nothing in the system knows which real name belongs where.
- Conversations and agents lose the thread: In the next message, "What did Maria ask for?" points to nothing, because Maria was never a stable entity to begin with.
Agents make the fourth problem worse. An agent might pull a CRM record, check an order, draft an email and log a ticket in one task, passing data between each step. If every step masks values afresh, the agent cannot match the customer in the CRM to the customer in the order system, and the workflow breaks halfway.
None of this is a bug in any specific tool. It is how masking is designed: once a value is hidden, it is gone for good.
Can masked PII be re-identified? What businesses should know
Masked PII can be re-identified when partial values or leftover context point to one person, whereas redacted data can never be linked back and tokenized data can be reversed only through a vault.
Why redacted data can never be linked back
Redaction deletes the value. There is no key, no mapping and no copy kept anywhere, so the business itself cannot recover it, let alone an attacker.
That is its strength and its limit in one. A redacted support transcript is safe to share, but if a manager later needs to know which customer complained, the answer is not in the data anymore.
How partially masked data still leaks identity
The name is hidden, but the rest of the record tells the story. Consider a masked HR note:
[PERSON], the only female VP of Engineering in the Denver office, joined in March 2021 and is on leave until June.
Anyone at that company knows exactly who this is. Details like these are called quasi-identifiers: fields that seem harmless alone but narrow a group down to one person when combined. Watch for these in any masked text:
- Job title and seniority
- ZIP code or office location
- Dates of birth, hire or admission
- Rare events, such as a specific incident or diagnosis
- The last four digits of a card or phone number
LLMs make this easier, not harder. A model is very good at combining small clues from a long document, so a prompt that looks masked to a regex rule can still describe a person clearly enough to identify them.
How tokenization allows controlled re-identification
Tokenization takes a different approach. Each value becomes a token such as [NAME: Ry0Ixd1], and the real value is kept in a separate, secured vault.
| Who | Can they see the real value? |
|---|---|
| The LLM provider | No, it only receives tokens |
| Logs, gateways and analytics tools | No, they only store tokens |
| An authorized application or user | Yes, through the vault, when the business allows it |
| An attacker who steals a log file | No, the token means nothing without vault access |
Re-identification becomes a decision rather than an accident. The business controls who can reverse a token, for how long the mapping exists and when it is deleted for good. Because a mapping exists, tokenized data is still treated as pseudonymous personal data under GDPR, so the business holding the vault keeps its obligations.
PII tokenization for LLMs: the modern alternative to masking
PII tokenization for LLMs is a process that replaces personal data with consistent tokens, keeps the real values in a secure vault and restores them in responses only for authorized systems.
Where masking hides a value and throws it away, tokenization hides it and keeps a secured record of what it was. The model sees a stand-in, the real value sits in an encrypted vault, and only systems the business authorizes can swap it back. Each token keeps its entity type, and labels that would give the value away are neutralized too:
Before: Update the account for John Smith (SSN: 123-45-6789).
After: Update the account for [NAME: Ry0Ixd1] (ID number: [GOV_ID: k8Lm2n]).
"SSN" becomes "ID number," so the model cannot tell what kind of ID the token hides, yet it still knows a person and an ID are involved.
That one change closes the gaps masking leaves open in live AI workflows:
- Context lost: each person and value gets its own token, so the model can still tell a sender from a receiver
- No way back: tokens in the model's reply are replaced with real values before anyone reads it
- Identity resets: the same value keeps the same token across messages, sessions and agent steps
- Uncontrolled re-identification: only vault access can reverse a token, and mappings can expire or be deleted on demand
Switching an existing app from raw prompts to tokenized ones takes one line of code:
# Before: PII goes straight to the LLM provider
client = OpenAI(api_key="sk-...")
# After: PII is tokenized before it leaves your environment and restored on the way back
client = OpenAI(
api_key="sk-...",
base_url="https://api.nopii.co"
)
Tokenization is the better fit for customer-facing chat, internal copilots and AI agents, where replies must reach the right person by name. Masking remains the simpler choice for one-way work such as analytics, logs and training data.
Masking vs tokenization in action: a sample LLM prompt
Masking vs tokenization in a sample prompt shows that masked text loses who the customer is, while tokenized text keeps each person distinct and still works in later follow-up prompts.
Here is one customer support request followed through the full round trip, first with masking and then with tokenization.
With masking
1. App sends: Customer Maria Lopez (maria.lopez@mail.com) says her husband
Daniel Lopez was charged twice for order 5521. Draft a reply
to Maria confirming the refund.
2. LLM receives: Customer [PERSON] ([EMAIL]) says her husband [PERSON] was
charged twice for order 5521. Draft a reply to [PERSON]
confirming the refund.
3. LLM returns: Hi [PERSON], we're sorry [PERSON] was charged twice for
order 5521. Your refund is on its way.
4. Agent sees: Hi [PERSON], we're sorry [PERSON] was charged twice for
order 5521. Your refund is on its way.
The model had to guess which [PERSON] was the customer, and nothing on the way back knows which name goes where. The agent rewrites the reply by hand, or it goes out with placeholders in it.
With tokenization
1. App sends: Customer Maria Lopez (maria.lopez@mail.com) says her husband
Daniel Lopez was charged twice for order 5521. Draft a reply
to Maria confirming the refund.
2. LLM receives: Customer [NAME: Ry0Ixd1] ([EMAIL: bN3dF5h]) says her husband
[NAME: Pq7Lk2] was charged twice for order 5521. Draft a
reply to [NAME: Ry0Ixd1] confirming the refund.
3. LLM returns: Hi [NAME: Ry0Ixd1], we're sorry [NAME: Pq7Lk2] was charged
twice for order 5521. Your refund is on its way.
4. Agent sees: Hi Maria Lopez, we're sorry Daniel Lopez was charged twice
for order 5521. Your refund is on its way.
Two people get two different tokens, so the model addresses the right one. On the way back, each token is swapped for its real value, and the agent gets a reply ready to send.
The difference grows in the next message. When the agent asks "Has Maria Lopez contacted us before?", tokenization sends "Has [NAME: Ry0Ixd1] contacted us before?" and the model links it to the earlier thread. With masking, the question arrives as "Has [PERSON] contacted us before?" and there is nothing to link it to.
Why deterministic tokens matter for accurate LLM workflows
Deterministic tokens are tokens that always map the same value to the same placeholder, so an LLM can follow each person across prompts and sessions and return answers that restore correctly.
Why the same token has to belong to the same person, every time
Picture an agent preparing a renewal for Daniel Kim. It finds him in the CRM, then in the billing system, then in last month's support emails. If each lookup turned his name into a fresh random token, the model would see three strangers, and the renewal summary would split his history across them.
Deterministic tokenization prevents that. The same input always produces the same token, so Daniel is [NAME: Hx4Rv8] in the CRM record, the invoice and the email thread alike.
Keeping context across prompts and sessions
Consistency pays off most when work spans more than one message:
- Multi-turn chats, where "What did she ask for last week?" must point to the right person
- Case files reviewed over several days by different team members
- Retrieval pipelines, where the same person appears in dozens of stored documents
- Reporting across sessions, such as counting how often one customer has contacted support
In each case, the token acts as a stable ID the model can track without ever seeing the real value.
Restoring real values in LLM responses
Because each token points to exactly one original value, the proxy knows precisely what to put back, even when a reply streams in word by word. The agent's draft leaves the model as "Hi [NAME: Hx4Rv8], your plan renews on 1 November" and reaches the account manager as "Hi Daniel, your plan renews on 1 November."
Per-request placeholders cannot do this reliably. When [PERSON_1] means one customer today and another tomorrow, there is no single answer to which real name belongs in the reply.
PII masking vs tokenization for LLMs: key differences explained
PII masking hides personal data permanently and strips context from LLM prompts, whereas tokenization swaps it for consistent tokens that keep context and can be restored by authorized systems.
| Factor | PII masking | PII tokenization |
|---|---|---|
| Reversibility | No, the original value is discarded | Yes, through the vault, for authorized systems |
| Context kept | Low to medium, depending on the technique | High, each value stays a distinct entity |
| Identity across prompts | Lost, or reset with each request | Kept, the same value gets the same token |
| Real values in responses | Not possible, placeholders remain | Restored automatically before the user sees them |
| Who controls re-identification | Nobody, the value is gone (or anyone who can read partial clues) | The business, through vault access and retention rules |
| Setup effort | Low, often a library or regex rules | Moderate, needs a vault or a managed proxy |
| Best use | Training data, analytics, logs, test data | Chat, support, agents and any workflow that needs real values back |
The setup row is worth a second look. Masking is cheaper to start, but a managed tokenization proxy has narrowed that gap to little more than a configuration change.
When to use PII masking vs tokenization for LLM workflows
Use PII masking when real values are never needed again, such as training data, analytics and logs, while tokenization suits chat, support and agent workflows that must return real values.
The deciding question is whether anyone needs the real value back. If not, mask. If yes, or if the model has to keep people apart across messages, tokenize.
Here is how that plays out in real workflows:
| Workflow | Better choice | Why |
|---|---|---|
| Fine-tuning a support model on last year's tickets | Masking (synthetic data) | The model should learn patterns, never real customers |
| Support copilot drafting replies to live tickets | Tokenization | Replies must address the right customer by name |
| Weekly sentiment report across all conversations | Masking (placeholders) | Only the trend matters, not who said what |
| Agent that updates CRM records and emails customers | Tokenization | The same person must be matched across tools and steps |
| Application logs shared with an outside vendor | Masking (redaction) | The vendor should never see the values, now or later |
| Contract review assistant for a legal team | Tokenization | Party names must stay distinct and return in the summary |
In practice, most teams run both. A common setup tokenizes live prompts so users get complete, accurate answers, then applies redaction or synthetic data when records are exported for analytics, testing or fine-tuning, where nobody needs the real values again.
How tokenization protects PII in ChatGPT and Claude workflows
Tokenization protects PII in ChatGPT and Claude workflows by routing API calls through a proxy that tokenizes personal data on the way out and restores real values in the reply on the way back.
The protection sits at the API layer, where apps built on OpenAI's GPT models or Anthropic's Claude send their requests. A tokenization proxy slots in between the app and the provider through the one-line base_url change shown in the tokenization section above.
From then on, every request makes the same round trip:
- The app sends its prompt as usual, now addressed to the proxy.
- The proxy detects personal data and swaps each value for a token.
- Only the tokenized prompt travels to OpenAI or Anthropic, while the real values stay in the vault.
- Tokens in the reply are replaced with real values, chunk by chunk for streamed responses, before the app shows it.
Retrieved documents and tool results reach the model inside those same API requests, so they pass through the proxy too. That covers all three doors described earlier: user input, RAG context and agent tool outputs.
A well-built proxy also adds safeguards that plain masking lacks:
| Safeguard | What it does |
|---|---|
| Context phrase neutralization | Rewrites labels such as "social security number" to "ID number" so the words around a token do not give away what it is |
| Fail-safe blocking | Stops the request with an error if detection or tokenization fails, instead of sending raw PII |
| Retention control | Deletes each token mapping after a set period, or on demand |
One boundary is worth stating clearly. This protects traffic that flows through the proxy, such as company apps, internal copilots and approved integrations. An employee pasting customer data into a personal ChatGPT or Claude account is outside that path and needs policy, training and an approved alternative.
How Enigma NoPII delivers this
Enigma NoPII is a PII protection layer for LLM APIs that replaces 35+ configurable entity types with deterministic, format-preserving tokens, blocks the request if detection fails and restores real values in every reply, including streamed ones.
It works with OpenAI, Anthropic, Gemini and six other providers, runs on PCI DSS Level 1 and SOC 2 Type II certified infrastructure, and needs only a base URL change to set up. Start free with NoPII and protect your first app in about five minutes.