Businesses cannot book a trip, process a payment or treat a patient without collecting personal information. A traveler emails a passport scan to confirm a tour. A patient types an insurance ID into an intake form. A shopper saves a card for the next order.
Each of these routine steps puts personally identifiable information into a company's systems. This guide explains what is PII, how it is classified, where it hides inside a business, the laws that govern it and the practical steps that keep it safe.
What is PII (personally identifiable information)?
PII (personally identifiable information) is any data that can identify a specific person, alone or combined with other details, such as a name, Social Security number, email or card number.
Personally identifiable information is collected every time someone opens an account, buys a product, fills out a form or visits a doctor. It covers obvious identifiers, such as a passport number, and less obvious ones, such as the IP address of the device used to place an order.
The US National Institute of Standards and Technology (NIST SP 800-122) splits PII into two groups:
- Distinguishing information: Data that tells one individual apart from everyone else, such as a Social Security number or fingerprint.
- Traceable information: Data that links activity back to one person, such as login history or location data.
Consider a single online order. The name, shipping address, email, card number and IP address all land in one record, and together they show who the customer is, where they live and how they pay. GDPR calls this "personal data," while California's CCPA uses "personal information," but the core question is the same: can this data be linked to a real person?
What are the main categories of PII? Direct vs. indirect identifiers
Direct identifiers are data that identify a person on their own, such as a passport number, whereas indirect identifiers identify a person only when combined with other data, such as a ZIP code.
Direct identifiers (linked data)
Direct identifiers point to one individual without any extra context. If a support agent sees a Social Security number in a ticket, there is no doubt whose record it belongs to. In most businesses, they show up as:
- Full name, in most contexts
- Social Security number or national ID number
- Passport and driver's license numbers
- Biometric data, such as fingerprints or face scans
- Personal email address and mobile number
Indirect identifiers (linkable data)
Indirect identifiers, also called quasi-identifiers, describe a person without naming them. A ZIP code, job title, date of birth, gender, IP address or cookie ID could each belong to thousands of people. Inside a customer database, though, they sit side by side in the same row.
How indirect identifiers become direct
Put a few of them together and the picture sharpens quickly. "Female, 34, pediatric nurse, lives in a town of 8,000" may describe only one person, even though no single detail names her. This is why removing names from a dataset does not make it anonymous.
Sensitive vs. non-sensitive PII: what's the difference?
Sensitive PII is data that can cause serious harm if exposed, such as an SSN, card number or medical record, whereas non-sensitive PII, such as a name or ZIP code, carries low risk on its own.
Sensitive PII
Sensitive PII could lead to identity theft, financial loss or discrimination if it reaches the wrong person. Social Security and passport numbers, card and bank account numbers, medical records and biometrics all belong here. A criminal with a stolen SSN and date of birth can open credit lines in someone else's name, so this data needs encryption or tokenization and strict access rules.
Non-sensitive PII
A first name, work email or business phone number rarely harms anyone on its own. Much of it already appears on company websites and social profiles, but attackers still use it to write convincing phishing emails.
When non-sensitive PII becomes sensitive
Context changes the risk. Names in a newsletter list are low-risk, yet the same names on a clinic's appointment list reveal who is receiving treatment. A phone number on a business card is harmless, but the same number used for two-factor codes becomes a target for SIM-swap fraud.
PII risk matrix: how identifiability and sensitivity work together
| Group | Examples | Protection level |
|---|---|---|
| Direct and sensitive | SSN, passport number, fingerprints | Highest: tokenize or encrypt, need-to-know access |
| Indirect and sensitive | Diagnosis, salary, precise location | High: encrypt and store apart from identifiers |
| Direct and non-sensitive | Full name, personal email | Moderate: role-based access |
| Indirect and non-sensitive | ZIP code, job title | Baseline: minimize, watch combinations |
A record takes the risk level of its most sensitive field, so a patient file holding a name, ZIP code and diagnosis needs top-tier protection.
What are common examples of PII?
PII examples range from full names and home addresses to Social Security numbers, passport numbers, card numbers, medical record numbers, fingerprints, IP addresses and online login details.
PII shows up in almost every system a business runs, from the checkout page to the HR portal.
| Category | PII examples |
|---|---|
| Identity | Full name, date of birth, photo, signature |
| Government IDs | Social Security number, passport, driver's license, tax ID |
| Financial | Card number (PAN), bank account and routing number |
| Contact | Home address, personal email, mobile number |
| Health | Medical record number, diagnosis, health insurance ID |
| Biometric | Fingerprints, face scans, voiceprints |
| Digital | IP address, device ID, cookie IDs, precise geolocation |
| Employment | Employee ID, salary, background check results |
These fields pile up quickly. One travel booking can hold a traveler's name, date of birth, passport number, address, phone number and payment card. Onboarding a new employee adds an SSN, payroll bank details and a driver's license scan, and that file then feeds payroll, benefits and IT systems.
What is not PII? Common misconceptions
Non-PII is data that cannot be used on its own to identify, contact or trace a specific person, such as aggregated statistics or anonymized records, while pseudonymous data still counts as PII.
When it is fully detached from any one person's record, the following usually falls outside PII:
- Aggregated data: Totals and trends, such as "500 visitors viewed this page today"
- Anonymized records: Data with every identifier removed and no way to restore it
- Generic device data: Browser type or operating system not tied to a user profile
- Business data: A company name, head office address or main phone line
The confusion starts with data that only looks anonymous:
| Misconception | Reality |
|---|---|
| "No name means no PII." | IP addresses, device IDs and location history can be traced to a person. |
| "ZIP code, birth date and gender are safe." | Together they can single out one individual. |
| "Public data isn't PII." | Names in public profiles are still protected under GDPR. |
| "Tokenized data isn't PII." | The vault can re-identify it, so it is pseudonymous personal data. |
PII vs. personal data vs. PHI: what's the difference?
PII is data that identifies a person under US frameworks, while personal data is GDPR's broader term, whereas PHI is identifiable health data held by HIPAA-covered entities and their partners.
These terms overlap, and teams often use them interchangeably. The difference matters, because it decides which rules apply to a record.
| Aspect | PII | Personal data | PHI |
|---|---|---|---|
| Main source | NIST, US federal and state laws | GDPR, UK GDPR | HIPAA |
| Scope | Data that identifies or links to a person | Any information about an identifiable person | Identifiable health data held by covered entities or business associates |
| Online identifiers | Depends on context | Yes, including cookie and device IDs | Yes, when linked to health data |
| Example | Name and SSN in a bank record | Cookie ID tied to a user profile | Diagnosis in a hospital record |
PHI is not simply "more sensitive PII." A name in a travel CRM is PII but not PHI. The same name on a treatment record at a covered clinic is PHI, while a medical note in an employer's HR file often is not, because HIPAA excludes employment records.
Where does PII live inside a business?
Data sprawl is the process by which PII spreads from one database into CRMs, inboxes, support tickets, shared drives, spreadsheets, analytics tools, logs, backups, vendor systems and AI prompts.
PII rarely stays where it was first collected. A customer's details might start in a CRM, get copied into a data warehouse for reporting, move into a test environment for developers and appear in a support ticket when the customer calls.
Everyday shortcuts add even more copies:
| Everyday shortcut | Where the PII ends up |
|---|---|
| "Email us your passport" | Mailboxes, forwards, phones, archives |
| Send it on WhatsApp | Personal phones, cloud backups, screenshots |
| Paste it into CRM notes | Broad staff access, exports, plugins |
| Paste it into an AI assistant | Third-party processing, logs, retention |
Every copy is another place an attacker can reach, and another place to find when a customer asks to be deleted under GDPR or CCPA.
How to identify and classify PII in your organization
PII classification is a process of discovering where personal data lives, sorting each field into a sensitivity tier and applying security controls that match the risk level of each tier.
A business cannot encrypt, restrict or delete data it does not know it has, which makes classification the starting point for every other control.
Map your data sources
List every system that collects or stores personal data: CRMs, billing and HR platforms, but also PDFs, spreadsheets, email threads, ticketing tools and test environments. Talk to each department, because shadow tools often hold more PII than approved systems.
Inventory PII fields
Record which fields each system holds, where they came from, where they are copied and who can read them. Automated discovery tools help find PII buried in free text, such as a card number typed into a support chat.
Classify by sensitivity
| Tier | Examples | Minimum controls |
|---|---|---|
| Restricted | SSNs, card numbers, medical records | Tokenization or encryption, MFA |
| Confidential | Home address, date of birth, IP address | Encryption, role-based access |
| Internal | Job titles, work email lists | Standard access control |
| Public | Company address, press releases | Approval before publishing |
Pay attention to combinations. A birth date, ZIP code and gender may each be Confidential, but stored together they carry Restricted-level risk.
Assign owners and retention rules
Every dataset needs a named owner who approves access and a deletion date tied to its purpose.
Keep labels and the inventory current
Tag columns and files with their tier so the label travels with the data, and revisit the inventory whenever a new system or vendor is added.
Which industries handle the most sensitive PII?
The industries handling the most sensitive PII are healthcare, financial services, insurance, legal, travel and HR, which collect medical records, account numbers, government IDs and passports.
Attackers target these sectors most. In the Identity Theft Resource Center's 2025 Annual Data Breach Report, financial services recorded the most compromises (739), followed by healthcare (534) and professional services (478).
| Industry | Sensitive PII they hold | Key rules |
|---|---|---|
| Healthcare | Medical records, insurance IDs | HIPAA |
| Financial services | Account numbers, SSNs, KYC documents | GLBA, PCI DSS |
| Insurance | Claims, medical and driving history | GLBA, state laws |
| Legal | Case files, IDs, financial statements | Client confidentiality |
| Travel and hospitality | Passports, payment cards | PCI DSS, GDPR |
| HR and staffing | SSNs, payroll, background checks | FCRA, state laws |
How these industries collect data adds to the risk. Law firms receive IDs before a client is formally engaged, and travel agencies take card and passport details over email or phone to book suppliers.
Why protecting PII matters: risks, costs, and real breaches
Protecting PII is essential because breaches are costly and common. IBM's 2026 report puts the average breach at $4.99 million globally and $11.5 million in the US, before fines and lost trust.
According to the Identity Theft Resource Center, US organizations reported 3,322 data compromises in 2025, the highest number on record, resulting in nearly 279 million victim notices.
The cost per incident is rising too. IBM's Cost of a Data Breach Report 2026 found that the global average reached $4.99 million, up 12% on the previous year, while the US average rose to $11.5 million. Verizon's 2026 DBIR found third parties involved in 48% of breaches.
Criminals want PII because it turns into money in several ways:
- Identity theft: Opening credit cards or loans in a victim's name
- Account takeover: Using leaked details to reset passwords
- Targeted phishing: Writing emails that reference real names and orders
- Extortion: Threatening to publish stolen records unless a ransom is paid
For the business, the fallout includes fines, lawsuits, notification costs, lost customers and higher insurance premiums.
Who is responsible for protecting PII in an organization?
PII protection is a shared responsibility: executives own the risk, security teams run the controls, privacy officers manage compliance, engineers build safeguards and employees handle data safely.
| Role | What they own |
|---|---|
| Executives and board | Accountability, budget and risk appetite |
| CISO and security team | Encryption, access control, monitoring, incident response |
| Privacy officer or DPO | Policies, impact assessments, data subject requests |
| Engineering and data teams | Minimization, tokenization and secure design |
| Legal and procurement | Vendor contracts and breach notices |
| Every employee | Keeping PII out of email, chat and AI tools |
Vendors that process PII also carry obligations, but outsourcing does not move the accountability. Under GDPR, the business that decides why data is processed (the controller) stays responsible, even when a vendor (the processor) handles it.
What laws and regulations govern PII?
PII regulation is a mix of regional and sector laws, including GDPR in the EU, CCPA/CPRA and HIPAA in the US, India's DPDP Act, China's PIPL, Brazil's LGPD and the PCI DSS standard for card data.
No single law answers what is PII the same way, so businesses with customers in several regions usually follow the strictest rule that applies.
United States
Protection comes from sector laws rather than one federal privacy law:
- HIPAA for health information held by providers, health plans and their business associates
- GLBA for customer financial data
- COPPA for data collected online from children under 13
- CCPA/CPRA for California residents, with rights to access, delete and opt out
States such as Virginia, Colorado and Texas have passed similar laws, and all 50 states require breach notification.
European Union and UK
GDPR and UK GDPR protect any data about an identifiable person. Fines can reach €20 million or 4% of global annual turnover, whichever is higher.
Asia-Pacific
| Country | Law | Main focus |
|---|---|---|
| India | DPDP Act | Consent, purpose limits |
| China | PIPL | Consent, cross-border transfers |
| Japan | APPI | Purpose of use, third-party sharing |
| Singapore | PDPA | Consent, breach notification |
Other regions
Canada's PIPEDA, Brazil's LGPD and South Africa's POPIA follow the same core principles of purpose, security and individual rights.
Industry standards
PCI DSS governs cardholder data, while SOC 2 and ISO/IEC 27001 give security programs a structure auditors recognize.
How to protect PII: best practices for businesses
PII protection is a process of collecting less data, encrypting and tokenizing sensitive fields, limiting access, monitoring use, deleting data on schedule, training staff and checking vendors.
Data minimization
The safest PII is the data a business never collects. A newsletter signup can usually work with a name and email, without a birth date or home address.
Encryption in transit and at rest
- TLS for traffic between apps, APIs and services
- AES-256 for databases, files and backups
- A separate key management system, with keys rotated on schedule
Tokenization
Replace SSNs, card numbers and bank details with tokens and keep the originals in a secured vault, so most systems never hold the real value.
Access control, least privilege and MFA
A support agent may only need the last four digits of a card, while billing staff need more. Grant access by role, require MFA for sensitive data and review permissions every quarter.
Monitoring and data loss prevention
Log every access to sensitive records and alert on behavior that looks wrong, such as one account exporting thousands of records. DLP tools can also block PII from leaving through email, cloud storage or AI prompts.
Audits and privacy impact assessments
Test controls regularly. Before launching a new product or vendor, complete a privacy impact assessment (a DPIA under GDPR).
Secure retention and deletion
Give every data type an end date, then delete it from live systems, backups and archives.
Employee training and limited sharing
Most leaks start with an ordinary mistake, such as a passport scan forwarded to the wrong address. Short, regular training on phishing and safe sharing does more than an annual slideshow.
Vendor management
Before any vendor receives PII, ask:
- Which certifications do they hold, such as SOC 2?
- How quickly will they report a breach?
- Can they delete your data on request?
How to protect PII across its lifecycle
The PII lifecycle is the path personal data follows from collection to storage, access, sharing, use and deletion, and each stage needs its own security control to keep the data protected.
Collect only what you need
Collection is where most sprawl begins. Ask only for required fields, record consent where the law requires it and capture sensitive data through secure intake forms instead of email attachments or PDFs.
Store it securely
Tokenize or encrypt sensitive fields so the main database never holds raw SSNs or card numbers. Backups need the same protection, because an unencrypted copy on a shared drive undoes the rest.
Limit access
Picture a customer calling about a billing issue. The agent sees a card ending in 4242, which is enough to help. If a supervisor needs the full number, the reveal is logged with their name and the time.
Share it safely
| Instead of | Use |
|---|---|
| Email attachments | Expiring, password-protected links |
| Copying full records into other tools | Tokens that point to the secure record |
| Informal transfers to vendors | Encrypted transfers under a data processing agreement |
Keep it out of logs and AI prompts
Developers log full requests while debugging, and staff paste customer details into AI tools. Tokenizing values first keeps both useful:
Before: payment failed for jane.doe@example.com
After: payment failed for [EMAIL: 7hQ2x]
Delete it on schedule
- Automate deletion when retention periods end.
- Respond to deletion requests within legal deadlines.
- Remove data from backups and vendor systems, and keep a record that it happened.
How tokenization and data privacy vaults protect PII
Tokenization is a process that replaces sensitive values with random tokens and keeps the originals in a secure vault, whereas encryption scrambles data but leaves it inside your own systems.
Encrypted PII still sits in the application database, and anyone who gets both the data and the key can read it. Tokenization moves the real value into a separate data privacy vault instead.
- A customer enters an SSN in a form.
- The vault encrypts and stores it, then returns a token such as tok_7Hq2Lx9.
- Your app, CRM and warehouse store only the token.
- Authorized services request the real value when needed, and every request is logged.
If an attacker breaches the app database or logs, they find tokens instead of real values. Deterministic tokens still allow search and matching, and deleting the original from the vault handles erasure requests in one place. Tokenized data remains pseudonymous, because the vault can map tokens back.
How to protect PII in ChatGPT, Claude, and other LLMs
Protecting PII in LLMs is a process of detecting personal data in prompts, replacing it with tokens before it reaches ChatGPT or Claude, and restoring the real values in the model's response.
Large language models add a new route for PII to leave a business. Support teams send tickets to AI assistants, legal teams summarize contracts and developers connect customer records to LLM APIs. Verizon's 2026 DBIR found that 45% of employees now use AI tools regularly at work.
| Stage | Prompt text |
|---|---|
| Original | "Refund John Smith, email john@acme.com" |
| Sent to the model | "Refund [NAME: Ry0Ixd1], email [EMAIL: m4Tz9]" |
| Returned to the agent | The reply, with the real name and email restored |
For company AI apps and LLM API calls, this step can run automatically before every request. Consumer chat tools need approved enterprise accounts and clear usage rules instead, because API-level controls do not reach a personal account.
What to do after a PII breach: a response checklist
A PII breach response is a process of containing the incident, scoping the exposed data, preserving evidence, notifying regulators and affected individuals on time and fixing the root cause.
- Contain the incident: Revoke compromised credentials and isolate affected systems.
- Scope the exposure: Identify which fields and people were affected.
- Preserve evidence: Keep logs and access records for forensic review.
- Report to regulators: GDPR requires notice within 72 hours; US state deadlines vary.
- Tell the people affected: Explain what happened and what they should do.
- Fix the root cause: Close the gap and update controls and training.
The checklist works best when it is written in advance, with named owners, legal contacts and draft notices tested in a tabletop exercise. Businesses that tokenize sensitive fields often have far less to report, because stolen tokens reveal nothing on their own.
How Enigma protects PII: across the vault and in AI
Enigma is a trust layer between your data and AI that keeps raw PII out of your systems by tokenizing cards, files and customer records, and by tokenizing PII before prompts reach LLM providers.
Enigma keeps the real values in a secure vault and gives your systems tokens instead: Data Vault encrypts any field while keeping it searchable, Card Vault lets you accept payments without storing card numbers, Customer Vault replaces email with secure intake links for IDs, cards and documents, and NoPII tokenizes PII in prompts before they reach LLM providers, all on PCI DSS Level 1 and SOC 2 Type II certified infrastructure.