Home Knowledge Base Protected Health Information (PHI) Detection

Protected Health Information (PHI) Detection is the specialized clinical NLP task of automatically identifying all 18 HIPAA-defined categories of personally identifiable health information in clinical text — enabling automated de-identification pipelines that make patient data available for research, AI training, and analytics while maintaining regulatory compliance with federal healthcare privacy law.

What Is PHI Detection?

PHI Detection vs. General NER

Standard NER (person, location, organization) is insufficient for PHI detection:

The i2b2 2014 De-Identification Gold Standard

The i2b2 2014 shared task is the definitive clinical PHI benchmark:

System Architectures

Rule-Based with Regex:

CRF + Clinical Lexicons:

BioBERT / ClinicalBERT NER:

Ensemble + Post-Processing:

Performance Results (i2b2 2014)

PHI CategoryBest RecallBest Precision
NAME98.9%97.4%
DATE99.8%99.5%
ID (MRN/SSN)99.2%98.7%
LOCATION97.6%95.3%
AGE (>89)96.1%93.8%
CONTACT98.4%97.1%
PROFESSION84.7%79.2%

Why PHI Detection Matters

PHI Detection is the privacy protection layer of clinical AI — the prerequisite NLP capability that makes all other healthcare AI innovation legally permissible by ensuring that patient-identifying information is identified, tracked, and appropriately protected before clinical text enters any data processing pipeline.

protected health information detectionphihealthcare ai

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.