Home Knowledge Base T-Closeness

T-Closeness is the privacy model requiring that the distribution of sensitive attribute values within each equivalence class be close to the distribution in the overall dataset — measured by a distance metric not exceeding threshold t, addressing both the homogeneity attack (defeated by l-diversity) and the skewness attack where the distribution within a group reveals information even when values are diverse.

What Is T-Closeness?

Why T-Closeness Matters

The Problem T-Closeness Solves

ScenarioGroup DistributionPopulation DistributionPrivacy
L-Diverse but SkewedCancer: 80%, Flu: 10%, Cold: 10%Cancer: 10%, Flu: 45%, Cold: 45%✗ Group membership strongly suggests cancer
T-CloseCancer: 12%, Flu: 43%, Cold: 45%Cancer: 10%, Flu: 45%, Cold: 45%✓ Group is similar to population

How T-Closeness Works

ComponentDescriptionImplementation
Equivalence ClassGroup of records sharing quasi-identifiersGrouped by generalized attributes
Class DistributionFrequency of sensitive values within groupCount sensitive values per group
Overall DistributionFrequency of sensitive values in full datasetCount sensitive values globally
EMD CalculationDistance between class and overall distributionsEarth Mover's Distance computation
Threshold CheckVerify EMD ≤ t for every equivalence classReject or modify groups exceeding t

Earth Mover's Distance (EMD)

Limitations

T-Closeness is the strongest member of the k-anonymity family of privacy models — closing the distributional loopholes in k-anonymity and l-diversity by ensuring that group membership reveals minimal information about sensitive attributes, providing near-population-level uncertainty for any identified individual.

t-closenessprivacy

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.