Richard Plant, Mario Valerio Giuffrida, Dimitra Gkatzia

Computer Speech & Language (2025)

Richard Plant, Valerio Giuffrida, Dimitra Gkatzia, (2025). You Are What You Write: Author re-identification privacy attacks in the era of pre-trained language models. Computer Speech & Language, Volume 90, https://doi.org/10.1016/j.csl.2024.101746.

wp-content/uploads/2020/10/tex.png
Get Paper
@article{PLANT2025101746,
title = {You Are What You Write: Author re-identification privacy attacks in the era of pre-trained language models},
journal = {Computer Speech & Language},
volume = {90},
pages = {101746},
year = {2025},
issn = {0885-2308},
doi = {https://doi.org/10.1016/j.csl.2024.101746},
url = {https://www.sciencedirect.com/science/article/pii/S0885230824001293},
author = {Richard Plant and Valerio Giuffrida and Dimitra Gkatzia},
keywords = {Language models, Privacy-preserving, Differential privacy, Adversarial learning, Re-identification attacks},
abstract = {The widespread use of pre-trained language models has revolutionised knowledge transfer in natural language processing tasks. However, there is a concern regarding potential breaches of user trust due to the risk of re-identification attacks, where malicious users could extract Personally Identifiable Information (PII) from other datasets. To assess the extent of extractable personal information on popular pre-trained models, we conduct the first wide coverage evaluation and comparison of state-of-the-art privacy-preserving algorithms on a large multi-lingual dataset for sentiment analysis annotated with demographic information (including location, age, and gender). Our results suggest a link between model complexity, pre-training data volume, and the efficacy of privacy-preserving embeddings. We found that privacy-preserving methods demonstrate greater effectiveness when applied to larger and more complex models, with improvements exceeding >20% over non-private baselines. Additionally, we observe that local differential privacy imposes serious performance penalties of ≈20% in our test setting, which can be mitigated using hybrid or metric-DP techniques.}
}


Abstract

The widespread use of pre-trained language models has revolutionised knowledge transfer in natural language processing tasks. However, there is a concern regarding potential breaches of user trust due to the risk of re-identification attacks, where malicious users could extract Personally Identifiable Information (PII) from other datasets. To assess the extent of extractable personal information on popular pre-trained models, we conduct the first wide coverage evaluation and comparison of state-of-the-art privacy-preserving algorithms on a large multi-lingual dataset for sentiment analysis annotated with demographic information (including location, age, and gender). Our results suggest a link between model complexity, pre-training data volume, and the efficacy of privacy-preserving embeddings. We found that privacy-preserving methods demonstrate greater effectiveness when applied to larger and more complex models, with improvements exceeding >20%
over non-private baselines. Additionally, we observe that local differential privacy imposes serious performance penalties of ~20%
in our test setting, which can be mitigated using hybrid or metric-DP techniques.