You Are What You Write: Author re-identification privacy attacks in the era of pre-trained language models
Richard Plant, Mario Valerio Giuffrida, Dimitra Gkatzia
Computer Speech & Language (2025)
Richard Plant, Valerio Giuffrida, Dimitra Gkatzia, (2025). You Are What You Write: Author re-identification privacy attacks in the era of pre-trained language models. Computer Speech & Language, Volume 90, https://doi.org/10.1016/j.csl.2024.101746.
@article{PLANT2025101746,
title = {You Are What You Write: Author re-identification privacy attacks in the era of pre-trained language models},
journal = {Computer Speech & Language},
volume = {90},
pages = {101746},
year = {2025},
issn = {0885-2308},
doi = {https://doi.org/10.1016/j.csl.2024.101746},
url = {https://www.sciencedirect.com/science/article/pii/S0885230824001293},
author = {Richard Plant and Valerio Giuffrida and Dimitra Gkatzia},
keywords = {Language models, Privacy-preserving, Differential privacy, Adversarial learning, Re-identification attacks},
abstract = {The widespread use of pre-trained language models has revolutionised knowledge transfer in natural language processing tasks. However, there is a concern regarding potential breaches of user trust due to the risk of re-identification attacks, where malicious users could extract Personally Identifiable Information (PII) from other datasets. To assess the extent of extractable personal information on popular pre-trained models, we conduct the first wide coverage evaluation and comparison of state-of-the-art privacy-preserving algorithms on a large multi-lingual dataset for sentiment analysis annotated with demographic information (including location, age, and gender). Our results suggest a link between model complexity, pre-training data volume, and the efficacy of privacy-preserving embeddings. We found that privacy-preserving methods demonstrate greater effectiveness when applied to larger and more complex models, with improvements exceeding >20% over non-private baselines. Additionally, we observe that local differential privacy imposes serious performance penalties of ≈20% in our test setting, which can be mitigated using hybrid or metric-DP techniques.}
}
Abstract
The widespread use of pre-trained language models has revolutionised knowledge transfer in natural language processing tasks. However, there is a concern regarding potential breaches of user trust due to the risk of re-identification attacks, where malicious users could extract Personally Identifiable Information (PII) from other datasets. To assess the extent of extractable personal information on popular pre-trained models, we conduct the first wide coverage evaluation and comparison of state-of-the-art privacy-preserving algorithms on a large multi-lingual dataset for sentiment analysis annotated with demographic information (including location, age, and gender). Our results suggest a link between model complexity, pre-training data volume, and the efficacy of privacy-preserving embeddings. We found that privacy-preserving methods demonstrate greater effectiveness when applied to larger and more complex models, with improvements exceeding >20% over non-private baselines. Additionally, we observe that local differential privacy imposes serious performance penalties of ~20% in our test setting, which can be mitigated using hybrid or metric-DP techniques.
Manage Cookie Consent
We use cookies to optimise our website and our service.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.