IP Library Granted Patent US 12,468,852
Granted Patent B2
US 12,468,852 · App. 17/653,817 · Granted Nov 11, 2025

Data anonymity protector

Inventors: Martin Gleize (Dublin, IE); Pierpaolo Tommasi (Dublin, IE); Yufang Hou (Dublin, IE); Debasis Ganguly (Dublin, IE)
Assignee: International Business Machines Corporation
G06F21/6263G06F21/31G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,852
App. No.
17/653,817
Granted
Nov 11, 2025
Kind
B2
Abstract

Embodiments for providing enhanced data anonymity protection by a processor are disclosed. Selected portions of data intended for distribution in a communication channel or currently distributed on one or more data sources having a potential for revealing identify of a user in a public domain (and/or private domain) may be identified, where an assessment is provided indicating a current status of an amount of data currently exposing the identity of the user in the public domain. The selected portions of the data may be transformed into anonymous data by applying a one or more data corrective operations to prevent further exposure of the identity of the user into the public domain.

Claims (46)

1 . A computer-implemented method comprising:

receiving, by one or more processors, credentials to access a user account for a publicly available social media platform;

parsing, by one or more processors, the user account for the publicly available social media platform, the parsing comprising:

using, by one or more processors, the credentials to access the user account for the publicly available social media platform; and

extracting, by one or more processors, information from the publicly available social media platform using the user account;

responsive to parsing the user account, identifying, by one or more processors, selected portions of data of the information previously posted to the publicly available social media platform as having a potential for revealing an identity of a user;

providing, by one or more processors, an assessment indicating a current status of an amount of data currently exposing the identity of the user by analyzing user profile information posted to the publicly available social media platform; and

transforming, by one or more processors, the selected portions of the data into anonymous data by, subsequent to accessing the user account using the credentials, editing and applying a data corrective operation to the selected portions of the data previously posted to the publicly available social media platform that reduces further exposure of the identity of the user, wherein transforming the selected portions of data into anonymous data comprises modifying the data such that the anonymous data is a more generic version of the data, the modifying the data comprising swapping user country origin with ability to speak a language typically associated with the country.

2 . The computer-implemented method of claim 1 , further comprising generating, by one or more processors, a report of the data indicating a potential for revealing the identity of the user based on the selected portions of data.

3 . The computer-implemented method of claim 1 , further comprising indicating, by one or more processors, the selected portions of the data are intended for distribution in a communication channel and expose the identity of the user.

4 . The computer-implemented method of claim 1 , further comprising identifying, by one or more processors, the selected portion of the data reveals the identity of the user based on using a machine learning operation.

5 . The computer-implemented method of claim 1 , further comprising initiating, by one or more processors, a machine learning operation to identify and learn: (i) all data, data attributes, and sensitive data associated with the user from prior communications to one or more communication channels, (ii) all data intended for distribution in the communication channel, and (iii) all data currently distributed on one or more data sources that exposes the identity of the user.

6 . The computer-implemented method of claim 1 , further comprising training, by one or more processors, one or more machine learning models to identify and transform the selected portions of data into the anonymous data according to a user profile.

7 . A computer system comprising:

a processor set;

one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media to cause the processor set to perform operations comprising:

receiving credentials to access a user account for a publicly available social media platform;

parsing the user account for the publicly available social media platform, the parsing comprising:

using the credentials to access the user account for the publicly available social media platform; and

extracting information from the publicly available social media platform using the user account;

responsive to parsing the user account, identifying selected portions of data of the information previously posted to the publicly available social media platform as having a potential for revealing an identity of a user;

providing an assessment indicating a current status of an amount of data currently exposing the identity of the user by analyzing user profile information posted to the publicly available social media platform; and

transforming the selected portions of the data into anonymous data by, subsequent to accessing the user account using the credentials, editing and applying a data corrective operation to the selected portions of the data previously posted to the publicly available social media platform that reduces further exposure of the identity of the user, wherein transforming the selected portions of data into anonymous data comprises modifying the data such that the anonymous data is a more generic version of the data, the modifying the data comprising swapping user country origin with ability to speak a language typically associated with the country.

8 . The computer system of claim 7 , wherein the operations further comprise generating a report of the data indicating a potential for revealing the identity of the user based on the selected portions of data.

9 . The computer system of claim 7 , wherein the operations further comprise indicating the selected portions of the data are intended for distribution in a communication channel and expose the identity of the user.

10 . The computer system of claim 7 , wherein the operations further comprise identifying the selected portion of the data reveals the identity of the user based on using a machine learning operation.

11 . The computer system of claim 7 , wherein the operations further comprise initiating a machine learning operation to identify and learn: (i) all data, data attributes, and sensitive data associated with the user from prior communications to one or more communication channels, (ii) all data intended for distribution in the communication channel, and (iii) all data currently distributed on one or more data sources that exposes the identity of the user.

12 . The computer system of claim 7 , wherein the operations further comprise training one or more machine learning models to identify and transform the selected portions of data into the anonymous data according to a user profile.

13 . A computer program product comprising:

one or more computer readable hardware storage media;

and program instructions stored on the one or more computer readable storage media to perform operations comprising:

receiving credentials to access a user account for a publicly available social media platform;

parsing the user account for the publicly available social media platform, the parsing comprising:

using the credentials to access the user account for the publicly available social media platform; and

extracting information from the publicly available social media platform using the user account:

responsive to parsing the user account, identifying selected portions of data of the information previously posted to the publicly available social media platform as having a potential for revealing an identity of a user;

providing an assessment indicating a current status of an amount of data currently exposing the identity of the user by analyzing user profile information posted to the publicly available social media platform; and

transforming the selected portions of the data into anonymous data by, subsequent to accessing the user account using the credentials, editing and applying a data corrective operation to the selected portions of the data previously posted to the publicly available social media platform that reduces further exposure of the identity of the user, wherein transforming the selected portions of data into anonymous data comprises modifying the data such that the anonymous data is a more generic version of the data, the modifying the data comprising swapping user country origin with ability to speak a language typically associated with the country.

14 . The computer program product of claim 13 , wherein the operations further comprise generating a report of the data indicating a potential for revealing the identity of the user based on the selected portions of data.

15 . The computer program product of claim 13 , wherein the operations further comprise indicating the selected portions of the data are intended for distribution in a communication channel and expose the identity of the user.

16 . The computer program product of claim 13 , wherein the operations further comprise identifying the selected portion of the data reveals the identity of the user based on using a machine learning operation.

17 . The computer program product of claim 13 , wherein the operations further comprise:

initiating a machine learning operation to identify and learn: (i) all data, data attributes, and sensitive data associated with the user from prior communications to one or more communication channels, (ii) all data intended for distribution in the communication channel, and (iii) all data currently distributed on one or more data sources that exposes the identity of the user; and

training one or more machine learning models to identify and transform the selected portions of data into the anonymous data according to a user profile.

18 . The computer-implemented method of claim 1 , wherein the data is currently distributed on the publicly available platform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2022
From: GLEIZE, MARTIN; TOMMASI, PIERPAOLO; HOU, YUFANG; GANGULY, DEBASIS
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059188/0806 →
Continuity (1)
Related Publication 20230281340A1 · Sep 7, 2023
References Cited (19)
US 8166313B2 · Fedtke · 2012 [cited by applicant]
US 8649552B2 · Balakrishnan · 2014 [cited by applicant]
US 10026021B2 · Stoop et al. · 2018 [cited by applicant]
US 10140321B2 · Hakkani-Tur · 2018 [cited by applicant]
US 10387511B2 · Evnine · 2019 [cited by applicant]
US 20070266079A1 · Criddle · 2007 [cited by examiner]
US 20100318489A1 · De Barros · 2010 [cited by examiner]
US 20150213288A1 · Bilodeau · 2015 [cited by examiner]
US 20170243162A1 · Gavrielides · 2017 [cited by examiner]
US 20170262477A1 · Carroll · 2017 [cited by examiner]
US 20190222602A1 · Linder · 2019 [cited by examiner]
Amreen, Sadika, “Methods of Disambiguating and De-anonymizing Authorship in Large Scale Operational Data”, PhD dissertation, University of Tennessee, 2019, https://trace.tennessee.edu/utk_graddiss/5453 (130 pages). [cited by applicant]
Umair, Amber, “Machine Learning based Information Forensics from Smart Sources”, Thesis, University of Technology Sydney, Australia, 2007, (118 pages). [cited by applicant]
Metcalf et al., “Where are human subjects in Big Data research? The emerging ethics divide”, Research article, Big Data & Society, May 2016, DOI: 10.1177/2053951716650211, (14 pages). [cited by applicant]
Chen, Zhipeng, “Inferencing User Characteristics and Behaviors from Social Media with Applications in Regulatory Science”, Electronic Dissertation, The University of Arizona, Jul. 2021, http://hdl.handle.net/10150/64578… [cited by applicant]
Breidbach et al., “Accountable algorithms? The ethical implications of data-driven business models”, Journal of Service Management, vol. 31, pp. 163-185, May 2020, DOI:10.1108/JOSM-03-2019-0073, (23 pages). [cited by applicant]
Garcia-Pablos, et. al., “Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT”, https://www.aclweb.org/anthology/2020.lrec-1.552.pdf, Proceedings of the 12th Conference on Language… [cited by applicant]
Ghazinour, et. al., “Detecting Health-Related Privacy Leaks in Social Networks Using Text Mining Tools”, https://link.springer.com/chapter/10.1007/978-3-642-38457-8_3, O. Zaïane and S. Zilles (Eds.): Canadian AI 2013, L… [cited by applicant]
Xu, et. al., “Personal Information Leakage Detection in Conversations”, https://aclanthology.org/2020.emnlp-main.532/, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pp. 6567-658… [cited by applicant]