IP Library Granted Patent US 12,412,087
Granted Patent B2
US 12,412,087 · App. 17/238,567 · Granted Sep 9, 2025

Classifying data from de-identified content

Inventors: Aswin Kannan (Chennai, IN); Balaji Ganesan (Bengaluru, IN); Shanmukha Chaitanya Guttula (Vijayawada, IN)
Assignee: International Business Machines Corporation
G06N3/08G06F16/93G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,087
App. No.
17/238,567
Granted
Sep 9, 2025
Kind
B2
Abstract

Methods, systems, and computer program products for classifying data from de-identified content are provided herein. A computer-implemented method includes applying one or more rules to identify one or more structural elements of a document; determining, based at least in part on the one or more structural elements, one or more pairs of words within the document having a hypernym relationship; extracting de-identified content within the document based on one or more de-identification techniques applied to the document; and applying a set of causal rules to the de-identified content and the one or more pairs of words to annotate at a least a portion of the de-identified content as belonging to a class of protected content.

Claims (59)

1. A computer-implemented method, the method comprising:

applying one or more rules to identify one or more structural elements of a document, wherein the one or more structural elements are indicative of at least one of an organization of text within the document and a visual presentation of text within the document;

determining, based at least in part on the one or more structural elements, one or more pairs of words within the document having a hypernym relationship;

extracting de-identified content within the document based on one or more de-identification techniques applied to the document;

applying a set of causal rules to the de-identified content and the one or more pairs of words by computing mutual dependence and/or mutual independence between the de-identified content and the one or more pairs of words;

annotating, based at least in part on the computed mutual dependence and/or mutual independence, at least a portion of the de-identified content as belonging to a class of protected content; and

training a machine learning model on a set of training data to automatically identify at least a portion of the de-identification techniques in at least one other document, wherein the set of training data comprises the annotated portion of the de-identified content;

wherein the method is carried out by at least one computing device.

2. The computer-implemented method of claim 1 , wherein determining the one or more pairs of words comprises determining one or more contexts in which a given one of the words is used within the document.

3. The computer-implemented method of claim 2 , wherein the one or more contexts correspond to the one or more structural elements.

4. The computer-implemented method of claim 2 , wherein the determining the one or more contexts comprises computing a set of context vectors for the given word, wherein the set of context vectors comprises:

a first vector indicating the contexts in which the given word is used as a relational hypernym;

a second vector indicating the contexts in which the given word is used as a relational hyponym; and

a third vector indicating the contexts in which the given word is used as a general hypernym.

5. The computer-implemented method of claim 4 , wherein computing the set of context vectors comprises at least one of:

determining a number of times the given word appears in the document;

determining a number of times the given word precedes each of the one or more structural elements; and

determining a number of times the given word succeeds each of the one or more structural elements.

6. The computer-implemented method of claim 1 , wherein the one or more structural elements relate to at least one of:

at least one of a type and a style of font;

a type of spacing;

a type of indention;

at least one of: a shape or a position of a shape within the document;

one or more symbols; and

one or more keywords.

7. The computer-implemented method of claim 1 , wherein the de-identification techniques comprise at least one of: one or more anonymization techniques and one or more pseudonymization techniques.

8. The computer-implemented method of claim 7 , wherein the machine learning model comprises a deep neural network, and wherein the annotated portion of the de-identified content comprises one or more labeled samples to identify at least a portion of the de-identification techniques.

9. The computer-implemented method of claim 1 , wherein the set of causal rules comprises at least one of a CCU causality rule and a CCC causality rule.

10. The computer-implemented method of claim 1 , wherein

the mutual independence and/or the mutual dependence between the de-identified content and the one or more pairs of words is computed by applying one or more thresholds.

11. The computer-implemented method of claim 1 , wherein the class of protected content is defined by a regulatory document.

12. The computer-implemented method of claim 1 , wherein software is provided as a service in a cloud environment.

13. A computer program product comprising:

a computer readable storage medium; program instructions embodied therewithstored in the computer readable storage medium, for causing a computing device to perform the following computer operations:

apply one or more rules to identify one or more structural elements of a document, wherein the one or more structural elements are indicative of at least one of an organization of text within the document and a visual presentation of text within the document;

determine, based at least in part on the one or more structural elements, one or more pairs of words within the document having a hypernym relationship;

extract de-identified content within the document based on one or more de-identification techniques applied to the document;

apply a set of causal rules to the de-identified content and the one or more pairs of words by computing mutual dependence and/or mutual independence between the de-identified content and the one or more pairs of words;

annotate, based at least in part on the computed mutual dependence and/or mutual independence, at least a portion of the de-identified content as belonging to a class of protected content; and

train a machine learning model on a set of training data to automatically identify at least a portion of the de-identification techniques in at least one other document, wherein the set of training data comprises the annotated portion of the de-identified content.

14. The computer program product of claim 13 , wherein determining the one or more pairs of words comprises determining one or more contexts in which a given one of the words is used within the document.

15. The computer program product of claim 14 , wherein the one or more contexts correspond to the one or more structural elements.

16. The computer program product of claim 14 , wherein the determining the one or more contexts comprises computing a set of context vectors for the given word, wherein the set of context vectors comprises:

a first vector indicating the contexts in which the given word is used as a relational hypernym;

a second vector indicating the contexts in which the given word is used as a relational hyponym; and

a third vector indicating the contexts in which the given word is used as a general hypernym.

17. The computer program product of claim 13 , wherein the de-identification techniques comprise at least one of: one or more anonymization techniques and one or more pseudonymization techniques.

18. The computer program product of claim 13 , wherein the set of causal rules comprises at least one of a CCU causality rule and a CCC causality rule.

19. The computer program product of claim 13 , wherein

the mutual independence and/or the mutual dependence between the de-identified content and the one or more pairs of words is computed by applying one or more thresholds.

20. A system comprising:

a memory configured to store program instructions;

a processor operatively coupled to the memory to execute the program instructions to:

apply one or more rules to identify one or more structural elements of a document, wherein the one or more structural elements are indicative of at least one of an organization of text within the document and a visual presentation of text within the document;

determine, based at least in part on the one or more structural elements, one or more pairs of words within the document having a hypernym relationship;

extract de-identified content within the document based on one or more de-identification techniques applied to the document;

apply a set of causal rules to the de-identified content and the one or more pairs of words by computing mutual dependence and/or mutual independence between the de-identified content and the one or more pairs of words;

annotate, based at least in part on the computed mutual dependence and/or mutual independence, at least a portion of the de-identified content as belonging to a class of protected content; and

train a machine learning model on a set of training data to automatically identify at least a portion of the de-identification techniques in at least one other document, wherein the set of training data comprises the annotated portion of the de-identified content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2021
From: KANNAN, ASWIN; GANESAN, BALAJI; GUTTULA, SHANMUKHA CHAITANYA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056035/0592 →
Continuity (1)
Related Publication 20220343151A1 · Oct 27, 2022
References Cited (33)
US 7630981B2 · Xu et al. · 2009 [cited by applicant]
US 7853555B2 · Peoples et al. · 2010 [cited by applicant]
US 8024339B2 · Barker et al. · 2011 [cited by applicant]
US 8635107B2 · Chang et al. · 2014 [cited by applicant]
US 9558179B1 · Jurca et al. · 2017 [cited by applicant]
US 9754128B2 · Golic · 2017 [cited by applicant]
US 10223586B1 · Leibovitz · 2019 [cited by examiner]
US 10878124B1 · Sitaraman · 2020 [cited by examiner]
US 20080027895A1 · Combaz · 2008 [cited by examiner]
US 20100076957A1 · Staddon · 2010 [cited by examiner]
US 20130166303A1 · Chang · 2013 [cited by examiner]
US 20140059011A1 · Bostick et al. · 2014 [cited by applicant]
US 20150235143A1 · Eder · 2015 [cited by examiner]
US 20170300565A1 · Calapodescu · 2017 [cited by examiner]
US 20200152302A1 · Co · 2020 [cited by examiner]
US 20200334381A1 · Yarowsky · 2020 [cited by examiner]
US 20210256160A1 · Hachey · 2021 [cited by examiner]
US 20220188567A1 · Ganesan et al. · 2022 [cited by applicant]
Espinosa-Anke, L., et al.,, Finding and Expanding Hypernymic Relations In The Music Domain, in 19th International Conference of the Catalan Association for Artificial Intelligence (CCIA), Barcelona, Spain, Oct. 19, 2016… [cited by applicant]
Snow, R., et al., Learning Syntactic Patterns for Automatic Hypernym Discovery, in Advances in Neural Information Processing Systems 17, L. K. Saul, Y. Weiss, and L. Bottou, eds., MIT Press, 2005, pp. 1297-1304. [cited by applicant]
Wang, Chenguang, et al., Towards Re-Defining Relation Under-Standing in Financial Domain. Proceedings of the 3rd International Workshop on Data Science for Macro-Modeling with Financial and Economic Datasets. 8, ACM. (2… [cited by applicant]
V. Baisa and V. Suchomel, Corpus Based Extraction of Hypernyms in Terminological Thesaurus for Land Surveying Domain, in Ninth Workshop on Recent Advances in Slavonic Natural Language Processing, Brno, 2015, Tribun EU, … [cited by applicant]
Wintergerst, Luca et al., Protecting GDPR Personal Data with Pseudonymization, published in elastic.co, available at https://www.elastic.co/blog/gdpr-personal-data-pseudonymization-part-1, Mar. 27, 2018, 14 pgs. [cited by applicant]
Eder, Elisabeth, et al., De-Identification of Emails: Pseudonymizing Privacy-Sensitive Data in a German Email Corpus, Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP … [cited by applicant]
Alfalahi, A., Brissman, S., & Dalianis, H. (2012). Pseudonymisation Of Personal Names and Other PHIs in an Annotated Clinical Swedish Corpus. In Third Workshop on Building and Evaluating Resources for Biomedical Text Mi… [cited by applicant]
Apache UIMA Ruta (Rule-based Text Annotation), published in Unstructured Information Management Architecture, available at https://uima.apache.org/ruta.html, last accessed Apr. 11, 2021, 3 pgs. [cited by applicant]
Five Ways Imperva Helps You with GDPR Compliance, Imperva, available at https://www.imperva.com/resources/datasheets/Imperva-Five-Ways-Imperva-HelpsWith-GDPR-2020.pdf, last accessed Apr. 11, 2021, 6 pgs. [cited by applicant]
Whitelegg, Dave, Minimizing application privacy risk, IBM, https://developer.ibm.com/solutions/security/articles/s-gdpr3/, published May 25, 2018, 19 pgs. [cited by applicant]
Mouna Kamel, et al., Extracting hypernym relations from Wikipedia disambiguation pages: comparing symbolic and machine learning approaches, International Conference on Computational Semantics (IWCS 2017), Montpellier, F… [cited by applicant]
Mell, Peter, et al., The NIST Definition of Cloud Computing, National Institute of Standards and Technology, U.S. Department of Commerce, NIST Special Publication 800-145, Sep. 2011, 7 pgs. [cited by applicant]
Kannan et al., “Document Structure Measure for Hypernym discovery”, arXiv:1811.12728, Nov. 30, 2018, 07 pages, https://doi.org/10.48550/arXiv.1811.12728. [cited by applicant]
Klimt et al., “The Enron Corpus: A New Dataset for Email Classification Research”, ECML 2004, Lecture Notes in Computer Science, vol. 3201, Sep. 2004, pp. 217-226, DOI:10.1007/978-3-540-30115-8_22. [cited by applicant]
Vannur et al., “Data Augmentation for Personal Knowledge Base Population”, arXiv:2002.10943, Aug. 18, 2020, 08 pages, https://doi.org/10.48550/arXiv.2002.10943. [cited by applicant]