IP Library Granted Patent US 11,755,839
Granted Patent B2
US 11,755,839 · App. 17/324,212 · Granted Sep 12, 2023

Low resource named entity recognition for sensitive personal information

Inventors: Youngja Park (Princeton, NJ); Jatin Arora (Urbana, IL)
Assignee: International Business Machines Corporation
G06F40/295G06F40/166G06F40/284G06F40/30G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,755,839
App. No.
17/324,212
Granted
Sep 12, 2023
Kind
B2
Abstract

Natural language processing (NLP) methodologies and mechanisms are provided that include a named entity recognition (NER) computer model augmented to operate on an entity pattern embedding input feature in addition to other embedding input features. The mechanisms tokenize natural language content (NLC) to generate tokens and process a selected token in accordance with a predetermined entity pattern embedding technique to generate an entity pattern embedding input feature for the selected token. The entity pattern embedding input feature specifies a pattern of characters present in the selected token. The mechanisms process the NLC to generate the other embedding input features in accordance with other embedding techniques, and process, by the NER computer model, the other embedding input features and the entity pattern embedding input feature for the selected token to generate a predicted tag for the selected token. The predicted tag specifies a named entity type classification for the selected token.

Claims (47)

1. A method, in a natural language processing (NLP) computing system comprising a named entity recognition (NER) computer model augmented to operate on an entity pattern embedding input feature in addition to one or more other embedding input features, the method comprising:

tokenizing natural language content to generate one or more tokens, wherein each token represents a subset of text in the natural language content;

processing a selected token, in the one or more tokens, in accordance with a predetermined entity pattern embedding technique to generate an entity pattern embedding input feature for the selected token, wherein the entity pattern embedding input feature specifies a pattern of characters present in the selected token;

processing the natural language content to generate the one or more other embedding input features in accordance with one or more other embedding techniques;

processing, by the NER computer model, the one or more other embedding input features and the entity pattern embedding input feature for the selected token to generate a predicted tag for the selected token, wherein the predicted tag specifies a named entity type classification for the selected token; and

performing, by the NLP computing system, an operation based on the predicted tag.

2. The method of claim 1 , wherein the one or more other embedding input features comprise at least one of a word embedding input feature or a character embedding input feature.

3. The method of claim 1 , wherein the one or more other embedding input features comprise a semantic category embedding input feature that specifies a context pattern of one or more high-resource named entity types present in the natural language content.

4. The method of claim 3 , further comprising:

generating the semantic category embedding input feature at least by processing tokens in the natural language content via one or more high resource named entity computer models or matching the tokens to entries in one or more knowledge resources to generate named entity type classifications for one or more other tokens in the natural language content; and

determining a semantic category embedding based on a combination of the named entity type classifications for the one or more other tokens in the natural language content.

5. The method of claim 4 , wherein processing the one or more other embedding input features and the entity pattern embedding input feature for the selected token to generate the predicted tag for the selected token comprises:

processing, by the NER computer model, the combination of named entity type classifications based on learned association patterns specifying entity types that appear together in natural language content to generate a first prediction of a tag for the selected token;

processing, by the NER computer model, the entity pattern embedding input feature to map the entity pattern embedding input feature to one or more second predictions of the tag for the selected token; and

combining, by the NER computer model, the first prediction and the one or more second predictions to generate the predicted tag.

6. The method of claim 1 , wherein the operation is a machine learning operation that trains the NER computer model to recognize named entities in natural language content, and wherein the machine learning operation trains the NER computer model to recognize selected named entities by learning a correlation of the one or more other embedding input features and the entity pattern embedding input feature to a correct tag, wherein the correct tag specifies a correct named entity type classification for a selected named entity corresponding to the one or more other embedding input features and the entity pattern embedding input feature.

7. The method of claim 6 , wherein the selected named entity is one of a sensitive personal information entity or a domain specific entity, for which there is less than a predetermined number of labeled instances in training data used to train the NER computer model.

8. The method of claim 7 , wherein the sensitive personal information entity is at least one of a government identification entity, a membership identifier entity, a biomedical entity, or a cybersecurity entity.

9. The method of claim 1 , wherein the entity pattern embedding technique comprises replacing letters with first entity pattern elements corresponding to a type of the letter and replacing numerical characters with second entity pattern elements specifying the numerical characters to be digits, and maintaining symbol characters as third entity pattern elements.

10. The method of claim 1 , wherein the entity pattern embedding technique comprises replacing first alphanumeric characters of a first character type with a first entity pattern element specifying the first character type and followed by a first numerical value specifying a number of the first alphanumeric characters of the first character type, replacing second alphanumeric characters of a second character type, different from the first character type, with a second entity pattern element specifying the second character type followed by a second numerical value specifying a number of the second alphanumeric characters of the second character type, and maintaining symbol characters as third entity pattern elements.

11. A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a data processing system, configures the data processing system to implement a natural language processing (NLP) computing system having a named entity recognition (NER) computer model, wherein the data processing system is further configured to:

tokenize natural language content to generate one or more tokens, wherein each token represents a subset of text in the natural language content;

process a selected token, in the one or more tokens, in accordance with a predetermined entity pattern embedding technique to generate an entity pattern embedding input feature for the selected token, wherein the entity pattern embedding input feature specifies a pattern of characters present in the selected token;

process the natural language content to generate the one or more other embedding input features in accordance with one or more other embedding techniques;

process, by the NER computer model, the one or more other embedding input features and the entity pattern embedding input feature for the selected token to generate a predicted tag for the selected token, wherein the predicted tag specifies a named entity type classification for the selected token; and

perform, by the NLP computing system, an operation based on the predicted tag.

12. The computer program product of claim 11 , wherein the one or more other embedding input features comprise at least one of a word embedding input feature or a character embedding input feature.

13. The computer program product of claim 11 , wherein the one or more other embedding input features comprise a semantic category embedding input feature that specifies a context pattern of one or more high-resource named entity types present in the natural language content.

14. The computer program product of claim 13 , wherein the computer readable program further causes the data processing system to:

generate the semantic category embedding input feature at least by processing tokens in the natural language content via one or more high resource named entity computer models or matching the tokens to entries in one or more knowledge resources to generate named entity type classifications for one or more other tokens in the natural language content; and

determine a semantic category embedding based on a combination of the named entity type classifications for the one or more other tokens in the natural language content.

15. The computer program product of claim 14 , wherein the computer readable program further causes the data processing system to process the one or more other embedding input features and the entity pattern embedding input feature for the selected token to generate the predicted tag for the selected token at least by:

processing, by the NER computer model, the combination of named entity type classifications based on learned association patterns specifying entity types that appear together in natural language content to generate a first prediction of a tag for the selected token; and

processing, by the NER computer model, the entity pattern embedding input feature to map the entity pattern embedding input feature to one or more second predictions of the tag for the selected token; and

combining, by the NER computer model, the first prediction and the one or more second predictions to generate the predicted tag.

16. The computer program product of claim 11 , wherein the operation is a machine learning operation that trains the NER computer model to recognize named entities in natural language content, and wherein the machine learning operation trains the NER computer model to recognize selected named entities by learning a correlation of the one or more other embedding input features and the entity pattern embedding input feature to a correct tag, wherein the correct tag specifies a correct named entity type classification for a selected named entity corresponding to the one or more other embedding input features and the entity pattern embedding input feature.

17. The computer program product of claim 16 , wherein the selected named entity is one of a sensitive personal information entity or a domain specific entity, for which there is less than a predetermined number of labeled instances in training data used to train the NER computer model.

18. The computer program product of claim 17 , wherein the sensitive personal information entity is at least one of a government identification entity, a membership identifier entity, a biomedical entity, or a cybersecurity entity.

19. The computer program product of claim 11 , wherein the entity pattern embedding technique comprises replacing letters with first entity pattern elements corresponding to a type of the letter and replacing numerical characters with second entity pattern elements specifying the numerical characters to be digits, and maintaining symbol characters as third entity pattern elements.

20. An apparatus comprising:

at least one processor; and

at least one memory coupled to the at least one processor, wherein the at least one memory comprises instructions which, when executed by the at least one processor, configures the a least one processor to implement a natural language processing (NLP) computing system having a named entity recognition (NER) computer model, wherein the at least one processor is further configured to:

tokenize natural language content to generate one or more tokens, wherein each token represents a subset of text in the natural language content;

process a selected token, in the one or more tokens, in accordance with a predetermined entity pattern embedding technique to generate an entity pattern embedding input feature for the selected token, wherein the entity pattern embedding input feature specifies a pattern of characters present in the selected token;

process the natural language content to generate the one or more other embedding input features in accordance with one or more other embedding techniques;

process, by the NER computer model, the one or more other embedding input features and the entity pattern embedding input feature for the selected token to generate a predicted tag for the selected token, wherein the predicted tag specifies a named entity type classification for the selected token; and

perform, by the NLP computing system, an operation based on the predicted tag.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: PARK, YOUNGJA; ARORA, JATIN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056284/0156 →
Continuity (1)
Related Publication 20220374602A1 · Nov 24, 2022
Cited By (1)
US 12,488,191