IP Library › Granted Patent US 11,663,407
Granted Patent B2
US 11,663,407 · App. 17/109,196 · Granted May 30, 2023

Management of text-item recognition systems

Inventors: Francesco Fusco (Zurich, CH); Abderrahim Labbi (Gattikon, CH); Peter Willem Jan Staar (Wädenswil, CH)
Assignee: International Business Machines Corporation
G06F40/284G06F18/24147G06F40/169G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,663,407
App. No.
17/109,196
Granted
May 30, 2023
Kind
B2
Abstract

A tool for managing text-item recognition systems such as NER (Named Entity Recognition) systems. The tool applies the system to a text corpus containing instances of text items, such as named entities, to be recognized by the system, and selecting from the text corpus a set of instances of text items which the system recognized. The tool tokenizes the text corpus such that each instance in the aforementioned set is encoded as a single token and processing the tokenized text via a word embedding scheme to generate a word embedding matrix. The tool, responsive to selecting a seed token corresponding to an instance in the aforementioned set, performs a nearest-neighbor search of the embedding space to identify a set of neighboring tokens for the seed token, and identifies the text corresponding to each neighboring token as a potential instance of a text item to be annotated.

Claims (61)

1. A computer-implemented method for managing a text-item recognition system, the method comprising:

applying said system to a text corpus containing instances of text items to be recognized by the system;

selecting from the text corpus a set of instances of text items which the system recognized;

tokenizing the text corpus such that each instance in said set is encoded as a single token;

processing the tokenized text via a word embedding scheme to generate a word embedding matrix comprising vectors which indicate locations of respective tokens in a word embedding space;

in response to selection of a seed token corresponding to an instance in said set, performing a nearest-neighbor search of an embedding space to identify a set of neighboring tokens for the seed token;

for at least a subset of the neighboring tokens, identifying the text corresponding to each neighboring token as a potential instance of a text item to be annotated; and

targeting an annotation effort on a particular area of interest based on a set of identified potential instances of text items within the text corpus that are likely to improve said system, wherein the annotation effort includes a manual annotation of each instance of the set of identified potential instances of text items, and wherein said system is a rule-based text-item recognition system.

2. The computer-implemented method as claimed in claim 1 wherein said system provides a confidence value for each instance of a text item recognized by the system, the method including selecting instances of text items having confidence values above a threshold for inclusion in said set of instances.

3. The computer-implemented method as claimed in claim 1 wherein said system comprises a named-entity recognition system and said instances of text items comprise instances of named entities.

4. The computer-implemented method as claimed in claim 1 wherein said system comprises a machine learning model, the method including, for text identified as a said potential instance of a text item, storing at least one fragment of the text corpus which includes that text as a text sample to be annotated.

5. The computer-implemented method as claimed in claim 4 including, in response to annotation of a set of said text samples, training the model on the set of annotated text samples.

6. The computer-implemented method as claimed in claim 1 including:

extracting from the text corpus a plurality of seed text fragments each comprising a fragment of the text corpus which includes an instance corresponding to said seed token;

generating a vector representation of each seed text fragment based on a word-embedding of the fragment in an embedding space;

extracting from the text corpus a plurality of candidate text fragments each comprising a fragment of the text corpus which includes text corresponding to a neighboring token;

generating a said vector representation of each candidate text fragment;

computing, for each candidate text fragment, a distance value indicative of distance between the vector representation of that candidate text fragment and the vector representations of the seed text fragments; and

identifying the text corresponding to the neighboring token in each candidate text fragment with less than a threshold distance value as a said potential instance of a text item to be annotated.

7. The computer-implemented method as claimed in claim 6 wherein said system provides a confidence value for each instance of a text item recognized by the system, the method including:

selecting instances of text items having a confidence value above a first threshold for inclusion in said set of instances; and

selecting said candidate text fragments from text fragments which include text, corresponding to a neighboring token, that the system recognized as an instance of a text item with a confidence value between the first threshold and a second, lower threshold.

8. The computer-implemented method as claimed in claim 6 wherein said system comprises a machine learning model, the method including storing each candidate text fragment with less than a threshold distance value as a text sample to be annotated.

9. The computer-implemented method as claimed in claim 8 including, in response to annotation of a set of said text samples, training the model on the set of annotated text samples.

10. The computer-implemented method as claimed in claim 6 including:

generating said vector representation of each seed text fragment after removing the instance corresponding to the seed token in that fragment; and

generating said vector representation of each candidate text fragment after removing the text corresponding to the neighboring token in that fragment.

11. The computer-implemented method as claimed in claim 6 including generating said vector representation of a fragment by supplying that fragment to a pretrained word-embedding model.

12. The computer-implemented method as claimed in claim 1 including providing a user interface for user-selection of said seed token via selection of an instance from said set of instances of text items which the system recognized.

13. The computer-implemented method as claimed in claim 12 including, in response to user-selection of a plurality of seed tokens via the user interface, performing said nearest-neighbor search, and identifying text corresponding to at least said subset of the neighboring tokens identified by that search, for each of those seed tokens.

14. The computer-implemented method as claimed in claim 12 including, after performing the nearest-neighbor search for the seed token to identify said set of neighboring tokens, displaying in the interface the text corresponding to each neighboring token for user-selection of a restricted set of neighboring tokens.

15. A computer program product for managing a text-item recognition system, the computer program product comprising a computer readable storage medium having program instructions embodied therein, the program instructions being executable by a computing apparatus to cause the computing apparatus to:

apply said system to a text corpus containing instances of text items to be recognized by the system;

select from the text corpus a set of instances of text items which the system recognized;

tokenize the text corpus such that each instance in said set is encoded as a single token;

process the tokenized text via a word embedding scheme to generate a word embedding matrix comprising vectors which indicate locations of respective tokens in a word embedding space;

in response to selection of a seed token corresponding to an instance in said set, perform a nearest-neighbor search of the embedding space to identify a set of neighboring tokens for the seed token;

for at least a subset of the neighboring tokens, identify the text corresponding to each neighboring token as a potential instance of a text item to be annotated for improving operation of the system; and

target an annotation effort on a particular area of interest based on a set of identified potential instances of text items within the text corpus that are likely to improve said system, wherein the annotation effort includes a manual annotation of each instance of the set of identified potential instances of text items, and wherein said system is a rule-based text-item recognition system.

16. The computer program product as claimed in claim 15 wherein said system provides a confidence value for each instance of a text item recognized by the system, said program instructions being adapted to cause the computing apparatus to select instances of text items having confidence values above a threshold for inclusion in said set of instances.

17. The computer program product as claimed in claim 15 wherein said program instructions are adapted to cause the computing apparatus to:

extract from the text corpus a plurality of seed text fragments each comprising a fragment of the text corpus which includes an instance corresponding to said seed token;

generate a vector representation of each seed text fragment based on a word-embedding of the fragment in an embedding space;

extract from the text corpus a plurality of candidate text fragments each comprising a fragment of the text corpus which includes text corresponding to a neighboring token;

generate a said vector representation of each candidate text fragment;

compute, for each candidate text fragment, a distance value indicative of distance between the vector representation of that candidate text fragment and the vector representations of the seed text fragments; and

identify the text corresponding to the neighboring token in each candidate text fragment with less than a threshold distance value as a said potential instance of a text item to be annotated.

18. The computer program product as claimed in claim 17 wherein said system provides a confidence value for each instance of a text item recognized by the system, said program instructions being further adapted to cause the computing apparatus to:

select instances of text items having a confidence value above a first threshold for inclusion in said set of instances; and

select said candidate text fragments from text fragments which include text, corresponding to a neighboring token, that the system recognized as an instance of a text item with a confidence value between the first threshold and a second, lower threshold.

19. The computer program product as claimed in claim 17 wherein said system comprises a machine learning model, said program instructions being further adapted to cause the computing apparatus to:

store each candidate text fragment with less than a threshold distance value as a text sample to be annotated; and

in response to annotation of a set of said text samples, train the model on the set of annotated text samples.

20. A computing apparatus for managing a text-item recognition system, the apparatus comprising memory for storing a text corpus containing instances of text items to be recognized by the system, and control logic adapted to:

apply said system to the text corpus;

select from the text corpus a set of instances of text items which the system recognized;

tokenize the text corpus such that each instance in said set is encoded as a single token;

process the tokenized text via a word embedding scheme to generate a word embedding matrix comprising vectors which indicate locations of respective tokens in a word embedding space;

in response to selection of a seed token corresponding to an instance in said set, perform a nearest-neighbor search of the embedding space to identify a set of neighboring tokens for the seed token;

for at least a subset of the neighboring tokens, identify the text corresponding to each neighboring token as a potential instance of a text item to be annotated for improving operation of the system; and

target an annotation effort on a particular area of interest based on a set of identified potential instances of text items within the text corpus that are likely to improve said system, wherein the annotation effort includes a manual annotation of each instance of the set of identified potential instances of text items, and wherein said system is a rule-based text-item recognition system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2020
From: FUSCO, FRANCESCO; LABBI, ABDERRAHIM; STAAR, PETER WILLEM JAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054511/0494 →
Continuity (1)
Related Publication 20220171931A1 · Jun 2, 2022
Cited By (1)
US 12,333,835