IP Library Granted Patent US 11,687,719
Granted Patent B2
US 11,687,719 · App. 17/188,362 · Granted Jun 27, 2023

Post-filtering of named entities with machine learning

Inventors: Christian Schäfer (Berlin, DE); Michael Kieweg (Berlin, DE); Florian Kuhlmann (Berlin, DE)
Assignee: LEVERTON HOLDING LLC
G06F40/295G06F17/16G06F18/2185G06F18/295G06N3/08G06V10/82G06V30/19173G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,687,719
App. No.
17/188,362
Granted
Jun 27, 2023
Kind
B2
Abstract

A method for identifying errors associated with named entity recognition includes recognizing a candidate named entity within a text and extracting a chunk from the text containing the candidate named entity. The method further includes creating a feature vector associated with the chunk and analyzing the feature vector for an indication of an error associated with the candidate named entity. The method also includes correcting the error associated with the candidate named entity.

Claims (48)

1. A computer-implemented method, comprising:

recognizing a candidate named entity within a text;

extracting a first subset of the text that contains the candidate named entity;

analyzing, with a classifier, one or more features of the first subset of the text to identify an error associated with the candidate named entity; and

correcting the error associated with the candidate named entity.

2. The method of claim 1 , further comprising:

storing a document image in a memory; and

recognizing the text from the document image.

3. The method of claim 1 , wherein the error associated with the candidate named entity is that the candidate named entity is not a named entity and correcting the error associated with the candidate named entity includes removing the candidate named entity as a potential named entity in the text.

4. The method of claim 1 , wherein the classifier analyzes the one or more features using a first machine learning model.

5. The method of claim 4 , wherein the first machine learning model includes one or more of a recurrent neural network, a convolutional neural network, a conditional random field model, and a Markov model.

6. The method of claim 4 , further comprising:

receiving a labeled training text comprising (i) a candidate training named entity, (ii) a subset of the training text associated with the candidate training named entity, and (iii) a labeling output indicating whether the candidate training named entity is a named entity;

creating a training feature vector, wherein the training feature vector includes one or more features of the training subset of the subset of the training text;

analyzing the training feature vector using the first machine learning model to generate an indication of whether the first machine learning model identified an error associated with the candidate training named entity;

comparing the indication with the labeling output to identify one or more errors in the indication; and

updating one or more parameters of the first machine learning model based on the one or more errors in the indication.

7. The method of claim 6 , wherein the first machine learning model is initially configured to identify errors associated with candidate named entities recognized from a first document type and updating one or more parameters of the first machine learning model enables the first machine learning model to identify errors associated with candidate named entities recognized from a second document type.

8. The method of claim 1 , wherein the candidate named entity is recognized using a second machine learning model.

9. The method of claim 1 , wherein the one or more features include at least one of a named entity label associated with the candidate named entity, a recognition accuracy prediction of the candidate named entity, a distance measure between the first subset of the text and a previous subset of the text, a distance measure between the first subset of the text and a subsequent subset of the text, an embedding vector associated with the first subset of the text, semantics of the first subset of the text, and/or a similarity of the candidate named entity contained within the first subset of the text and a named entity and/or a candidate named entity contained within a second subset of the text.

10. The method of claim 1 , wherein removing the candidate named entity improves the accuracy of named entities recognized within the text.

11. The method of claim 1 , wherein the steps of the method are performed on a plurality of candidate named entities recognized within the text.

12. The method of claim 1 , wherein the recognizing a candidate named entity within the text further comprises recognizing a plurality of candidate named entities from a plurality of texts, and wherein performing the method corrects errors associated with at least a subset of the plurality of candidate named entities from the plurality of texts.

13. A system comprising:

a processor; and

a memory storing instructions which, when executed by the processor, cause the processor to:

recognize a candidate named entity within a text;

extract a first subset of the text that contains the candidate named entity;

analyze, with a classifier, one or more features of the first subset of the text to identify an error associated with the candidate named entity; and

correct the error associated with the candidate named entity.

14. The system of claim 13 , wherein the instructions further cause the processor to:

storing a document image in a memory; and

recognizing the text from the document image.

15. The system of claim 13 , wherein the error associated with the candidate named entity is that the candidate named entity is not a named entity and correcting the error associated with the candidate named entity includes removing the candidate named entity as a potential named entity in the text.

16. The system of claim 13 , wherein the classifier analyzes the one or more features using a first machine learning model.

17. The system of claim 16 , wherein the instructions further cause the processor to:

receive a labeled training text comprising (i) a candidate training named entity, (ii) a subset of the training text associated with the candidate training named entity, and (iii) a labeling output indicating whether the candidate training named entity is a named entity;

create a training feature vector, wherein the training feature vector includes one or more features of the training subset of the subset of the training text;

analyze the training feature vector using the first machine learning model to generate an indication of whether the first machine learning model identified an error associated with the candidate training named entity;

compare the indication with the labeling output to identify one or more errors in the indication; and

update one or more parameters of the first machine learning model based on the one or more errors in the indication.

18. The system of claim 17 , wherein the first machine learning model is initially configured to identify errors associated with candidate named entities recognized from a first document type and updating one or more parameters of the first machine learning model enables the first machine learning model to identify errors associated with candidate named entities recognized from a second document type.

19. The system of claim 13 , wherein the one or more features include at least one of a named entity label associated with the candidate named entity, a recognition accuracy prediction of the candidate named entity, a distance measure between the first subset of the text and a previous subset of the text, a distance measure between the first subset of the text and a subsequent subset of the text, an embedding vector associated with the first subset of the text, semantics of the first subset of the text, and/or a similarity of the candidate named entity contained within the first subset of the text and a named entity and/or a candidate named entity contained within a second subset of the text.

20. A non-transitory, computer readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to:

recognize a candidate named entity within a text;

extract a first subset of the text that contains the candidate named entity;

analyze, with a classifier, one or more features of the first subset of the text to identify an error associated with the candidate named entity; and

correct the error associated with the candidate named entity.

Assignments (3)
SECURITY INTEREST Recorded Oct 2, 2025
From: LEVERTON HOLDING, LLC
To: GOLUB CAPITAL MARKETS LLC, AS COLLATERAL AGENT
Reel/Frame 072447/0265 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2023
From: SCHAFER, CHRISTIAN; KIEWEG, MICHAEL; KUHLMANN, FLORIAN
To: LEVERTON GMBH
Reel/Frame 062705/0626 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2023
From: LEVERTON GMBH
To: LEVERTON HOLDING LLC
Reel/Frame 062705/0875 →
Continuity (3)
Continuation 16416827 · May 20, 2019
Provisional Application 62674312 · May 21, 2018
Related Publication 20210182494A1 · Jun 17, 2021