IP Library Granted Patent US 12,056,945
Granted Patent B2
US 12,056,945 · App. 17/098,902 · Granted Aug 6, 2024

Method and system for extracting information from a document image

Inventor: Andrii Matiukhov (Moraga, CA)
Assignee: KYOCERA Document Solutions Inc.
G06V30/19G06F18/214G06F18/217G06F18/285G06F18/40G06N20/00G06V10/40G06V30/416G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,056,945
App. No.
17/098,902
Granted
Aug 6, 2024
Kind
B2
Abstract

A method performed by a computing system includes receiving, by a document data extraction system (DDES), image data associated with a document. The DDES extracts, via optical character recognition (OCR) logic of the DDES, metadata from the image data. The metadata specifies sequences of text content items and text content item features associated with each text content item of the sequences of text content items. A machine learning logic (MLL) module of the DDES determines, based on the sequences of text content items and the text content item features, one or more text content items associated with a key. The DDES communicates information that specifies the key and a corresponding value that is associated with the one or more text content items that are associated with the key to a terminal.

Claims (58)

1. A method performed by a computing system, the method comprising:

receiving, by a document data extraction system (DDES), image data associated with a document;

extracting, by optical character recognition (OCR) logic of the DDES, metadata from the image data, wherein the metadata specifies sequences of text content items and text content item features associated with each text content item of the sequences of text content items;

determining, by a machine learning logic module of the DDES and based on the sequences of text content items and the text content item features, one or more text content items associated with a key; and

communicating, by the DDES and to a terminal, information that specifies the key and a corresponding value that is associated with the one or more text content items that are associated with the key.

2. The method according to claim 1 , wherein each text content item feature specifies an amount of area in the image data occupied by a corresponding text content item, and a distance from an origin of the image data to the corresponding text content item.

3. The method according to claim 2 , further comprising:

determining, by the DDES, additional text content item features associated with each text content item, wherein the additional text content item features specify one or more of: a shape, morphological pattern, syntactic dependency, presence-of-hyphen indication, stop word indication, and style associated with a corresponding text content item.

4. The method according to claim 1 , wherein determining the one or more text content items associated with a key further comprises:

receiving, by first logic of the machine learning logic module that includes a recurrent neural network layer, the sequences of text content items;

receiving, by second logic of the machine learning logic module that includes a multilayer perceptron, the text content item features; and

combining, by third logic of the machine learning logic module that includes a fully connected layer, an output of the first logic and an output of the second logic, wherein for each of a plurality of keys, the fully connected layer outputs a probability that the one or more text content items is associated with a particular key of the plurality.

5. The method according to claim 1 , further comprising:

determining a document type; and

selecting, based on the document type, a machine learning logic module, from a plurality of machine learning logic modules, that is configured to determine one or more text content items associated with a key for a document of the document type.

6. The method according to claim 5 , further comprising:

responsive to determining that a machine learning logic module for the document type does not exist, generating a user interface that facilitates training a machine learning logic module to determine one or more text content items associated with a key for a document of the document type; and

associating a trained machine learning logic module with the document type.

7. The method according to claim 6 , wherein the user interface facilitates

mapping one or more text content items of a training document to corresponding one or more keys associated with the one or more text content items.

8. The method according to claim 6 , further comprising:

responsive to determining that a prediction accuracy of a particular machine learning logic module is below an accuracy threshold, training the machine learning logic module with another document of the document type.

9. A document data extraction system (DDES) comprising:

a memory that stores instruction code; and

a processor in communication with the memory, wherein the instruction code is executable by the processor to perform operations comprising:

receiving, by the document data extraction system (DDES), image data associated with a document;

extracting, by optical character recognition (OCR) logic of the DDES, metadata from the image data, wherein the metadata specifies sequences of text content items and text content item features associated with each text content item of the sequences of text content items;

determining, by a machine learning logic module of the DDES and based on the sequences of text content items and the text content item features, one or more text content items associated with a key; and

communicating, by the DDES and to a terminal, information that specifies the key and a corresponding value that is associated with the one or more text content items that are associated with the key.

10. The system according to claim 9 , wherein each text content item feature specifies an amount of area in the image data occupied by a corresponding text content item, and a distance from an origin of the image data to the corresponding text content item.

11. The system according to claim 10 , wherein the operations further comprise:

determining, by the DDES, additional text content item features associated with each text content item, wherein the additional text content item features specify one or more of: a shape, morphological pattern, syntactic dependency, presence-of-hyphen indication, stop word indication, and style associated with a corresponding text content item.

12. The system according to claim 9 , wherein in determining the one or more text content items associated with a key, the operations further comprise:

receiving, by first logic of the machine learning logic module that includes a recurrent neural network layer, the sequences of text content items;

receiving, by second logic of the machine learning logic module that includes a multilayer perceptron, the text content item features; and

combining, by third logic of the machine learning logic module that includes a fully connected layer, an output of the first logic and an output of the second logic, wherein for each of a plurality of keys, the fully connected layer outputs a probability that the one or more text content items is associated with a particular key of the plurality.

13. The system according to claim 9 , wherein the operations further comprise:

determining a document type; and

selecting, based on the document type, a machine learning logic module, from a plurality of machine learning logic modules, that is configured to determine one or more text content items associated with a key for a document of the document type.

14. The system according to claim 13 , wherein the operations further comprise:

responsive to determining that a machine learning logic module for the document type does not exist, generating a user interface that facilitates training a machine learning logic module to determine one or more text content items associated with a key for a document of the document type; and

associating a trained machine learning logic module with the document type.

15. The system according to claim 14 , wherein the user interface facilitates

mapping one or more text content items of a training document to corresponding one or more keys associated with the one or more text content items.

16. The system according to claim 14 , wherein the operations further comprise:

responsive to determining that a prediction accuracy of a particular machine learning logic module is below an accuracy threshold, training the machine learning logic module with another document of the document type.

17. A non-transitory computer-readable medium having stored thereon instruction code, wherein the instruction code is executable by a processor for causing the processor to perform operations comprising:

receiving, by a document data extraction system (DDES), image data associated with a document;

extracting, by optical character recognition (OCR) logic of the DDES, metadata from the image data, wherein the metadata specifies sequences of text content items and text content item features associated with each text content item of the sequences of text content items;

determining, by a machine learning logic module of the DDES and based on the sequences of text content items and the text content item features, one or more text content items associated with a key; and

communicating, by the DDES and to a terminal, information that specifies the key and a corresponding value that is associated with the one or more text content items that are associated with the key.

18. The non-transitory computer-readable medium according to claim 17 , wherein each text content item feature specifies an amount of area in the image data occupied by a corresponding text content item, and a distance from an origin of the image data to the corresponding text content item.

19. The non-transitory computer-readable medium according to claim 18 , wherein the operations further comprise:

determining, by the DDES, additional text content item features associated with each text content item, wherein the additional text content item features specify one or more of: a shape, morphological pattern, syntactic dependency, presence-of-hyphen indication, stop word indication, and style associated with a corresponding text content item.

20. The non-transitory computer-readable medium according to claim 17 , wherein in determining the one or more text content items associated with a key, the operations further comprise:

receiving, by first logic of the machine learning logic module that includes a recurrent neural network layer, the sequences of text content items;

receiving, by second logic of the machine learning logic module that includes a multilayer perceptron, the text content item features; and

combining, by third logic of the machine learning logic module that includes a fully connected layer, an output of the first logic and an output of the second logic, wherein for each of a plurality of keys, the fully connected layer outputs a probability that the one or more text content items is associated with a particular key of the plurality.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2020
From: MATIUKHOV, ANDRII
To: KYOCERA DOCUMENT SOLUTIONS INC.
Reel/Frame 054378/0456 →
Continuity (1)
Related Publication 20220156490A1 · May 19, 2022
Cited By (1)
US 12,525,000