IP Library Granted Patent US 12,354,156
Granted Patent B2
US 12,354,156 · App. 17/966,521 · Granted Jul 8, 2025

System and method for data extraction

Inventors: Mihir Joshi (Atlanta, GA); Kyle Wade (Dallas, TX); Brian A. Klotzman (Lewisville, TX); Eric Smith (Grapevine, TX)
Assignee: NCR Voyix Corporation
G06Q30/0641G06F40/103G06F40/166G06V30/414G06Q50/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,156
App. No.
17/966,521
Granted
Jul 8, 2025
Kind
B2
Abstract

A system and method of data extraction is disclosed. An image file of a scan of a printed information list is received at a server via a network connection. When not in portable document format, the received image file is processed with an optical character recognition (OCR) engine at the server to identify all text therein and then the processed image file is stored in a memory. When in portable document format, the image file is processed using metadata and positional data at the server to generate sentences; process the sentences to identify prices, descriptions, items, and categories; link each identified item to an associated price, description, and category; and extract and link all modifiers from identified description, and then all identified and extracted information is stored in memory. A user interface is provided to a user via the server for graphically visualizing and editing a stored processed file.

Claims (41)

1. A method of data extraction, comprising:

receiving an image file of a scan of a printed information list at a server via a network connection;

determining, at the server, if the image file is in portable document format;

when the image file is not in portable document format, processing the received image file with an optical character recognition (OCR) engine at the server to identify all text therein and storing the processed image file in a memory coupled to the server;

when the image is in portable document format, processing the image file using metadata and positional data at the server to:

generate sentences,

process the sentences to identify prices, descriptions, items, and categories,

link each identified item to an associated price, description, and category, and

extract and link all modifiers from each identified description; and

storing all identified and extracted information as a processed portable document format file in the memory; and

providing a user interface via the server for graphically visualizing and editing a selected one of a stored processed image file or a stored processed portable document format file, the user interface including a first pane for displaying the scan of the printed information list, a second pane for displaying a status of the identified and extracted data from the printed information list, and a set of commands for editing the identified and extracted information.

2. The method of claim 1 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to generate sentences by processing text information character-by-character, from left to right, and designating a new sentence when a predetermined number of blank spaces are found, a new font is found, or a new font size is found.

3. The method of claim 1 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify prices by stripping each currency symbol and designating each number having a value less than one thousand as a price.

4. The method of claim 3 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify items by identifying one or more fonts that appear most frequently but which are not used for a description or a price.

5. The method of claim 4 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify categories by identifying fonts which occur less frequently than fonts for items and which are not used for a description or a price.

6. The method of claim 4 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to link each item among the identified items to an associated price by identifying a price among the identified prices closest to and immediately to a right of or below that item.

7. The method of claim 1 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify descriptions by calculating a ratio between a separator count and a total character count for each different font found, and designating all text having the associated font as descriptions responsive to the calculated ratio being above a predetermined threshold.

8. The method of claim 7 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify modifiers by extracting all noun phrases from each of the identified descriptions.

9. The method of claim 1 , wherein the user interface provides adaptive threshold text selection in response to a box drawn by a user by identifying any text which is within the drawn box by at least a predetermined threshold.

10. The method of claim 1 , wherein the user interface provides dynamic line detection by identifying a line for each word within a box drawn by a user by determining if that word overlaps an existing line or a first word in the box by at least a predetermined threshold.

11. A system for data extraction, comprising:

a processor and a first memory, the first memory storing instructions which, when executed by the processor, cause the processor to perform the following steps:

receive an image file of a scan of a printed information list via a network connection;

determine if the image file is in portable document format;

when the image file is not in portable document format, process the received image file with an optical character recognition (OCR) engine to identify all text therein and store the processed image file in a second memory;

when the image is in portable document format, process the image file using metadata and positional data at the server to:

generate sentences,

process the sentences to identify prices, descriptions, items, and categories,

link each identified item to an associated price, description, and category, and

extract and link all modifiers from each identified description; and

store all identified and extracted information as a processed portable document format file in the memory; and

provide a user interface via the server for graphically visualizing and editing a selected one of a stored processed image file or a stored processed portable document format file, the user interface including a first pane for displaying the scan of the printed information list, a second pane for displaying a status of the identified and extracted data from the printed information list, and a set of commands for editing the identified and extracted information.

12. The system of claim 11 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to generate sentences by processing text information character-by-character, from left to right, and designating a new sentence when a predetermined number of blank spaces are found, a new font is found, or a new font size is found.

13. The system of claim 11 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify prices by stripping each currency symbol and designating each number having a value less than one thousand as a price.

14. The system of claim 13 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify items by identifying one or more fonts that appear most frequently but which are not used for a description or a price.

15. The system of claim 14 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify categories by identifying fonts which occur less frequently than fonts for items and which are not used for a description or a price.

16. The system of claim 14 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to link each item among the identified items to an associated price by identifying a price among the identified prices closest to and immediately to a right of or below that item.

17. The system of claim 11 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify descriptions by calculating a ratio between a separator count and a total character count for each different font found, and designating all text having the associated font as descriptions responsive to the calculated ratio being above a predetermined threshold.

18. The system of claim 17 , wherein when the image is in portable document format, the image file is processed using metadata and positional data at the server to identify modifiers by extracting all noun phrases from each of the identified descriptions.

19. The system of claim 11 , wherein the user interface provides adaptive threshold text selection in response to a box drawn by a user by identifying any text which is within the drawn box by at least a predetermined threshold.

20. The system of claim 11 , wherein the user interface provides dynamic line detection by identifying a line for each word within a box drawn by a user by determining if that word overlaps an existing line or a first word in the box by a predetermined threshold.

Assignments (1)
CHANGE OF NAME Recorded Feb 29, 2024
From: NCR CORPORATION
To: NCR VOYIX CORPORATION
Reel/Frame 066602/0539 →