IP Library › Granted Patent US 11,361,565
Granted Patent B2
US 11,361,565 · App. 16/711,619 · Granted Jun 14, 2022

Natural language processing (NLP) pipeline for automated attribute extraction

Inventors: Sema Ustuntas (Redmond, WA); Kris Fantin Nunes (Bellevue, WA); Krishna Srinivasmurthy (Bothell, WA)
Assignee: THE BOEING COMPANY
G06V30/413G06V10/22G06V20/62G06V30/224
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,361,565
App. No.
16/711,619
Granted
Jun 14, 2022
Kind
B2
Abstract

A method for training a filter-based text recognition system for cataloging image portions associated with files using text from the image portions, the method comprising: receiving a first set of text represented in a first image portion associated with a first file; classifying the first image portion into a predetermined group, wherein the classifying is based at least in part on the first set of text; extracting a first set of features from the first set of text; harmonizing existing data in the predetermined group with the first set of text to modify the first set of features; categorizing the first set of text; and determining analytics-based rules based at least in part on the first set of features.

Claims (51)

1. A method for training a filter-based text recognition system for cataloging image portions associated with files using text from the image portions, the method comprising:

receiving a first set of text represented in a first image portion associated with a first file;

classifying the first image portion into a predetermined group, wherein the classifying is based at least in part on the first set of text;

extracting a first set of features from the first set of text;

harmonizing existing data in the predetermined group with the first set of text to modify the first set of features;

categorizing the first set of text;

determining analytics-based rules based at least in part on the first set of features;

aggregating the first set of text with existing data;

harmonizing the aggregated data and a second set of text into a second set of features, wherein the second set of text is represented in a second image portion classified into the predetermined group;

analyzing the second set of features; and

adapting at least one of the classifying step and the harmonizing step based on the analyzed second set of features.

2. The method of claim 1 , further comprising sorting the first set of text into predetermined structures, wherein the predetermined structures comprise at least one of the group consisting of words, numbers, dates, sentences, paragraphs, tables, and alphanumeric codes.

3. The method of claim 2 , wherein the classification of the first image portion is based at least in part on the predetermined structures into which the first set of text is sorted.

4. The method of claim 1 , further comprising receiving a second set of text, wherein the second set of text is not obtained from the first image portion.

5. The method of claim 1 , further comprising at least one of adding and modifying at least one feature for the first set of features.

6. The method of claim 1 , wherein the analytics-based rules are used at least in part to correct errors in the first set of text.

7. The method of claim 1 , wherein the method is performed on a large area network.

8. The method of claim 1 , further comprising initially analyzing the first set of text using optical character recognition.

9. The method of claim 1 , wherein classifying the first image portion into the predetermined group based at least in part on the first set of text comprises classifying the first image portion based on orders of words in the predetermined group.

10. A filter-based text recognition system for cataloging image portions associated with files using text from the image portions, the system comprising:

a memory; and

processing circuitry coupled with the memory, wherein the processing circuitry is operable to:

receive a first set of text represented in a first image portion associated with a first file;

classify the first image portion into a predetermined group, wherein the classifying is based at least in part on the first set of text;

extract a first set of features from the first set of text;

harmonize existing data in the predetermined group with the first set of text to modify the first set of features;

categorize the first set of text; and

determine analytics-based rules based at least in part on the first set of features and correcting the first set of text by appending additional data based on the rules.

11. The system of claim 10 , wherein the processing circuitry is further operable to aggregate the first set of text with existing data.

12. The system of claim 11 , wherein the processing circuitry is further operable to:

harmonize the aggregated data and a second set of text into a second set of features, wherein the second set of text is represented in a second image portion classified into the predetermined group;

analyze the second set of features; and

adapt at least one of the classifying step and the harmonizing step based on the analyzed second set of features.

13. The system of claim 10 , wherein the processing circuitry is further operable to sort the first set of text into predetermined structures, the predetermined structures comprise at least one of the group consisting of words, numbers, dates, sentences, paragraphs, tables, and alphanumeric codes.

14. The system of claim 13 , wherein the classification of the first image portion is based at least in part on the predetermined structures into which the first set of text is sorted.

15. The system of claim 10 , wherein the processing circuitry is further operable to receive a second set of text, wherein the second set of text is not obtained from the first image portion.

16. The system of claim 10 , wherein the first set of text is obtained using an optical character recognition technique.

17. The system of claim 10 , wherein the processing circuitry is further operable to at least one of add and modify at least one feature for the first set of features.

18. The system of claim 10 , wherein the system is a large area network.

19. The system of claim 10 , wherein the analytics-based rules comprise searching for a particular word in the first image portion.

20. A non-transitory computer-readable storage medium for cataloging image portions associated with files using text from the image portions, the computer-readable storage medium being non-transitory and having computer-readable program code portions stored therein that in response to execution by a processing circuitry, cause an apparatus to at least:

receive a first set of text represented in a first image portion associated with a first file;

classify the first image portion into a predetermined group, wherein the classifying is based at least in part on the first set of text;

extract a first set of features from the first set of text;

harmonize existing data in the predetermined group with the first set of text to modify the first set of features;

categorize the first set of text;

determine analytics-based rules based at least in part on the first set of features;

aggregate the first set of text with existing data;

harmonize the aggregated data and a second set of text into a second set of features, wherein the second set of text is represented in a second image portion classified into the predetermined group;

analyze the second set of features; and

adapt at least one of the classifying step, and the harmonizing step based on the analyzed second set of features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2019
From: USTUNTAS, SEMA; NUNES, KRIS FANTIN; SRINIVASMURTHY, KRISHNA
To: THE BOEING COMPANY
Reel/Frame 051261/0303 →
Continuity (1)
Related Publication 20210182549A1 · Jun 17, 2021