IP Library › Granted Patent US 11,562,591
Granted Patent B2
US 11,562,591 · App. 17/132,114 · Granted Jan 24, 2023

Computer vision systems and methods for information extraction from text images using evidence grounding techniques

Inventors: Khoi Nguyen (Corvallis, OR); Maneesh Kumar Singh (Princeton, NJ)
Assignee: Insurance Services Office, Inc.
G06V30/413G06K9/6276G06K9/6297G06V30/414G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,591
App. No.
17/132,114
Granted
Jan 24, 2023
Kind
B2
Abstract

Computer vision systems and methods for text classification are provided. The system detects a plurality of text regions in an image and generates a bounding box for each detected text region. The system utilizes a neural network to recognize text present within each bounding box and classifies the recognized text, based on at least one extracted feature of each bounding box and the recognized text present within each bounding box, according to a plurality of predefined tags. The system can associate a key with a value and return a key-value pair for each predefined tag.

Claims (49)

1. A computer vision system for text classification comprising:

a memory; and

a processor in communication with the memory, the processor:

detecting a plurality of text regions in an image,

generating a bounding box for each detected text region,

recognizing text present within each bounding box using a neural network,

classifying the recognized text based on at least one extracted feature of each bounding box and the recognized text present within each bounding box according to a plurality of predefined tags, and

associating a key with a value and returning a key-value pair for each predefined tag,

wherein the processor:

extracts a positional feature of each bounding box,

extracts a textual feature of the recognized text present within each bounding box, the textual feature being indicative of a length of the recognized text, and

concatenates the extracted positional feature and the extracted textual feature to generate a fixed vector representation of the recognized text.

2. The system of claim 1 , wherein the processor detects the plurality of text regions in the image by applying an Efficient and Accurate Scene Text Detector (EAST) model to the image.

3. The system of claim 1 , wherein the neural network is a convolutional recurrent neural network.

4. The system of claim 1 , wherein the processor classifies the recognized text by utilizing a modified conditional random fields machine learning system implemented as a recurrent neural network and a modified graph attention network.

5. The system of claim 1 , wherein the processor:

extracts the positional feature of each bounding box based on coordinates of each bounding box and a width and a height of each bounding box relative to a width and a height of the image, and

extracts the textual feature of the recognized text present within each bounding box by utilizing a bidirectional long short-term memory (LSTM).

6. The system of claim 1 , wherein the processor generates a graph having an adjacency list by applying a k-nearest neighbors algorithm based on a Euclidean distance between center positions of each bounding box.

7. A method for text classification by a computer vision system, comprising the steps of:

detecting a plurality of text regions in an image;

generating a bounding box for each detected text region;

recognizing text present within each bounding box using a neural network;

classifying the recognized text, based on at least one extracted feature of each bounding box and the recognized text present within each bounding box, according to a plurality of predefined tags;

associating a key with a value and returning a key-value pair for each predefined tag; extracting a positional feature of each bounding box;

extracting a textual feature of the recognized text present within each bounding box, the textual feature being indicative of a length of the recognized text; and

concatenating the extracted positional feature and the extracted textual feature to generate a fixed vector representation of the recognized text.

8. The method of claim 7 , further comprising the step of detecting the plurality of text regions in the image by applying an Efficient and Accurate Scene Text Detector (EAST) model to the image.

9. The method of claim 7 , wherein the neural network is a convolutional recurrent neural network.

10. The method of claim 7 , further comprising the step of classifying the recognized text by utilizing a modified conditional random fields machine learning system implemented as a recurrent neural network and a modified graph attention network.

11. The method of claim 7 , further comprising the steps of:

extracting the positional feature of each bounding box based on coordinates of each bounding box and a width and a height of each bounding box relative to a width and a height of the image; and

extracting the textual feature of the recognized text present within each bounding box by utilizing a bidirectional long short-term memory (LSTM).

12. The method of claim 7 , further comprising the step of generating a graph having an adjacency list by applying a k-nearest neighbors algorithm based on a Euclidean distance between center positions of each bounding box.

13. A non-transitory computer readable medium having instructions stored thereon for text classification by a computer vision system which, when executed by a processor, causes the processor to carry out the steps of:

detecting a plurality of text regions in an image;

generating a bounding box for each detected text region;

recognizing text present within each bounding box using a neural network;

classifying the recognized text, based on at least one extracted feature of each bounding box and the recognized text present within each bounding box, according to a plurality of predefined tags;

associating a key with a value and returning a key-value pair for each predefined tag;

extracting a positional feature of each bounding box;

extracting a textual feature of the recognized text present within each bounding box, the textual feature being indicative of a length of the recognized text; and

concatenating the extracted positional feature and the extracted textual feature to generate a fixed vector representation of the recognized text.

14. The non-transitory computer readable medium of claim 13 , the processor further carrying out the step of detecting the plurality of text regions in the image by applying an Efficient and Accurate Scene Text Detector (EAST) model to the image.

15. The non-transitory computer readable medium of claim 13 , wherein the neural network is a convolutional recurrent neural network.

16. The non-transitory computer readable medium of claim 13 , the processor further carrying out the step of classifying the recognized text by utilizing a modified conditional random fields machine learning system implemented as a recurrent neural network and a modified graph attention network.

17. The non-transitory computer readable medium of claim 13 , the processor further carrying out the steps of:

extracting the positional feature of each bounding box based on coordinates of each bounding box and a width and a height of each bounding box relative to a width and a height of the image; and

extracting the textual feature of the recognized text present within each bounding box by utilizing a bidirectional long short-term memory (LSTM).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: NGUYEN, KHOI; SINGH, MANEESH KUMAR
To: INSURANCE SERVICES OFFICE, INC.
Reel/Frame 059518/0099 →
Continuity (2)
Provisional Application 62952749 · Dec 23, 2019
Related Publication 20210192201A1 · Jun 24, 2021