IP Library › Patent Application 18789940
Patent Application
App. No. 18/789,940

Computer Vision Systems and Methods for Information Extraction from Text Images Using Evidence Grounding Techniques

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/789,940
Abstract

Computer vision systems and methods for text classification are provided. The system detects a plurality of text regions in an image and generates a bounding box for each detected text region. The system utilizes a neural network to recognize text present within each bounding box and classifies the recognized text, based on at least one extracted feature of each bounding box and the recognized text present within each bounding box, according to a plurality of predefined tags. The system can associate a key with a value and return a key-value pair for each predefined tag.

Claims (31)

1 . A computer vision system for text classification comprising:

a memory; and

a processor in communication with the memory, the processor:

detecting a plurality of text regions in an image,

generating a bounding box for each detected text region,

recognizing text present within each bounding box using a neural network,

classifying the recognized text based on at least one extracted feature of each bounding box and the recognized text present within each bounding box according to a plurality of predefined tags, and

associating a key with a value and returning a key-value pair for each predefined tag.

2 . The system of claim 1 , wherein the processor detects the plurality of text regions in the image by applying an Efficient and Accurate Scene Text Detector (EAST) model to the image.

3 . The system of claim 1 , wherein the neural network is a convolutional recurrent neural network.

4 . The system of claim 1 , wherein the processor classifies the recognized text by utilizing a modified conditional random fields machine learning system implemented as a recurrent neural network and a modified graph attention network.

5 . The system of claim 1 , wherein the processor generates a graph having an adjacency list by applying a k-nearest neighbors algorithm based on a Euclidean distance between center positions of each bounding box.

6 . A method for text classification by a computer vision system, comprising the steps of:

detecting a plurality of text regions in an image;

generating a bounding box for each detected text region;

recognizing text present within each bounding box using a neural network;

classifying the recognized text, based on at least one extracted feature of each bounding box and the recognized text present within each bounding box, according to a plurality of predefined tags; and

associating a key with a value and returning a key-value pair for each predefined tag;

7 . The method of claim 6 , further comprising the step of detecting the plurality of text regions in the image by applying an Efficient and Accurate Scene Text Detector (EAST) model to the image.

8 . The method of claim 6 , wherein the neural network is a convolutional recurrent neural network.

9 . The method of claim 6 , further comprising the step of classifying the recognized text by utilizing a modified conditional random fields machine learning system implemented as a recurrent neural network and a modified graph attention network.

10 . The method of claim 6 , further comprising the step of generating a graph having an adjacency list by applying a k-nearest neighbors algorithm based on a Euclidean distance between center positions of each bounding box.

11 . A non-transitory computer readable medium having instructions stored thereon for text classification by a computer vision system which, when executed by a processor, causes the processor to carry out the steps of:

detecting a plurality of text regions in an image;

generating a bounding box for each detected text region;

recognizing text present within each bounding box using a neural network;

classifying the recognized text, based on at least one extracted feature of each bounding box and the recognized text present within each bounding box, according to a plurality of predefined tags; and

associating a key with a value and returning a key-value pair for each predefined tag.

12 . The non-transitory computer readable medium of claim 11 , the processor further carrying out the step of detecting the plurality of text regions in the image by applying an Efficient and Accurate Scene Text Detector (EAST) model to the image.

13 . The non-transitory computer readable medium of claim 11 , wherein the neural network is a convolutional recurrent neural network.

14 . The non-transitory computer readable medium of claim 11 , the processor further carrying out the step of classifying the recognized text by utilizing a modified conditional random fields machine learning system implemented as a recurrent neural network and a modified graph attention network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2024
From: NGUYEN, KHOI; SINGH, MANEESH KUMAR
To: INSURANCE SERVICES OFFICE, INC.
Reel/Frame 068134/0299 →