IP Library › Granted Patent US 11,410,445
Granted Patent B2
US 11,410,445 · App. 17/060,103 · Granted Aug 9, 2022

System and method for obtaining documents from a composite file

Inventors: Akshay Uppal (Bengaluru, IN); Shreyas Sudeendra (Bengaluru, IN)
Assignee: Infrrd Inc.
G06V30/413G06F16/93G06F40/106G06F40/114G06V10/464G06F40/149
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,445
App. No.
17/060,103
Granted
Aug 9, 2022
Kind
B2
Abstract

A system for obtaining documents from a composite file comprising a stream of multiple pages is provided. The system may comprise one or more processors configured to receive the composite file comprising the multiple pages and split the composite file to obtain individual pages of the composite file, wherein image of each of the individual pages and image vector for each of the individual pages from the image of the respective page may be obtained. The processor may further obtain text present in each of the individual pages and text vector for each of the individual pages from the text of the respective page. The processor may further determine continuity pattern between pages that are consecutive based on the image vector and the text vector of the consecutive pages and may categorize the consecutive pages as belonging to the same document in case the determined continuity pattern between the consecutive pages indicate that the consecutive pages belong to the same document.

Claims (27)

1. A system for obtaining documents from a composite file comprising a stream of multiple pages, the system comprising one or more processors configured to:

receive the composite file comprising the multiple pages;

split the composite file to obtain individual pages of the composite file;

obtain image of each of the individual pages;

feed at least one of the split individual pages, into a deep learning model;

identify whether the individual page fed into the deep learning model comprises multiple documents;

if the individual page fed into the deep learning model comprises multiple documents, coordinates of at least one of the documents within the individual page are determined;

extract the identified document based on the determined coordinates;

obtain image vector for each of the individual pages from the image of the respective page;

obtain text present in each of the individual pages;

obtain text vector for each of the individual pages from the text of the respective page;

combine the image vector and the text vector of at least one of the individual page and image vector and text vector of its consecutive page to generate a page vector of the at least one of the individual pages and its consecutive page;

combine the page vector of the at least one of the individual pages and its consecutive page to generate a document vector;

determine continuity pattern between pages that are consecutive by processing the generated document vector to determine the continuity pattern across the generated document vector; and

categorize the consecutive pages as belonging to the same document in case the determined continuity pattern between the consecutive pages indicates that the consecutive pages belong to the same document.

2. The system of claim 1 , wherein the processor is configured to combine the consecutive pages on determining that the consecutive pages belong to the same document.

3. The system of claim 2 , wherein the processor is configured to:

feed at least one of the combined documents into a classification module; and

classify the document based on the text and the image layout.

4. The system of claim 1 , wherein the processor is configured to:

feed at least one of the extracted document into a classification module; and

classify the extracted document based on the text.

5. The system of claim 1 , wherein the processor is configured to reduce the dimension of the image before obtaining image vector.

6. The system of claim 1 , wherein the processor is configured to obtain the image vector of 1×256 vector size.

7. The system of claim 6 , wherein the processor is configured to obtain the text vector of 1×256 vector size.

8. The system of claim 7 , wherein the processor is configured to combine the image vector and the text vector to obtain a vector of 1×512 vector size.

9. The system of claim 8 , wherein the document vector is of a size 1×1024.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2020
From: UPPAL, AKSHAY; SUDEENDRA, SHREYAS
To: INFRRD INC
Reel/Frame 054108/0366 →
Continuity (1)
Related Publication 20210019512A1 · Jan 21, 2021