IP Library Granted Patent US 11,789,990
Granted Patent B1
US 11,789,990 · App. 17/733,581 · Granted Oct 17, 2023

Automated splitting of document packages and identification of relevant documents

Inventors: Meena AbdelMaseeh Adly Fouad (Mississauga, CA); Zhihong Zeng (Acton, MA); Anirudh Prabakaran (Atlanta, GA); Samriddhi Shakya (Washington, DC); Tom Sebastian (Bangalore, IN); Tallam Sai Teja (Andhra Pradesh, IN); Simon Ioffe (Boston, MA); Narasimha Goli (Boston, MA)
Assignee: Iron Mountain Incorporated
G06F16/35G06F16/383G06N3/045G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,789,990
App. No.
17/733,581
Granted
Oct 17, 2023
Kind
B1
Abstract

A process for document processing (e.g., automated package splitting) may involve producing, for each document page of an ordered plurality of document pages, an image of the document page and a representation of text from the document page; generating, for each document page of the ordered plurality, and based on the image of the document page and the representation of text from the document page, an embedding of the document page; and generating, for each document page among a subset of the ordered plurality, a label for the document page that indicates whether the document page is a document first page, based on the embedding of the document page, the embedding of each of at least one document page that precedes the document page in the ordered plurality, and the embedding of each of at least one document page that follows the document page in the ordered plurality.

Claims (44)

1. A computer-implemented method for processing an ordered plurality of document pages, the method comprising:

producing, for each document page of the ordered plurality of document pages, an image of the document page and a representation of text from the document page;

generating, for each document page of the ordered plurality of document pages, and based on the image of the document page and the representation of text from the document page, an embedding of the document page; and

generating, for each document page among a subset of the ordered plurality of document pages, a label for the document page that indicates whether the document page is a document first page,

wherein, for each document page among the subset of the ordered plurality of document pages, generating the label for the document page is based on the embedding of the document page, the embedding of each of at least one document page that precedes the document page in the ordered plurality of document pages, and the embedding of each of at least one document page that follows the document page in the ordered plurality of document pages.

2. The computer-implemented method of claim 1 , wherein the subset of the ordered plurality of document pages does not include an initial page of the ordered plurality of document pages and does not include a final page of the ordered plurality of document pages.

3. The computer-implemented method of claim 1 , wherein, for each document page among the subset of the ordered plurality of document pages, generating the label for the document page is based on:

the embedding of each of a plurality of document pages that precede the document page in a document package, and

the embedding of each of a plurality of document pages that follow the document page in the document package.

4. The computer-implemented method of claim 1 , wherein generating, for each document page of the ordered plurality of document pages, the embedding of the document page is based on an embedding of a context of recognized words from the document page.

5. The computer-implemented method of claim 1 , wherein generating, for each document page of the ordered plurality of document pages, the embedding of the document page is based on a graph of word bounding boxes in the page.

6. The computer-implemented method of claim 1 , wherein generating, for each document page of the ordered plurality of document pages, the embedding of the document page is based on an embedding of a formatting of text within the document page.

7. The computer-implemented method of claim 1 , wherein generating, for each document page of the ordered plurality of document pages, the embedding of the document page is based on a statistical distribution of pixel values of the image of the document page.

8. The computer-implemented method of claim 1 , wherein generating the label for the document page includes:

generating, by a first recurrent neural network and based on the embedding of each of the at least one document page that precedes the document page in the ordered plurality of document pages, a backward prediction; and

generating, by a second recurrent neural network and based on the embedding of each of the at least one document page that follows the document page in the ordered plurality of document pages, and

wherein the label is based on the backward prediction and on a forward prediction.

9. The computer-implemented method of claim 1 , wherein generating, for each document page among the subset of the ordered plurality of document pages, the label for the document page comprises generating a vector, each element of the vector corresponding to a respective document page in the ordered plurality and indicating the label of the document page.

10. The computer-implemented method of claim 1 , wherein the labels for the document pages among the subset of the ordered plurality of document pages indicate a division of the ordered plurality of document pages into a plurality of individual documents.

11. The computer-implemented method of claim 10 , wherein the method further comprises, for each of the plurality of individual documents, classifying the document based on the page embeddings of document pages of the document.

12. The computer-implemented method of claim 1 , the method further comprising, based on the page embeddings of the document pages of the ordered plurality of document pages and on the labels for the document pages among the subset of the ordered plurality of document pages, classifying each of a plurality of documents within the ordered plurality of document pages.

13. A document package processing system, the system comprising:

one or more processing devices; and

one or more non-transitory computer-readable media communicatively coupled to the one or more processing devices, wherein the one or more processing devices are configured to execute program code stored in the non-transitory computer-readable media and thereby perform operations for processing an ordered plurality of document pages, the operations comprising:

producing, for each document page of the ordered plurality of document pages, an image of the document page and a representation of text from the document page;

generating, for each document page of the ordered plurality of document pages, and based on the image of the document page and the representation of text from the document page, an embedding of the document page; and

generating, for each document page among a subset of the ordered plurality of document pages, a label for the document page that indicates whether the document page is a document first page,

wherein, for each document page among the subset of the ordered plurality of document pages, generating the label for the document page is based on the embedding of the document page, the embedding of each of at least one document page that precedes the document page in the ordered plurality of document pages, and the embedding of each of at least one document page that follows the document page in the ordered plurality of document pages.

14. The document package processing system of claim 13 , wherein, for each document page among the subset of the ordered plurality of document pages, generating the label for the document page is based on:

the embedding of each of a plurality of document pages that precede the document page in a document package, and

the embedding of each of a plurality of document pages that follow the document page in the document package.

15. The document package processing system of claim 13 , wherein generating, for each document page of the ordered plurality of document pages, the embedding of the document page is based on a graph of word bounding boxes in the page.

16. The document package processing system of claim 13 , wherein generating the label for the document page includes:

generating, by a first recurrent neural network and based on the embedding of each of the at least one document page that precedes the document page in the ordered plurality of document pages, a backward prediction; and

generating, by a second recurrent neural network and based on the embedding of each of the at least one document page that follows the document page in the ordered plurality of document pages, and

wherein the label is based on the backward prediction and on a forward prediction.

17. The document package processing system of claim 13 , wherein the labels for the document pages among the subset of the ordered plurality of document pages indicate a division of the ordered plurality of document pages into a plurality of individual documents.

18. The document package processing system of claim 17 , wherein the operations further comprise, for each of the plurality of individual documents, classifying the document based on the page embeddings of document pages of the document.

19. One or more non-transitory computer-readable media storing computer-executable instructions to cause a computer to perform operations for processing an ordered plurality of document pages, the operations comprising:

producing, for each document page of the ordered plurality of document pages, an image of the document page and a representation of text from the document page;

generating, for each document page of the ordered plurality of document pages, and based on the image of the document page and the representation of text from the document page, an embedding of the document page; and

generating, for each document page among a subset of the ordered plurality of document pages, a label for the document page that indicates whether the document page is a document first page,

wherein, for each document page among the subset of the ordered plurality of document pages, generating the label for the document page is based on the embedding of the document page, the embedding of each of at least one document page that precedes the document page in the ordered plurality of document pages, and the embedding of each of at least one document page that follows the document page in the ordered plurality of document pages.

20. The one or more non-transitory computer-readable media of claim 19 , the operations further comprising, based on the page embeddings of the document pages of the ordered plurality of document pages and on the labels for the document pages among the subset of the ordered plurality of document pages, classifying each of a plurality of documents within the ordered plurality of document pages.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2023
From: FOUAD, MEENA ABDELMASEEH ADLY; ZENG, ZHIHONG; PRABAKARAN, ANIRUDH; SHAKYA, SAMRIDDHI; SEBASTIAN, TOM; TEJA, TALLAM SAI; IOFFE, SIMON; GOLI, NARASIMHA
To: IRON MOUNTAIN INCORPORATED
Reel/Frame 062449/0339 →
Cited By (3)
US 12,265,787 US 12,437,569 US 12,608,537