IP Library Granted Patent US 11,960,816
Granted Patent B2
US 11,960,816 · App. 17/577,793 · Granted Apr 16, 2024

Automatic document generation and segmentation system

Inventors: James M. Kukla (Ellicott City, MD); Maryam Esmaeilkhanian (Ellicott City, MD)
Assignee: RedShred LLC
G06F40/103G06F40/166G06N3/08G06V30/412
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,960,816
App. No.
17/577,793
Granted
Apr 16, 2024
Kind
B2
Abstract

Methods and systems are provided for generating a corpus of documents from an original document or document corpus. Original documents are processed to extract document layout and elements. Elements are clustered and processed by neural networks, such as GANs, to generate additional elements that are then combined with the extracted layout to produce new documents, such as for training automated document processing systems.

Claims (46)

1. A method of generating a document corpus, the method comprising:

receiving an original document, wherein the original document is a restricted-access document having security requirements that restrict access to the original document;

separating the original document into a plurality of pages;

for a first page of the plurality of pages, performing a first sub-process comprising:

converting the first page to an image;

determining a layout of the first page, the layout indicating a location of a first element of the first page;

assigning an element type to the first element; and

based upon the element type, assigning the first element to an element cluster;

based upon the element cluster, generating a plurality of generated elements of the same type as the first element;

based upon the layout and the plurality of generated elements, generating a plurality of generated documents that include the generated elements, wherein the plurality of generated documents are not restricted-access documents and wherein no generated document duplicates the entirety of the original document; and

training a machine learning system based upon the plurality of generated documents.

2. The method of claim 1 , wherein the layout indicates locations of a plurality of elements of the first page.

3. The method of claim 1 , further comprising performing the sub-process for a set of pages of the plurality of pages.

4. The method of claim 3 , wherein the set of pages includes all pages in the plurality of pages.

5. The method of claim 3 , wherein the plurality of generated documents is based on a plurality of layouts, each of the plurality of layouts determined for a page in the set of pages.

6. The method of claim 1 , wherein the element type is represented by a color.

7. The method of claim 1 , wherein each of the plurality of generated elements is generated by a generative adversarial network (GAN).

8. The method of claim 7 , wherein a separate GAN is used to generate elements of each type of element identified in the document.

9. The method of claim 1 , wherein the plurality of generated documents are generated based upon outputs of the separate GANs.

10. A system for generating a document corpus, the system comprising:

a computer-readable data storage storing a plurality of documents; and

a computer processor configured to:

obtain an original document from the plurality of documents, wherein the original document is a restricted-access document having security requirements that restrict access to the original document;

separate the original document into a plurality of pages;

for a first page of the plurality of pages, perform a first sub-process comprising:

converting the first page to an image;

determining a layout of the first page, the layout indicating a location of a first element of the first page;

assigning an element type to the first element; and

based upon the element type, assigning the first element to an element cluster;

based upon the element cluster, generate a plurality of generated elements of the same type as the first element;

based upon the layout and the plurality of generated elements, generate a plurality of generated documents that include the generated elements, wherein the plurality of generated documents are not restricted-access documents and no generated document duplicates the entirety of the original document; and

training a machine learning system based upon the plurality of generated documents.

11. The system of claim 10 , the processor further configured to train a machine learning system based upon the plurality of generated documents.

12. The system of claim 10 , wherein the layout indicates locations of a plurality of elements of the first page.

13. The system of claim 10 , wherein the plurality of generated documents is based on a plurality of layouts, each of the plurality of layouts determined for a page in the set of pages.

14. A non-transitory computer-readable medium storing a plurality of instructions which, when executed by a computer processor, cause the processor to perform a method comprising:

receiving an original document, wherein the original document is a restricted-access document having security requirements that restrict access to the original document;

separating the original document into a plurality of pages;

for a first page of the plurality of pages, performing a first sub-process comprising:

converting the first page to an image;

determining a layout of the first page, the layout indicating a location of a first element of the first page;

assigning an element type to the first element; and

based upon the element type, assigning the first element to an element cluster;

based upon the element cluster, generating a plurality of generated elements of the same type as the first element;

based upon the layout and the plurality of generated elements, generating a plurality of generated documents that include the generated elements, wherein the plurality of generated documents are not restricted-access documents and wherein no generated document duplicates the entirety of the original document; and

training a machine learning system based upon the plurality of generated documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2024
From: KUKLA, JAMES; ESMAEILKHANIAN, MARYAM
To: REDSHRED LLC
Reel/Frame 066217/0749 →
Continuity (2)
Provisional Application 63138074 · Jan 15, 2021
Related Publication 20220229969A1 · Jul 21, 2022
Cited By (1)
US 12,585,414