IP Library › Granted Patent US 11,392,763
Granted Patent B2
US 11,392,763 · App. 16/986,136 · Granted Jul 19, 2022

Cross-document intelligent authoring and processing, including format for semantically-annotated documents

Inventors: Andrew Paul Begun (Redmond, WA); Steven DeRose (Silver Spring, MD); Taqi Jaffri (Kirkland, WA); Luis Marti Orosa (Las Condes, CL); Michael Palmer (Edmonds, WA); Jean Paoli (Kirkland, WA); Christina Pavlopoulou (Emeryville, CA); Elena Pricoiu (Issaquah, WA); Swagatika Sarangi (Bellevue, WA); Marcin Sawicki (Kirkland, WA); Manar Shehadeh (Kirkland, WA); Michael Taron (Seattle, WA); Bhaven Toprani (Cupertino, CA); Zubin Rustom Wadia (Chappaqua, NY); David Watson (Seattle, WA); Eric White (San Luis Obispo, CA); Joshua Yongshin Fan (Bellevue, WA); Kush Gupta (Seattle, WA); Andrew Minh Hoang (Olympia, WA); Zhanlin Liu (Seattle, WA); Jerome George Paliakkara (Seattle, WA); Zhaofeng Wu (Seattle, WA); Yue Zhang (St Paul, MN); Xiaoquan Zhou (Bellevue, WA)
Assignee: DOCUGAMI, INC.
G06F40/186G06F16/248G06F16/2457G06F16/93G06F40/106G06F40/117G06F40/169G06F40/289G06F40/295G06F40/30G06N20/00G06V30/414G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,392,763
App. No.
16/986,136
Granted
Jul 19, 2022
Kind
B2
Abstract

Machine learning, artificial intelligence, and other computer-implemented methods are used to identify various semantically important chunks in documents, automatically label them with appropriate datatypes and semantic roles, and use this enhanced information to assist authors and to support downstream processes. Chunk locations, datatypes, and semantic roles can often be automatically determined from what is here called “context”, to wit, the combination of their formatting, structure, and content; those of adjacent or nearby content; overall patterns of occurrence in a document, and similarities of all these things across documents (mainly but not exclusively among documents in the same document set). Similarity is not limited to exact or fuzzy string or property comparisons, but may include similarity of natural language grammatical structure, ML (machine learning) techniques such as measuring similarity of word, chunk, and other embeddings, and the datatypes and semantic roles of previously-identified chunks.

Claims (47)

1. A method implemented on a computer system executing instructions for processing documents, the method comprising:

processing a document set that contains a plurality of documents to identify chunks in the documents and to generate corresponding annotations, comprising stages of:

processing images of the documents to identify visual chunks that comprise visually distinct regions of the images of the documents; and generating first annotations that specify spacing and formatting of the visual chunks, the first annotations including signatures of bitmap tiles of the images of the documents;

processing the visual chunks and first annotations to identify structural chunks that contain content from structures within the visual chunks; and generating second annotations that specify layout of the structural chunks;

processing the structural chunks and second annotations to identify topic-level chunks based on a grouping of content in structural chunks according to topic; and generating third annotations that specify topics of the topic-level chunks; and

processing the topic-level chunks and third annotations to identify field chunks comprising content suitable for use as fields in document templates; and generating fourth annotations that specify the fields of the field chunks; and

wherein at least some of the processing is performed in parallel by the computer system to identify the chunks and annotations; and

arbitrating among the identified chunks to produce a well-formed hierarchy of chunks, wherein arbitrating among the chunks comprises modifying some chunks to avoid overlap of the chunks within the well-formed hierarchy; and

generating representations of the processed documents in a format comprising the well-formed hierarchy of chunks that includes the field chunks and at least some of the other identified chunks from the documents and corresponding annotations for the chunks; and

making the representations in the format available for use by any of a plurality of software applications in downstream processes.

2. A method implemented on a computer system executing instructions for processing documents, the method comprising:

processing a document set that contains a plurality of documents to identify chunks in the documents and to generate corresponding annotations, comprising stages of:

processing images of the documents to identify visual chunks that comprise visually distinct regions of the images of the documents; and generating first annotations that specify spacing and formatting of the visual chunks, the first annotations including signatures of bitmap tiles of the images of the documents;

processing the visual chunks and first annotations to identify structural chunks that contain content from structures within the visual chunks; and generating second annotations that specify layout of the structural chunks, the structural chunks including hyperlines of text;

generating signatures of the hyperlines within the structural chunks and renesting the structural chunks based on the signatures to create a nested structure of structural chunks; and

processing the nested structure of structural chunks and second annotations to identify topic-level chunks based on a grouping of content in structural chunks according to topic; and generating third annotations that specify topics of the topic-level chunks; and

processing the topic-level chunks and third annotations to identify field chunks comprising content suitable for use as fields in document templates; and generating fourth annotations that specify the fields of the field chunks;

generating representations of the processed documents in a format comprising the field chunks and at least some of the other identified chunks from the documents and corresponding annotations for the chunks; and

making the representations in the format available for use by any of a plurality of software applications in downstream processes.

3. A method implemented on a computer system executing instructions for processing documents, the method comprising:

processing a document set that contains a plurality of documents to identify chunks in the documents and to generate corresponding annotations, comprising stages of:

processing images of the documents to identify visual chunks that comprise visually distinct regions of the images of the documents; and generating first annotations that specify spacing and formatting of the visual chunks, the first annotations including signatures of bitmap tiles of the images of the documents;

processing the visual chunks and first annotations to identify structural chunks that contain content from structures within the visual chunks; and generating second annotations that specify layout of the structural chunks;

processing the structural chunks and second annotations to identify topic-level chunks based on a grouping of content in structural chunks according to topic; and generating third annotations that specify topics of the topic-level chunks; and

processing the topic-level chunks and third annotations to identify field chunks comprising content suitable for use as fields in document templates; and generating fourth annotations that specify the fields of the field chunks;

generating representations of the processed documents in a format comprising the field chunks and at least some of the other identified chunks from the documents and corresponding annotations for the chunks; and

making the representations in the format available for use by any of a plurality of software applications in downstream processes.

4. The computer-implemented method of claim 3 , wherein the representations of the processed documents comprise all of the chunks identified in processing the documents and all of the corresponding annotations generated in processing the documents.

5. The computer-implemented method of claim 3 , wherein each of the stages of processing the documents uses machine learning, artificial intelligence and/or natural language processing.

6. The computer-implemented method of claim 3 , wherein each of the stages of processing the documents identifies chunks with less than 100% confidence.

7. The computer-implemented method of claim 6 , wherein the representations of the processed documents further comprise annotations specifying confidence levels for the identification of chunks.

8. The computer-implemented method of claim 6 , further comprising:

receiving user corrections for incorrectly identified chunks; and

improving the stages of automatically identifying chunks in response to the user corrections.

9. The computer-implemented method of claim 3 , wherein the stages of processing visual chunks, processing structural chunks and processing topic-level chunks is performed recursively for visual chunks contained within other visual chunks.

10. The computer-implemented method of claim 3 , wherein the representations of the processed documents further comprise annotations for the datatypes and semantic role labels of a plurality of the chunks, wherein the semantic role labels are descriptive of the semantic roles played by the chunks.

11. The computer-implemented method of claim 3 , wherein some higher-level chunks contain other lower-level chunks as sub-chunks, and the representations of the processed documents further comprise annotations specifying containment of lower-level chunks in higher-level chunks.

12. The computer-implemented method of claim 3 , wherein some chunks have a hierarchical relationship, and the representations of the processed documents further comprise annotations specifying hierarchical relationships between chunks.

13. The computer-implemented method of claim 3 , wherein the chunks in the representations of the processed documents comprise a plurality of sections, headings, lists, items, markers, and/or Named Entities at multiple different levels.

14. The computer-implemented method of claim 3 , wherein the plurality of documents in the document set are all a same document type.

15. The computer-implemented method of claim 3 , further comprising:

assembling the document set by clustering documents into the document set based on similarity of content and/or layout.

16. The computer-implemented method of claim 3 , wherein representations of the processed documents are in an XML format.

17. The computer-implemented method of claim 3 , wherein the representations of the processed documents further comprise annotations for locations of chunks implemented using digital signatures.

18. The computer-implemented method of claim 3 , wherein the documents have original layouts, and the representations of the processed documents contain sufficient information to reconstruct the documents with the original layouts.

19. The computer-implemented method of claim 3 , wherein the plurality of software applications comprise software applications with a user interface for a user to create, edit and/or review the representations of the processed documents.

20. The computer-implemented method of claim 3 , wherein the format is a standardized, published format.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2020
From: BEGUN, ANDREW PAUL; DEROSE, STEVEN; JAFFRI, TAQI; OROSA, LUIS MARTI; PALMER, MICHAEL B.; PAOLI, JEAN; PAVLOPOULOU, CHRISTINA; PRICOIU, ELENA; SARANGI, SWAGATIKA; SAWICKI, MARCIN; SHEHADEH, MANAR; TARON, MICHAEL; TOPRANI, BHAVEN; WADIA, ZUBIN RUSTOM; WATSON, DAVID; WHITE, ERIC; FAN, JOSHUA YONGSHIN; GUPTA, KUSH; HOANG, ANDREW MINH; LIU, ZHANLIN; PALIAKKARA, JEROME GEORGE; WU, ZHAOFENG; ZHANG, YUE; ZHOU, XIAOQUAN
To: DOCUGAMI, INC.
Reel/Frame 053505/0153 →
Continuity (3)
Continuation PCTUS2020043606 · Jul 24, 2020
Provisional Application 62900793 · Sep 16, 2019
Related Publication 20210081601A1 · Mar 18, 2021
Cited By (2)
US 12,602,376 US 12,602,378