IP Library › Granted Patent US 11,514,238
Granted Patent B2
US 11,514,238 · App. 16/986,142 · Granted Nov 29, 2022

Automatically assigning semantic role labels to parts of documents

Inventors: Andrew Paul Begun (Redmond, WA); Steven DeRose (Silver Spring, MD); Taqi Jaffri (Kirkland, WA); Luis Marti Orosa (Las Condes, CL); Michael Palmer (Edmonds, WA); Jean Paoli (Kirkland, WA); Christina Pavlopoulou (Emeryville, CA); Elena Pricoiu (Issaquah, WA); Swagatika Sarangi (Bellevue, WA); Marcin Sawicki (Kirkland, WA); Manar Shehadeh (Kirkland, WA); Michael Taron (Seattle, WA); Bhaven Toprani (Cupertino, CA); Zubin Rustom Wadia (Chappaqua, NY); David Watson (Seattle, WA); Eric White (San Luis Obispo, CA); Joshua Yongshin Fan (Bellevue, WA); Kush Gupta (Seattle, WA); Andrew Minh Hoang (Olympia, WA); Zhanlin Liu (Seattle, WA); Jerome George Paliakkara (Seattle, WA); Zhaofeng Wu (Seattle, WA); Yue Zhang (St Paul, MN); Xiaoquan Zhou (Bellevue, WA)
Assignee: Docugami, Inc.
G06F40/186G06F16/248G06F16/2457G06F16/93G06F40/106G06F40/117G06F40/169G06F40/289G06F40/295G06F40/30G06N20/00G06V30/414G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,238
App. No.
16/986,142
Filed
Aug 5, 2020
Granted
Nov 29, 2022
Kind
B2
Art Unit
2176
USPC
704/9
Abstract

Machine learning, artificial intelligence, and other computer-implemented methods are used to identify various semantically important chunks in documents, automatically label them with appropriate datatypes and semantic roles, and use this enhanced information to assist authors and to support downstream processes. Chunk locations, datatypes, and semantic roles can often be automatically determined from what is here called “context”, to wit, the combination of their formatting, structure, and content; those of adjacent or nearby content; overall patterns of occurrence in a document, and similarities of all these things across documents (mainly but not exclusively among documents in the same document set). Similarity is not limited to exact or fuzzy string or property comparisons, but may include similarity of natural language grammatical structure, ML (machine learning) techniques such as measuring similarity of word, chunk, and other embeddings, and the datatypes and semantic roles of previously-identified chunks.

Claims (60)

1. A method implemented on a computer system executing instructions for analyzing and improving documents, the method comprising:

accessing a document set that contains a plurality of documents, wherein the document set also identifies chunks within the individual documents of the document set;

automatically assigning semantic role labels to a plurality of the chunks, wherein the semantic role labels are descriptive of the semantic roles played by the chunks in a transaction described by the document; and automatically assigning semantic role labels to the chunks (a) comprises using machine learning and/or natural language processing methods to determine semantic roles for chunks; and (b) is also based on patterns of occurrence of counterpart chunks in different documents across the document set, wherein the counterpart chunks are different chunks in different documents that play the same semantic role within their respective documents; and

using the chunks and their semantic role labels in further processing of documents in the document set.

2. The computer-implemented method of claim 1 , wherein the plurality of documents in the document set are all a same document type.

3. The computer-implemented method of claim 1 , wherein the chunks in the document set comprise:

field chunks that contain content within the documents suitable for use as fields in document templates, wherein some of the field chunks are hierarchical and contain other chunks as sub-chunks; and

structural chunks that contain content comprising structures within the layout of the documents.

4. The computer-implemented method of claim 1 , wherein the document set contains legal documents; and the semantic roles comprise (a) roles played by parties to the legal documents, and (b) roles played by dates, time periods or other expressions of time.

5. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

automatically extracting some of the semantic role labels from chunks; and

assigning the extracted semantic role labels to chunks.

6. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

using machine learning to automatically extract semantic role labels from chunks (a) based on the content, layout and contexts of chunks in individual documents; (b) based on patterns of content, layout and contexts of chunks across the documents in the document set; and (c) based on datatypes of chunks; and

assigning the extracted semantic role labels to chunks.

7. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

using auto-encoder machine learning techniques to automatically extract some of the semantic role labels; and

assigning the extracted semantic role labels to chunks.

8. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

automatically extracting candidate semantic role labels from the chunks;

using machine learning to refine the candidate semantic role labels; and

assigning the extracted semantic role labels to chunks.

9. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

automatically extracting some of the semantic role labels from chunks based on similarity of content, layout and/or context of chunks from different documents in the document set; and

assigning the extracted semantic role labels to chunks.

10. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

assigning candidate semantic role labels to chunks;

grouping chunks into clusters based on similarity of the semantic roles played by the chunks;

standardizing the candidate semantic role labels among the chunks in clusters; and

assigning the standardized semantic role labels to chunks.

11. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

assigning candidate semantic role labels to chunks;

grouping chunks into chunk clusters based on similarity of size and text embedding of the chunks;

grouping candidate semantic role labels into label clusters based on similarity of text embedding of the candidate semantic role labels;

standardizing the candidate semantic role labels based on the chunk clusters and the label clusters; and

assigning the standardized semantic role labels to chunks.

12. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises:

assigning candidate semantic role labels to chunks that comprise sections of documents, wherein the candidate semantic role labels are based on headings of the sections;

grouping the chunks into clusters based on similarity of the content in the sections;

standardizing the candidate semantic role labels by selecting the most common candidate semantic role label as the semantic role label for all chunks in a cluster; and

assigning the standardized semantic role labels to chunks.

13. The computer-implemented method of claim 1 , wherein the semantic role labels are chosen from a predetermined set of semantic role labels.

14. The computer-implemented method of claim 1 , wherein the semantic role labels comprise labels recognized by a software application used for the further processing of documents in the document set.

15. The computer-implemented method of claim 1 , wherein automatically assigning semantic role labels to chunks comprises at least one of: (a) using machine learning to determine semantic roles for chunks based on other chunks that are nearby or based on containing chunks that contain said chunks, or (b) using natural language processing methods based on grammatical structures of nearby chunks to determine semantic roles for chunks.

16. The computer-implemented method of claim 1 , wherein some of the chunks are Named Entity References, such chunks are labeled with semantic role labels for the semantic roles played by the those chunks in the documents, and such chunks are also labeled with a datatype of the chunk.

17. The computer-implemented method of claim 1 , wherein some of the chunks are multi-paragraph structures in the documents, and such chunks are labeled with semantic role labels for the semantic roles played by those chunks in the documents.

18. The computer-implemented method of claim 1 , further comprising:

estimating a confidence level for the automatically assigned semantic role labels;

based on the estimated confidence level, presenting some assignments to a user for confirmation;

receiving user feedback for the automatically assigned semantic role labels; and

improving the machine learning and/or natural language processing methods in response to the user feedback.

19. A non-transitory computer-readable storage medium storing executable computer program instructions for analyzing and improving documents, the instructions executable by a computer system and causing the computer system to perform a method comprising:

accessing a document set that contains a plurality of documents, wherein the document set also identifies chunks within the individual documents of the document set;

automatically assigning semantic role labels to a plurality of the chunks, wherein the semantic role labels are descriptive of the semantic roles played by the chunks in a transaction described by the document; and automatically assigning semantic role labels to the chunks (a) comprises using machine learning and/or natural language processing methods to determine semantic roles for chunks; and (b) is also based on patterns of occurrence of counterpart chunks in different documents across the document set, wherein the counterpart chunks are different chunks in different documents that play the same semantic role within their respective documents; and

making the chunks and their semantic role labels available for further processing of documents in the document set.

20. A computer system for analyzing and improving documents, the computer system comprising:

a storage medium for receiving and storing a document set that contains a plurality of documents, wherein the document set also identifies chunks within the individual documents of the document set; and

a processor system having access to the storage medium and executing an application program for analyzing and improving documents, wherein the processor system executing the application program:

automatically assigns semantic role labels to a plurality of the chunks, wherein the semantic role labels are descriptive of the semantic roles played by the chunks in a transaction described by the document; and automatically assigning semantic role labels to the chunks (a) comprises using machine learning and/or natural language processing methods to determine semantic roles for chunks; and (b) is also based on patterns of occurrence of counterpart chunks in different documents across the document set, wherein the counterpart chunks are different chunks in different documents that play the same semantic role within their respective documents; and

makes the chunks and their semantic role labels available for further processing of documents in the document set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2020
From: BEGUN, ANDREW PAUL; DEROSE, STEVEN; JAFFRI, TAQI; OROSA, LUIS MARTI; PALMER, MICHAEL B.; PAOLI, JEAN; PAVLOPOULOU, CHRISTINA; PRICOIU, ELENA; SARANGI, SWAGATIKA; SAWICKI, MARCIN; SHEHADEH, MANAR; TARON, MICHAEL; TOPRANI, BHAVEN; WADIA, ZUBIN RUSTOM; WATSON, DAVID; WHITE, ERIC; FAN, JOSHUA YONGSHIN; GUPTA, KUSH; HOANG, ANDREW MINH; LIU, ZHANLIN; PALIAKKARA, JEROME GEORGE; WU, ZHAOFENG; ZHANG, YUE; ZHOU, XIAOQUAN
To: DOCUGAMI, INC.
Reel/Frame 053505/0153 →
Continuity (3)
Continuation PCTUS2020043606 · Jul 24, 2020
Provisional Application 62900793 · Sep 16, 2019
Related Publication 20210081613A1 · Mar 18, 2021
Cited By (3)
US 12,405,970 US 12,511,925 US 12,525,047