IP Library › Granted Patent US 12,657,379
Granted Patent B2
US 12,657,379 · App. 18/609,740 · Granted Jun 16, 2026

Cross-document intelligent authoring and processing

Inventors: Andrew Paul Begun (Redmond, WA); Steven DeRose (Silver Spring, MD); Taqi Jaffri (Kirkland, WA); Luis Marti Orosa (Las Condes, CL); Michael B. Palmer (Edmonds, WA); Jean Paoli (Kirkland, WA); Christina Pavlopoulou (Emeryville, CA); Elena Pricoiu (Issaquah, WA); Swagatika Sarangi (Bellevue, WA); Marcin Sawicki (Kirkland, WA); Manar Shehadeh (Kirkland, WA); Michael Taron (Seattle, WA); Bhaven Toprani (Cupertino, CA); Zubin Rustom Wadia (Chappaqua, NY); David Watson (Seattle, WA); Eric White (San Luis Obispo, CA); Joshua Yongshin Fan (Bellevue, WA); Kush Gupta (Seattle, WA); Andrew Minh Hoang (Olympia, WA); Zhanlin Liu (Seattle, WA); Jerome George Paliakkara (Seattle, WA); Zhaofeng Wu (Seattle, WA); Yue Zhang (St Paul, MN); Xiaoquan Zhou (Bellevue, WA)
Assignee: Docugami, Inc.
G06F40/186G06F16/2457G06F16/248G06F16/93G06F40/106G06F40/117G06F40/169G06F40/289G06F40/295G06F40/30G06N20/00G06V30/414G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,379
App. No.
18/609,740
Granted
Jun 16, 2026
Kind
B2
Abstract

Machine learning, artificial intelligence, and other computer-implemented methods are used to identify various semantically important chunks in documents, automatically label them with appropriate datatypes and semantic roles, and use this enhanced information to assist authors and to support downstream processes. Chunk locations, datatypes, and semantic roles can often be automatically determined from what is here called “context”, to wit, the combination of their formatting, structure, and content; those of adjacent or nearby content; overall patterns of occurrence in a document, and similarities of all these things across documents (mainly but not exclusively among documents in the same document set). Similarity is not limited to exact or fuzzy string or property comparisons, but may include similarity of natural language grammatical structure, ML (machine learning) techniques such as measuring similarity of word, chunk, and other embeddings, and the datatypes and semantic roles of previously-identified chunks.

Claims (45)

1 . A method implemented on a computer system executing instructions for labeling chunks in different types of documents, the method comprising:

receiving a plurality of documents containing text, the documents uploaded according to an original file organization;

grouping the plurality of documents into sets of different types of documents;

creating a second file organization based on the grouping into sets, wherein the second file organization is different than the original file organization;

for multiple sets of documents, processing the documents in the sets using information from the multiple sets and from both the original file organization and the second file organization, the processing further comprising:

performing linguistic analysis of the text of the documents;

identifying structures and patterns in the text within documents and in the text across documents, based on the linguistic analysis;

extracting semantic role labels from the text of the documents, based on the identified structures and patterns; and

annotating chunks of text from the documents with the semantic role labels;

wherein the processing is automatically adjusted for different sets of documents based on the type of documents in that set.

2 . The method of claim 1 , wherein performing linguistic analysis and identifying structures and patterns in the text is different for different sets of documents.

3 . The method of claim 2 , wherein the extracted semantic role labels are different for different sets of documents.

4 . The method of claim 1 , wherein chunks of text are annotated with semantic role labels extracted from nearby text.

5 . The method of claim 1 , further comprising: improving the grouping of documents into sets, based on the extracted semantic role labels and chunks for different documents and sets.

6 . The method of claim 1 , wherein improving the grouping of documents into sets is based on the extracted semantic role labels and chunks but ignores the particular contents of chunks.

7 . The method of claim 1 , wherein improving the grouping of documents comprises: re-grouping the plurality of documents into sets of different types of documents.

8 . The method of claim 7 , wherein grouping the plurality of documents into sets comprises: clustering the documents based on previously detected text content, layout information, and structural information.

9 . The method of claim 1 , wherein grouping the plurality of documents into sets comprises:

automatically grouping the documents into sets; and

receiving user feedback to check the grouping of documents into sets.

10 . The method of claim 1 , wherein grouping the plurality of documents into sets of different types of documents comprises: automatically naming the different sets based on the type of document in that set.

11 . The method of claim 1 , wherein grouping the plurality of documents into sets facilitates machine learning and/or reasoning about the documents.

12 . The method of claim 1 , wherein grouping the plurality of documents into sets facilitates machine learning and/or reasoning about differences between documents in the same set.

13 . A computer system for labeling chunks in different types of documents, the computer system comprising a storage medium for storing computer program instructions; and a processor system having access to the storage medium, wherein executing the computer program instructions causes the processor system to:

receive a plurality of documents containing text, the documents uploaded according to an original file organization;

group the plurality of documents into sets of different types of documents;

create a second file organization based on the grouping into sets, wherein the second file organization is different than the original file organization;

for multiple sets of documents, process the documents in the sets using information from the multiple sets and from both the original file organization and the second file organization, the processing further comprising:

performing linguistic analysis of the text of the documents;

identifying structures and patterns in the text within documents and in the text across documents, based on the linguistic analysis;

extracting semantic role labels from the text of the documents, based on the identified structures and patterns; and

annotating chunks of text from the documents with the semantic role labels;

wherein the processing is automatically adjusted for different sets of documents based on the type of documents in that set.

14 . The method of claim 1 , wherein the processing further comprises:

arbitrating among the chunks to produce hierarchies of the chunks that are well-formed without any partial overlap of the chunks.

15 . A non-transitory computer-readable storage medium storing executable computer program instructions for labeling chunks in different types of documents, the instructions executable by a computer system and causing the computer system to:

receive a plurality of documents containing text, the documents uploaded according to an original file organization;

group the plurality of documents into sets of different types of documents;

create a second file organization based on the grouping into sets, wherein the second file organization is different than the original file organization;

for multiple sets of documents, process the documents in the sets using information from the multiple sets and from both the original file organization and the second file organization, the processing further comprising:

performing linguistic analysis of the text of the documents;

identifying structures and patterns in the text within documents and in the text across documents, based on the linguistic analysis;

extracting semantic role labels from the text of the documents, based on the identified structures and patterns; and

annotating chunks of text from the documents with the semantic role labels;

wherein the processing is automatically adjusted for different sets of documents based on the type of documents in that set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2024
From: BEGUN, ANDREW PAUL; DEROSE, STEVEN; JAFFRI, TAQI; OROSA, LUIS MARTI; PALMER, MICHAEL B.; PAOLI, JEAN; PAVLOPOULOU, CHRISTINA; PRICOIU, ELENA; SARANGI, SWAGATIKA; SAWICKI, MARCIN; SHEHADEH, MANAR; TARON, MICHAEL; TOPRANI, BHAVEN; WADIA, ZUBIN RUSTOM; WATSON, DAVID; WHITE, ERIC; FAN, JOSHUA YONGSHIN; GUPTA, KUSH; HOANG, ANDREW MINH; LIU, ZHANLIN; PALIAKKARA, JEROME GEORGE; WU, ZHAOFENG; ZHANG, YUE; ZHOU, XIAOQUAN
To: DOCUGAMI, INC.
Reel/Frame 066841/0386 →
Continuity (5)
Continuation 17724934 · Apr 20, 2022
Continuation 16986136 · Aug 5, 2020
Continuation PCTUS2020043606 · Jul 24, 2020
Provisional Application 62900793 · Sep 16, 2019
Related Publication 20240232518A1 · Jul 11, 2024
References Cited (20)
US 9880988B2 · Vanderwende · 2018 [cited by examiner]
US 20050060643A1 · Glass · 2005 [cited by examiner]
US 20130021346A1 · Terman · 2013 [cited by applicant]
US 20140324808A1 · Sandhu · 2014 [cited by examiner]
US 20150088888A1 · Brennan · 2015 [cited by examiner]
US 20150178853A1 · Byron et al. · 2015 [cited by applicant]
US 20150261743A1 · Sengupta et al. · 2015 [cited by applicant]
US 20160232204A1 · Zholudev · 2016 [cited by examiner]
US 20200175041A1 · Ramachandra · 2020 [cited by examiner]
CN 101681348A · 2010 [cited by applicant]
CN 106295706A · 2017 [cited by applicant]
CN 109582949B · 2019 [cited by applicant]
JP 2005266903A · 2005 [cited by applicant]
JP 2007094855A · 2007 [cited by applicant]
JP 2017004074A · 2017 [cited by applicant]
Chinese Patent Office, Office Action, CN Patent Application No. 202080064610.1, Jan. 10, 2025, 7 pages. [cited by applicant]
Chinese Patent Office, Office Action, CN Patent Application No. 202080064610.1, Jun. 23, 2025, 28 pages. [cited by applicant]
European Patent Office, European Search Report and Opinion, European Patent Application No. 20864772.7, Jan. 22, 2025, seven pages. [cited by applicant]
Japanese Patent Office, Office Action, JP Patent Application No. 2022-542307, Mar. 11, 2025, six pages. [cited by applicant]
Rahman, M et al., “Understanding the Logical and Semantic Structure of Large Documents”, Arxiv.org, Cornell University Library, Sep. 3, 2017, pp. 1-10. [cited by applicant]