Cross-Document Intelligent Authoring and Processing
Machine learning, artificial intelligence, and other computer-implemented methods are used to identify various semantically important chunks in documents, automatically label them with appropriate datatypes and semantic roles, and use this enhanced information to assist authors and to support downstream processes. Chunk locations, datatypes, and semantic roles can often be automatically determined from what is here called “context”, to wit, the combination of their formatting, structure, and content; those of adjacent or nearby content; overall patterns of occurrence in a document, and similarities of all these things across documents (mainly but not exclusively among documents in the same document set). Similarity is not limited to exact or fuzzy string or property comparisons, but may include similarity of natural language grammatical structure, ML (machine learning) techniques such as measuring similarity of word, chunk, and other embeddings, and the datatypes and semantic roles of previously-identified chunks.
1 . A method implemented on a computer system, the method comprising:
accessing a document set that contains a plurality of documents, chunks within the individual documents of the document set, and semantic role labels for some of the chunks, wherein the semantic role labels are descriptive of the semantic roles played by the chunks within their respective documents;
performing linguistic analysis of the document set;
providing a user interface for a user to provide user instructions to develop a target document; and
automatically generating suggestions to develop the target document based on the linguistic analysis of the document set, and according to the user instructions provided through the user interface.
2 . The method of claim 1 , wherein automatically generating suggestions to develop the target document comprises:
generating embeddings for chunks within the target document;
identifying semantic role labels for the chunks within the target document based on similarities between the embeddings for the chunks within the target document and embeddings of the chunks within the document set; and
automatically generating the suggestions based on the identified semantic role labels.
3 . The method of claim 2 , wherein identifying semantic role labels for the chunks within the target document comprises:
clustering chunks based on their embeddings; and
ranking candidate semantic role labels for the clusters using a weighted PageRank algorithm.
4 . The method of claim 1 , wherein the chunks are identified based on extracting n-grams from documents.
5 . The method of claim 1 , wherein automatically generating the suggestions comprises:
identifying counterpart chunks missing in the target document; and
automatically generating suggestions for the missing counterpart chunks, based on counterpart chunks in other documents.
6 . The method of claim 1 , wherein automatically generating the suggestions comprises:
automatically generating suggestions based on text from other documents, but replacing some sub-chunks with sub-chunks specific to the target document.
7 . The method of claim 1 , wherein the linguistic analysis comprises using embeddings and clustering to produce semantic role labels for headings in a document.
8 . A system comprising:
a data store configured to store a document set that contains a plurality of documents, chunks within the individual documents of the document set, and semantic role labels for some of the chunks, wherein the semantic role labels are descriptive of the semantic roles played by the chunks within their respective documents;
a user interface configured for a user to provide user instructions for developing a target document; and
one or more learning models configured to perform linguistic analysis of the document set and of the target document and to automatically generate suggestions to develop the target document based on the linguistic analysis and on the user instructions; wherein the user interface is further configured to present the suggestions to the user.
9 . The system of claim 8 wherein the one or more learning models comprises:
a question-answering model that is trained to predict semantic role labels for chunks based on contents of the chunks.
10 . The system of claim 8 wherein the one or more learning models comprises:
a document language model that is trained on textual content and visual characteristics of documents including at least one of font, geometry, spacing, and structural information.
11 . The system of claim 8 wherein the one or more learning models comprises:
a sequence-to-sequence neural model to reconstruct a hierarchical structure from a flat sequence of text units.
12 . The system of claim 11 wherein, wherein the sequence-to-sequence neural model is pretrained on a publicly available dataset and then fine-tuned based on user feedback.
13 . The system of claim 8 wherein at least one of the learning models is fine-tuned based on user feedback.
14 . The system of claim 13 wherein at least one of the learning models is fine-tuned based on user feedback using a few-shot learning technique.
15 . The system of claim 8 wherein the one or more learning models comprises:
a neural network model that learns from sentence parses what texts are likely semantic role labels.
16 . The system of claim 8 wherein the one or more learning models comprises:
an autoencoder that identifies visual events in document layouts.
17 . The system of claim 8 further comprising:
a dispatcher component that receives user feedback and provides the user feedback to the relevant learning model for training.
18 . The system of claim 8 wherein the one or more learning models comprises:
a language model of n-grams trained on expected words, wherein the language model identifies n-grams in the target document that do not fit the language model.
19 . The system of claim 8 wherein the system is used by multiple groups of users, and the system keeps private confidential data and models of each group, but shares among the groups general learning based on non-confidential data and models.
20 . The system of claim 19 wherein the system supports fleet querying.