Cross-document intelligent authoring and processing
Machine learning, artificial intelligence, and other computer-implemented methods are used to identify various semantically important chunks in documents, automatically label them with appropriate datatypes and semantic roles, and use this enhanced information to assist authors and to support downstream processes. Chunk locations, datatypes, and semantic roles can often be automatically determined from what is here called “context”, to wit, the combination of their formatting, structure, and content; those of adjacent or nearby content; overall patterns of occurrence in a document, and similarities of all these things across documents (mainly but not exclusively among documents in the same document set). Similarity is not limited to exact or fuzzy string or property comparisons, but may include similarity of natural language grammatical structure, ML (machine learning) techniques such as measuring similarity of word, chunk, and other embeddings, and the datatypes and semantic roles of previously-identified chunks.
1 . A method implemented on a computer system executing instructions for labeling chunks in different types of documents, the method comprising:
receiving a plurality of documents containing text, the documents uploaded according to an original file organization;
grouping the plurality of documents into sets of different types of documents;
creating a second file organization based on the grouping into sets, wherein the second file organization is different than the original file organization;
for multiple sets of documents, processing the documents in the sets using information from the multiple sets and from both the original file organization and the second file organization, the processing further comprising:
performing linguistic analysis of the text of the documents;
identifying structures and patterns in the text within documents and in the text across documents, based on the linguistic analysis;
extracting semantic role labels from the text of the documents, based on the identified structures and patterns; and
annotating chunks of text from the documents with the semantic role labels;
wherein the processing is automatically adjusted for different sets of documents based on the type of documents in that set.
2 . The method of claim 1 , wherein performing linguistic analysis and identifying structures and patterns in the text is different for different sets of documents.
3 . The method of claim 2 , wherein the extracted semantic role labels are different for different sets of documents.
4 . The method of claim 1 , wherein chunks of text are annotated with semantic role labels extracted from nearby text.
5 . The method of claim 1 , further comprising: improving the grouping of documents into sets, based on the extracted semantic role labels and chunks for different documents and sets.
6 . The method of claim 1 , wherein improving the grouping of documents into sets is based on the extracted semantic role labels and chunks but ignores the particular contents of chunks.
7 . The method of claim 1 , wherein improving the grouping of documents comprises: re-grouping the plurality of documents into sets of different types of documents.
8 . The method of claim 7 , wherein grouping the plurality of documents into sets comprises: clustering the documents based on previously detected text content, layout information, and structural information.
9 . The method of claim 1 , wherein grouping the plurality of documents into sets comprises:
automatically grouping the documents into sets; and
receiving user feedback to check the grouping of documents into sets.
10 . The method of claim 1 , wherein grouping the plurality of documents into sets of different types of documents comprises: automatically naming the different sets based on the type of document in that set.
11 . The method of claim 1 , wherein grouping the plurality of documents into sets facilitates machine learning and/or reasoning about the documents.
12 . The method of claim 1 , wherein grouping the plurality of documents into sets facilitates machine learning and/or reasoning about differences between documents in the same set.
13 . A computer system for labeling chunks in different types of documents, the computer system comprising a storage medium for storing computer program instructions; and a processor system having access to the storage medium, wherein executing the computer program instructions causes the processor system to:
receive a plurality of documents containing text, the documents uploaded according to an original file organization;
group the plurality of documents into sets of different types of documents;
create a second file organization based on the grouping into sets, wherein the second file organization is different than the original file organization;
for multiple sets of documents, process the documents in the sets using information from the multiple sets and from both the original file organization and the second file organization, the processing further comprising:
performing linguistic analysis of the text of the documents;
identifying structures and patterns in the text within documents and in the text across documents, based on the linguistic analysis;
extracting semantic role labels from the text of the documents, based on the identified structures and patterns; and
annotating chunks of text from the documents with the semantic role labels;
wherein the processing is automatically adjusted for different sets of documents based on the type of documents in that set.
14 . The method of claim 1 , wherein the processing further comprises:
arbitrating among the chunks to produce hierarchies of the chunks that are well-formed without any partial overlap of the chunks.
15 . A non-transitory computer-readable storage medium storing executable computer program instructions for labeling chunks in different types of documents, the instructions executable by a computer system and causing the computer system to:
receive a plurality of documents containing text, the documents uploaded according to an original file organization;
group the plurality of documents into sets of different types of documents;
create a second file organization based on the grouping into sets, wherein the second file organization is different than the original file organization;
for multiple sets of documents, process the documents in the sets using information from the multiple sets and from both the original file organization and the second file organization, the processing further comprising:
performing linguistic analysis of the text of the documents;
identifying structures and patterns in the text within documents and in the text across documents, based on the linguistic analysis;
extracting semantic role labels from the text of the documents, based on the identified structures and patterns; and
annotating chunks of text from the documents with the semantic role labels;
wherein the processing is automatically adjusted for different sets of documents based on the type of documents in that set.