System and method for automated file reporting
A document index generating system and method are provided. The system has at least one processor and a memory storing a sequence of instructions which when executed by the at least one processor configures the at least one processor to perform the method. The method preprocesses a plurality of pages into a collection of data structures, classifies each preprocessed page into at least one document type, segments groups of classified pages into documents, and generates a page and document index for the plurality of pages based on the classified pages and documents. Each data structure has a representation of data for a page of the plurality of pages. The representation has at least one region on the page.
1 . A document index generating system comprising:
at least one processor; and
a memory storing a sequence of instructions which when executed by the at least one processor configure the at least one processor to:
preprocess a plurality of pages into a collection of data structures, each data structure comprising a representation of data for a page of the plurality of pages, the representation comprising at least one region on the page;
classify each preprocessed page into at least one document type, including determining candidate document types for each page in the collection of data structures;
segment groups of classified pages into separate documents, said segmenting based on contiguous pages having at least one of:
similar document types;
similar titles; or
sequential page numbers located at a same region; and
generate a page and document index for the plurality of pages based on the classified pages and the documents;
obtain a document among the separate documents;
divide the document into chunks of content;
encode each chunk of content;
cluster each encoded chunk of content;
determine at least one central encoded chunk in each centroid of clustered encoded chunks; and
generate a summary for the document based on the at least one central encoded chunk for each of the clustered encoded chunks,
wherein to preprocess the plurality of pages into the collection of data structures, the at least one processor is configured to determine regions on that page based on the location of the region in relation to other regions on the page; and
wherein to determine candidate document types for the page, the at least one processor is configured to determine confidence score values for each candidate document type based on content in a summary block of the page.
2 . The document index generating system as claimed in claim 1 , wherein to preprocess the plurality of pages into a collection of data structures, the at least one processor is configured to:
for each page in the plurality of pages:
convert that page to a bit map file format;
determine regions on that page based on at least one of:
the location of a region on the page; or
the content in the region;
convert each region of that page into a machine-encoded content;
collect the regions and corresponding content for that page into a data structure for that page; and
merge the data structure for each page into the collection of data structures.
3 . The document index generating system as claimed in claim 2 , wherein to determine regions on that page, the at least one processor is configured to:
search sections of the page for text or other items, each section of the page comprising at least one of:
a top third of the page;
a middle third of the page;
a bottom third of the page;
a top quadrant of the page;
a bottom 15 percent of the page;
a bottom right corner of the page;
a top right corner of the page; or
the full page.
4 . The document index generating system as claimed in claim 1 , wherein to determine candidate document types for each page, the at least one processor is configured to:
determine confidence score values for each candidate document type based on at least one of:
a presence of a combination of regions on the page; or
content in at least one of:
a region category type for each region on the page;
a title of the page;
an origin of the page;
a date of the page; or
a page identifier.
5 . The document index generating system as claimed in claim 1 , wherein to segment groups of classified pages into documents, the at least one processor is configured to:
cluster contiguous pages based on at least one of:
similar document types;
similar document titles; or
sequential page numbers.
6 . The document index generating system as claimed in claim 1 , wherein the at least one processor is configured to:
analyze characteristics of the classified pages and the documents to update missing information in the page and document index.
7 . The document index generating system as claimed in claim 1 , wherein the at least one processor is further configured to:
determine a similarity score between a ground truth graph associated with the document and a predicted graph associated with the document, the at least one processor configured to:
obtain ground truth data with manually applied labels;
generate a graph for the ground truth data with manually applied labels;
generate a predicted graph using predicted attributes associated with the document, the at least one processor is configured to:
obtain classified pages and unclassified pages from the document using a known document classifier;
extract known attributes from the classified pages using a document type classifier; and
extract the predicted attributes from the unclassified pages using a page classifier; and
determine a graph edit distance between the generated graph for the ground truth data and the predicted graph.
8 . The document index generating system as claimed in claim 7 , wherein the at least one processor is configured to:
tokenize the document into sentences; and
group the sentences into the chunks of content.
9 . The document index generating system as claimed in claim 7 , wherein the at least one processor is configured to:
determine a corresponding vector for each chunk of content and determine a similarity score for all pairs of such vectors;
build a graph associated with the document, wherein nodes in the graph comprise the vectors and edges in the graph comprise connections between the vectors;
cluster the nodes in the graph into groupings;
determine at least one influential node in each grouping; and
generate the summary based on the at least one influential node in each grouping.
10 . The document index generating system as claimed in claim 7 , wherein the at least one processor is configured to:
determine a corresponding vector for each chunk of the chunks of content and determine a similarity score for all pairs of such vectors;
build a graph associated with the document, wherein nodes in the graph comprise the vectors and edges in the graph comprise connections between the vectors;
cluster the nodes in the graph into groupings;
determine at least one influential node in each grouping; and
generate the summary based on the at least one influential node in each grouping.
11 . A computer-implemented method of generating an index of a document, the method comprising:
preprocessing a plurality of pages into a collection of data structures, each data structure comprising a representation of data for a page of the plurality of pages, the representation comprising at least one region on the page;
classifying each preprocessed page into at least one document type, including determining candidate document types for each page in the collection of data structures;
segmenting groups of classified pages into separate documents, said segmenting based on contiguous pages having at least one of:
similar document types;
similar titles; or
sequential page numbers located at a same region; and
generating a page and document index for the plurality of pages based on the classified pages and the documents;
obtaining a document among the separate documents;
dividing the document into chunks of content;
encoding each chunk of content;
clustering each encoded chunk of content;
determining at least one central encoded chunk in each centroid of clustered encoded chunks;
generating a summary for the document based on the at least one central encoded chunk for each of the clustered encoded chunks,
wherein preprocessing the plurality of pages into the collection of data structures comprises determining regions on that page based on the location of the region in relation to other regions on the page; and
wherein determining candidate document types for the page comprises determining confidence score values for each candidate document type based on content in a summary block of the page.
12 . The computer-implemented method as claimed in claim 11 , wherein preprocessing the plurality of pages into a collection of data structures comprises:
for each page in the plurality of pages:
converting that page to a bit map file format;
determining regions on that page based on at least one of:
the location of a region on the page;
the content in the region; or
the location of the region in relation to other regions on the page;
converting each region of that page into a machine-encoded content;
collecting the regions and corresponding content for that page into a data structure for that page; and
merging the data structure for each page into the collection of data structures.
13 . The computer-implemented method as claimed in claim 12 , wherein determining regions on that page comprises:
searching sections of the page for text or other items, each section of the page comprising at least one of:
a top third of the page;
a middle third of the page;
a bottom third of the page;
a top quadrant of the page;
a bottom 15 percent of the page;
a bottom right corner of the page;
a top right corner of the page; or
the full page.
14 . The computer-implemented method as claimed in claim 11 , wherein determining candidate document types for each page comprises:
determining confidence score values for each candidate document type based on at least one of:
a presence of a combination of regions on the page; or
content in at least one of:
a region category type for each region on the page;
a title of the page;
an origin of the page;
a date of the page; or
a page identifier.
15 . The computer-implemented method as claimed in claim 11 , wherein segmenting groups of classified pages into documents comprises:
clustering contiguous pages based on at least one of:
similar document types;
similar document titles; or
sequential page numbers.
16 . The computer-implemented method as claimed in claim 11 , comprising:
analyzing characteristics of the classified pages and the documents to update missing information in the page and document index.
17 . The computer-implemented method as claimed in claim 11 , further comprising:
determining a similarity score between a ground truth graph associated with the document and a predicted graph associated with the document by:
obtaining ground truth data with manually applied labels;
generating a graph for the ground truth data with manually applied labels;
generating a predicted graph using predicted attributes associated with the document, said predicted attributes obtained by:
obtaining classified pages and unclassified pages from the document using a known document classifier;
extracting known attributes from the classified pages using a document type classifier; and
extracting the predicted attributes from the unclassified pages using a page classifier; and
determining a graph edit distance between the generated graph for the ground truth data and the predicted graph.
18 . The computer-implemented method as claimed in claim 17 , comprising:
tokenizing the document into sentences; and
grouping the sentences into the chunks of content.
19 . The computer-implemented method as claimed in claim 17 , comprising:
determining a corresponding vector for each chunk of content and determining a similarity score for all pairs of such vectors;
building a graph associated with the document, wherein nodes in the graph comprise the vectors and edges in the graph comprise connections between the vectors;
clustering the nodes in the graph into groupings;
determining at least one influential node in each grouping; and
generating the summary based on the at least one influential node in each grouping.