IP Library › Granted Patent US 11,379,534
Granted Patent B2
US 11,379,534 · App. 16/687,848 · Granted Jul 5, 2022

Document feature repository management

Inventors: Kun Yan Yin (Ningbo, CN); Qi Y Wang (Ningbo, CN); Wen Wang (Beijing, CN); Hai Ji (Beijing, CN); Rui W W Wang (Beijing, CN)
Assignee: International Business Machines Corporation
G06F16/9038G06F16/5846G06F16/9035G06F16/93G06V10/462G06V30/153G06V30/414G06V30/418
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,379,534
App. No.
16/687,848
Granted
Jul 5, 2022
Kind
B2
Abstract

A plurality of documents is received. From one or more documents within the plurality, a set of image features is extracted. A document feature repository is generated by storing one or more sets of image features, by document. A document search query is received. The document search query is pre-processed to generate a set of variable document search queries. The document feature repository is searched using the set of variable document search queries. The search results are presented to a user.

Claims (85)

1. A method for generating and searching a document feature repository, the method comprising:

receiving a plurality of documents;

extracting a set of image features from one or more documents within the plurality of documents, wherein extracting the set of image features comprises:

parsing the one or more documents into a set of page images, by document;

generating one or more sets of character blocks for each page image within the set of page images;

extracting a plurality of character block features for each character block;

generating the document feature repository by storing the one or more sets of image features, by document;

receiving a document search query; pre-processing the document search query to generate a set of variable document search queries;

searching the document feature repository using the set of variable document search queries wherein searching the document feature repository using the set of variable document search queries further comprises generating a set of match scores between each set within the plurality of character block features and the set of variable document search queries, wherein generating a set of match scores between the one or more documents and the set of variable document search queries further comprises:

generating, using Fast Library for Approximate Nearest Neighbors (FLANN), a set of weighted match scores between the plurality of character block features for each of the one or more documents and the set of variable document search queries; and

combining the set of weighted match scores to generate the match score for each of the one or more documents; and

presenting the search results to a user.

2. The method of claim 1 , wherein extracting the set of image features further comprises:

storing the plurality of character block features in the document feature repository, by page image.

3. The method of claim 2 , wherein generating one or more sets of character blocks further comprises:

performing image binarization on the document; and

performing morphological dilation on the binarized document to generate the one or more sets of character blocks.

4. The method of claim 1 , wherein pre-processing the document search query further comprises:

parsing the document search query into a set of query subcomponents;

modifying the set of query subcomponents by replacing a first query subcomponent with a second query subcomponent, wherein the first and second query subcomponents are similar; and

outputting the modified set of query subcomponents as a first variable document search query within the set of variable document search queries.

5. The method of claim 2 , wherein the plurality of character block features includes a set of character features, a set of word features, and a set of sentence features.

6. The method of claim 5 , wherein searching the document feature repository using the set of variable document search queries further comprises:

determining a best match, according to the set of match scores; and

presenting the document associated with the best match score to the user.

7. The method of claim 1 , wherein generating the set of weighted match scores includes adjusting a weight and a bias of one or more neural network edges.

8. A computer program product for generating and searching a document feature repository, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a device to cause the device to:

receive a plurality of documents;

extract a set of image features from one or more documents within the plurality of documents, wherein extracting the set of image features comprises:

parsing the one or more documents into a set of page images, by document;

generating one or more sets of character blocks for each page image within the set of page images;

extracting a plurality of character block features for each character block;

generate the document feature repository by storing the one or more sets of image features, by document;

receive a document search query;

pre-process the document search query to generate a set of variable document search queries;

search the document feature repository using the set of variable document search queries, wherein searching the document feature repository using the set of variable document search queries further comprises generating a set of match scores between each set within the plurality of character block features and the set of variable document search queries, wherein generating a set of match scores between the one or more documents and the set of variable document search queries further comprises:

generating, using Fast Library for Approximate Nearest Neighbors (FLANN), a set of weighted match scores between the plurality of character block features for each of the one or more documents and the set of variable document search queries; and

combining the set of weighted match scores to generate the match score for each of the one or more documents; and

present the search results to a user.

9. The computer program product of claim 8 , wherein extracting the set of image features further comprises:

storing the plurality of character block features in the document feature repository, by page image.

10. The computer program product of claim 9 , wherein generating one or more sets of character blocks further comprises:

performing image binarization on the document; and

performing morphological dilation on the binarized document to generate the one or more sets of character blocks.

11. The computer program product of claim 8 , wherein pre-processing the document search query further comprises:

parsing the document search query into a set of query subcomponents;

modifying the set of query subcomponents by replacing a first query subcomponent with a second query subcomponent, wherein the first and second query subcomponents are similar; and

outputting the modified set of query subcomponents as a first variable document search query within the set of variable document search queries.

12. The computer program product of claim 9 , wherein the plurality of character block features includes a set of character features, a set of word features, and a set of sentence features.

13. The computer program product of claim 12 , wherein searching the document feature repository using the set of variable document search queries further comprises:

determining a best match, according to the set of match scores; and

presenting the document associated with the best match score to the user.

14. A system for generating and searching a document feature repository, comprising:

a memory with program instructions included thereon; and

a processor in communication with the memory, wherein the program instructions cause the processor to:

receive a plurality of documents;

extract a set of image features from one or more documents within the plurality of documents, wherein extracting the set of image features comprises:

parsing the one or more documents into a set of page images, by document

generating one or more sets of character blocks for each page image within the set of page images;

extracting a plurality of character block features for each character block;

generate the document feature repository by storing the one or more sets of image features, by document;

receive a document search query;

pre-process the document search query to generate a set of variable document search queries;

search the document feature repository using the set of variable document search queries, wherein searching the document feature repository using the set of variable document search queries further comprises generating a set of match scores between each set within the plurality of character block features and the set of variable document search queries, wherein generating a set of match scores between the one or more documents and the set of variable document search queries further comprises:

generating, using Fast Library for Approximate Nearest Neighbors (FLANN), a set of weighted match scores between the plurality of character block features for each of the one or more documents and the set of variable document search queries; and

combining the set of weighted match scores to generate the match score for each of the one or more documents; and

present the search results to a user.

15. The system of claim 14 , wherein extracting the set of image features comprises:

parsing the one or more documents into a set of page images, by document;

generating one or more sets of character blocks for each page image within the set of page images;

extracting a plurality of character block features for each character block; and

storing the plurality of character block features in the document feature repository, by page image.

16. The system of claim 15 , wherein generating one or more sets of character blocks further comprises:

performing image binarization on the document; and

performing morphological dilation on the binarized document to generate the one or more sets of character blocks.

17. The system of claim 14 , wherein pre-processing the document search query further comprises:

parsing the document search query into a set of query subcomponents;

modifying the set of query subcomponents by replacing a first query subcomponent with a second query subcomponent, wherein the first and second query subcomponents are similar; and

outputting the modified set of query subcomponents as a first variable document search query within the set of variable document search queries.

18. The system of claim 15 , wherein the plurality of character block features includes a set of character features, a set of word features, and a set of sentence features.

19. The system of claim 15 , wherein extracting the set of image features further comprises:

storing the plurality of character block features in the document feature repository, by page image.

20. The system of claim 15 , wherein searching the document feature repository using the set of variable document search queries further comprises:

determining a best match, according to the set of match scores; and

presenting the document associated with the best match score to the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2019
From: YIN, KUN YAN; WANG, QI Y; WANG, WEN; JI, HAI; WANG, RUI WW
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051046/0582 →
Continuity (1)
Related Publication 20210149962A1 · May 20, 2021
Cited By (1)
US 12,210,556