IP Library Granted Patent US 9,798,724
Granted Patent B2
US 9,798,724 · App. 14/588,194 · Granted Oct 24, 2017

Document discovery strategy to find original electronic file from hardcopy version

Inventor: Kirk Steven Tecu (Longmont, CO)
Assignee: Konica Minolta Laboratory U.S.A., Inc.
G06F17/30011G06F17/30247G06K9/00483G06K9/2063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,798,724
App. No.
14/588,194
Granted
Oct 24, 2017
Kind
B2
Abstract

A method for document discovery includes receiving a scan of a physical copy of a document with a non-text object, determining a tag for the non-text object defining a portion of the non-text object in an original file, and generating, based on the tag, non-text object metadata with composition information of the non-text object. The method further includes searching, using the non-text object metadata, electronic documents stored in a data repository, where each of the electronic documents has an object and searchable metadata associated with the object, comparing the non-text object metadata with the searchable metadata, and providing a location of the original file to a user when the non-text object metadata matches the searchable metadata.

Claims (59)

1. A method for document discovery, comprising:

receiving a scan of a physical copy of a document comprising a non-text object;

determining a first tag for the non-text object by comparing the non-text object with a plurality of templates comprising a plurality of tags, wherein the first tag defines a portion of the non-text object in an original file and specifies a type of the non-text object and a formatting attribute of the non-text object;

generating, based on the first tag, non-text object metadata comprising composition information comprising the type and the formatting attribute for the non-text object;

searching a plurality of electronic documents stored in a data repository with a search query comprising the non-text object metadata, wherein each of the plurality of electronic documents comprises searchable metadata;

comparing the non-text object metadata with the searchable metadata; and

providing a location of the original file to a user when the non-text object metadata in the search query matches the searchable metadata of the original file.

2. The method of claim 1 , further comprising:

processing an electronic document from the plurality of electronic documents stored in the data repository by:

extracting a second tag for an object in the electronic document;

generating the searchable metadata based on the second tag, wherein the searchable metadata of the electronic document describes the object; and

storing the searchable metadata in the electronic document associated with the object.

3. The method of claim 1 , wherein the original file is an Office Open XML file, and wherein the original file is one of the plurality of electronic documents stored in the data repository.

4. The method of claim 1 , further comprising:

determining whether the user has authorization to access the original file, wherein the location is provided only when the user is determined to have authorization to access the original file.

5. The method of claim 1 , wherein the location is provided in an e-mail to the user.

6. The method of claim 1 , wherein the location is provided by displaying the location on a display of a scanner.

7. The method of claim 1 , wherein the data repository is part of an enterprise content management (ECM) system.

8. The method of claim 1 , wherein the searching further comprises using standard text discovered in the document through Optical Character Recognition (OCR).

9. A system for document discovery, comprising:

a data repository storing a plurality of electronic documents, wherein each of the plurality of electronic documents comprises searchable metadata;

a computer processor connected to the data repository that:

receives a scan of a physical copy of a document comprising a non-text object;

determines a first tag for the non-text object by comparing the non-text object with a plurality of templates comprising a plurality of tags, wherein the first tag defines a portion of the non-text object in an original file and specifies a type of the non-text object and a formatting attribute of the non-text object;

generates, based on the first tag, non-text object metadata comprising composition information comprising the type and the formatting attribute for the non-text object;

searches the plurality of electronic documents stored in the data repository with a search query comprising the non-text object metadata;

compares the non-text object metadata with the searchable metadata; and

provides a location of the original file to a user when the non-text object metadata in the search query matches the searchable metadata of the original file.

10. The system of claim 9 , wherein the computer processor also:

processes an electronic document from the plurality of electronic documents stored in the data repository by:

extracting a second tag for an object in the electronic document;

generating the searchable metadata based on the second tag, wherein the searchable metadata of the electronic document describes the object; and

storing the searchable metadata in the electronic document associated with the object.

11. The system of claim 9 , wherein the original file is an Office Open XML file, and wherein the original file is one of the plurality of electronic documents stored in the data repository.

12. The system of claim 9 , wherein the computer processor also:

determines whether the user has authorization to access the original file, wherein the location is provided only when the user is determined to have authorization to access the original file.

13. The system of claim 9 , wherein the location is provided in an e-mail to the user.

14. The system of claim 9 , wherein the location is provided by displaying the location on a display of a scanner.

15. The system of claim 9 , wherein the data repository is part of an enterprise content management (ECM) system.

16. The system of claim 9 , wherein the searching further comprises using standard text discovered in the document through Optical Character Recognition (OCR).

17. A non-transitory computer readable medium comprising instructions for document discovery, the instructions, when executed, are configured to:

receive a scan of a physical copy of a document comprising a non-text object;

determine a first tag for the non-text object by comparing the non-text object with a plurality of templates comprising a plurality of tags, wherein the first tag defines a portion of the non-text object in an original file and specifies a type of the non-text object and a formatting attribute for the non-text object;

generate, based on the first tag, non-text object metadata comprising composition information comprising the type and formatting attribute for the non-text object;

search, using the non-text object metadata, a plurality of electronic documents stored in a data repository with a search query comprising the non-text object metadata, wherein each of the plurality of electronic documents comprises searchable metadata;

compare the non-text object metadata with the searchable metadata; and

provide a location of the original file to a user when the non-text object metadata in the search query matches the searchable metadata of the original file.

18. The non-transitory computer readable medium of claim 17 , the instructions further configured to:

process an electronic document from the plurality of electronic documents stored in the data repository by:

extracting a second tag for an object in the electronic document;

generating the searchable metadata based on the second tag, wherein the searchable metadata of the electronic document describes the object; and

storing the searchable metadata in the electronic document associated with the object.

19. The non-transitory computer readable medium of claim 17 , wherein the original file is an Office Open XML file, and wherein the original file is one of the plurality of electronic documents stored in the data repository.

20. The non-transitory computer readable medium of claim 17 , the instructions further configured to:

determine whether the user has authorization to access the original file, wherein the location is provided only when the user is determined to have authorization to access the original file.

21. The non-transitory computer readable medium of claim 17 , wherein the location is provided in an e-mail to the user.

22. The non-transitory computer readable medium of claim 17 , wherein the location is provided by displaying the location on a display of a scanner.

23. The non-transitory computer readable medium of claim 17 , wherein the data repository is part of an enterprise content management (ECM) system.

24. The non-transitory computer readable medium of claim 17 , wherein the searching further comprises using standard text discovered in the document through Optical Character Recognition (OCR).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2015
From: TECU, KIRK STEVEN
To: KONICA MINOLTA LABORATORY U.S.A., INC.
Reel/Frame 034647/0554 →
Continuity (1)
Related Publication 20160188580A1 · Jun 30, 2016