IP Library Granted Patent US 8,335,789
Granted Patent B2
US 8,335,789 · App. 11/461,286 · Granted Dec 18, 2012

Method and system for document fingerprint matching in a mixed media environment

Assignee: Ricoh Co., Ltd.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,335,789
App. No.
11/461,286
Granted
Dec 18, 2012
Kind
B2
Abstract

A Mixed Media Reality (MMR) system and associated techniques are disclosed. The MMR system provides mechanisms for forming a mixed media document that includes media of at least two types (e.g., printed paper as a first medium and digital content and/or web link as a second medium). In one particular embodiment, the MMR system provides for document fingerprint matching.

Claims (99)

1. A method of image matching, comprising:

receiving an image of at least part of a first media type;

generating a horizontal profile from the image, the horizontal profile identifying words in the image;

generating a plurality of bounding boxes, each bounding box surrounding a word in the horizontal profile;

horizontally classifying the plurality of bounding boxes in the image;

vertically classifying the plurality of bounding boxes in the image;

determining at least one spatial relationship between the plurality of bounding boxes by associating a first length of a first word with a second length of a second word in the horizontal profile and combining the horizontal and vertical classifications;

generating at least one horizontal grouping of bounding boxes and at least one vertical grouping of bounding boxes based on the spatial relationship;

generating a list of documents from a database of one or more documents, the list of documents including at least one common bounding box comprising an overlap of the at least one horizontal grouping of bounding boxes and the at least one vertical grouping of bounding boxes at a location in each document in the list;

determining a number of votes for each document in the list based on a number of common bounding boxes; and

identifying a matching document with a most number of votes from the list as a document containing the image.

2. The method of claim 1 , wherein the first media type is a paper document.

3. The method of claim 1 , further comprising:

identifying a second media type dependent on a location of the image in the matching document.

4. The method of claim 3 , wherein the second media type comprises at least one selected from the group consisting of a data structure, a command, text, audio, video, an image, a digital photograph, web link text, an application file, updated information, and services.

5. The method of claim 1 , wherein the first length of the first word in the image includes a count of characters in the first word and wherein the second word has a location that is at least one of below and above the first word.

6. The method of claim 1 , wherein determining the at least one spatial relationship further comprises:

representing content within a first word region with at least one of a first value; and

representing a space between the first word region and a second word region with at least one of a second value.

7. The method of claim 1 , further comprising:

identifying a first feature point in the image;

identifying a second feature point in the image; and

wherein determining the at least one spatial relationship includes determining at least a distance or an angular relationship between the first feature point and the second feature point.

8. The method of claim 1 , wherein the horizontal grouping of bounding boxes comprises a trigram of three words.

9. The method of claim 1 , further comprising:

determining an x-y location of the common bounding box; and

identifying the x-y location as the location of the image within the matching document.

10. A system for image matching, comprising:

a database operable to store one or more documents;

a processor; and

a feature extraction module stored on a memory and executable by the processor, the feature extraction module operable to:

receive an image of at least part of a first media type,

generate a horizontal profile from the image, the horizontal profile identifying words in the image,

generate a plurality of bounding boxes, each bounding box surrounding a word in the horizontal profile,

horizontally classify the plurality of bounding boxes in the image,

vertically classify the plurality of bounding boxes in the image,

determine at least one spatial relationship between the plurality of bounding boxes by an association of a first length of a first word with a second length of a second word in the horizontal profile and a combination of the horizontal and vertical classifications,

generate at least one horizontal grouping of bounding boxes and at least one vertical grouping of bounding boxes based on the spatial relationship,

generate a list of documents from the database, the list of documents including at least one common bounding box comprising an overlap of the at least one horizontal grouping of bounding boxes and the at least one vertical grouping of bounding boxes at a location in each document in the list,

determine a number of votes for each document in the list based on a number of common bounding boxes and

identify a matching document with a most number of votes from the list as a document containing the image.

11. The system of claim 10 , further comprising an image capture device that comprises the feature extraction module.

12. The system of claim 10 , wherein the database comprises the feature extraction module.

13. The system of claim 10 , wherein the first media type is a paper document.

14. The system of claim 10 , wherein a second media type is communicated to the feature extraction module based on identification of a location of the image in the matching document.

15. The system of claim 14 , wherein the second media type comprises at least one selected from the group consisting of a data structure, a command, text, audio, video, an image, a digital photograph, web link text, an application file, updated information, and services.

16. The system of claim 10 , wherein the first length of the first word in the image includes a count of characters in the first word and wherein the second word has a location that is at least one of below and above the first word.

17. The system of claim 10 , wherein the feature extraction module is further operable to:

represent content within a first word region with at least one of a first value; and

represent a space between the first word region and a second word region with at least one of a second value.

18. The system of claim 10 , wherein the feature extraction module is further operable to:

identify a first feature point in the image;

identify a second feature point in the image; and

determine a distance or an angular relationship between the first feature point and the second feature point to determine the spatial relationship.

19. The system of claim 10 , wherein the horizontal grouping of bounding boxes comprises a trigram of three words.

20. The system of claim 10 , wherein the feature extraction module is further operable to:

determine an x-y location of the common bounding box; and

identify the x-y location as the location of the image within the matching document.

21. The system of claim 10 , further comprising an image capture device that captures the image and transmits the image to the feature extraction module, the image capture device comprising one selected from the group consisting of a cellular camera phone, a Personal Digital Assistant (PDA) device, a digital camera, a barcode reader, a radio frequency identification (RFID) reader, a computer peripheral, a web camera, and a video card.

22. A method of providing interaction between a first media type and a second media type, comprising:

receiving an image of at least part of the first media type;

extracting a plurality of features from the image, including a plurality of words;

generating a horizontal profile from the image, the horizontal profile including the words in the image;

generating a plurality of bounding boxes, each bounding box surrounding a word in the horizontal profile;

horizontally classifying the plurality of bounding boxes in the image;

vertically classifying the plurality of bounding boxes in the image;

determining at least one spatial relationship between the plurality of bounding boxes by associating a first length of a first word with a second length of a second word in the horizontal profile and combining the horizontal and vertical classifications;

generating at least one horizontal grouping of bounding boxes and at least one vertical grouping of bounding boxes based on the spatial relationship;

generating a list of documents from a database of one or more documents, the list of documents including at least one symbolic representation comprising an overlap of the at least one horizontal grouping of bounding boxes and the at least one vertical grouping of bounding boxes at a location in each document in the list;

determining a number of votes for each document in the list based on a number of symbolic representations;

identifying a matching document with a most number of votes as a document containing the image; and

providing the document containing the image as the second media type.

23. The method of claim 22 , wherein the first media type is a paper document.

24. The method of claim 22 , wherein the second media type comprises at least one selected from the group consisting of a data structure, a command, text, audio, video, an image, a digital photograph, web link text, an application file, updated information, and services.

25. The method of claim 22 , further comprising:

performing an action based on the identification, wherein the action comprises at least one selected from the group consisting of retrieving information, placing an order, retrieving a video, retrieving a sound, storing information, creating a new document, printing a document or image, displaying a document or image, searching information, and presenting information.

26. The method of claim 22 , wherein the first length of the first word in the image includes a count of characters in the first word and wherein the second word has a location that is at least one of below and above the first word.

27. The method of claim 22 , wherein determining the at least one spatial relationship further comprises:

representing content within a first word region with at least one of a first value; and

representing a space between the first word region and a second word region with at least one of a second value.

28. The method of claim 22 , further comprising:

identifying a first feature point in the image;

identifying a second feature point in the image; and

wherein determining the at least one spatial relationship includes determining at least a distance or an angular relationship between the first feature point and the second feature point.

29. The method of claim 22 , wherein the horizontal grouping of bounding boxes comprises a trigram of three words.

30. The method of claim 22 , further comprising:

determining an x-y location of the common bounding box; and

identifying the x-y location as the location of the image within the matching document.

31. A computer program product comprising a computer readable non-transitory storage medium including a computer readable program, wherein the computer readable program when executed on a computer causes the computer to:

receive an image of at least part of a first media type;

generate a horizontal profile from the image, the horizontal profile identifying words in the image;

generate a plurality of bounding boxes, each bounding box surrounding a word in the horizontal profile;

horizontally classify the plurality of bounding boxes in the image;

vertically classify the plurality of bounding boxes in the image;

determine at least one spatial relationship between the plurality of bounding boxes by an association of a first length of a first word with a second length of a second word in the horizontal profile and a combination the horizontal and vertical classifications;

generate at least one horizontal grouping of bounding boxes and at least one vertical grouping of bounding boxes based on the spatial relationship;

generate a list of documents from a database of one or more documents, the list of documents including at least one common bounding box comprising an overlap of the at least one horizontal grouping of bounding boxes and the at least one vertical grouping of bounding boxes at a location in each document in the list;

determine a number of votes for each document in the list based on a number of common bounding boxes; and

identify a matching document page with a most number of votes from the list as a document containing the image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2006
From: EROL, BERNA; HART, PETER E.; LEE, DAR-SHYANG; PIERSOL, KURT; HULL, JONATHAN J.
To: RICOH CO., LTD.
Reel/Frame 018037/0238 →
Continuity (5)
Continuation In Part 10957080 · Oct 1, 2004
Provisional Application 60710767 · Aug 23, 2005
Provisional Application 60792912 · Apr 17, 2006
Provisional Application 60807654 · Jul 18, 2006
Related Publication 20060285172A1 · Dec 21, 2006