IP Library › Granted Patent US 9,020,966
Granted Patent B2
US 9,020,966 · App. 12/340,124 · Granted Apr 28, 2015

Client device for interacting with a mixed media reality recognition system

Inventors: Berna Erol (San Jose, CA); Jorge Moraleda (Menlo Park, CA); Jonathan J. Hull (San Carlos, CA)
Assignee: Ricoh Co., Ltd.
G06F17/30247G06F17/30026G06F17/30038G06F17/30047G06F17/30876G06K9/00463G06K9/00979G06K9/036G06K9/6217G06K9/6262
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,020,966
App. No.
12/340,124
Granted
Apr 28, 2015
Kind
B2
Abstract

The mobile device includes a client that has a number of modules, and the MMR Gateway and MMR matching unit are implemented as a server that has a number of modules. The implementation of the MMR system as a client and a server is advantageous because the modules may be distributed among the client and the server in a variety of configurations. The present invention includes a capture module, a preprocessing module, a feature extraction module, a retrieval module, a send message module, an action module, a prediction module, a feedback module, a sending module, an MMR database, a streaming module, an e-mail module, a voice recognition system and an audio database. These modules and systems are operational upon the client or the server.

Claims (55)

1. A method for generating and processing a retrieval request for a visual recognition system, the method comprising:

receiving an image and generating an image query from the image;

receiving audio data associated with the image, the audio data specifying at least one of a location within a document to be retrieved, at least one type of recognition algorithm for recognizing the image, and an order of the at least one type of recognition algorithm that determines a next algorithm used to recognize the image if a previous algorithm fails;

performing command and data recognition on the audio data to produce audio recognition results for improving image recognition, wherein the audio recognition results include a keyword;

performing retrieval of the document from a database of documents based on the image query and the audio recognition results to produce a retrieval result including a document identification, a portion of the document and an x-y location of the image on the portion of the document, wherein performing retrieval of the document includes performing image recognition on the image based on the image query to produce image recognition results from the database of documents, generating confidence scores associated with the image recognition results, modifying the confidence scores using the keyword and identifying the document based on the modified confidence scores; and

providing the document to a user or performing an action based on the document.

2. The method of claim 1 , further comprising receiving metadata associated with the image.

3. The method of claim 1 , wherein the audio recognition results are used to specify a subject of the document.

4. The method of claim 1 , wherein the audio recognition results are used to improve an accuracy of retrieval.

5. The method of claim 1 , wherein the audio recognition results are also used to specify how retrieval is performed.

6. The method of claim 1 , wherein the audio recognition results are also used to specify the action.

7. The method of claim 1 , wherein the audio recognition results are also used for biometric verification of the user that transmitted the image.

8. The method of claim 1 , wherein

modifying the confidence scores using the keyword includes increasing a confidence score by a certain increment for each keyword found in an image recognition result associated with the confidence score.

9. The method of claim 1 , further comprising:

receiving metadata associated with the image; and

wherein modifying the confidence scores is also based on the metadata.

10. The method of claim 1 , further comprising using the audio recognition results to select one of a plurality of databases to use for performing the retrieval.

11. The method of claim 1 , wherein the audio data is from a video.

12. A method for generating and processing a retrieval request in a distributed visual recognition system, the method comprising:

receiving an image and generating an image query from the image;

receiving audio data associated with the image, the audio data specifying at least one of a location within a document to be retrieved, at least one type of recognition algorithm for recognizing the image, and an order of the at least one type of recognition algorithm that determines a next algorithm used to recognize the image if a previous algorithm fails;

performing command and data recognition on the audio data to produce audio recognition results for improving image recognition, wherein the audio recognition results include a keyword;

performing retrieval of the document from a database of documents based on the image query and the audio recognition results, wherein performing retrieval of the document includes performing image recognition on the image based on the image query to produce image recognition results from the database of documents, generating confidence scores associated with the image recognition results, modifying the confidence scores using the keyword and identifying the document based on the modified confidence scores; and

generating and sending a first message including a document identification, a portion identification and an x-y location of the image on the portion of the document.

13. The method of claim 12 further comprising:

performing an action based on the document identification, the portion identification and the x-y location.

14. The method of claim 12 wherein retrieval of the document is performed on one from the group of a mobile device and a hardware server.

15. The method of claim 13 wherein performing the action includes:

generating and sending a second message, the second message including hotspot data.

16. The method of claim 12 wherein performing retrieval of the document comprises:

performing feature extraction on the image query to produce extracted features; and

querying the database of documents using the extracted features to generate the image recognition results including the document identification, the portion identification and the x-y location.

17. A method for generating and processing a retrieval request for a visual recognition system, the method comprising:

receiving an image and generating an image query from the image;

receiving audio data and metadata associated with the image, the audio data specifying at least one of a location within a document to be retrieved, at least one type of recognition algorithm for recognizing the image, and an order of the at least one type of recognition algorithm that determines a next algorithm used to recognize the image if a previous algorithm fails;

performing command and data recognition on the audio data to produce audio recognition results for improving image recognition, wherein the audio recognition results include a keyword;

performing retrieval of the document from a database of documents based on the image query, the audio recognition results and the metadata to produce a retrieval result including a document identification, a portion of the document and an x-y location of the image on the portion of the document, wherein performing retrieval of the document includes performing image recognition on the image based on the image query to produce image recognition results from the database of documents, generating confidence scores associated with the image recognition results, modifying the confidence scores using the keyword and identifying the document based on the modified confidence scores; and

performing an action based on the document.

18. The method of claim 17 , wherein receiving the image and receiving the audio data and the metadata are performed by a capture device and wherein the capture device generates a first message that includes the image, the audio data and the metadata in a multimedia messaging service (MMS) format.

19. The method of claim 17 , wherein the metadata includes one from the group of an email address, a user identification, a preference, global positioning system (GPS) information, device settings, metadata about a query image, a query image location, an audio file, a query image document name and information about the action.

20. The method of claim 17 , wherein the action is one from the group of sending a return confirmation message, sending the document, sending the portion of the document, sending the x-y location, sending a thumbnail of the document, sending a video overview of the document, sending a message about the action performed, and sending instructions about receiving the image.

21. The method of claim 17 , wherein performing retrieval of the document and performing the action are performed by a server coupled to a mobile device.

22. A system for generating and processing a retrieval request for a visual recognition, the system comprising:

a processor;

a send message module stored on a memory and executable by the processor, the send message module for receiving an image query and audio data associated with an image, the audio data specifying at least one of a location within a document to be retrieved, at least one type of recognition algorithm for recognizing the image, and an order of the at least one type of recognition algorithm that determines a next algorithm used to recognize the image if a previous algorithm fails;

a voice recognition module coupled to the send message module, the voice recognition module for performing command and data recognition on the audio data to produce audio recognition results for improving image recognition, wherein the audio recognition results include a keyword; and

a retrieval module coupled to the send message module, the retrieval module for performing retrieval of the document from a database of documents based on the image query and the audio recognition results to produce a retrieval result including a document identification, a portion of the document and an x-y location of the image on the portion of the document, wherein performing retrieval of the document includes performing image recognition on the image based on the image query to produce image recognition results from the database of documents, generating confidence scores associated with the image recognition results, modifying the confidence scores using the keyword and identifying the document based on the modified confidence scores;

wherein the send message module provides the document to a user or performs an action based on the document retrieved by the retrieval module.

23. A non-transitory computer readable storage medium comprising a computer readable program, wherein the computer readable program when executed on a computer causes the computer to:

receive an image and generate an image query from the image;

receive audio data associated with the image, the audio data specifying at least one of a location within a document to be retrieved, at least one type of recognition algorithm for recognizing the image, and an order of the at least one type of recognition algorithm that determines a next algorithm used to recognize the image if a previous algorithm fails;

perform command and data recognition on the audio data to produce audio recognition results for improving image recognition, wherein the audio recognition results include a keyword;

perform retrieval of the document from a database of documents based on the image query and the audio recognition results to produce a retrieval result including a document identification, a portion of the document and an x-y location of the image on the portion of the document, wherein performing retrieval of the document includes performing image recognition on the image based on the image query to produce image recognition results from the database of documents, generating confidence scores associated with the image recognition results, modifying the confidence scores using the keyword and identifying the document based on the modified confidence scores; and

providing the document to a user or performing an action based on the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2008
From: EROL, BERNA; MORALEDA, JORGE; HULL, JONATHAN J.
To: RICOH CO., LTD.
Reel/Frame 022009/0969 →
Continuity (40)
Continuation In Part 11461017 · Jul 31, 2006
Continuation In Part 11461279 · Jul 31, 2006
Continuation In Part 11461286 · Jul 31, 2006
Continuation In Part 11461294 · Jul 31, 2006
Continuation In Part 11461300 · Jul 31, 2006
Continuation In Part 11461126 · Jul 31, 2006
Continuation In Part 11461143 · Jul 31, 2006
Continuation In Part 11461268 · Jul 31, 2006
Continuation In Part 11461272 · Jul 31, 2006
Continuation In Part 11461064 · Jul 31, 2006
Continuation In Part 11461075 · Jul 31, 2006
Continuation In Part 11461090 · Jul 31, 2006
Continuation In Part 11461037 · Jul 31, 2006
Continuation In Part 11461085 · Jul 31, 2006
Continuation In Part 11461091 · Jul 31, 2006
Continuation In Part 11461095 · Jul 31, 2006
Continuation In Part 11466414 · Aug 22, 2006
Continuation In Part 11461147 · Jul 31, 2006
Continuation In Part 11461164 · Jul 31, 2006
Continuation In Part 11461024 · Jul 31, 2006
Continuation In Part 11461032 · Jul 31, 2006
Continuation In Part 11461049 · Jul 31, 2006
Continuation In Part 11461109 · Jul 31, 2006
Continuation In Part 11827530 · Jul 11, 2007
Continuation In Part 12060194 · Mar 31, 2008
Continuation In Part 12059583 · Mar 31, 2008
Continuation In Part 12060198 · Mar 31, 2008
Continuation In Part 12060200 · Mar 31, 2008
Continuation In Part 12060206 · Mar 31, 2008
Continuation In Part 12121275 · May 15, 2008
Continuation In Part 11776510 · Jul 11, 2007
Continuation In Part 11776520 · Jul 11, 2007
Continuation In Part 11776530 · Jul 11, 2007
Continuation In Part 11777142 · Jul 12, 2007
Continuation In Part 11624466 · Jan 18, 2007
Continuation In Part 12210511 · Sep 15, 2008
Continuation In Part 12210519 · Sep 15, 2008
Continuation In Part 12210532 · Sep 15, 2008
Continuation In Part 12210540 · Sep 15, 2008
Related Publication 20090100050A1 · Apr 16, 2009