IP Library Granted Patent US 8,655,657
Granted Patent B1
US 8,655,657 · App. 13/768,232 · Granted Feb 18, 2014

Identifying media content

Inventors: Matthew Sharifi (Zurich, CH); Gheorghe Postelnicu (Zurich, CH)
Assignee: Google Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,655,657
App. No.
13/768,232
Granted
Feb 18, 2014
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving (i) audio data that encodes a spoken natural language query, and (ii) environmental audio data, obtaining a transcription of the spoken natural language query, determining a particular content type associated with one or more keywords in the transcription, providing at least a portion of the environmental audio data to a content recognition engine, and identifying a content item that has been output by the content recognition engine, and that matches the particular content type.

Claims (38)

1. A computer-implemented method comprising:

receiving data including (i) audio data that encodes a spoken natural language query, and (ii) sensor data;

obtaining a transcription of the spoken natural language query;

determining a particular content type associated with one or more keywords in the transcription;

providing at least a portion of the sensor data to a content recognition engine; and

identifying a content item that (i) has been output by the content recognition engine, and (ii) matches the particular content type associated with the one or more keywords in the transcription.

2. The computer-implemented method of claim 1 , wherein the sensor data includes image data.

3. The computer-implemented method of claim 2 , wherein receiving the data further comprises receiving the data from a mobile computing device.

4. The computer-implemented method of claim 2 , wherein receiving the data further comprises receiving environmental audio data.

5. The computer-implemented method of claim 2 , wherein the image data includes environmental image data.

6. The computer-implemented method of claim 2 , wherein the image data is generated within a predetermined period of time prior to the spoken natural language query.

7. The computer-implemented method of claim 2 , wherein determining the particular content type further includes identifying the one or more keywords using one or more databases that, for each of multiple content types, maps at least one of the keywords to at least one of the multiple content types.

8. The computer-implemented method of claim 7 , wherein the multiple content types includes the particular content type, and wherein mapping further includes mapping at least one of the keywords to the particular content type.

9. The computer-implemented method of claim 2 , wherein providing further includes providing data identifying the particular content type to the content recognition engine, and

wherein identifying the content item further includes receiving data identifying the content item from the content recognition engine.

10. The computer-implemented method of claim 2 , further including receiving two or more content recognition candidates from the content recognition system, and

wherein identifying the content item further includes selecting a particular content recognition candidate based on the particular content type.

11. The computer-implemented method of claim 10 , wherein each of the two or more content recognition candidates is associated with a ranking score, the method further including adjusting the ranking scores of the two or more content recognition candidates based on the particular content type.

12. The computer-implemented method of claim 11 , further including ranking the two or more content recognition candidates based on the adjusted ranking scores.

13. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving data including (i) audio data that encodes a spoken natural language query, and (ii) sensor data;

obtaining a transcription of the spoken natural language query;

determining a particular content type associated with one or more keywords in the transcription;

providing at least a portion of the sensor data to a content recognition engine; and

identifying a content item that (i) has been output by the content recognition engine, and (ii) matches the particular content type associated with the one or more keywords in the transcription.

14. The system of claim 13 , wherein the sensor data includes image data.

15. The system of claim 14 , wherein receiving the data further comprises receiving the data from a mobile computing device.

16. The system of claim 14 , wherein receiving the data further comprises receiving environmental audio data.

17. The system of claim 14 , wherein the image data includes environmental image data.

18. The system of claim 14 , wherein the image data is generated within a predetermined period of time prior to the spoken natural language query.

19. The system of claim 14 , wherein determining the particular content type further includes identifying the one or more keywords using one or more databases that, for each of multiple content types, maps at least one of the keywords to at least one of the multiple content types.

20. A computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving data including (i) audio data that encodes a spoken natural language query, and (ii) sensor data;

obtaining a transcription of the spoken natural language query;

determining a particular content type associated with one or more keywords in the transcription;

providing at least a portion of the sensor data to a content recognition engine; and

identifying a content item that (i) has been output by the content recognition engine, and (ii) matches the particular content type associated with the one or more keywords in the transcription.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0299 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2013
From: SHARIFI, MATTHEW; POSTELNICU, GHEORGHE
To: GOOGLE INC.
Reel/Frame 030027/0930 →
Continuity (2)
Continuation 13626351 · Sep 25, 2012
Provisional Application 61698949 · Sep 10, 2012