IP Library Granted Patent US 11,847,124
Granted Patent B2
US 11,847,124 · App. 17/531,332 · Granted Dec 19, 2023

Contextual search on multimedia content

Inventors: Gökhan Hasan Bakir (Zurich, CH); Károly Csalogány (Zurich, CH); Behshad Behzadi (Zurich, CH)
Assignee: GOOGLE LLC
G06F16/24575G06F16/43G06F16/951G06F16/9538
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,847,124
App. No.
17/531,332
Granted
Dec 19, 2023
Kind
B2
Abstract

Techniques for contextual search on multimedia content are provided. An example method includes extracting entities associated with multimedia content, wherein the entities include values characterizing one or more objects represented in the multimedia content, generating one or more query rewrite candidates based on the extracted entities and one or more terms in a query related to the multimedia content, providing the one or more query rewrite candidates to a search engine, scoring the one or more query rewrite candidates, ranking the scored one or more query rewrite candidates based on their respective scores, rewriting the query related to the multimedia content based on a particular ranked query rewrite candidate and providing for display, responsive to the query related to the multimedia content, a result set from the search engine based on the rewritten query.

Claims (55)

1. A method implemented by one or more processors, the method comprising:

receiving a voice input of a user during streaming of multimedia content at a client device of the user, wherein the voice input includes a plurality of terms; and

in response to receiving the voice input:

extracting an entity associated with the multimedia content, wherein the entity includes at least one value characterizing at least one object represented in the multimedia content at a time of receiving the voice input;

extracting an additional entity associated with the multimedia content, wherein the additional entity includes at least one additional value characterizing at least one additional object represented in the multimedia content at the time of receiving the voice input, and wherein the additional entity is in addition to the entity;

generating, based on one or more of the plurality of terms of the voice input and the entity extracted from the multimedia content, a query;

generating, based on one or more of the plurality of terms of the voice input and the additional entity extracted from the multimedia content, an additional query, wherein the additional query is in addition to the query;

executing, based on the query, a search over one or more databases;

executing, based on the additional query, an additional search over one or more of the databases;

identifying, based on the search executed over one or more of the databases, a result that is responsive to the query;

identifying, based on the additional search executed over one or more of the databases, an additional result that is responsive to the additional query, wherein the additional result that is responsive to the additional query is in addition to the result that is responsive to the query;

selecting, from among the result that is responsive to the query and the additional result that is responsive to the additional query, a given result to be provided for presentation to the user responsive to receiving the voice input; and

causing the given result to be provided for presentation to the user responsive to receiving the voice input and without interrupting consumption of the multimedia content during the streaming of the multimedia content.

2. The method of claim 1 , wherein generating the query is performed based on the voice input being received from the user during the streaming of the multimedia content.

3. The method of claim 1 , wherein the multimedia content is a video.

4. The method of claim 1 , wherein causing the result to be provided for presentation to the user comprises causing the result to be provided for audible presentation to the user via a speaker of the client device.

5. The method of claim 1 , wherein causing the result to be provided for presentation to the user comprises causing the result to be provided for visual presentation to the user via a display of the client device.

6. The method of claim 1 , wherein extracting the entity associated with the multimedia content comprises:

extracting the entity from metadata associated with the multimedia content.

7. The method of claim 1 , wherein the at least one value characterizing the at least one object represented in the multimedia content is one of: a name of a person represented in the multimedia content, a name of a person associated with the multimedia content, or a name of an object represented in the multimedia content.

8. A system comprising:

memory comprising instructions; and

at least one processor configured to execute the instructions to:

receive a voice input of a user during streaming of multimedia content at a client device of the user, wherein the voice input includes a plurality of terms; and

in response to receiving the voice input:

extract an entity associated with the multimedia content, wherein the entity includes at least one value characterizing at least one object represented in the multimedia content at a time of receiving the voice input;

extract an additional entity associated with the multimedia content, wherein the additional entity includes at least one additional value characterizing at least one additional object represented in the multimedia content at the time of receiving the voice input, and wherein the additional entity is in addition to the entity

generate, based on one or more of the plurality of terms of the voice input and the entity extracted from the multimedia content, a query;

generate, based on one or more of the plurality of terms of the voice input and the additional entity extracted from the multimedia content, an additional query, wherein the additional query is in addition to the query;

execute, based on the query, a search over one or more databases;

execute, based on the additional query, an additional search over one or more of the databases;

identify, based on the search executed over one or more of the databases, a result that is responsive to the query;

identify, based on the additional search executed over one or more of the databases, an additional result that is responsive to the additional query, wherein the additional result that is responsive to the additional query is in addition to the result that is responsive to the query;

select, from among the result that is responsive to the query and the additional result that is responsive to the additional query, a given result to be provided for presentation to the user responsive to receiving the voice input; and

cause the given result to be provided for presentation to the user responsive to receiving the voice input and without interrupting consumption of the multimedia content during the streaming of the multimedia content.

9. The system of claim 8 , wherein generating the query is performed based on the voice input being received from the user during the streaming of the multimedia content.

10. The system of claim 8 , wherein the multimedia content is a video.

11. The system of claim 8 , wherein the instruction to cause the result to be provided for presentation to the user comprise instructions to cause the result to be provided for audible presentation to the user via a speaker of the client device.

12. The system of claim 8 , wherein the instruction to cause the result to be provided for presentation to the user comprise instructions to cause the result to be provided for visual presentation to the user via a display of the client device.

13. The system of claim 8 , wherein the instructions to extract the entity associated with the multimedia content comprise instruction to:

extract the entity from metadata associated with the multimedia content.

14. The system of claim 8 , wherein the at least one value characterizing the at least one object represented in the multimedia content is one of: a name of a person represented in the multimedia content, a name of a person associated with the multimedia content, or a name of an object represented in the multimedia content.

15. A non-transitory machine-readable medium comprising instructions stored therein, which when executed by a processor, causes the processor to perform operations comprising:

receiving a voice input of a user during streaming of multimedia content at a client device of the user, wherein the voice input includes a plurality of terms; and

in response to receiving the voice input:

extracting an entity associated with the multimedia content, wherein the entity includes at least one value characterizing at least one object represented in the multimedia content at a time of receiving the voice input;

extracting an additional entity associated with the multimedia content, wherein the additional entity includes at least one additional value characterizing at least one additional object represented in the multimedia content at the time of receiving the voice input, and wherein the additional entity is in addition to the entity;

generating, based on one or more of the plurality of terms of the voice input and the entity extracted from the multimedia content, a query;

generating, based on one or more of the plurality of terms of the voice input and the additional entity extracted from the multimedia content, an additional query, wherein the additional query is in addition to the query;

executing, based on the query, a search over one or more databases;

executing, based on the additional query, an additional search over one or more of the databases;

identifying, based on the search executed over one or more of the databases, a result that is responsive to the query;

identifying, based on the additional search executed over one or more of the databases, an additional result that is responsive to the additional query, wherein the additional result that is responsive to the additional query is in addition to the result that is responsive to the query;

selecting, from among the result that is responsive to the query and the additional result that is responsive to the additional query, a given result to be provided for presentation to the user responsive to receiving the voice input; and

causing the given result to be provided for presentation to the user responsive to receiving the voice input and without interrupting consumption of the multimedia content during the streaming of the multimedia content.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: BAKIR, GÖKHAN HASAN; CSALOGÁNY, KÁROLY; BEHZADI, BEHSHAD
To: GOOGLE INC.
Reel/Frame 059484/0706 →
CHANGE OF NAME Recorded Mar 23, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 059484/0861 →
Continuity (3)
Continuation 15815349 · Nov 16, 2017
Continuation 14312630 · Jun 23, 2014
Related Publication 20220075787A1 · Mar 10, 2022