IP Library Granted Patent US 11,688,191
Granted Patent B2
US 11,688,191 · App. 17/941,971 · Granted Jun 27, 2023

Contextually disambiguating queries

Inventors: Ibrahim Badr (Zurich, CH); Nils Grimsmo (Adliswil, CH); Gokhan H. Bakir (Zurich, CH); Kamil Anikiej (Lachen, CH); Aayush Kumar (Zurich, CH); Viacheslav Kuznetsov (Rüschlikon, IN)
Assignee: GOOGLE LLC
G06V30/262G06F16/5866G06F16/9032G06V10/768G06V20/63G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,688,191
App. No.
17/941,971
Granted
Jun 27, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for contextually disambiguating queries are disclosed. In an aspect, a method includes receiving an image being presented on a display of a computing device and a transcription of an utterance spoken by a user of the computing device, identifying a particular sub-image that is included in the image, and based on performing image recognition on the particular sub-image, determining one or more first labels that indicate a context of the particular sub-image. The method also includes, based on performing text recognition on a portion of the image other than the particular sub-image, determining one or more second labels that indicate the context of the particular sub-image, based on the transcription, the first labels, and the second labels, generating a search query, and providing, for output, the search query.

Claims (62)

1. A method implemented by one or more processors, the method comprising:

determining, by a client device of a user, to generate a search query for the user based on an image capturing screen content displayed by the client device at a particular time and based on a voice input of the user received subsequent to the particular time;

processing the image of the screen content displayed by the client device at the particular time to identify a particular sub-image of a plurality of disparate sub-images included in the image of the screen content displayed by the client device at the particular time;

processing a plurality of separate portions of the particular sub-image to generate a plurality of labels that each correspond to at least one of the separate portions of the particular sub-image included in the image of the screen content displayed by the client device at the particular time;

receiving, by the client device and subsequent to the particular time, audio data including the voice input of the user;

selecting a particular subset of the plurality of labels based on a transcription of the voice input of the user and based on identifying a screen content type associated with a particular portion of the particular sub-image, of the plurality of separate portions of the particular sub-image, that is associated with the particular subset of the plurality of labels;

generating the search query for the user based on the transcription of the voice input of the user and the particular selected subset of the plurality of labels that correspond to the particular portion of the particular sub-image; and

providing, for display at the client device of the user, one or more search results obtained responsive to the search query that was generated for the user.

2. The method of claim 1 , wherein generating the search query for the user based on the transcription of the voice input of the user and the particular selected subset of the plurality of labels that correspond to the particular disparate portion of the particular sub-image includes generating the search query to include at least one first term included in transcription of the voice input of the user and at least one second term associated with the particular selected subset of the plurality of labels.

3. The method of claim 1 , wherein processing the image of the screen content displayed by the client device at the particular time to identify the particular sub-image of the plurality of disparate sub-images included in the image further includes processing the image of the screen content to identify an additional particular sub-image of the plurality of disparate sub-images included in the image, and further comprising:

processing a plurality of additional separate portions of the additional particular sub-image to generate a plurality of additional labels that each correspond to at least one of the additional separate portions of the additional particular sub-image included in the image of the screen content displayed by the client device at the particular time;

selecting, for use in generating the search query for the user, at least one additional label that corresponds to at least one of the additional separate portions of the additional particular sub-image based on identifying a type of screen content respective screen content types associated with the at least one additional separate portion of the additional particular sub-image.

4. The method of claim 3 , wherein generating the search query for the user using the at least one additional label that corresponds to the at least one of the additional separate portions of the additional particular sub-image includes:

generating the transcription of the voice input of the user based on the at least one additional label; and

generating the search query for the user based on the transcription of the voice input.

5. The method of claim 1 , wherein generating the search query for the user includes:

generating a plurality of candidate search queries;

comparing the plurality of candidate search queries to a plurality of recent search queries associated with a plurality of users; and

selecting a candidate search query, of the plurality of candidate search queries, to be the search query for the user based on a frequency of each of the candidate search queries of the plurality appears in the plurality of recent search queries.

6. The method of claim 1 , wherein the screen content displayed by the client device of the user at the particular time includes video content.

7. A system, comprising:

one or more processors; and

memory storing instructions that, when executed by one or more of the processors, cause the one or more processors to perform operations comprising:

determining, by a client device of a user, to generate a search query for the user based on an image capturing screen content displayed by the client device at a particular time and based on a voice input of the user received subsequent to the particular time;

processing the image of the screen content displayed by the client device at the particular time to identify a particular sub-image of a plurality of disparate sub-images included in the image of the screen content displayed by the client device at the particular time;

processing a plurality of separate portions of the particular sub-image to generate a plurality of labels that each correspond to at least one of the separate portions of the particular sub-image included in the image of the screen content displayed by the client device at the particular time;

receiving, by the client device and subsequent to the particular time, audio data including the voice input of the user;

selecting a particular subset of the plurality of labels based on a transcription of the voice input of the user and based on identifying a screen content type associated with a particular portion of the particular sub-image, of the plurality of separate portions of the particular sub-image, that is associated with the particular subset of the plurality of labels;

generating the search query for the user based on the transcription of the voice input of the user and the particular selected subset of the plurality of labels that correspond to the particular portion of the particular sub-image; and

providing, for display at the client device of the user, one or more search results obtained responsive to the search query that was generated for the user.

8. The system of claim 7 , wherein generating the search query for the user based on the transcription of the voice input of the user and the particular selected subset of the plurality of labels that correspond to the particular disparate portion of the particular sub-image includes generating the search query to include at least one first term included in transcription of the voice input of the user and at least one second term associated with the particular selected subset of the plurality of labels.

9. The system of claim 7 , wherein processing the image of the screen content displayed by the client device at the particular time to identify the particular sub-image of the plurality of disparate sub-images included in the image further includes processing the image of the screen content to identify an additional particular sub-image of the plurality of disparate sub-images included in the image, and the operations further comprising:

processing a plurality of additional separate portions of the additional particular sub-image to generate a plurality of additional labels that each correspond to at least one of the additional separate portions of the additional particular sub-image included in the image of the screen content displayed by the client device at the particular time;

selecting, for use in generating the search query for the user, at least one additional label that corresponds to at least one of the additional separate portions of the additional particular sub-image based on identifying a type of screen content respective screen content types associated with the at least one additional separate portion of the additional particular sub-image.

10. The system of claim 9 , wherein generating the search query for the user using the at least one additional label that corresponds to the at least one of the additional separate portions of the additional particular sub-image includes:

generating the transcription of the voice input of the user based on the at least one additional label; and

generating the search query for the user based on the transcription of the voice input.

11. The system of claim 7 , wherein generating the search query for the user includes:

generating a plurality of candidate search queries;

comparing the plurality of candidate search queries to a plurality of recent search queries associated with a plurality of users; and

selecting a candidate search query, of the plurality of candidate search queries, to be the search query for the user based on a frequency of each of the candidate search queries of the plurality appears in the plurality of recent search queries.

12. The system of claim 7 , wherein the screen content displayed by the client device of the user at the particular time includes video content.

13. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

determining, by a client device of a user, to generate a search query for the user based on an image capturing screen content displayed by the client device at a particular time and based on a voice input of the user received subsequent to the particular time;

processing the image of the screen content displayed by the client device at the particular time to identify a particular sub-image of a plurality of disparate sub-images included in the image of the screen content displayed by the client device at the particular time;

processing a plurality of separate portions of the particular sub-image to generate a plurality of labels that each correspond to at least one of the separate portions of the particular sub-image included in the image of the screen content displayed by the client device at the particular time;

receiving, by the client device and subsequent to the particular time, audio data including the voice input of the user;

selecting a particular subset of the plurality of labels based on a transcription of the voice input of the user and based on identifying a screen content type associated with a particular portion of the particular sub-image, of the plurality of separate portions of the particular sub-image, that is associated with the particular subset of the plurality of labels;

generating the search query for the user based on the transcription of the voice input of the user and the particular selected subset of the plurality of labels that correspond to the particular portion of the particular sub-image; and

providing, for display at the client device of the user, one or more search results obtained responsive to the search query that was generated for the user.

14. The one or more non-transitory computer-readable storage media of claim 13 , wherein generating the search query for the user based on the transcription of the voice input of the user and the particular selected subset of the plurality of labels that correspond to the particular disparate portion of the particular sub-image includes generating the search query to include at least one first term included in transcription of the voice input of the user and at least one second term associated with the particular selected subset of the plurality of labels.

15. The one or more non-transitory computer-readable storage media of claim 13 , wherein processing the image of the screen content displayed by the client device at the particular time to identify the particular sub-image of the plurality of disparate sub-images included in the image further includes processing the image of the screen content to identify an additional particular sub-image of the plurality of disparate sub-images included in the image, and the operations further comprising:

processing a plurality of additional separate portions of the additional particular sub-image to generate a plurality of additional labels that each correspond to at least one of the additional separate portions of the additional particular sub-image included in the image of the screen content displayed by the client device at the particular time;

selecting, for use in generating the search query for the user, at least one additional label that corresponds to at least one of the additional separate portions of the additional particular sub-image based on identifying a type of screen content respective screen content types associated with the at least one additional separate portion of the additional particular sub-image.

16. The one or more non-transitory computer-readable storage media of claim 15 , wherein generating the search query for the user using the at least one additional label that corresponds to the at least one of the additional separate portions of the additional particular sub-image includes:

generating the transcription of the voice input of the user based on the at least one additional label; and

generating the search query for the user based on the transcription of the voice input.

17. The one or more non-transitory computer-readable storage media of claim 13 , wherein generating the search query for the user includes:

generating a plurality of candidate search queries;

comparing the plurality of candidate search queries to a plurality of recent search queries associated with a plurality of users; and

selecting a candidate search query, of the plurality of candidate search queries, to be the search query for the user based on a frequency of each of the candidate search queries of the plurality appears in the plurality of recent search queries.

18. The one or more non-transitory computer-readable storage media of claim 13 , wherein the screen content displayed by the client device of the user at the particular time includes video content.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2022
From: BADR, IBRAHIM; GRIMSMO, NILS; BAKIR, GOKHAN H.; ANIKIEJ, KAMIL; KUMAR, AAYUSH; KUZNETSOV, VIACHESLAV
To: GOOGLE INC.
Reel/Frame 061073/0043 →
CHANGE OF NAME Recorded Sep 13, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 061419/0424 →
Continuity (3)
Continuation 16731786 · Dec 31, 2019
Continuation 15463018 · Mar 20, 2017
Related Publication 20230004597A1 · Jan 5, 2023