IP Library › Granted Patent US 12,347,430
Granted Patent B2
US 12,347,430 · App. 17/892,803 · Granted Jul 1, 2025

Automated assistant adaptation of a response to an utterance and/or of processing of the utterance, based on determined interaction measure

Inventors: Victor Carbune (Zurich, CH); Matthew Sharifi (Kilchberg, CH)
Assignee: GOOGLE LLC
G10L15/22G06F3/167G06F16/3344G06N20/00G10L2015/223G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,430
App. No.
17/892,803
Granted
Jul 1, 2025
Kind
B2
Abstract

Implementations set forth herein relate to an automated assistant that provides a response for certain user queries based on a level of interaction of the user with respect to the automated assistant. Interaction can be characterized by sensor data, which can be processed using one or more trained machine learning models in order to identify parameters for generating a response. In this way, the response can be limited to preserve computational resources and/or ensure that the response is more readily understood given the amount of interaction exhibited by the user. In some instances, a response that embodies information that is supplemental, to an otherwise suitable response, can be provided when a user is exhibiting a particular level of interaction. In other instances, such supplemental information can be withheld when the user is not exhibiting that particular level of interaction, at least in order to preserve computational resources.

Claims (70)

1. A method implemented by one or more processors, the method comprising:

determining that a user has provided a user input directed to an automated assistant via a client device of the user;

determining a first level of interaction of a user at one or more instances of time;

generating, based on the first level of interaction, automated assistant output data that includes content that is responsive to the user input and that is also responsive to one or more queries that are associated with a context of the user at the one or more instances of time,

wherein the context of the user at the one or more instances of time is determined based on a physical activity being performed by the user, in an environment of the user, at the one or more instances of time,

wherein the physical activity being performed by the user, in the environment of the user, at the one or more instances of time is determined based on sensor data generated by the client device or an additional client device of the user that is in addition to the client device, and

wherein the one or more queries that are associated with the context of the user at the one or more instances of time are determined based on the physical activity being performed by the user, in the environment of the user, at the one or more instances of time;

causing, based on the automated assistant output data, the automated assistant to render a first responsive output for presentation to the user via the client device; and

subsequent to the automated assistant rendering the first responsive output:

determining that the user has provided an additional user input directed to the automated assistant via the client device or the additional client device of the user;

determining a second level of interaction for the user at one or more subsequent instances of time that are subsequent to the one or more instances of time;

generating, based on the second level of interaction, additional automated assistant output data that includes other content that is responsive to the additional user input but is not responsive to any queries associated with a subsequent context of the user at the one or more subsequent instances of time,

wherein the subsequent context of the user at the one or more subsequent instances of time is determined based on a subsequent physical activity being performed by the user, in an environment of the user, at the one or more subsequent instances of time,

wherein the subsequent physical activity being performed by the user, in the environment of the user, at the one or more subsequent instances of time is determined based on subsequent sensor data generated by the client device or the additional client device, and

wherein no queries are associated with the subsequent context of the user at the one or more subsequent instances of time based on the subsequent physical activity being performed by the user, in the environment of the user, at the one or more subsequent instances of time; and

causing, based on the additional automated assistant output data, the automated assistant to render a second responsive output for presentation to the user via the client device or the additional client device,

wherein the second responsive output is different from the first responsive output.

2. The method of claim 1 , wherein the first level of interaction is a metric that characterizes an estimated amount of attention the user is providing to the automated assistant at the one or more instances of time.

3. The method of 2 , wherein the second level of interaction is an additional metric that characterizes the estimated amount of attention the user is providing to the automated assistant at the one or more subsequent instances of time.

4. The method of claim 1 , wherein the user input is an instance of a spoken utterance directed to the automated assistant, and wherein the additional user input is an additional instance of the spoken utterance directed to the automated assistant.

5. The method of 4 , wherein generating the additional automated assistant output data includes:

selecting, based on the second level of interaction, a voice filter for filtering audio that is different from a voice of the user, and

processing the audio data using the voice filter.

6. The method of 4 , wherein generating the additional automated assistant output data includes:

processing, based on the second level of interaction, the audio data using an automatic speech recognition (ASR) process that is biased towards speech of the user.

7. The method of claim 1 , wherein the user input is an instance of typed input or touch input directed to the automated assistant, and wherein the additional user input is an additional instance of typed input or touch input directed to the automated assistant.

8. The method of claim 1 , wherein the additional user input is directed to the automated assistant via the additional client device.

9. The method of claim 1 , wherein the other content that is responsive to the additional user input is also responsive to one or more additional queries that are associated with an additional context of the user at the one or more subsequent instances of time.

10. The method of claim 9 , wherein the additional context of the user at the one or more subsequent instances of time includes one or more additional environmental features of an additional environment of the user and/or one or more additional user features of the user, wherein the one or more additional environmental features of the additional environment of the user and/or the one or more user features of the user are determined based on additional sensor data generated by the one or more sensors of the client device and/or the one or more additional sensors of the additional client device, wherein the one or more additional queries that are associated with the additional context of the user at the one or more subsequent instances of time are generated based on the one or more additional environmental features of the environment of the user and/or the one or more additional user features of the user, and wherein the one or more additional queries differ from the one or more queries.

11. A method implemented by one or more processors, the method comprising:

determining that a user has provided a user input directed to an automated assistant via a client device of the user;

causing an automated assistant output that is responsive to the user input to be provided for presentation to the user,

wherein the automated assistant output is rendered for presentation to the user at one or more interfaces of the client device;

detecting a gaze of the user that is directed to the automated assistant output and while the automated assistant is providing the automated assistant output for presentation to the user;

determining, based on detecting the gaze of the user that is directed to the automated assistant and while the automated assistant is providing the automated assistant output for presentation to the user, that a change in a level of interaction of the user has occurred; and

in response to determining that the change in the level of interaction of the user has occurred and without the user having provided an additional user input directed to the automated assistant:

identifying, based on detecting the gaze of the user that is directed to the automated assistant and while the automated assistant is providing the automated assistant output for presentation to the user, a portion of the automated assistant output that was being actively rendered when the gaze of the user is directed to the automated assistant;

generating, based on the portion of the automated assistant output that was being actively rendered when the gaze of the user is directed to the automated assistant, additional content that is associated with the portion of the automated assistant output that was being actively rendered when the change in the level of interaction occurred; and

causing, based on the additional content, an additional automated assistant output to be provided for presentation to the user,

wherein the additional automated assistant output is rendered for presentation to the user at one or more of the interfaces of the client device or one or more additional interfaces of an additional client device of the user, and

wherein the additional automated assistant output is different from the automated assistant output.

12. The method of claim 11 , wherein the automated assistant output is based on a first document, and the additional content is selected from a second document that is different from the first document.

13. The method of claim 11 , wherein the automated assistant output is responsive to a first query that was included in the user input, and wherein the additional content is responsive to a second query that is different from the first query and that was not previously provided by the user to the automated assistant.

14. The method of claim 11 ,

wherein providing the automated assistant output to the user includes rendering the automated assistant output in response to the user providing the first query to the automated assistant while the user is exhibiting one or more user characteristics; and

wherein determining that the change in the level of interaction of the user has occurred includes determining that the one or more user characteristics changed in response to the automated assistant providing the portion of the automated assistant output.

15. The method of claim 11 , wherein generating the additional content includes selecting one or more features of the additional content based on the change in the level of interaction, and wherein the automated assistant renders the additional automated assistant output as an audio output or a video output that embodies the one or more features.

16. The method of claim 11 , further comprising:

in response to determining that no change in the level of interaction of the user has occurred and without the user having provided the additional user input directed to the automated assistant:

refraining from generating the additional content.

17. The method of claim 11 , wherein the additional user input is directed to the automated assistant via the additional client device.

18. A system comprising:

at least one processor; and

memory storing instructions that, when executed, cause the at least one processor to:

determine that a user has provided a user input directed to an automated assistant via a client device of the user;

determine a first level of interaction of a user at one or more instances of time;

generate, based on the first level of interaction, automated assistant output data that includes content that is responsive to the user input and that is also responsive to one or more queries that are associated with a context of the user at the one or more instances of time,

wherein the context of the user at the one or more instances of time is determined based on a physical activity being performed by the user, in an environment of the user, at the one or more instances of time,

wherein the physical activity being performed by the user, in the environment of the user, at the one or more instances of time is determined based on sensor data generated by the client device or an additional client device of the user that is in addition to the client device, and

wherein the one or more queries that are associated with the context of the user at the one or more instances of time are determined based on the physical activity being performed by the user, in the environment of the user, at the one or more instances of time;

cause, based on the automated assistant output data, the automated assistant to render a first responsive output for presentation to the user via the client device; and

subsequent to the automated assistant rendering the first responsive output:

determine that the user has provided an additional user input directed to the automated assistant via the client device or the additional client device of the user that is in addition to the client device;

determine a second level of interaction for the user at one or more subsequent instances of time that are subsequent to the one or more instances of time;

generate, based on the second level of interaction, additional automated assistant output data that includes other content that is responsive to the additional instance of the user input but is not responsive to any queries associated with a subsequent context of the user at the one or more subsequent instances of time,

wherein the subsequent context of the user at the one or more subsequent instances of time is determined based on a subsequent physical activity being performed by the user, in an environment of the user, at the one or more subsequent instances of time,

wherein the subsequent physical activity being performed by the user, in the environment of the user, at the one or more subsequent instances of time is determined based on subsequent sensor data generated by the client device or the additional client device, and

wherein no queries are associated with the subsequent context of the user at the one or more subsequent instances of time based on the subsequent physical activity being performed by the user, in the environment of the user, at the one or more subsequent instances of time; and

cause, based on the additional automated assistant output data, the automated assistant to render a second responsive output for presentation to the user via the client device or the additional client device,

wherein the second responsive output is different than the first responsive output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2022
From: CARBUNE, VICTOR; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 060864/0951 →
Continuity (3)
Continuation 16947513 · Aug 4, 2020
Provisional Application 63057099 · Jul 27, 2020
Related Publication 20220392449A1 · Dec 8, 2022
References Cited (17)
US 10192551B2 · Carbune et al. · 2019 [cited by applicant]
US 10594757B1 · Shevchenko · 2020 [cited by examiner]
US 10706873B2 · Tsiartas · 2020 [cited by examiner]
US 10783876B1 · Hoover · 2020 [cited by examiner]
US 10937420B2 · Park · 2021 [cited by examiner]
US 11016968B1 · Hoover · 2021 [cited by examiner]
US 20020097848A1 · Wesemann · 2002 [cited by examiner]
US 20200193264A1 · Zavesky · 2020 [cited by examiner]
US 20200327895A1 · Gruber · 2020 [cited by examiner]
US 20200379787A1 · Martin · 2020 [cited by examiner]
US 20200380980A1 · Shum · 2020 [cited by examiner]
US 20210020177A1 · Oh · 2021 [cited by examiner]
US 20210043208A1 · Luan · 2021 [cited by examiner]
US 20210118450A1 · Gorsica · 2021 [cited by examiner]
US 20210304756A1 · Iwase · 2021 [cited by examiner]
US 20220028379A1 · Carbune et al. · 2022 [cited by applicant]
WO 2019236372 · 2019 [cited by applicant]