IP Library › Granted Patent US 11,170,772
Granted Patent B2
US 11,170,772 · App. 16/269,275 · Granted Nov 9, 2021

Multi-modal interaction between users, automated assistants, and other computing services

Inventors: Ulas Kirazci (Mountain View, CA); Adam Coimbra (Cupertino, CA); Abraham Lee (Belmont, CA); Wei Dong (Union City, CA); Thushan Amarasiriwardena (Mountain View, CA); Yudong Sun (Cupertino, CA); Xiao Gao (Fremont, CA)
Assignee: GOOGLE LLC
G10L15/22G06F3/0485G06F3/16G10L13/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,170,772
App. No.
16/269,275
Granted
Nov 9, 2021
Kind
B2
Abstract

Techniques are described herein for multi-modal interaction between users, automated assistants, and other computing services. In various implementations, a user may engage with the automated assistant in order to further engage with a third party computing service. In some implementations, the user may advance through dialog state machines associated with third party computing service using both verbal input modalities and input modalities other than verbal modalities, such as visual/tactile modalities.

Claims (40)

1. A method implemented using one or more processors, comprising:

receiving, by a computing service implemented at least in part by the one or more processors, from a server associated with an automated assistant, data indicative of an intent of a user of a computing device in communication with the automated assistant as part of a human-to-computer dialog session between the user and the automated assistant,

wherein the automated assistant executes in part on the server and in part on an automated assistant application executing on the computing device of the user, and

wherein the intent of the user is not resolvable by the automated assistant;

resolving, by the computing service, the intent of the user to generate resolution information,

wherein resolving the intent of the user to generate resolution information includes at least one of:

controlling or configuring a smart home device of the user, or

procuring a physical item for the user;

updating a display context maintained for a visual dialog state machine of the computing service in association with the human-to-computer dialog session, wherein the updating is based at least in part on one or both of the intent and the resolution information; and

providing, by the computing service and over one or more communication networks via an application programming interface (“API”), data indicative of the display context to the server associated with the automated assistant, wherein the data indicative of the display context is provided by the server associated with the automated assistant to the automated assistant application executing on the computing device,

wherein the data indicative of the display context causes the automated assistant application executing on the computing device to invoke a function embedded in a markup language document already rendered on a graphical user interface of the automated assistant application,

wherein invocation of the function triggers a touchless interaction between the user and the graphical user interface of the automated assistant application, and

wherein the touchless interaction between the user and the graphical user interface of the automated assistant application visually represents the updating of the display context maintained for the visual dialog state machine of the computing service.

2. The method of claim 1 , wherein the graphical user interface comprises a web browser embedded into the assistant application.

3. The method of claim 2 , wherein the embedded web browser renders graphics based on the markup language document provided by the computing service to the automated assistant application executing on the computing device.

4. The method of claim 1 , wherein the embedded function comprises a JavaScript function call.

5. The method of claim 1 , wherein the data indicative of the intent of the user comprises speech recognition output of vocal free form input provided by the user at the computing device.

6. The method of claim 5 , further comprising determining, by the computing service, the intent of the user based on the speech recognition output.

7. The method of claim 1 , wherein the touchless interaction comprises operation of a selectable element of the graphical user interface.

8. The method of claim 1 , wherein the touchless interaction comprises scrolling to a particular position of a document rendered in the graphical user interface, or zooming in on a portion of the graphical user interface.

9. A system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to:

receive, by a computing service implemented at least in part by the one or more processors, from a server associated with an automated assistant, data indicative of an intent of a user of a computing device in communication with the automated assistant as part of a human-to-computer dialog session between the user and the automated assistant,

wherein the automated assistant executes in part on the server and in part on an automated assistant application executing on the computing device of the user, and

wherein the intent of the user is not resolvable by the automated assistant;

resolve, by the computing service, the intent of the user to generate resolution information,

wherein resolving the intent of the user to generate resolution information includes at least one of:

controlling or configuring a smart home device of the user, or

procuring a physical item for the user;

update a display context maintained for a visual dialog state machine of the computing service in association with the human-to-computer dialog session, wherein the updating is based at least in part on one or both of the intent and the resolution information; and

provide, by the computing service and over one or more communication networks via an application programming interface (“API”), data indicative of the display context to the server associated with the automated assistant, wherein the data indicative of the display context is provided by the server associated with the automated assistant to the automated assistant application executing on computing device,

wherein the data indicative of the display context causes the automated assistant application executing on the computing device to invoke a function embedded in a markup language document already rendered on a graphical user interface of the automated assistant application,

wherein invocation of the function triggers a touchless interaction between the user and the graphical user interface of the automated assistant application and

wherein the touchless interaction between the user and the graphical user interface of the automated assistant application visually represents the updating of the display context maintained for the visual dialog state machine of the computing service.

10. The system of claim 9 , wherein the graphical user interface comprises a web browser embedded into the assistant application.

11. The system of claim 9 , wherein the embedded function comprises a JavaScript function call.

12. The system of claim 9 , wherein the data indicative of the intent of the user comprises speech recognition output of vocal free form input provided by the user at the computing device.

13. The system of claim 12 , further comprising instructions to, by the computing service, the intent of the user based on the speech recognition output.

14. The system of claim 9 , wherein the touchless interaction comprises operation of a selectable element of the graphical user interface.

15. The system of claim 9 , wherein the touchless interaction comprises scrolling to a particular position of a document rendered in the graphical user interface, or zooming in on a portion of the graphical user interface.

16. The system of claim 10 , wherein the embedded web browser renders graphics based on the markup language document provided by the computing service to the automated assistant application executing on the computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2019
From: KIRAZCI, ULAS; COIMBRA, ADAM; LEE, ABRAHAM; DONG, WEI; AMARASIRIWARDENA, THUSHAN; SUN, YUDONG; GAO, XIAO
To: GOOGLE LLC
Reel/Frame 048255/0784 →
Continuity (2)
Continuation In Part 15774950
Related Publication 20190341040A1 · Nov 7, 2019