IP Library › Granted Patent US 11,735,182
Granted Patent B2
US 11,735,182 · App. 17/192,230 · Granted Aug 22, 2023

Multi-modal interaction between users, automated assistants, and other computing services

Inventors: Ulas Kirazci (Mountain View, CA); Adam Coimbra (Cupertino, CA); Abraham Lee (Belmont, CA); Wei Dong (Union City, CA); Thushan Amarasiriwardena (Mountain View, CA)
Assignee: GOOGLE LLC
G10L15/22G06F3/167G06F9/4498G10L13/027G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,735,182
App. No.
17/192,230
Granted
Aug 22, 2023
Kind
B2
Abstract

Techniques are described herein for multi-modal interaction between users, automated assistants, and other computing services. In various implementations, a user may engage with the automated assistant in order to further engage with a third party computing service. In some implementations, the user may advance through dialog state machines associated with third party computing service using both verbal input modalities and input modalities other than verbal modalities, such as visual/tactile modalities.

Claims (36)

1. A method implemented using one or more processors, comprising:

receiving, by a third party computing service implemented at least in part by the one or more processors, first data transmitted over one or more computer networks from an automated assistant, wherein the first data is indicative of touch input provided by a user of a computing device in communication with the automated assistant as part of a human-to-computer dialog session between the user and the automated assistant;

generating resolution information based on the touch input;

updating a visual dialog state machine maintained for the third party computing service in association with the human-to-computer dialog session, wherein the updating is based at least in part on one or both of the touch input and the resolution information;

based on one or more links between one or more visual dialog states of the visual dialog state machine and one or more verbal dialog states of a verbal dialog state machine maintained for the third party computing service in parallel with the visual dialog state machine, automatically updating the verbal dialog state machine to a particular verbal dialog state;

receiving, by the third party computing service, second data transmitted over one or more of the computer networks from the automated assistant, wherein the second data is indicative of speech input provided by the user of the computing device;

generating additional resolution information based on the speech input and the updated verbal dialog state machine;

updating the visual dialog state machine based on the additional resolution information; and

transmitting third data to the automated assistant over one or more of the computer networks, wherein the third data is indicative of the updated visual dialog state machine and causes an assistant application executing on the computing device to trigger a touchless interaction between the user and a graphical user interface of the assistant application.

2. The method of claim 1 , wherein the graphical user interface comprises a web browser embedded into the assistant application.

3. The method of claim 1 , wherein the touchless interaction comprises operation of a selectable element of the graphical user interface.

4. The method of claim 1 , wherein the touchless interaction comprises scrolling to a particular position of a document rendered in the graphical user interface.

5. The method of claim 1 , wherein the touchless interaction comprises zooming in on a portion of the graphical user interface.

6. A system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to:

receive, by a third party computing service implemented at least in part by the one or more processors, first data transmitted over one or more computer networks from an automated assistant, wherein the first data is indicative of touch input provided by a user of a computing device in communication with the automated assistant as part of a human-to-computer dialog session between the user and the automated assistant;

generate resolution information based on the touch input;

update a visual dialog state machine maintained for the third party computing service in association with the human-to-computer dialog session, wherein the updating is based at least in part on one or both of the touch input and the resolution information;

based on one or more links between one or more visual dialog states of the visual dialog state machine and one or more verbal dialog states of a verbal dialog state machine maintained for the third party computing service in parallel with the visual dialog state machine, automatically updating the verbal dialog state machine to a particular verbal dialog state;

receiving, by the third party computing service, second data transmitted over one or more of the computer networks from the automated assistant, wherein the second data is indicative of speech input provided by the user of the computing device;

generating additional resolution information based on the speech input and the updated verbal dialog state machine;

updating the visual dialog state machine based on the additional resolution information; and

transmit third data to the automated assistant over one or more of the computer networks, wherein the third second data is indicative of the updated visual dialog state machine and causes an assistant application executing on the computing device to trigger a touchless interaction between the user and a graphical user interface of the assistant application.

7. The system of claim 6 , wherein the graphical user interface comprises a web browser embedded into the assistant application.

8. The system of claim 6 , wherein the touchless interaction comprises operation of a selectable element of the graphical user interface.

9. The system of claim 6 , wherein the touchless interaction comprises scrolling to a particular position of a document rendered in the graphical user interface.

10. The system of claim 6 , wherein the touchless interaction comprises zooming in on a portion of the graphical user interface.

11. A method implemented using one or more processors, comprising:

receiving, by a third party computing service implemented at least in part by the one or more processors, first data transmitted over one or more computer networks from an automated assistant, wherein the first data is indicative of touch input provided by a user of a computing device in communication with the automated assistant as part of a human-to-computer dialog session between the user and the automated assistant;

generating resolution information based on the touch input;

updating a display context maintained for the third party computing service in association with the human-to-computer dialog session, wherein the updating is based at least in part on one or both of the intent and the resolution information;

based on one or more links between one or more visual dialog states of the visual dialog state machine and one or more verbal dialog states of a verbal dialog state machine maintained for the third party computing service in parallel with the visual dialog state machine, automatically updating the verbal dialog state machine to a particular verbal dialog state;

receiving, by the third party computing service, second data transmitted over one or more of the computer networks from the automated assistant, wherein the second data is indicative of speech input provided by the user of the computing device;

generating additional resolution information based on the speech input and the updated verbal dialog state machine;

updating the verbal dialog state machine based on the additional resolution information; and

transmitting third data to the automated assistant over one or more of the computer networks, wherein the third data is indicative of the updated verbal dialog state machine.

12. The method of claim 11 , wherein the graphical user interface comprises a web browser embedded into the assistant application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: KIRAZCI, ULAS; COIMBRA, ADAM; LEE, ABRAHAM; DONG, WEI; AMARASIRIWARDENA, THUSHAN
To: GOOGLE LLC
Reel/Frame 056206/0779 →
Continuity (2)
Continuation 15774950
Related Publication 20210193146A1 · Jun 24, 2021
Cited By (1)
US 12,243,550