IP Library › Granted Patent US 12,271,408
Granted Patent B2
US 12,271,408 · App. 18/400,826 · Granted Apr 8, 2025

Streaming real-time dialog management

Inventors: David Elson (Port Washington, NY); Christa Wimberley (Mountain View, CA); Benjamin Ross (Mountain View, CA); David Eisenberg (New York, NY); Sudeep Gandhe (Mountain View, CA); Kevin Chavez (Mountain View, CA); Raj Agarwal (Mountain View, CA)
Assignee: GOOGLE LLC
G06F16/3329G06F16/00G06F16/3344G06F40/30G06N20/00G06Q10/10G10L15/22G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,408
App. No.
18/400,826
Granted
Apr 8, 2025
Kind
B2
Abstract

Systems and methods provides for dialog management in real-time rather than turn taking. An example method included generating first candidate responses to triggering event. The triggering event may be receipt of a live stream chunk for the dialog or receipt of a backend response to a previous backend request for a dialog shema. The method also includes updating a list of candidate responses that are accepted or pending with at least on of the first candidate responses, and determining, for the triggering event, whether the list of candidate responses includes a candidate response that has a confidence score that meets a triggering threshold. The method also includes waiting for a next triggering event without providing a candidate response when the list does not include a candidate response that has a confidence score that meets the triggering threshold.

Claims (48)

1. A method implemented by one or more processors, the method comprising:

receiving, at an assistant device, a first input from a user during a real-time dialog between the user and the assistant device, wherein the first input from the user is received at one or more of a graphical interface, an audio interface, or a haptic interface of the assistant device;

transmitting, by the assistant device, features of the first input to a remote device;

receiving, at the assistant device, one or more predicted responses generated in response to the transmitting of the features of the first input by the assistant device to the remote device, wherein one or more of the predicted responses correspond to a particular schema of one or more schemas;

storing, during the real-time dialog between the user and the assistant device, the one or more predicted responses and data indicating that the one or more predicted responses correspond to the particular schema;

receiving, at the assistant device and subsequent to receiving the first input, a second input from the user during the real-time dialog between the user and the assistant device, wherein the second input is different from the first input, and wherein the second input from the user is received at one or more of the graphical interface, the audio interface, or the haptic interface of the assistant device, and/or one or more of a graphical interface, a audio interface, or a haptic interface of another assistant device;

in response to receiving the second input:

identifying, based on processing the second input, that the second input also corresponds to the particular schema;

determining, based on processing the data indicating that the one or more predicted responses correspond to the particular schema and based on identifying that the second input also corresponds to the particular schema to provide a predicted response, of the one or more predicted responses that correspond to the particular schema, to the user; and

causing, based on determining to provide the predicted response, one or more of the assistant device or the another assistant device to render a prompt for the user regarding the predicted response.

2. The method of claim 1 , wherein determining whether to provide the predicted response is based on how much time has passed between receiving the first input and receiving the second input.

3. The method of claim 1 , wherein determining whether to provide the predicted response is based on a duration of the real-time dialog.

4. The method of claim 1 , wherein determining whether to provide the predicted response is based on an intonation associated with the first input and/or the second input.

5. The method of claim 1 , wherein the one or more predicted responses include an expression that indicates attention or comprehension.

6. The method of claim 5 , wherein the one or more predicted responses include “uh-huh”, “hmm”, or “right”.

7. The method of claim 1 , wherein the first input and/or the second input include natural language input (NLI).

8. The method of claim 1 , wherein causing prompting regarding the predicted response comprises rendering, by the assistant device, the predicted response as natural language output (NLO).

9. A system comprising:

one or more computers comprising one or more processors, and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more processors to perform operations comprising:

receiving, at an assistant device, a first input from a user during a real-time dialog between the user and the assistant device, wherein the first input from the user is received at one or more of a graphical interface, an audio interface, or a haptic interface of the assistant device;

transmitting, by the assistant device, features of the first input to a remote device;

receiving, at the assistant device, one or more predicted responses generated in response to the transmitting of the features of the first input by the assistant device to the remote device, wherein one or more of the predicted responses correspond to a particular schema of one or more schemas;

storing, during the real-time dialog between the user and the assistant device, the one or more predicted responses and data indicating that the one or more predicted responses correspond to the particular schema;

receiving, at the assistant device and subsequent to receiving the first input, a second input from the user during the real-time dialog between the user and the assistant device, wherein the second input is different from the first input, and wherein the second input from the user is received at one or more of the graphical interface, the audio interface, or the haptic interface of the assistant device, and/or one or more of a graphical interface, an audio interface, or a haptic interface of another assistant device;

in response to receiving the second input:

identifying, based on processing the second input, that the second input also corresponds to the particular schema;

determining, based on processing the data indicating that the one or more predicted responses correspond to the particular schema and based on identifying that the second input also corresponds to the particular schema to provide a predicted response, of the one or more predicted responses that correspond to the particular schema, to the user; and

causing, based on determining to provide the predicted response, one or more of the assistant device or the other assistant device to render a prompt for the user regarding the predicted response.

10. The system of claim 9 , wherein determining whether to provide the predicted response is based on how much time has passed between receiving the first input and receiving the second input.

11. The system of claim 9 , wherein determining whether to provide the predicted response is based on a duration of the real-time dialog.

12. The system of claim 9 , wherein determining whether to provide the predicted response is based on an intonation associated with the first input and/or the second input.

13. The system of claim 9 , wherein the one or more predicted responses include an expression that indicates attention or comprehension.

14. The system of claim 13 , wherein the one or more predicted responses include “uh-huh”, “hmm”, or “right”.

15. The system of claim 9 , wherein the first input and/or the second input include natural language input (NLI).

16. The system of claim 9 , wherein causing prompting regarding the predicted response comprises rendering, by the assistant device, the predicted response as natural language output (NLO).

17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, at an assistant device, a first input from a user during a real-time dialog between the user and the assistant device, wherein the first input from the user is received at one or more of a graphical interface, an audio interface, or a haptic interface of the assistant device;

transmitting, by the assistant device, features of the first input to a remote device;

receiving, at the assistant device, one or more predicted responses generated in response to the transmitting of the features of the first input by the assistant device to the remote device, wherein one or more of the predicted responses correspond to a particular schema of one or more schemas;

storing, during the real-time dialog between the user and the assistant device, the one or more predicted responses and data indicating that the one or more predicted responses correspond to the particular schema;

receiving, at the assistant device and subsequent to receiving the first input, a second input from the user during the real-time dialog between the user and the assistant device, wherein the second input is different from the first input, and wherein the second input from the user is received at one or more of the graphical interface, the audio interface, or the haptic interface of the assistant device, and/or a graphical interface, an audio interface, or a haptic interface of another assistant device;

in response to receiving the second input:

identifying, based on processing the second input, that the second input also corresponds to the particular schema;

determining, based on processing the data indicating that the one or more predicted responses correspond to the particular schema and based on identifying that the second input also corresponds to the particular schema to provide a predicted response, of the one or more predicted responses that correspond to the particular schema, to the user; and

causing, based on determining to provide the predicted response, one or more of the assistant device or the other assistant device to render a prompt for the user regarding the predicted response.

18. The non-transitory computer-readable medium of claim 17 , wherein determining whether to provide the predicted response is based on how much time has passed between receiving the first input and receiving the second input.

19. The non-transitory computer-readable medium of claim 17 , wherein determining whether to provide the predicted response is based on a duration of the real-time dialog.

20. The non-transitory computer-readable medium of claim 17 , wherein determining whether to provide the predicted response is based on an intonation associated with the first input and/or the second input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2024
From: ELSON, DAVID; WIMBERLEY, CHRISTA; ROSS, BENJAMIN; EISENBERG, DAVID; GANDHE, SUDEEP; CHAVEZ, KEVIN; AGARWAL, RAJ
To: GOOGLE LLC
Reel/Frame 066068/0175 →
Continuity (5)
Continuation 18088270 · Dec 23, 2022
Continuation 17114350 · Dec 7, 2020
Continuation 15783290 · Oct 13, 2017
Provisional Application 62459820 · Feb 16, 2017
Related Publication 20240134893A1 · Apr 25, 2024
References Cited (31)
US 8478584B1 · Grove · 2013 [cited by examiner]
US 9158974B1 · Laska · 2015 [cited by examiner]
US 10319042B2 · Arvapally et al. · 2019 [cited by applicant]
US 10504521B1 · Taubman et al. · 2019 [cited by applicant]
US 10860628B2 · Elson et al. · 2020 [cited by applicant]
US 11537646B2 · Elson et al. · 2022 [cited by applicant]
US 20100036667A1 · Byford · 2010 [cited by examiner]
US 20120078889A1 · Chu-Carroll et al. · 2012 [cited by applicant]
US 20140006012A1 · Zhou et al. · 2014 [cited by applicant]
US 20140163959A1 · Hebert et al. · 2014 [cited by applicant]
US 20140310001A1 · Kalns et al. · 2014 [cited by applicant]
US 20140365209A1 · Evermann · 2014 [cited by applicant]
US 20160171114A1 · Whipp et al. · 2016 [cited by applicant]
US 20160188565A1 · Robichaud et al. · 2016 [cited by applicant]
US 20160196110A1 · Yehoshua et al. · 2016 [cited by applicant]
US 20170009831A1 · Iwasaki et al. · 2017 [cited by applicant]
US 20170032791A1 · Elson et al. · 2017 [cited by applicant]
US 20210089565A1 · Elson et al. · 2021 [cited by applicant]
US 20230132020A1 · Elson et al. · 2023 [cited by applicant]
CN 103631853 · 2014 [cited by applicant]
CN 106227779 · 2016 [cited by applicant]
EP 1343144 · 2003 [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 17794470.9; 49 pages; dated Jul. 23, 2024. [cited by applicant]
European Patent Office; Communication pursuant to Article 94(3) EPC issued in Application No. 17794470.9; 5 pages; dated Jan. 30, 2023. [cited by applicant]
China Intellectual Property; Notice of Allowance issued in Application No. 201711035649.7; 5 pages; dated Nov. 25, 2021. [cited by applicant]
Deutsches Patent Office; Examination Report issued in Application No. 102017125001.8; 18 pages; dated May 14, 2021. [cited by applicant]
China Intellectual Property; Notice of Office Action issued in Application No. 201711035649.7; 5 pages; dated Jun. 2, 2021. [cited by applicant]
European Patent Office; International Preliminary Report on Patentability of PCT Ser. No. PCT/US2017/056720; 19 pages; dated Mar. 14, 2019 Mar. 14, 2019. [cited by applicant]
International Search Report and Written Opinion of PCT Ser. No. PCT/US2017/056720; 14 pages Jan. 19, 2018. [cited by applicant]
United Kingdom Intellectual Property Office; Examination Report issued in Application No. 1717421.0 dated Mar. 13, 2018 Mar. 13, 2018. [cited by applicant]
United Kingdom Intellectual Property Office; Examination Report issued in Application No. GB1717421.0; 6 pages; dated Mar. 24, 2021. [cited by applicant]