IP Library › Granted Patent US 11,537,646
Granted Patent B2
US 11,537,646 · App. 17/114,350 · Granted Dec 27, 2022

Streaming real-time dialog management

Inventors: David Elson (Port Washington, NY); Christa Wimberley (Mountain View, CA); Benjamin Ross (New York, NY); David Eisenberg (New York, NY); Sudeep Gandhe (Sunnyvale, CA); Kevin Chavez (Palo Alto, CA); Raj Agarwal (Fremont, CA)
Assignee: GOOGLE LLC
G06F16/3329G06F16/00G06F16/3344G06F40/30G06N20/00G06Q10/10G10L15/22G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,537,646
App. No.
17/114,350
Granted
Dec 27, 2022
Kind
B2
Abstract

Systems and methods provides for dialog management in real-time rather than turn taking. An example method included generating first candidate responses to triggering event. The triggering event may be receipt of a live stream chunk for the dialog or receipt of a backend response to a previous backend request for a dialog shema. The method also includes updating a list of candidate responses that are accepted or pending with at least on of the first candidate responses, and determining, for the triggering event, whether the list of candidate responses includes a candidate response that has a confidence score that meets a triggering threshold. The method also includes waiting for a next triggering event without providing a candidate response when the list does not include a candidate response that has a confidence score that meets the triggering threshold.

Claims (53)

1. A computing device configured to manage a real-time dialog with a user, the computing device comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the computing device to:

receive an initial chunk or a next chunk of a series of chunks, the series of chunks representing a live-stream of speech from a user;

generate a first candidate responses based on receiving the chunk;

update, with at least one of the first candidate responses, a ranked list of candidate responses that, at the time of receipt of the chunk, are accepted or pending, the updated ranked list of candidate responses including one or more backend requests for one or more dialog schemas and one or more system responses generated by one or more of the dialog schemas;

execute a backend request to two or more of the dialog schemas based on the updated ranked list; and

subsequent to executing the backend request:

generate backend responses based on information, for the two or more dialog schemas, obtained responsive to the backend request;

derive a composite candidate response from the two or more dialog schemas based on the generated backend responses;

further update the updated ranked list based on the composite candidate response; and

prune the further updated ranked list based on updated ranks of the candidate responses of the further updated ranked list.

2. The computing device of claim 1 , wherein each candidate response in the ranked list of candidate responses has a corresponding dialog state and is assigned to a path in a dialog beam, the dialog beam including at least two paths.

3. The computing device of claim 2 , the operations further comprising:

determining that the further updated ranked list includes a system response candidate with a confidence score satisfying a triggering threshold; and

triggering that system response.

4. The computing device of claim 3 , wherein each candidate response in the ranked list has respective annotations and a respective dialog state and ranking the response candidates includes:

processing, by a machine learning model, the annotations and the candidate responses to determine a confidence score for each candidate response in the ranked list.

5. The computing device of claim 3 , wherein each backend request of the updated ranked list of candidate responses is associated with a provisional dialog state, and wherein at least one of the generated backend responses is generated based on the provisional dialog state.

6. A method comprising:

receiving an initial chunk or a next chunk of a series of chunks, the series of chunks representing a live-stream of speech from a user;

generating a first candidate responses based on receiving the chunk;

updating, with at least one of the first candidate responses, a ranked list of candidate responses that, at the time of receipt of the chunk, are accepted or pending, the updated ranked list of candidate responses including one or more backend requests for one or more dialog schemas and one or more system responses generated by one or more of the dialog schemas;

executing a backend request to two or more of the dialog schemas based on the updated ranked list; and

subsequent to executing the backend request:

generating backend responses based on information, for the two or more dialog schemas, obtained responsive to the backend request;

deriving a composite candidate response from the two or more dialog schemas based on the generated backend responses;

further updating the updated ranked list based on the composite candidate response; and

pruning the further updated ranked list based on updated ranks of the candidate responses of the further updated ranked list.

7. The method of claim 6 , wherein each candidate response in the ranked list of candidate responses has a corresponding dialog state and is assigned to a path in a dialog beam, the dialog beam including at least two paths.

8. The method of claim 7 , further comprising:

determining that the further updated ranked list includes a system response candidate with a confidence score satisfying a triggering threshold; and

triggering that system response.

9. The method of claim 8 , wherein each candidate response in the ranked list has respective annotations and a respective dialog state and ranking the response candidates includes:

processing, by a machine learning model, the annotations and the candidate responses to determine a confidence score for each candidate response in the ranked list.

10. The method of claim 8 , wherein each backend request of the updated ranked list of candidate responses is associated with a provisional dialog state, and wherein at least one of the generated backend responses is generated based on the provisional dialog state.

11. A non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to:

receive an initial chunk or a next chunk of a series of chunks, the series of chunks representing a live-stream of speech from a user;

generate a first candidate responses based on receiving the chunk;

update, with at least one of the first candidate responses, a ranked list of candidate responses that, at the time of receipt of the chunk, are accepted or pending, the updated ranked list of candidate responses including one or more backend requests for one or more dialog schemas and one or more system responses generated by one or more of the dialog schemas;

execute a backend request to two or more of the dialog schemas based on the updated ranked list; and

subsequent to executing the backend request:

generate backend responses based on information, for the two or more dialog schemas, obtained responsive to the backend request;

derive a composite candidate response from the two or more dialog schemas based on the generated backend responses;

further update the updated ranked list based on the composite candidate response; and

prune the further updated ranked list based on updated ranks of the candidate responses of the further updated ranked list.

12. The non-transitory computer-readable storage medium of claim 11 , wherein each candidate response in the ranked list of candidate responses has a corresponding dialog state and is assigned to a path in a dialog beam, the dialog beam including at least two paths.

13. The non-transitory computer-readable storage medium of claim 12 , the operations further comprising:

determining that the further updated ranked list includes a system response candidate with a confidence score satisfying a triggering threshold; and

triggering that system response.

14. The non-transitory computer-readable storage medium of claim 13 , wherein each candidate response in the ranked list has respective annotations and a respective dialog state and ranking the response candidates includes:

processing, by a machine learning model, the annotations and the candidate responses to determine a confidence score for each candidate response in the ranked list.

15. The non-transitory computer-readable storage medium of claim 13 , wherein each backend request of the updated ranked list of candidate responses is associated with a provisional dialog state, and wherein at least one of the generated backend responses is generated based on the provisional dialog state.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2021
From: ELSON, DAVID; WIMBERLEY, CHRISTA; ROSS, BENJAMIN; EISENBERG, DAVID; GANDHE, SUDEEP; CHAVEZ, KEVIN; AGARWAL, RAJ
To: GOOGLE LLC
Reel/Frame 055426/0112 →
Continuity (3)
Continuation 15783290 · Oct 13, 2017
Provisional Application 62459820 · Feb 16, 2017
Related Publication 20210089565A1 · Mar 25, 2021
Cited By (1)
US 12,271,408