IP Library › Granted Patent US 11,714,870
Granted Patent B2
US 11,714,870 · App. 17/122,906 · Granted Aug 1, 2023

Using frames for action dialogs

Inventors: David P. Whipp (San Jose, CA); David Kliger Elson (Brooklyn, NY); Shir Judith Yehoshua (San Francisco, CA)
Assignee: GOOGLE LLC
G06F16/9535G06F3/167G06F16/243G06F16/3329G06F16/953G06F16/9536G10L15/1822G10L15/22H04M3/4936G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,870
App. No.
17/122,906
Granted
Aug 1, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using frames for performing tasks. One of the methods includes receiving a first request to perform a task, the first request comprising user speech identifying the task; generating a frame associated with the task, wherein the frame comprises one or more types of values necessary to perform the task, and wherein each type of value can be satisfied by a respective value; receiving a second request to provide information related to a question, the second request comprising user speech identifying the question; providing information identifying the question to a search engine, and receiving a response identifying one or more terms; determining that at least one term can satisfy a type of value necessary to perform the task; and storing the at least one term in the frame.

Claims (79)

1. A method performed by one or more computers of an automated spoken dialog system, the method comprising:

receiving, from a user device of a user, a first request to perform a task, wherein the first request comprises user speech identifying the task;

in response to receiving the first request:

generating a frame associated with the task, wherein the frame specifies one or more types of values used to perform the task; and

causing output to be provided for presentation to the user, via the user device, that solicits at least one value for one or more of the types of values used to perform the task;

receiving, from the user device, a second request including a question that requests information, the second request comprising user speech identifying the question;

determining, based on processing the user speech identifying the question, that the question is a user request by the user for a search system to access user data of the user device;

in response to determining that the question is the user request by the user for the search system to access the user data of the user device:

causing the search system to execute a search over the user data based on the question;

receiving, responsive to the question indicated by the second request, a response from the search system that identifies a plurality of results from the user data; and

determining, subsequent to receiving the response from the search system that identifies the plurality of results from the user data, that one or more of the plurality of results that are responsive to the question included in the second request provides at least one value corresponding to one or more of the types of values specified by the frame associated with the first request;

storing, in the frame associated with the task requested by the first request and responsive to determining that the one or more of the plurality of results received responsive to the second request provides at least one value corresponding to the one or more types of values specified by the frame associated with the first request, the at least one value that was provided by the one or more of the plurality of results; and

using the stored value in the frame to carry out the task requested by the first request.

2. The method of claim 1 , wherein the response identified respective terms related to a respective question of two or more questions, each question including a different meaning of one or more terms in the question, and wherein the method further comprises:

causing a disambiguation question to be provided for presentation to the user, via the user device, wherein the disambiguation question comprises audio data identifying the different meanings; and

receiving user speech identifying a selection of a particular meaning responsive to the disambiguation question.

3. The method of claim 1 , wherein the one or more types of values include at least one of: a date, a time, a location, a phone number, a user name, a web address, or a person's name.

4. The method of claim 1 , wherein the task comprises: setting a reminder, scheduling a calendar event, providing information, or providing directions.

5. The method of claim 1 , further comprising:

obtaining information identifying a question template for the at least one value;

causing the question template to be provided for presentation to the user via the user device; and

receiving a confirmation of the at least one value, the confirmation comprising user speech that is responsive to the question template.

6. The method of claim 1 , further comprising:

determining that the one or more of the plurality of results provides at least one additional value corresponding to one or more of the types of values specified by the frame;

determining respective confidence levels for the at least one value and the at least one additional value based on one or more rules; and

selecting the at least one value based on the respective confidence levels.

7. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations, the operations comprising:

receiving, from a user device of a user, a first request to perform a task, wherein the first request comprises user speech identifying the task;

in response to receiving the first request:

generating a frame associated with the task, wherein the frame specifies one or more types of values used to perform the task; and

causing output to be provided for presentation to the user, via the user device, that solicits at least one value for one or more of the types of values used to perform the task;

receiving, from the user device, a second request including a question that requests information, the second request comprising user speech identifying the question;

determining, based on processing the user speech identifying the question, that the question is a user request by the user for a search system to access user data of the user device;

in response to determining that the question is the user request by the user for the search system to access the user data of the user device:

causing the search system to execute a search over the user data based on the question;

receiving, responsive to the question indicated by the second request, a response from the search system that identifies a plurality of results from the user data; and

determining, subsequent to receiving the response from the search system that identifies the plurality of results from the user data, that one or more of the plurality of results that are responsive to the question included in the second request provides at least one value corresponding to one or more of the types of values specified by the frame associated with the first request;

storing, in the frame associated with the task requested by the first request and responsive to determining that the one or more of the plurality of results received responsive to the second request provides at least one value corresponding to the one or more types of values specified by the frame associated with the first request, the at least one value that was provided by the one or more of the plurality of results; and

using the stored value in the frame to carry out the task requested by the first request.

8. The system of claim 7 , wherein the response identified respective terms related to a respective question of two or more questions, each question including a different meaning of one or more terms in the question, and wherein the operations further comprising:

causing a disambiguation question to be provided for presentation to the user, via the user device, wherein the disambiguation question comprises audio data identifying the different meanings; and

receiving user speech identifying a selection of a particular meaning responsive to the disambiguation question.

9. The system of claim 7 , wherein the one or more types of values include at least one of: a date, a time, a location, a phone number, a user name, a web address, or a person's name.

10. The system of claim 7 , wherein the task comprises: setting a reminder, scheduling a calendar event, providing information, or providing directions.

11. The system of claim 7 , the operations further comprising:

obtaining information identifying a question template for the at least one value;

causing the question template to be provided for presentation to the user via the user device; and

receiving a confirmation of the at least one value, the confirmation comprising user speech that is responsive to the question template.

12. The system of claim 7 , the operations further comprising:

determining that the one or more of the plurality of results provides at least one additional value corresponding to one or more of the types of values specified by the frame;

determining respective confidence levels for the at least one value and the at least one additional value based on one or more rules; and

selecting the at least one value based on the respective confidence levels.

13. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving, from a user device of a user, a first request to perform a task, wherein the first request comprises user speech identifying the task;

in response to receiving the first request:

generating a frame associated with the task, wherein the frame specifies one or more types of values used to perform the task; and

causing output to be provided for presentation to the user, via the user device, that solicits at least one value for one or more of the types of values used to perform the task;

receiving, from the user device, a second request including a question that requests information, the second request comprising user speech identifying the question;

determining, based on processing the user speech identifying the question, that the question is a user request by the user for a search system to access user data of the user device;

in response to determining that the question is the user request by the user for the search system to access the user data of the user device:

causing the search system to execute a search over the user data based on the question;

receiving, responsive to the question indicated by the second request, a response from the search system that identifies a plurality of results from the user data; and

determining, subsequent to receiving the response from the search system that identifies the plurality of results from the user data, that one or more of the plurality of results that are responsive to the question included in the second request provides at least one value corresponding to one or more of the types of values specified by the frame associated with the first request;

storing, in the frame associated with the task requested by the first request and responsive to determining that the one or more of the plurality of results received responsive to the second request provides at least one value corresponding to the one or more types of values specified by the frame associated with the first request, the at least one value that was provided by the one or more of the plurality of results; and

using the stored value in the frame to carry out the task requested by the first request.

14. The one or more non-transitory computer-readable storage media of claim 13 , wherein the response identified respective terms related to a respective question of two or more questions, each question including a different meaning of one or more terms in the question and wherein the operations further comprising:

causing a disambiguation question to be provided for presentation to the user, via the user device, wherein the disambiguation question comprises audio data identifying the different meanings; and

receiving user speech identifying a selection of a particular meaning responsive to the disambiguation question.

15. The one or more non-transitory computer-readable storage media of claim 13 , wherein the one or more types of values include at least one of: a date, a time, a location, a phone number, a user name, a web address, or a person's name.

16. The one or more non-transitory computer-readable storage media of claim 13 , wherein the task comprises: setting a reminder, scheduling a calendar event, providing information, or providing directions.

17. The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising:

obtaining information identifying a question template for the at least one value;

causing the question template to be provided for presentation to the user via the user device; and

receiving a confirmation of the at least one value, the confirmation comprising user speech that is responsive to the question template.

18. The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising:

determining that the one or more of the plurality of results provides at least one additional value corresponding to one or more of the types of values specified by the frame;

determining respective confidence levels for the at least one value and the at least one additional value based on one or more rules; and

selecting the at least one value based on the respective confidence levels.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2021
From: WHIPP, DAVID P.; ELSON, DAVID KLIGER; YEHOSHUA, SHIR JUDITH
To: GOOGLE INC.
Reel/Frame 055071/0664 →
CHANGE OF NAME Recorded Jan 29, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 055174/0513 →
Continuity (3)
Continuation 14881778 · Oct 13, 2015
Provisional Application 62090093 · Dec 10, 2014
Related Publication 20210103624A1 · Apr 8, 2021