IP Library › Granted Patent US 11,699,441
Granted Patent B2
US 11,699,441 · App. 17/573,014 · Granted Jul 11, 2023

Contextual content for voice user interfaces

Inventors: Mark Conrad Kockerbeck (Irvine, CA); Muhammad Yahia (Anaheim, CA); Jordan Michael Hughes (San Diego, CA); Kevin Boehm (Seattle, WA); Rohit Sauhta (San Diego, CA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L15/1815G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,699,441
App. No.
17/573,014
Granted
Jul 11, 2023
Kind
B2
Abstract

The present disclosure describes techniques for dynamically determining when information is to be output to a user, as well as what information is to be output to a user. A natural language processing system may receive, from a first device, first data representing information to be output at a first point during a skill session. The natural language processing system may also receive, from a second device, second data representing a natural language input. The natural language processing system may determine a skill component is to execute with respect to the natural language input. The natural language processing system may send, to the skill component, second data representing the natural language input. The natural language processing system may receive, from the skill component, an indication that an ongoing first skill session with the second device has reached the first point. After receiving the indication and based at least in part on system usage data associated with at least one user, the natural language processing system may determine third data representing a prompt corresponding to the information and send, to the second device, the third data for output.

Claims (44)

1. A computer-implemented method comprising:

determining first data corresponding to timing for presentation of information during an exchange with a natural language processing system;

receiving, from a first device and after determining the first data, second data representing a first natural language input;

determining a first application is to execute with respect to the first natural language input;

sending, to the first application, third data representing the first natural language input;

processing at least the first data and dialog data using a first component to determine a first indication that an ongoing first dialog corresponding to the first natural language input has reached a first point corresponding to the presentation of information;

determining first information to be output at the first point; and

causing, based at least in part on receiving the first indication, the first device to output the first information.

2. The computer-implemented method of claim 1 , further comprising:

determining context data corresponding to a user identifier associated with the ongoing first dialog,

wherein determining the first data is based at least in part on the context data.

3. The computer-implemented method of claim 2 , wherein the context data corresponds to a frequency a user has interacted with the first application.

4. The computer-implemented method of claim 2 , wherein the context data corresponds to a geographic location.

5. The computer-implemented method of claim 1 , further comprising:

determining the first point corresponds to an end of the first dialog.

6. The computer-implemented method of claim 1 , wherein the first information corresponds to a purchase offer.

7. The computer-implemented method of claim 1 , wherein the first information corresponds to a second application different from the first application.

8. The computer-implemented method of claim 1 , further comprising:

after causing the first device to output the first information, resuming the first dialog.

9. The computer-implemented method of claim 1 , wherein determining the first information comprises using a trained model to determine the first information.

10. The computer-implemented method of claim 1 , wherein the first component comprises a trained model.

11. A system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

determine first data corresponding to timing for presentation of information during an exchange with a natural language processing system;

receive, from a first device and after determining the first data, second data representing a first natural language input;

determine a first application is to execute with respect to the first natural language input;

send, to the first application, third data representing the first natural language input;

process at least the first data and dialog data using a first component to determine a first indication that an ongoing first dialog corresponding to the first natural language input has reached a first point corresponding to the presentation of information;

determine first information to be output at the first point; and

cause, based at least in part on receiving the first indication, the first device to output the first information.

12. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine context data corresponding to a user identifier associated with the ongoing first dialog,

wherein determination of the first data is based at least in part on the context data.

13. The system of claim 12 , wherein the context data corresponds to a frequency a user has interacted with the first application.

14. The system of claim 12 , wherein the context data corresponds to a geographic location.

15. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine the first point corresponds to an end of the first dialog.

16. The system of claim 11 , wherein the first information corresponds to a purchase offer.

17. The system of claim 11 , wherein the first information corresponds to a second application different from the first application.

18. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

after causing the first device to output the first information, resume the first dialog.

19. The system of claim 11 , wherein determination of the first information comprises using a trained model to determine the first information.

20. The system of claim 11 , wherein the first component comprises a trained model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2022
From: KOCKERBECK, MARK CONRAD; YAHIA, MUHAMMAD; HUGHES, JORDAN MICHAEL; BOEHM, KEVIN; SAUHTA, ROHIT
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 058618/0987 →
Continuity (2)
Continuation 16455530 · Jun 27, 2019
Related Publication 20220130389A1 · Apr 28, 2022
Cited By (1)
US 12,254,881