IP Library › Granted Patent US 12,307,163
Granted Patent B1
US 12,307,163 · App. 18/461,671 · Granted May 20, 2025

Multiple results presentation

Inventors: Chaitanya Krishna Reddy Konda (Santa Clara, CA); Nicholas Adam Cummings (Lynnwood, WA)
Assignee: Amazon Technologies, Inc.
G06F3/167G10L13/02G10L15/18G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,163
App. No.
18/461,671
Granted
May 20, 2025
Kind
B1
Abstract

In some disclosed embodiments, first input data corresponding to a first natural language input may be received and processed to determine at least a first natural language understanding (NLU) hypothesis for the first natural language input. First session data identifying a first skill corresponding to the first NLU hypothesis may be determined and used to obtain first visual content corresponding to the first skill. Second session data identifying a second skill may also be determined in response to the input data and be used to obtain second visual content corresponding to the second skill. The device may output a first graphical user interface (GUI) element including the first visual content and a second GUI element including the second visual content. Second input data corresponding to a second input may be received from the device and used to determine, using the second session data, that the second input corresponds to an intent to invoke the second skill.

Claims (100)

1. A computer-implemented method, comprising:

receiving first input audio data corresponding to a first utterance detected by a device;

performing speech processing using the first input audio data to determine at least a first natural language understanding (NLU) hypothesis and a second NLU hypothesis for the first utterance;

determining that the first NLU hypothesis corresponds to a first intent;

determining that the second NLU hypothesis corresponds to a second intent;

determining, that the first NLU hypothesis is more likely accurate than the second NLU hypothesis;

storing first session data including at least a first identifier of a first resource of the device and a second identifier corresponding to a first skill component associated with a first skill corresponding to the first intent;

based at least in part on the first session data, obtaining first visual content corresponding to the first intent;

causing the device to use the first resource to present a first graphical user interface (GUI) element including the first visual content;

determining result data corresponding to the first intent;

performing speech synthesis using the result data to determine output audio data responsive to the first utterance;

causing the device to present output audio corresponding to the output audio data;

in response to the first input audio data, storing second session data including at least a second identifier of a second resource of the device and a fourth identifier corresponding to a second skill component associated with a second skill;

based at least in part on the second session data, obtaining second visual content corresponding to the second skill;

causing the device to use the second resource to present, together with the first GUI element, a second GUI element including the second visual content;

receiving, from the device, input data corresponding to a second input;

determining, based at least in part on the second session data, that the second input corresponds to a command to invoke the second skill;

using at least a portion of the second session data to obtain output content corresponding to the second skill and responsive to the second input; and

causing the device to present the output content.

2. The computer-implemented method of claim 1 , further comprising:

determining a user profile corresponding to the device;

determining, based at least in part on the user profile and the first input audio data, a probability that an action is likely to be requested following the first utterance;

based at least in part on the probability, determining third session data including at least a fourth identifier of a third resource of the device and a fourth identifier corresponding to a third skill component associated with a third skill;

based at least in part on the third session data, obtaining third visual content corresponding to the third skill; and

causing the device to use the third resource to present, together with the first GUI element and the second GUI element, a third GUI element that includes the third visual content and is selectable to cause execution of the action.

3. The computer-implemented method of claim 1 , further comprising:

obtaining at least the second session data in response to receipt of the input data; and

determining that the second input corresponds to the command at least in part by determining that at least a first portion of the input data corresponds to a second portion of the second session data.

4. The computer-implemented method of claim 1 , further comprising:

generating the second session data to identify a subset of resources that the second skill component is permitted to use to perform an action corresponding to the second session data, wherein the subset of resources does not include any resource for generating audio data.

5. A computer-implemented method, comprising:

receiving, from a device, first input data corresponding to a first natural language input;

processing the first input data to determine at least a first natural language understanding (NLU) hypothesis for the first natural language input;

determining first session data identifying a first skill corresponding to the first NLU hypothesis;

based at least in part on the first session data, obtaining first visual content corresponding to the first skill;

in response to the first input data, determining second session data identifying a second skill;

based at least in part on the second session data, obtaining second visual content corresponding to the second skill;

causing the device to output a first graphical user interface (GUI) element including the first visual content and a second GUI element including the second visual content;

receiving, from the device, second input data corresponding to a second input;

determining, based at least in part on the second session data, that the second input corresponds to an intent to invoke the second skill; and

causing the device to output content corresponding to the second skill in response to the second input.

6. The computer-implemented method of claim 5 , further comprising:

based at least in part on the first session data, causing the device to present output audio corresponding to the first skill.

7. The computer-implemented method of claim 6 , further comprising:

based at least in part on the second session data, refraining from causing the device to present output audio corresponding to the second skill.

8. The computer-implemented method of claim 5 , further comprising:

processing the first input data to determine a second NLU hypothesis for the first natural language input; and

determining that the second NLU hypothesis corresponds to the second skill.

9. The computer-implemented method of claim 8 , further comprising:

determining a first confidence value corresponding to the first NLU hypothesis;

determining a second confidence value corresponding to the second NLU hypothesis;

determining that the first confidence value and the second confidence value are outside a first range of similarity corresponding to performing disambiguation before presenting output results; and

determining that the first confidence value and the second confidence value are within a second range of similarity corresponding to outputting results corresponding to the first NLU hypothesis while presenting information corresponding to the second NLU hypothesis.

10. The computer-implemented method of claim 5 , further comprising:

obtaining at least the second session data in response to receipt of the second input data; and

determining that the second input corresponds to the intent at least in part by determining that at least a first portion of the second input data corresponds to a second portion of the second session data.

11. The computer-implemented method of claim 5 , further comprising:

determining a user profile corresponding to the device;

determining, based at least in part on the user profile and the first input data, a predicted action likely to be requested following the first natural language input;

determining third session data identifying a third skill corresponding to the predicted action;

based at least in part on the third session data, obtaining third visual content corresponding to the third skill; and

causing the device to present, together with the first GUI element and the second GUI element, a third GUI that is selectable to cause execution of the predicted action.

12. The computer-implemented method of claim 5 , further comprising:

generating the second session data to identify a subset of resources that a skill component is permitted to use to perform an action corresponding to the second session data, wherein the subset of resources does not include any resource for generating audio data.

13. A system, comprising:

at least one processor; and

at least one non-transitory computer-readable medium encoded with instructions which, when executed by the at least one processor, cause the system to:

receive, from a device, first input data corresponding to a first natural language input;

process the first input data to determine at least a first natural language understanding (NLU) hypothesis for the first natural language input;

determine first session data identifying a first skill corresponding to the first NLU hypothesis;

based at least in part on the first session data, obtain first visual content corresponding to the first skill;

in response to the first input data, determine second session data identifying a second skill;

based at least in part on the second session data, obtain second visual content corresponding to the second skill;

cause the device to output a first graphical user interface (GUI) element including the first visual content and a second GUI element including the second visual content;

receive, from the device, second input data corresponding to a second input;

determine, based at least in part on the second session data, that the second input corresponds to an intent to invoke the second skill; and

cause the device to output content corresponding to the second skill in response to the second input.

14. The system of claim 13 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

based at least in part on the first session data, cause the device to present output audio corresponding to the first skill.

15. The system of claim 14 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

based at least in part on the second session data, refrain from causing the device to present output audio corresponding to the second skill.

16. The system of claim 13 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

process the first input data to determine a second NLU hypothesis for the first natural language input; and

determine that the second NLU hypothesis corresponds to the second skill.

17. The system of claim 16 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

determine a first confidence value corresponding to the first NLU hypothesis;

determine a second confidence value corresponding to the second NLU hypothesis;

determine that the first confidence value and the second confidence value are outside a first range of similarity corresponding to performing disambiguation before presenting output results; and

determine that the first confidence value and the second confidence value are within a second range of similarity corresponding to outputting results corresponding to the first NLU hypothesis while presenting information corresponding to the second NLU hypothesis.

18. The system of claim 13 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

obtain at least the second session data in response to receipt of the second input data; and

determine that the second input corresponds to the intent at least in part by determining that at least a first portion of the second input data corresponds to a second portion of the second session data.

19. The system of claim 13 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

determine a user profile corresponding to the device;

determine, based at least in part on the user profile and the first input data, a predicted action likely to be requested following the first natural language input;

determine third session data identifying a third skill corresponding to the predicted action;

based at least in part on the third session data, obtain third visual content corresponding to the third skill; and

cause the device to present, together with the first GUI element and the second GUI element, a third GUI that is selectable to cause execution of the predicted action.

20. The system of claim 13 , wherein the at least one non-transitory computer-readable medium is further encoded with additional instructions which, when executed by the at least one processor, further cause the system to:

generate the second session data to identify a subset of resources that a skill component is permitted to use to perform an action corresponding to the second session data, wherein the subset of resources does not include any resource for generating audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2023
From: KONDA, CHAITANYA KRISHNA REDDY; CUMMINGS, NICHOLAS ADAM
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064812/0758 →
References Cited (6)
US 11693622B1 · Elders · 2023 [cited by examiner]
US 11900072B1 · Bossio · 2024 [cited by examiner]
US 20190347068A1 · Khaitan · 2019 [cited by examiner]
US 20220310080A1 · Qiu · 2022 [cited by examiner]
US 20240029720A1 · Qiu · 2024 [cited by examiner]
US 20240203406A1 · Khorram · 2024 [cited by examiner]
Cited By (1)
US 12,444,411