IP Library › Granted Patent US 11,741,965
Granted Patent B1
US 11,741,965 · App. 16/913,139 · Granted Aug 29, 2023

Configurable natural language output

Inventors: Ramsey Abou-Zaki Opp (Santa Barbara, CA); Anantdeep Gill (San Jose, CA); Angela Liu (Santa Barbara, CA); Anisha Jain (Santa Barbara, CA); Justin Maxwell Bollag (Santa Barbara, CA); Nathan Yeazel (Santa Barbara, CA); Sara Renee Bilich (Santa Barbara, CA); Spencer B Baker (Santa Barbara, CA)
Assignee: Amazon Technologies, Inc.
G10L15/26G10L13/00G10L13/086G10L15/005G10L15/18G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,965
App. No.
16/913,139
Granted
Aug 29, 2023
Kind
B1
Abstract

A system is provided for determining a natural language output, responsive to a user input, using different speech personality profiles. The system may determine to user a particular language generation profile based at least in part on data relating to the user input and data corresponding to the response to the user input. The language generation profile may include different attributes that are used to determine the natural language output, such as, prosody, replacement words, injected words, sentence structure, etc.

Claims (115)

1. A computer-implemented method comprising:

receiving, from a first device, first audio data representing a first utterance;

determining first output data responsive to the first utterance, the first output data being a first natural language output including a first plurality of words;

determining the first device corresponds to a first location;

identifying, using a model configured to determine a language generation profile, a first language generation profile associated with the first location, wherein the model was trained using location data and a plurality of language generation profiles;

using a natural language generation (NLG) component, processing the first output data to determine second output data representing a second natural language output, wherein processing the first output data comprises:

determining the first language generation profile represents a first word to be inserted in the first natural language output,

determining the first language generation profile represents a position indicating that the first word is to be inserted after the first plurality of words, and

determining the second output data to include the first plurality of words followed by the first word;

processing, using text-to-speech (TTS) processing, the second output data to determine first output audio data representing first synthesized speech; and

sending the first output audio data to the first device.

2. The computer-implemented method of claim 1 , further comprising:

receiving, from a second device, second audio data representing a second utterance;

performing spoken language understanding (SLU) processing on the second audio data to determine entity data corresponding to the second utterance;

determining third output data responsive to the second utterance, the third output data being a third natural language output including a second plurality of words;

identifying, using the model, a second language generation profile associated with an entity type associated with the entity data, wherein the model was further trained using entity types and the plurality of language generation profiles;

using the NLG component, processing the third output data to determine fourth output data representing a fourth natural language output, wherein processing the third output data comprises:

determining the second language generation profile represents a second word to be inserted before the third natural language output,

determining the second language generation profile indicates a second speech prosody to be applied to the second word,

determining a second SSML tag based on the second speech prosody,

determining the fourth output data to include the second word followed by the second plurality of words, and

associating the second SSML tag with a portion of the fourth output data, the portion corresponding to the second word;

processing the fourth output data to determine second output audio data representing second synthesized speech; and

sending the second output audio data to the second device.

3. The computer-implemented method of claim 1 , further comprising:

receiving second audio data representing a second utterance, the second audio data being associated with a user profile;

determining third output data responsive to the second utterance, the third output data being a third natural language output;

identifying, using the model, a second language generation profile associated with the user profile, wherein the model was further trained using user profile data and the plurality of language generation profiles;

using the NLG component, processing the third output data to determine fourth output data representing a fourth natural language output, wherein processing the third output data comprises:

determining a topic category corresponding to the third output data, and

determining, based on the topic category, to disregard the second language generation profile in determining the fourth natural language output; and

processing the fourth output data to determine second output audio data representing second synthesized speech.

4. The computer-implemented method of claim 1 , further comprising:

receiving second audio data representing a second utterance, the second audio data associated with a user profile;

determining third output data responsive to the second utterance, the third output data being a third natural language output including a second plurality of words;

identifying, using the model, a second language generation profile associated with the user profile, wherein the model was further trained using user profile data and the plurality of language generation profiles;

using the NLG component, processing the third output data to determine fourth output data representing a fourth natural language output, wherein processing the third output data comprises:

determining the second language generation profile represents a second word that is to replace a third word of the second plurality of words,

determining the second language generation profile indicates an arrangement for the second plurality of words, and

determining the fourth output data to include a third plurality of words comprising of the second word and words other than the third word of the second plurality of words, the third plurality of words arranged according to the arrangement; and

processing the fourth output data to determine second output audio data representing second synthesized speech.

5. A computer-implemented method comprising:

receiving, from a device, first data corresponding to a first user input;

receiving second data representing a first natural language output responsive to the first user input, the first natural language output including at least a first word;

processing, using a model configured to determine a language generation profile, at least the first data to determine a first language generation profile to be used to respond to the first user input, the first language generation profile including first word data, wherein the model was trained using a plurality of language generation profiles;

selecting, based on the second data, at least a second word from the first word data wherein the first word data includes third data representing an insertion position of the at least second word;

determining, using the second data, first output data representing a second natural language output, wherein the second natural language output includes the at least first word and the at least second word and wherein the at least second word is positioned, based on the third data, relative to the at least first word; and

processing, using text-to-speech (TTS) processing, the first output data to determine output audio data.

6. The computer-implemented method of claim 5 , further comprising:

determining that a third word in the first natural language output corresponds to an entity;

determining that the first language generation profile is associated with an entity type associated with the entity, the first word data representing a fourth word to be inserted after the third word; and

determining the second natural language output to include the at least first word, the second word, the third word, and the fourth word after the third word.

7. The computer-implemented method of claim 5 , wherein processing the first data further comprises:

determining a device identifier associated with the device that received the first user input, the first data including the device identifier; and

determining that the first language generation profile is associated with the device identifier.

8. The computer-implemented method of claim 5 , wherein determining the first output data further comprises:

determining the first word data represents the second word to be inserted in the first natural language output after the at least first word; and

determining the second natural language output to include the at least first word followed by the second word.

9. The computer-implemented method of claim 8 , wherein determining the first output data further comprises:

determining the first language generation profile representing a speech prosody to be applied to the second word;

determining a speech synthesis markup language (SSML) tag based on the speech prosody; and

determining the first output data to include the SSML tag associated with the second word.

10. The computer-implemented method of claim 5 , wherein determining the first output data further comprises:

determining the first word data represents at least a third word that is to replace the at least first word;

determining the first language generation profile indicates an arrangement of words to be used; and

determining the first output data to include a plurality of words comprising the third word instead of the first word and at least a fourth word included in the first natural language output, the plurality of words being arranged according to the arrangement.

11. The computer-implemented method of claim 5 , further comprising:

based on the first data and the second data, requesting, from a spoken language understanding (SLU) processing component, additional information to include in the second natural language output; and

receiving fourth data representing first additional information to include in the second natural language output,

wherein determining the first output data further comprises including at least a third word representing the fourth data.

12. The computer-implemented method of claim 5 , further comprising:

receiving fourth data corresponding to a second user input;

receiving fifth data representing a third natural language output responsive to the second user input;

processing, using the model, the fourth data to determine a second language generation profile to be used to respond to the second user input;

determining a topic category corresponding to the fifth data;

determining, based on the topic category, to disregard the second language generation profile in determining second output data; and

processing the fifth data to determine second output audio data.

13. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive, from a device, first data corresponding to a first user input;

receive second data representing a first natural language output responsive to the first user input, the first natural language output including at least a first word;

process, using a model configured to determine a language generation profile, at least the first data to determine a first language generation profile to be used to respond to the first user input, the first language generation profile including first word data, wherein the model was trained using a plurality of language generation profiles;

select, based on the second data, at least a second word from the first word data wherein the first word data includes third data representing an insertion position of the at least second word;

determine, using the second data, first output data representing a second natural language output, wherein the second natural language output includes the at least first word and the at least second word and wherein the at least second word is positioned, based on the third data, relative to the at least first word; and

process, using text-to-speech (TTS) processing, the first output data to determine output audio data.

14. The system of claim 13 , wherein the instructions that cause the system to process the first data further causes the system to:

determine that a third word in the first natural language output corresponds to an entity;

determine that the first language generation profile is associated with an entity type associated with the entity, the first word data representing a fourth word to be inserted after the third word; and

determine the second natural language output to include the at least first word, the second word, the third word, and the fourth word after the third word.

15. The system of claim 13 , wherein the instructions that cause the system to process the first data further causes the system to:

determine a device identifier associated with the device that received the first user input, the first data including the device identifier; and

determine that the first language generation profile is associated with the device identifier.

16. The system of claim 13 , wherein the instructions that cause the system to determine the first output data further causes the system to:

determine the first word data represents the second word to be inserted in the first natural language output after the at least first word; and

determine the second natural language output to include the at least first word followed by the second word.

17. The system of claim 16 , wherein the instructions that cause the system to determine the first output data further causes the system to:

determine the first language generation profile representing a speech prosody to be applied to the second word;

determine a speech synthesis markup language (SSML) tag based on the speech prosody; and

determine the first output data to include the SSML tag associated with the second word.

18. The system of claim 13 , wherein the instructions that cause the system to determine the first output data further causes the system to:

determine the first word data represents at least a third word that is to replace the at least first word;

determine the first language generation profile indicates an arrangement of words to be used; and

determine the first output data include a plurality of words comprising the third word instead of the first word and at least a fourth word included in the first natural language output, the plurality of words being arranged according to the arrangement.

19. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

based on the first data and the second data, request, from a spoken language understanding (SLU) processing component, additional information to include in the second natural language output; and

receive fourth data representing first additional information to include in the second natural language output,

wherein the first output data includes at least a third word representing the fourth data.

20. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

receive fourth data corresponding to a second user input;

receive fifth data representing a third natural language output responsive to the second user input;

process, using the model, the fourth data to determine a second language generation profile to be used to respond to the second user input;

determine a topic category corresponding to the fifth data;

determine, based on the topic category, to disregard the second language generation profile in determining second output data; and

process the fifth data to determine second output audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2020
From: OPP, RAMSEY ABOU-ZAKI; GILL, ANANTDEEP; LIU, ANGELA; JAIN, ANISHA; BOLLAG, JUSTIN MAXWELL; YEAZEL, NATHAN; BILICH, SARA RENEE; BAKER, SPENCER B
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 053128/0632 →
Cited By (2)
US 12,417,756 US 12,444,418