Stylizing text-to-speech (TTS) voice response for assistant systems
In one embodiment, a method includes receiving a voice input having first audio features at a client system, generating a text response corresponding to the voice input, wherein the text response is associated with style features, generating an output audio waveform of the text response by a text-to-speech model on the client system, wherein the output audio waveform is generated based on the first audio features and the style features, wherein the output audio waveform comprises second audio features, and rendering the output audio waveform at the client system in response to the voice input.
1 . A method comprising:
receiving, at a head-wearable device, a first voice input from a user, the first voice input having one or more first audio features;
determining a first voice input emotion of the first voice input based on the one or more first audio features;
generating a first text response corresponding to the first voice input;
in accordance with a determination, based on the first voice input emotion, that the one or more first audio features are appropriate for responding to the user, generating, by a text-to-speech model on the head-wearable device, a first output audio waveform of the first text response, wherein the first output audio waveform has the one or more first audio features;
in response to the first voice input, presenting, at the head-wearable device, the first output audio waveform;
receiving, at the head-wearable device, a second voice input from the user, the second voice input having one or more second audio features distinct from the first audio features;
determining a background noise level;
determining a second voice input emotion, distinct from the first voice input emotion, of the second voice input based on the one or more second audio features and the background noise level;
generating a second text response corresponding to the second voice input;
in accordance with a determination that the backgrounds noise level does not exceed a background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are not appropriate for responding to the user;
generating, by the text-to-speech model on the head-wearable device, a second output audio waveform of the second text response, wherein the second output audio waveform has one or more other audio features, distinct from the one or more second audio features; and
in response to the second voice input, presenting, at the head-wearable device, the second output audio waveform;
receiving, at the head-wearable device, a third voice input from the user, the third voice input having the one or more second audio features distinct from the first audio features;
determining a second background noise level;
determining the second voice input emotion based on the one or more second audio features and the second background noise level;
generating a third text response corresponding to the third voice input;
in accordance with a determination that the second background noise level exceeds the background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are appropriate for responding to the user;
generating, by the text-to-speech model on the head-wearable device, a third output audio waveform of the third text response, wherein the third output audio waveform has the one or more second audio features; and
in response to the third voice input, presenting, at the head-wearable device, the third output audio waveform.
2 . The method of claim 1 , wherein:
the one or more first audio features include one or more of a first tone, a first speed, a first volume, a first emphasis, a first pronunciation, a first accent, a first frequency, and a first prosody;
the one or more second audio features include one or more of a second tone, a second speed, a second volume, a second emphasis, a second pronunciation, a second accent, a second frequency, and a second prosody; and
the one or more other audio features include one or more of another tone, another speed, another volume, another emphasis, another pronunciation, another accent, another frequency, and another prosody.
3 . The method of claim 1 , further comprising:
before generating the second output audio waveform of the second text response, determining the one or more other audio features based on at least the second voice input emotion and the second voice input.
4 . The method of claim 3 , wherein the one or more other audio features are further based on one or more of:
environmental features associated with the user;
a time of receipt associated with the second voice input; and
one or more user settings associated with the user.
5 . The method of claim 3 , wherein the determining the first voice input emotion, the determining the second voice input emotion, and the determining the one or more other audio features are performed by a machine-learning model.
6 . The method of claim 1 , wherein determining the one or more other audio features based on at least the second voice input emotion and the second voice input is in accordance with the determination, based on the second voice input emotion, that the one or more other audio features are more appropriate than the one or more second audio features for responding to the user.
7 . The method of claim 1 , wherein:
the determination that the one or more first audio features are appropriate for responding to the user includes a determination that the first voice input emotion is one of a positive emotion and a neutral emotion; and
the determination that the one or more second audio features are not appropriate for responding to the user includes a determination that the second voice input emotion is a negative emotion; and
the determination that the one or more second audio features are appropriate for responding to the user includes a determination that the second voice input emotion is the negative emotion.
8 . The method of claim 1 , the method further comprising:
after receiving the first voice input, automatically transcribing the first voice input into a first input text, wherein generating the first text response corresponding to the first voice input is further based on the first input text;
after receiving the second voice input, automatically transcribing the second voice input into a second input text, wherein generating the second text response corresponding to the second voice input is further based on the second input text; and
after receiving the third voice input, automatically transcribing the third voice input into a third input text, wherein generating the third text response corresponding to the third voice input is further based on the third input text.
9 . The method of claim 1 , wherein the text-to-speech model is based on a prosody model and an acoustic model.
10 . The method of claim 9 , wherein the prosody model determines:
a first duration and a first frequency of the first output audio waveform;
a second duration and a second frequency of the second output audio waveform; and
a third duration and a third frequency of the third output audio waveform.
11 . The method of claim 10 , wherein generating the second output audio waveform of the second text response comprises:
accessing, by the prosody model, a style embedding corresponding to the one or more other audio features; and
setting one or more of the second duration and the second frequency of the second output audio waveform based on the style embedding to generate response prosody features of the second output audio waveform.
12 . The method of claim 1 , wherein:
the head-wearable device includes one or more speakers;
presenting the first output audio waveform includes playing the first output audio waveform at the one or more speakers;
presenting the second output audio waveform includes playing the second output audio waveform at the one or more speakers; and
presenting the third output audio waveform includes playing the third output audio waveform at the one or more speakers.
13 . A system comprising one or more processors and a memory coupled to the one or more processors, the memory including executable instructions that, when executed by the one or more processors, cause the one or more processors to:
receive a first voice input from a user of a head-wearable device, the first voice input having one or more first audio features;
determine a first voice input emotion of the first voice input based on the one or more first audio features;
generate a first text response corresponding to the first voice input;
in accordance with a determination, based on the first voice input emotion, that the one or more first audio features are appropriate for responding to the user of the head-wearable device, generate, by a text-to-speech model, a first output audio waveform of the first text response, wherein the first output audio waveform has the one or more first audio features;
in response to the first voice input, cause the first output audio waveform to be presented to the user of the head-wearable device;
receive a second voice input from the user of the head-wearable device, the second voice input having one or more second audio features distinct from the first audio features;
determine a background noise level;
determine a second voice input emotion, distinct from the first voice input emotion, of the second voice input based on the one or more second audio features and the background noise level;
generate a second text response corresponding to the second voice input;
in accordance with a determination that the background noise level does not exceed a background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are not appropriate for responding to the user of the head-wearable device:
generate, by the text-to-speech model, a second output audio waveform of the second text response, wherein the second output audio waveform has one or more other audio features, distinct from the one or more second audio features; and
in response to the second voice input, cause the second output audio waveform to be presented to the user of the head-wearable device:
receive a third voice input from the user of the head-wearable device, the third voice input having the one or more second audio features distinct from the first audio features;
determine a second background noise level;
determine the second voice input emotion based on the one or more second audio features and the second background noise level;
generate a third text response corresponding to the third voice input;
in accordance with a determination that the second background noise level exceeds the background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are appropriate for responding to the user of the head-wearable device:
generate, by the text-to-speech model, a third output audio waveform of the third text response, wherein the third output audio waveform has the one or more second audio features; and
in response to the third voice input, cause the third output audio waveform to be presented to the user of the head-wearable device.
14 . The system of claim 13 , wherein:
the one or more first audio features include one or more of a first tone, a first speed, a first volume, a first emphasis, a first pronunciation, a first accent, a first frequency, and a first prosody;
the one or more second audio features include one or more of a second tone, a second speed, a second volume, a second emphasis, a second pronunciation, a second accent, a second frequency, and a second prosody; and
the one or more other audio features include one or more of another tone, another speed, another volume, another emphasis, another pronunciation, another accent, another frequency, and another prosody.
15 . The system of claim 13 , wherein the executable instructions further cause the one or more processors to:
before generating the second output audio waveform of the second text response, determine the one or more other audio features based on at least the second voice input emotion and the second voice input.
16 . The system of claim 13 , wherein:
the determination that the one or more first audio features are appropriate for responding to the user includes a determination that the first voice input emotion is one of a positive emotion and a neutral emotion;
the determination that the one or more second audio features are not appropriate for responding to the user includes a determination that the second voice input emotion is a negative emotion; and
the determination that the one or more second audio features are appropriate for responding to the user includes a determination that the second voice input emotion is the negative emotion.
17 . A computer-readable non-transitory storage medium including executable instructions that, when executed by one or more processors, cause the one or more processors to:
receive a first voice input from a user of a head-wearable device, the first voice input having one or more first audio features;
determine a first voice input emotion of the first voice input based on the one or more first audio features;
generate a first text response corresponding to the first voice input;
in accordance with a determination, based on the first voice input emotion, that the one or more first audio features are appropriate for responding to the user of the head-wearable device, generate, by a text-to-speech model, a first output audio waveform of the first text response, wherein the first output audio waveform has the one or more first audio features;
in response to the first voice input, cause the first output audio waveform to be presented to the user of the head-wearable device;
receive a second voice input from the user of the head-wearable device, the second voice input having one or more second audio features distinct from the first audio features;
determine a background noise level;
determine a second voice input emotion, distinct from the first voice input emotion, of the second voice input based on the one or more second audio features and the background noise level;
generate a second text response corresponding to the second voice input;
in accordance with a determination that the background noise level does not exceed a background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are not appropriate for responding to the user of the head-wearable device;
generate, by the text-to-speech model, a second output audio waveform of the second text response, wherein the second output audio waveform has one or more other audio features, distinct from the one or more second audio features; and
in response to the second voice input, cause the second output audio waveform to be presented to the user of the head-wearable device;
receive a third voice input from the user of the head-wearable device, the third voice input having the one or more second audio features distinct from the first audio features;
determine a second background noise level;
determine the second voice input emotion based on the one or more second audio features and the second background noise level;
generate a third text response corresponding to the third voice input;
in accordance with a determination that the second background noise level exceeds the background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are appropriate for responding to the user of the head-wearable device:
generate, by the text-to-speech model, a third output audio waveform of the third text response, wherein the third output audio waveform has the one or more second audio features; and
in response to the third voice input, cause the third output audio waveform to be presented to the user of the head-wearable device.
18 . The computer-readable non-transitory storage medium of claim 17 , wherein:
the one or more first audio features include one or more of a first tone, a first speed, a first volume, a first emphasis, a first pronunciation, a first accent, a first frequency, and a first prosody;
the one or more second audio features include one or more of a second tone, a second speed, a second volume, a second emphasis, a second pronunciation, a second accent, a second frequency, and a second prosody; and
the one or more other audio features include one or more of another tone, another speed, another volume, another emphasis, another pronunciation, another accent, another frequency, and another prosody.
19 . The computer-readable non-transitory storage medium of claim 17 , wherein the executable instructions further cause the one or more processors to:
before generating the second output audio waveform of the second text response, determine the one or more other audio features based on at least the second voice input emotion and the second voice input.
20 . The computer-readable non-transitory storage medium of claim 17 , wherein:
the determination that the one or more first audio features are appropriate for responding to the user includes a determination that the first voice input emotion is one of a positive emotion and a neutral emotion;
the determination that the one or more second audio features are not appropriate for responding to the user includes a determination that the second voice input emotion is a negative emotion; and
the determination that the one or more second audio features are appropriate for responding to the user includes a determination that the second voice input emotion is the negative emotion.