IP Library Granted Patent US 12688863
Granted Patent B2
US 12688863 · App. 17/539,405 · Granted Jul 21, 2026

Emotionally-aware voice response generation method and apparatus

Inventors: Subham Biswas (Thane, IN); Saurabh Tahiliani (Noida, IN)
Assignee: Verizon Patent and Licensing Inc.
G10L25/63G10L15/22G10L15/26G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688863
App. No.
17/539,405
Granted
Jul 21, 2026
Kind
B2
Abstract

Techniques for generating emotionally-aware audio, or voice, responses for a user interface of an application, such as an automated voice response application, are disclosed. In one embodiment, a method is disclosed comprising obtaining voice input from a user via an automated voice response user interface of an application, obtaining a textual representation of the voice input, using the textual representation of the voice input from a user to obtain a source emotion of the user, determining a response emotion using the source emotion, generating a response textual representation indicating textual content of the response, generating a frequency spectrum representation of the response in accordance with the response textual representation and the response emotion, using the frequency spectrum representation of the response to generate a voice response reflective of the textual content of the response and the response emotion, and communicating the response to the user via the user interface.

Claims (72)

1 . A method comprising:

obtaining, by a computing device, voice input from a user via an automated voice response user interface of an application;

obtaining, by the computing device, a textual representation of the voice input;

using, by the computing device, the textual representation of the voice input to obtain a source emotion of the user;

determining, by the computing device, a response emotion using the source emotion;

generating, by the computing device, a response textual representation using the textual representation of the voice input, the response textual representation indicating textual content of the response;

generating, by the computing device, a frequency spectrum representation of the response in accordance with the response textual representation and the response emotion, the frequency spectrum representation comprising a spectrogram that represents the response textual representation and the response emotion, the generation comprising:

determining, by the computing device, a set of audio embeddings using a first neural network trained to generate the set of audio embeddings using a phonemic representation of the response textual representation and the response emotion;

determining, by the computing device, a set of text embeddings using a second neural network trained to generate the set of text embeddings using the phonemic representation of the response textual representation; and

using, by the computing device, the set of audio embeddings and the set of text embeddings and a third neural network trained to generate the frequency spectrum representation in accordance with the response textual representation and the response emotion;

using, by the computing device, the frequency spectrum representation of the response to generate a voice response reflective of the textual content of the response and the response emotion, the generation of the voice response comprising converting the spectrogram to a waveform representation;

communicating, by the computing device, the voice response to the user via the automated voice response user interface of the application, the communication comprising a visualization of the voice response based on the frequency spectrum representation, such that a signal strength of the voice response is visually provided respective to a time period at a number of frequencies;

obtaining, by the computing device, subsequent voice input from the user via the automated voice response user interface of the application;

determining, by the computing device, a source emotion for the subsequent voice input;

determining, by the computing device, an effectiveness of the voice response based on the determined source emotion for the subsequent voice input, the effectiveness determination comprising a determination of whether the determined source emotion for the subsequent voice input corresponds to an emotion within a set of predetermined emotions designated as undesirable or to an emotion within a set of predetermined emotions designated as desirable; and

tuning or training, by the computing device, based on the determined effectiveness of the voice response, at least one of a response emotion identifier or a response content generator.

2 . The method of claim 1 , wherein the first neural network comprises a Long Short-Term Memory (LSTM) neural network, the second neural network comprises an attention-based neural network and the third neural network comprises an LSTM neural network.

3 . The method of claim 1 , wherein generating a response textual representation further comprises:

using, by the computing device, the response emotion and the textual representation of the voice input to generate the response textual representation indicating textual content of the response.

4 . The method of claim 1 , wherein obtaining a source emotion of the user further comprises:

using, by the computing device, a trained emotion classifier and the textual representation of the voice input to determine the source emotion of the user.

5 . The method of claim 1 , wherein obtaining a response emotion further comprises:

determining, by the computing device, the response emotion using a mapping from the source emotion to the response emotion.

6 . The method of claim 1 , wherein generating a response textual representation further comprises:

using, by the computing device, a model trained using a number of samples and a machine learning algorithm to generate the response textual representation, each sample of the number of samples comprising a first communication and a second communication, the second communication acting as a label for the sample indicating a response to the first communication.

7 . The method of claim 1 , further comprising:

obtaining, by the computing device, a phonemic representation of the response using the determined response textual representation.

8 . The method of claim 7 , wherein the phonemic representation of the response is used to generate the frequency spectrum representation of the response.

9 . The method of claim 1 , wherein a waveform generator and the frequency spectrum representation of the response are used to generate the voice response reflective of the textual content of the response and the response emotion.

10 . The method of claim 1 , wherein the generation of the frequency spectrum representation comprises stacking the first neural network, the second neural network and the third neural network.

11 . The method of claim 1 , wherein the first neural network, the second neural network and third neural network are selected from a group consisting of: a Long Short-Term Memory (LSTM) neural network, an attention-based neural network (ANN), a recurrent neural network (RNN), wherein the RNN comprises one or more LSTM inner layers and an embedding layer for audio embedding generation.

12 . A non-transitory computer-readable storage medium tangibly encoded with computer-executable instructions that when executed by a processor associated with a computing device perform a method comprising:

obtaining voice input from a user via an automated voice response user interface of an application;

obtaining a textual representation of the voice input;

using the textual representation of the voice input to obtain a source emotion of the user;

determining a response emotion using the source emotion;

generating a response textual representation using the textual representation of the voice input, the response textual representation indicating textual content of the response;

generating a frequency spectrum representation of the response in accordance with the response textual representation and the response emotion, the frequency spectrum representation comprising a spectrogram that represents the response textual representation and the response emotion, the generation comprising:

determining a set of audio embeddings using a first neural network trained to generate the set of audio embeddings using a phonemic representation of the response textual representation and the response emotion;

determining a set of text embeddings using a second neural network trained to generate the set of text embeddings using the phonemic representation of the response textual representation; and

using the set of audio embeddings and the set of text embeddings and a third neural network trained to generate the frequency spectrum representation in accordance with the response textual representation and the response emotion;

using the frequency spectrum representation of the response to generate a voice response reflective of the textual content of the response and the response emotion, the generation of the voice response comprising converting the spectrogram to a waveform representation;

communicating the voice response to the user via the automated voice response user interface of the application, the communication comprising a visualization of the voice response based on the frequency spectrum representation, such that a signal strength of the voice response is visually provided respective to a time period at a number of frequencies;

obtaining, by the computing device, subsequent voice input from the user via the automated voice response user interface of the application;

determining, by the computing device, a source emotion for the subsequent voice input;

determining, by the computing device, an effectiveness of the voice response based on the determined source emotion for the subsequent voice input, the effectiveness determination comprising a determination of whether the determined source emotion for the subsequent voice input corresponds to an emotion within a set of predetermined emotions designated as undesirable or to an emotion within a set of predetermined emotions designated as desirable; and

tuning or training, by the computing device, based on the determined effectiveness of the voice response, at least one of a response emotion identifier or a response content generator.

13 . The non-transitory computer-readable storage medium of claim 12 , the method further comprising:

obtaining a phonemic representation of the response using the determined response textual representation.

14 . The non-transitory computer-readable storage medium of claim 13 , the phonemic representation of the response being used to generate the frequency spectrum representation of the response.

15 . The non-transitory computer-readable storage medium of claim 12 , the method further comprising using a waveform generator and the frequency spectrum representation of the response to generate the voice response reflective of the textual content of the response and the response emotion.

16 . The non-transitory computer-readable storage medium of claim 12 , wherein the generation of the frequency spectrum representation comprises stacking the first neural network, the second neural network and the third neural network.

17 . A computing device comprising:

a processor, configured to:

obtain voice input from a user via an automated voice response user interface of an application;

obtain a textual representation of the voice input;

use the textual representation of the voice input to obtain a source emotion of the user;

determine a response emotion using the source emotion;

generate a response textual representation using the textual representation of the voice input, the response textual representation indicating textual content of the response;

generate a frequency spectrum representation of the response in accordance with the response textual representation and the response emotion, the frequency spectrum representation comprising a spectrogram that represents the response textual representation and the response emotion, the generation comprising:

determining a set of audio embeddings using a first neural network trained to generate the set of audio embeddings using a phonemic representation of the response textual representation and the response emotion;

determining a set of text embeddings using a second neural network trained to generate the set of text embeddings using the phonemic representation of the response textual representation; and

using the set of audio embeddings and the set of text embeddings and a third neural network trained to generate the frequency spectrum representation in accordance with the response textual representation and the response emotion;

use the frequency spectrum representation of the response to generate a voice response reflective of the textual content of the response and the response emotion, the generation of the voice response comprising converting the spectrogram to a waveform representation;

communicate the voice response to the user via the automated voice response user interface of the application, the communication causing visualization of the voice response based on the frequency spectrum representation, such that a signal strength of the voice response is visually provided respective to a time period at a number of frequencies;

obtain subsequent voice input from the user via the automated voice response user interface of the application;

determine a source emotion for the subsequent voice input;

determine an effectiveness of the voice response based on the determined source emotion for the subsequent voice input, the effectiveness determination comprising a determination of whether the determined source emotion for the subsequent voice input corresponds to an emotion within a set of predetermined emotions designated as undesirable or to an emotion within a set of predetermined emotions designated as desirable; and

tune or train, based on the determined effectiveness of the voice response, at least one of a response emotion identifier or a response content generator.

18 . The computing device of claim 17 , the processor further configured to:

obtain a phonemic representation of the response using the determined response textual representation, the phonemic representation of the response being used to generate the frequency spectrum representation of the response.

19 . The computing device of claim 17 , the processor further configured to use a waveform generator and the frequency spectrum representation of the response to generate the voice response reflective of the textual content of the response and the response emotion.