IP Library Granted Patent US 11,922,923
Granted Patent B2
US 11,922,923 · App. 16/863,855 · Granted Mar 5, 2024

Optimal human-machine conversations using emotion-enhanced natural speech using hierarchical neural networks and reinforcement learning

Inventors: Alan McCord (Wakatipu Queenstown, NZ); Ashley Unitt (Basingstoke, GB); Brian Galvin (Seabeck, WA)
Assignee: VONAGE BUSINESS LIMITED
G10L13/10G06N3/04G06N3/045G06N3/08G10L13/033G10L13/047G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,922,923
App. No.
16/863,855
Granted
Mar 5, 2024
Kind
B2
Abstract

A system and method for emotion-enhanced natural speech using dilated convolutional neural networks, wherein an audio processing server receives a raw audio waveform from a dilated convolutional artificial neural network, associates text-based emotion content markers with portions of the raw audio waveform to produce an emotion-enhanced audio waveform, and provides the emotion-enhanced audio waveform to the dilated convolutional artificial neural network for use as a new input data set.

Claims (23)

1. A system for emotion-enhanced natural speech audio generation using dilated convolutional neural networks, comprising:

a computing device comprising a memory and a processor;

a first dilated convolutional artificial neural network stored in the memory of, and operating on the processor of, the computing device;

a second dilated convolutional artificial neural network stored in the memory of, and operating on the processor of, the computing device;

a neural network trainer, comprising a first plurality of programming instructions stored in the memory of the computing device which, when operating on the processor of the computing device, causes the computing device to:

train the first dilated convolutional artificial neural network to recognize emotion in text-based content by processing a plurality of text-based training data through the first dilated convolutional artificial neural network;

receive a first set of output data from the first dilated convolutional artificial network, the first set of output data comprising probability-based associations of text with emotions;

train the second dilated convolutional artificial neural network to recognize emotion in audio-based content by providing a plurality of audio-based training-data through the second dilated convolutional artificial neural network, the audio-based training data corresponding to the text-based training data;

receive a second set of output data from the second dilated convolutional artificial network, the second set of output data comprising probability-based associations of modulations of sounds with emotions;

construct an emotion injection model by associating text from the first set of output data with sounds from the second set of output data based on the emotions associated with each; and

an automated emotion engine comprising a second plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the programming instructions, when operating on the processor, cause the computing device to:

receive text content;

process the text content through the first dilated convolutional artificial neural network to recognize emotional states in the text content; and

convert the text content to audio content using a text-to-speech engine, modulating the audio content with the modulations of sounds associated with the text from the emotion injection model.

2. A method for emotion-enhanced natural speech audio generation using dilated convolutional neural networks, comprising:

training a first dilated convolutional artificial neural network to recognize emotion in text-based content by processing a plurality of text-based training data through the first dilated convolutional artificial neural network;

receiving a first set of output data from the first dilated convolutional artificial network, the first set of output data comprising probability-based associations of text with emotions;

training a second dilated convolutional artificial neural network to recognize emotion in audio-based content by providing a plurality of audio-based training-data through the second dilated convolutional artificial neural network, the audio-based training data corresponding to the text-based training data;

receiving a second set of output data from the second dilated convolutional artificial network, the second set of output data comprising probability-based associations of modulations of sounds with emotions;

constructing an emotion injection model by associating text from the first set of output data with sounds from the second set of output data based on the emotions associated with each;

receiving text content;

processing the text content through the first dilated convolutional artificial neural network to recognize emotional states in the text content; and

converting the text content to audio content using a text-to-speech engine, and modulating the audio content with the modulations of sounds associated with the text from the emotion injection model.

Assignments (2)
CHANGE OF NAME Recorded Feb 3, 2022
From: NEWVOICEMEDIA LIMITED
To: VONAGE BUSINESS LIMITED
Reel/Frame 058879/0481 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2020
From: MCCORD, ALAN; UNITT, ASHLEY; GALVIN, BRIAN
To: NEWVOICEMEDIA LTD.
Reel/Frame 053185/0560 →
Continuity (6)
Continuation 15661341 · Jul 27, 2017
Continuation In Part 15442667 · Feb 25, 2017
Continuation In Part 15268611 · Sep 18, 2016
Provisional Application 62516672 · Jun 8, 2017
Provisional Application 62441538 · Jan 2, 2017
Related Publication 20200320974A1 · Oct 8, 2020
Cited By (1)
US 12,712,970