IP Library Patent Application 18883144
Patent Application
App. No. 18/883,144

SYSTEMS AND METHODS FOR GENERATING SYNTHESIZED SPEECH RESPONSES TO VOICE INPUTS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/883,144
Abstract

The system provides a synthesized speech response to a voice input, based on the emotion of the voice input. The system receives the voice input and determines the emotion of the voice input. The system identifies, based on the first emotion of the voice input, a second emotion, different from the first emotion, for a response to the voice input. The system generates a synthesized speech voice output of the response comprising prosodic characteristics corresponding to the second emotion. The system outputs the synthesized speech response.

Claims (44)

1 . A method comprising:

receiving a voice input by a voice assistant system;

based on the voice input, determining a first emotion of the voice input;

based on the first emotion of the voice input, identifying a second emotion, different from the first emotion, for a response to the voice input;

generating a synthesized speech voice output as the response to the voice input, wherein the synthesized speech voice output comprises prosodic characteristics corresponding to the second emotion; and

causing an output of the synthesized speech voice output.

2 . The method of claim 1 , wherein the identifying the second emotion comprises identifying an emotion that counteracts the first emotion as the second emotion.

3 . The method of claim 1 , wherein identifying the second emotion comprises:

determining a correlation between the first emotion and the second emotion; and

identifying the second emotion based on the correlation.

4 . The method of claim 1 , further comprising training a neural network model based on a plurality of voice input emotion metrics and a plurality of response emotion metrics, wherein identifying the second emotion is based on an output of the trained neural network model.

5 . The method of claim 1 , wherein the determining the first emotion for the voice input comprises accessing user profile information associated with a user, and comparing characteristics of the voice input to the user profile information.

6 . The method of claim 1 , wherein the synthesized speech voice output is further based on at least one selected from the group comprising user voice input history, user language, user characteristics, user location, user preferences, metadata tags associated with a user, and any combination thereof.

7 . The method of claim 1 , wherein the generating the synthesized speech voice output further comprises interpolation operations to affect transitions between words and sounds in the synthesized speech voice output to improve the prosodic characteristics of the synthesized speech voice output.

8 . The method of claim 1 , wherein determining the first emotion of the voice input includes determining an emotional indicator associated with a word, syllable thereof, or grouping of words.

9 . A system comprising:

input/output circuitry configured to:

receive a voice input by a voice assistant system;

control circuitry configured to:

based on the voice input, determine a first emotion of the voice input;

based on the first emotion of the voice input, identify a second emotion, different from the first emotion, for a response to the voice input; and

generate a synthesized speech voice output as the response to the voice input, wherein the synthesized speech voice output comprises prosodic characteristics corresponding to the second emotion;

wherein the input/output circuitry is further configured to:

cause an output of the synthesized speech voice output.

10 . The system of claim 9 , wherein the control circuitry configured to identify the second emotion is further configured to identify an emotion that counteracts the first emotion as the second emotion.

11 . The system of claim 9 , wherein the control circuitry configured to identify the second emotion is further configured to:

determine a correlation between the first emotion and the second emotion; and

identify the second emotion based on the correlation.

12 . The system of claim 9 , wherein the control circuitry is further configured to train a neural network model based on a plurality of voice input emotion metrics and a plurality of response emotion metrics; and wherein the control circuitry further configured to identify the second emotion based on an output of the trained neural network model.

13 . The system of claim 9 , wherein the control circuitry configured to determine the first emotion for the voice input is further configured to access user profile information associated with a user, and compare characteristics of the voice input to the user profile information.

14 . The system of claim 9 , wherein the synthesized speech voice output is further based on at least one selected from the group comprising user voice input history, user language, user characteristics, user location, user preferences, metadata tags associated with a user, and any combination thereof.

15 . The system of claim 9 , wherein the control circuitry configured to generate the synthesized speech voice output is further configured to perform interpolation operations to affect transitions between words and sounds in the synthesized speech voice output to improve the prosodic characteristics of the synthesized speech voice output.

16 . The system of claim 9 , wherein the control circuitry configured determine the first emotion of the voice input includes determining an emotional indicator associated with a word, syllable thereof, or grouping of words.

17 . A non-transitory computer-readable medium having instructions encoded thereon that, when executed by control circuitry, cause the control circuitry to:

receive a voice input by a voice assistant system;

based on the voice input, determine a first emotion of the voice input;

based on the first emotion of the voice input, identify a second emotion, different from the first emotion, for a response to the voice input;

generate a synthesized speech voice output as the response to the voice input, wherein the synthesized speech voice output comprises prosodic characteristics corresponding to the second emotion; and

cause an output of the synthesized speech voice output.

18 . The non-transitory computer-readable medium of claim 17 , wherein the control circuitry is further caused to identify the second emotion is further configured to identify an emotion that counteracts the first emotion as the second emotion.

19 . The non-transitory computer-readable medium of claim 17 , wherein the control circuitry is further caused to:

determine a correlation between the first emotion and the second emotion; and

identify the second emotion based on the correlation.

20 . The non-transitory computer-readable medium of claim 17 , wherein the control circuitry is further caused to train a neural network model based on a plurality of voice input emotion metrics and a plurality of response emotion metrics; and wherein the control circuitry further configured to identify the second emotion based on an output of the trained neural network model.

Assignments (3)
SECURITY INTEREST Recorded May 28, 2025
From: ADEIA INC. (F/K/A XPERI HOLDING CORPORATION); ADEIA HOLDINGS INC.; ADEIA MEDIA HOLDINGS INC.; ADEIA IMAGING LLC; ADEIA MEDIA LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA TECHNOLOGIES INC.; ADEIA GUIDES INC.; ADEIA SOLUTIONS LLC; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR INTELLECTUAL PROPERTY LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA PUBLISHING INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 071454/0343 →
CHANGE OF NAME Recorded May 6, 2025
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 071188/0648 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2024
From: AHER, ANKUR; ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 068575/0521 →