IP Library Granted Patent US 12,118,978
Granted Patent B2
US 12,118,978 · App. 18/387,211 · Granted Oct 15, 2024

Systems and methods for generating synthesized speech responses to voice inputs indicative of a user in a hurry

Inventors: Ankur Aher (Maharashtra, IN); Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Rovi Guides, Inc.
G10L13/0335G10L25/63G10L13/00G10L13/02G10L13/06G10L13/08G10L15/00G10L15/10G10L15/16G10L15/18G10L15/22G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,118,978
App. No.
18/387,211
Granted
Oct 15, 2024
Kind
B2
Abstract

The system provides a synthesized speech response to a voice input, based on the prosodic character of the voice input. The system receives the voice input and calculates at least one prosodic metric of the voice input. The at least one prosodic metric can be associated with a word, phrase, grouping thereof, or the entire voice input. The system also determines a response to the voice input, which may include the sequence of words that form the response. The system generates the synthesized speech response, by determining prosodic characteristics based on the response, and on the prosodic character of the voice input. The system outputs the synthesized speech response, which includes a more natural, relevant, or both answer to the call of the voice input. The prosodic character of the voice input and/or response may include pitch, note, duration, prominence, timbre, rate, and rhythm, for example.

Claims (50)

1. A method comprising:

receiving a voice input by a voice assistant system;

determining, based on the voice input, that the voice input is indicative of a user in a hurry;

based on determining that the voice input is indicative of the user in the hurry:

adjusting a default setting of the voice assistant system to increase a play speed setting for responses of the voice assistant system that is higher than a play speed for responses when the voice assistant system operates in accordance with a default setting; and

generating a synthesized speech response that has a verbosity lesser than a verbosity of the voice assistant system when the voice assistant system operates in accordance with the default setting; and

causing to be output the synthesized speech response that plays at the increased play speed that is higher than when the voice assistant system operates in accordance with the default setting, and that is less verbose than when the voice assistant system operates in accordance with the default setting.

2. The method of claim 1 , comprising:

predicting a set of audio features to be applied for each phrase and word of the synthesized speech response.

3. The method of claim 2 , comprising:

processing the predicted set of audio features using an interpolation operation to improve a prosodic characteristic and a prosodic transition in the synthesized speech response.

4. The method of claim 3 , comprising:

determining whether adjacent words and phrases of the predicted set of audio features have disparate, unmatched, and conflicting values; and

in response to determining that the adjacent words and phrases of the predicted set of audio features have disparate, unmatched, and conflicting values, performing the interpolation operation affecting the transitions in the synthesized speech response.

5. The method of claim 1 , wherein the determining based on the voice input that the voice input is indicative of the user in the hurry includes:

accessing user profile information associated with a user, and comparing characteristics of the voice input to the user profile information.

6. The method of claim 1 , wherein the generating the synthesized speech response includes:

interpolation operations to affect transitions between words and sounds in the synthesized speech response to improve a prosodic characteristic of the synthesized speech response.

7. The method of claim 1 , wherein the generating the synthesized speech response includes:

interpolating, normalizing, and modifying a prosodic characteristic of the synthesized speech response to enhance prosody.

8. The method of claim 1 , comprising:

accessing interpolation metrics among word transitions from the voice input;

accessing a model trained to determine a desired prosody based on the interpolation metrics; and

based on the model, generating the desired prosody in the synthesized speech response.

9. A system comprising:

a communications circuitry configured to:

receive a voice input by a voice assistant system; and

control circuitry configured to:

determine, based on the voice input, that the voice input is indicative of a user in a hurry;

based on determining that the voice input is indicative of the user in the hurry:

adjust a default setting of the voice assistant system to increase a play speed setting for responses of the voice assistant system that is higher than a play speed for responses when the voice assistant system operates in accordance with a default setting; and

generate a synthesized speech response that has a verbosity lesser than a verbosity of the voice assistant system when the voice assistant system operates in accordance with the default setting; and

cause to be output the synthesized speech response that plays at the increased play speed that is higher than when the voice assistant system operates in accordance with the default setting, and that is less verbose than when the voice assistant system operates in accordance with the default setting.

10. The system of claim 9 , wherein the control circuitry is configured to:

predict a set of audio features to be applied for each phrase and word of the synthesized speech response.

11. The system of claim 10 , wherein the control circuitry is configured to:

process the predicted set of audio features using an interpolation operation to improve a prosodic characteristic and a prosodic transition in the synthesized speech response.

12. The system of claim 11 , wherein the control circuitry is configured to:

determine whether adjacent words and phrases of the predicted set of audio features have disparate, unmatched, and conflicting values; and

in response to determining that the adjacent words and phrases of the predicted set of audio features have disparate, unmatched, and conflicting values, perform the interpolation operation affecting the transitions in the synthesized speech response.

13. The system of claim 9 , wherein the control circuitry configured to determine based on the voice input that the voice input is indicative of the user in the hurry is configured to:

access user profile information associated with a user, and comparing characteristics of the voice input to the user profile information.

14. The system of claim 9 , wherein the control circuitry configured to generate the synthesized speech response is configured to:

perform interpolation operations to affect transitions between words and sounds in the synthesized speech response to improve a prosodic characteristic of the synthesized speech response.

15. The system of claim 9 , wherein the control circuitry configured to generate the synthesized speech response is configured to:

interpolate, normalize, and modify a prosodic characteristic of the synthesized speech response to enhance prosody.

16. The system of claim 9 , wherein the control circuitry is configured to:

access interpolation metrics among word transitions from the voice input;

access a model trained to determine a desired prosody based on the interpolation metrics; and

based on the model, generate the desired prosody in the synthesized speech response.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2023
From: AHER, ANKUR; ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 065471/0401 →
Priority Claims (1)
IN 202041015653 · Apr 9, 2020 · national
Continuity (3)
Continuation 17882289 · Aug 5, 2022
Continuation 15931074 · May 13, 2020
Related Publication 20240153483A1 · May 9, 2024