IP Library Granted Patent US 12,062,358
Granted Patent B2
US 12,062,358 · App. 18/136,515 · Granted Aug 13, 2024

Systems and methods for adjusting dubbed speech based on context of a scene

Inventors: Mario Sanchez (San Jose, CA); Ashleigh Miller (Denver, CO); Paul T. Stathacopoulos (San Carlos, CA)
Assignee: Rovi Guides, Inc.
G10L13/033G06V40/174G10L15/1815G10L15/22G10L21/02G10L25/51H04N9/802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,062,358
App. No.
18/136,515
Granted
Aug 13, 2024
Kind
B2
Abstract

Systems and methods are disclosed herein for detecting dubbed speech in a media asset and receiving metadata corresponding to the media asset. The systems and methods may determine a plurality of scenes in the media asset based on the metadata, retrieve a portion of the dubbed speech corresponding to the first scene, and process the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene. Further, the systems and methods may determine whether the speech characteristic of the character featured in the first scene matches the context of the first scene, and if the match fails, perform a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene matches the context of the first scene.

Claims (65)

1. A method comprising:

detecting a first text in a media asset and a second text in the media asset;

identifying a first portion of the media asset and a second portion of the media asset;

selecting a first emotion associated with the first text;

selecting a second emotion associated with the second text;

modifying a first audio for the first portion of the media asset based on the first text and the first emotion by:

retrieving a first set of vocal characteristics corresponding to the first portion of the media asset and a second set of vocal characteristics corresponding to the second portion of the media asset; and

in response to identifying that a first vocal characteristic from the first set of vocal characteristics does not match a second vocal characteristic from a second set of vocal characteristics, adjusting the first vocal characteristic from the first set of vocal characteristics to match the second vocal characteristic from the second set of vocal characteristics; and

modifying a second audio for the second portion of the media asset based on the second text and the second emotion.

2. The method of claim 1 , wherein the selecting the first emotion comprises:

generating for display on a user device a first plurality of emotions associated with the first text;

receiving a user interface selection of the first emotion from the first plurality of emotions;

and wherein the selecting the second emotion comprises:

generating for display on a user device a second plurality of emotions associated with the second text; and

receiving a user selection of the first emotion from the second plurality of emotions.

3. The method of claim 1 , wherein the selecting of the first emotion with the first text and the selecting of the second emotion with the second text comprises:

receiving a first input from an input interface, wherein the first input indicates the first selection of the first emotion; and

receiving a second input from an input interface, wherein the second input indicates the second selection of the second emotion.

4. The method of claim 1 , wherein the modifying of the first audio for the first portion of the media asset and the second audio for the second portion of the media asset further comprises:

generating the first audio for the first portion of the media asset based on the first text and the second audio for the second portion of the media asset based on the second text; and

modifying the first audio by inserting the first set of vocal characteristics into the first audio; and

modifying the second audio by inserting the second set of vocal characteristics into the second audio.

5. The method of claim 4 , wherein the first vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof, and wherein the second vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof.

6. The method of claim 1 , wherein the first emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof, and wherein the second emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof.

7. The method of claim 1 , wherein the modifying of the first audio for the first portion of the media asset based on the first text and the first emotion comprises:

retrieving a first personality metadata corresponding to the first portion of the media asset;

identifying a first speech characteristic from the first personality metadata; and

adjusting the audio for the first portion of the media asset from the identified first speech characteristic.

8. The method of claim 1 , wherein the modifying of the second audio for the second portion of the media asset based on the second text and the second emotion comprise:

retrieving a second personality metadata corresponding to the second portion of the media asset;

identifying a second speech characteristic from the second personality metadata; and

adjusting the second audio for the second portion of the media asset from the identified second speech characteristic.

9. A system comprising:

a control circuitry configured to:

detect a first text in a media asset and a second text in the media asset;

identify a first portion of the media asset and a second portion of the media asset;

select a first emotion associated with the first text;

select a second emotion associated with the second text;

modify a first audio for the first portion of the media asset based on the first text and the first emotion by:

retrieving a first set of vocal characteristics corresponding to the first portion of the media asset and a second set of vocal characteristics corresponding to the second portion of the media asset; and

in response to identifying that a first vocal characteristic from the first set of vocal characteristics does not match a second vocal characteristic from a second set of vocal characteristics, the control circuitry further configured to adjust the first vocal characteristic from the first set of vocal characteristics to match the second vocal characteristic from the second set of vocal characteristics; and

modify a second audio for the second portion of the media asset based on the second text and the second emotion.

10. The system of claim 9 , wherein the control circuitry is configured to select the first emotion by:

generating for display on a user device a first plurality of emotions associated with the first text;

receiving a user interface selection of the first emotion from the first plurality of emotions;

and wherein the control circuitry is configured to select the second emotion by:

generating for display on a user device a second plurality of emotions associated with the second text; and

receiving a user selection of the first emotion from the second plurality of emotions.

11. The system of claim 9 , wherein the control circuitry is configured to select the first emotion associated with the first text and the second emotion associated with the second text by:

receiving a first input from an input interface, wherein the first input indicates the first selection of the first emotion; and

receiving a second input from an input interface, wherein the second input indicates the second selection of the second emotion.

12. The system of claim 11 , wherein the first vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof, and wherein the second vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof.

13. The system of claim 9 , wherein the control circuitry is configured to modify the first audio for the first portion of the media asset and the second audio for the second portion of the media asset by:

generate the first audio for the first portion of the media asset based on the first text and the second audio for the second portion of the media asset based on the second text; and

modifying the first audio by inserting the first set of vocal characteristics into the first audio; and

modifying the second audio by inserting the second set of vocal characteristics into the second audio.

14. The system of claim 9 , wherein the first emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof, and wherein the second emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof.

15. The system of claim 9 , wherein the control circuitry is configured to modify the first audio for the first portion of the media asset based on the first text and the first emotion by:

retrieving a first personality metadata corresponding to the first portion of the media asset;

identifying a first speech characteristic from the first personality metadata; and

adjusting the first audio for the first portion of the media asset from the identified first speech characteristic.

16. The system of claim 9 , wherein the control circuitry is configured to modify the second audio for the second portion of the media asset based on the second text and the second emotion by:

retrieving a second personality metadata corresponding to the second portion of the media asset;

identifying a second speech characteristic from the second personality metadata; and

adjusting the second audio for the second portion of the media asset from the identified second speech characteristic.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069086/0199 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2023
From: SANCHEZ, MARIO; MILLER, ASHLEIGH; STATHACOPOULOS, PAUL T.
To: ROVI GUIDES, INC.
Reel/Frame 063379/0793 →
Continuity (4)
Continuation 17480550 · Sep 21, 2021
Continuation 16934230 · Jul 21, 2020
Continuation 16610225
Related Publication 20230343320A1 · Oct 26, 2023
Cited By (1)
US 12,614,540