IP Library Granted Patent US 11,670,284
Granted Patent B2
US 11,670,284 · App. 17/480,550 · Granted Jun 6, 2023

Systems and methods for adjusting dubbed speech based on context of a scene

Inventors: Mario Sanchez (San Jose, CA); Ashleigh Miller (Denver, CO); Paul T. Stathacopoulos (San Carlos, CA)
Assignee: Rovi Guides, Inc.
G10L13/033G06V40/174G10L15/1815G10L15/22G10L21/02G10L25/51H04N9/802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,670,284
App. No.
17/480,550
Granted
Jun 6, 2023
Kind
B2
Abstract

Systems and methods are disclosed herein for detecting dubbed speech in a media asset and receiving metadata corresponding to the media asset. The systems and methods may determine a plurality of scenes in the media asset based on the metadata, retrieve a portion of the dubbed speech corresponding to the first scene, and process the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene. Further, the systems and methods may determine whether the speech characteristic of the character featured in the first scene matches the context of the first scene, and if the match fails, perform a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene snatches the context of the first scene.

Claims (67)

1. A method comprising:

analyzing metadata of a media content portion to determine a context of the media content portion, the media content portion comprising audio and video;

analyzing the media content portion to identify modifications to the audio in the media content portion to cause the audio in the media content portion to match the context of the media content portion;

modifying the audio in the media content portion to match the context of the media content portion using the identified modifications to the audio in the media content portion; and

generating for output the media content portion with the modified audio.

2. The method of claim 1 , further comprising:

determining that the audio in the media content portion does not match with the context; and

in response to the determining that the audio in the media content portion does not match with the context, performing the modifying the audio in the media content portion to match with the context.

3. The method of claim 1 , further comprising:

retrieving a first set of vocal characteristics for the audio in the media content portion; and

retrieving a second set of vocal characteristics that match the context;

adjusting a particular characteristic of the first set of vocal characteristics to match a corresponding characteristic of the second set of vocal characteristics.

4. The method of claim 3 , wherein the particular characteristic is one or more of a pitch, pause, rate, or rhythm.

5. The method of claim of claim 3 , further comprising:

determining that the particular characteristic of the first set of vocal characteristics does not match the corresponding characteristic of the second set of vocal characteristics; and

performing the adjusting of the particular characteristic in response to determining that the particular characteristic does not match the corresponding characteristic.

6. The method of claim 1 , further comprising:

determining a speech characteristic corresponding to the context of the media content portion;

wherein modifying the audio in the media content portion comprises modifying speech audio to match the speech characteristic.

7. The method of claim 6 , wherein the speech characteristic comprises one or more of angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, or accusing.

8. The method of claim 6 , further comprising:

retrieving a speech template corresponding to the context, the speech template comprising the speech characteristic;

determining that the speech audio in the media content portion does not match the speech characteristic;

performing the modifying the audio in response to determining that the speech audio in the media content portion does not match the speech characteristic of the speech template.

9. A system comprising:

a memory storing a media content portion;

control circuitry configured to:

analyze metadata of the media content portion to determine a context of the media content portion, the media content portion comprising audio and video;

analyze the media content portion to identify modifications to the audio in the media content portion to cause the audio in the media content portion to match the context of the media content portion;

modify the audio in the media content portion to match the context of the media content portion using the identified modifications to the audio in the media content portion; and

generate for output the media content portion with the modified audio.

10. The system of claim 9 , wherein the control circuitry is further configured to:

determine that the audio in the media content portion does not match with the context; and

in response to the determining that the audio in the media content portion does not match with the context, performing the modifying the audio in the media content portion to match with the context.

11. The system of claim 9 , wherein the control circuitry is further configured to:

retrieve a first set of vocal characteristics for the audio in the media content portion; and

retrieve a second set of vocal characteristics that match the context;

adjust a particular characteristic of the first set of vocal characteristics to match a corresponding characteristic of the second set of vocal characteristics.

12. The system of claim 11 , wherein the particular characteristic is one or more of a pitch, pause, rate, or rhythm.

13. The system of claim 11 , wherein the control circuitry is further configured to:

determine that the particular characteristic of the first set of vocal characteristics does not match the corresponding characteristic of the second set of vocal characteristics; and

perform the adjusting of the particular characteristic in response to determining that the particular characteristic does not match the corresponding characteristic.

14. The system of claim 9 , wherein the control circuitry is further configured to:

determine a speech characteristic corresponding to the context of the media content portion;

wherein modifying the audio in the media content portion comprises modifying speech audio to match the speech characteristic.

15. The system of claim 14 , wherein the speech characteristic comprises one or more of angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, or accusing.

16. The system of claim 14 , wherein the control circuitry is further configured to:

retrieve a speech template corresponding to the context, the speech template comprising the speech characteristic;

determine that the speech audio in the media content portion does not match the speech characteristic;

perform the modifying the audio in response to determining that the speech audio in the media content portion does not match the speech characteristic of the speech template.

17. A non-transitory computer-readable medium having instructions encoded thereon to cause control circuitry to:

analyze metadata of a media content portion to determine a context of the media content portion, the media content portion comprising audio and video;

analyze the media content portion to identify modifications to the audio in the media content portion to cause the audio in the media content portion to match the context of the media content portion;

modify the audio in the media content portion to match the context of the media content portion using the identified modifications to the audio in the media content portion; and

generate for output the media content portion with the modified audio.

18. The non-transitory computer-readable medium of claim 17 , wherein the instructions cause the control circuitry to:

determine that the audio in the media content portion does not match with the context; and

in response to the determining that the audio in the media content portion does not match with the context, performing the modifying the audio in the media content portion to match with the context.

19. The non-transitory computer-readable medium of claim 17 , wherein the instructions cause the control circuitry to:

retrieve a first set of vocal characteristics for the audio in the media content portion; and

retrieve a second set of vocal characteristics that match the context;

adjust a particular characteristic of the first set of vocal characteristics to match a corresponding characteristic of the second set of vocal characteristics;

wherein the particular characteristic is one or more of a pitch, pause, rate, or rhythm.

20. The non-transitory computer-readable medium of claim 17 , wherein the instructions cause the control circuitry to:

determine a speech characteristic corresponding to the context of the media content portion;

wherein modifying the audio in the media content portion comprises modifying speech audio to match the speech characteristic;

wherein the speech characteristic comprises one or more of angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, or accusing.

Assignments (3)
CHANGE OF NAME Recorded Oct 2, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069086/0199 →
SECURITY INTEREST Recorded May 19, 2023
From: ADEIA GUIDES INC.; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063707/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2021
From: SANCHEZ, MARIO; MILLER, ASHLEIGH; STATHACOPOULOS, PAUL T.
To: ROVI GUIDES, INC.
Reel/Frame 057558/0703 →
Continuity (3)
Continuation 16934230 · Jul 21, 2020
Continuation 16610225
Related Publication 20220005455A1 · Jan 6, 2022
Cited By (1)
US 12,614,540