IP Library Granted Patent US 11,151,980
Granted Patent B2
US 11,151,980 · App. 16/934,230 · Granted Oct 19, 2021

Systems and methods for adjusting dubbed speech based on context of a scene

Inventors: Mario Sanchez (San Jose, CA); Ashleigh Miller (Denver, CO); Paul T. Stathacopoulos (San Carlos, CA)
Assignee: Rovi Guides, Inc.
G10L13/033G06K9/00302G10L15/1815G10L15/22G10L21/02G10L25/51H04N9/802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,151,980
App. No.
16/934,230
Granted
Oct 19, 2021
Kind
B2
Abstract

Systems and methods are disclosed herein for detecting dubbed speech in a media asset and receiving metadata corresponding to the media asset. The systems and methods may determine a plurality of scenes in the media asset based on the metadata, retrieve a portion of the dubbed speech corresponding to the first scene, and process the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene. Further, the systems and methods may determine whether the speech characteristic of the character featured in the first scene matches the context of the first scene, and if the match fails, perform a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene matches the context of the first scene.

Claims (59)

1. A method comprising:

determining a genre of a media content portion;

identifying a genre emotion based on the determined genre;

analyzing audio in the media content portion to determine that audio emotion corresponding to the audio in the media content portion does not match with the identified genre emotion of the media content portion;

in response to the determining that the audio emotion does not match with the identified genre emotion of the media content portion, modifying the audio in the media content portion to match with the identified genre emotion; and

generating for output the media content portion with the modified audio.

2. The method of claim 1 , wherein the analyzing comprises:

retrieving a portion of speech of a character in the media content portion; and

identifying a speech template among a plurality of templates that corresponds to the retrieved portion of the speech of the character, wherein each of the plurality of speech templates correspond to a different speech characteristic of the character in the media content portion.

3. The method of claim 2 , wherein the analyzing comprises:

retrieving a speech characteristic of the character in the portion of the speech in the media content portion; and

comparing the retrieved speech characteristic of the character in the media content portion with the speech characteristic of the character in the identified speech template.

4. The method of claim 2 wherein retrieving the portion of speech comprises retrieving a dialog that occurs in the media content portion.

5. The method of claim 2 wherein the analyzing further comprises:

retrieving a first set of vocal characteristics corresponding to the retrieved portion of the speech;

retrieving a second set of vocal characteristics corresponding to the identified speech template; and

comparing a first vocal characteristic among the first set of vocal characteristics to a corresponding second vocal characteristic among the second set of vocal characteristics.

6. The method of claim 5 wherein the vocal characteristics is one of a pitch, pause, rate, and rhythm or combinations thereof.

7. The method of claim 3 wherein the speech characteristic is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof.

8. A system comprising:

a control circuitry configured to:

determine a genre of a media content portion;

identify a genre emotion based on the determined genre;

analyze audio in the media content portion to determine that audio emotion corresponding to the audio in the media content portion does not match with the identified genre emotion of the media content portion;

in response to the determining that the audio emotion does not match with the identified genre emotion of the media content portion, modify the audio in the media content portion to match with the identified genre emotion; and

generating for output the media content portion with the modified audio.

9. The system of claim 8 , wherein the control circuitry is further configured to analyze audio in the media content portion by:

retrieve a portion of speech of a character in the media content portion; and

identify a speech template among a plurality of templates that corresponds to the retrieved portion of the speech of the character, wherein each of the plurality of speech templates correspond to a different speech characteristic of the character in the media content portion.

10. The system of claim 9 , wherein the control circuitry is further configured to analyze audio in the media content portion by:

retrieve a speech characteristic of the character in the portion of the speech in the media content portion; and

compare the retrieved speech characteristic of the character in the media content portion with the speech characteristic of the character in the identified speech template.

11. The system of claim 9 , wherein the control circuitry is further configured to retrieve the portion of speech by:

retrieve a dialog that occurs in the media content portion.

12. The system of claim 9 , wherein the control circuitry is further configured to analyze audio in the media content portion by:

retrieve a first set of vocal characteristics corresponding to the retrieved portion of the speech;

retrieve a second set of vocal characteristics corresponding to the identified speech template; and

compare a first vocal characteristic among the first set of vocal characteristics to a corresponding second vocal characteristic among the second set of vocal characteristics.

13. The system of claim 12 , wherein the vocal characteristics is one of a pitch, pause, rate, and rhythm or combinations thereof.

14. The system of claim 9 , wherein the speech characteristic is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof.

15. A non-transitory computer-readable medium having instructions encoded thereon to cause the control circuitry to:

determine a genre of a media content portion;

identify a genre emotion based on the determined genre;

analyze audio in the media content portion to determine that audio emotion corresponding to the audio in the media content portion does not match with the identified genre emotion of the media content portion;

in response to the determining that the audio emotion does not match with the identified genre emotion of the media content portion, modify the audio in the media content portion to match with the identified genre emotion; and

generating for output the media content portion with the modified audio.

16. The non-transitory computer-readable medium of claim 15 , wherein the instructions cause the control circuitry to analyze audio in the media content portion by:

retrieve a portion of speech of a character in the media content portion; and

identify a speech template among a plurality of templates that corresponds to the retrieved portion of the speech of the character, wherein each of the plurality of speech templates correspond to a different speech characteristic of the character in the media content portion.

17. The non-transitory computer-readable medium of claim 16 , wherein the instructions cause the control circuitry to analyze audio in the media content by:

retrieve a speech characteristic of the character in the portion of the speech in the media content portion; and

compare the retrieved speech characteristic of the character in the media content portion with the speech characteristic of the character in the identified speech template.

18. The non-transitory computer-readable medium of claim 16 , wherein the instructions cause the control circuitry to retrieve a portion of the speech by:

retrieve a dialog that occurs in the media content portion.

19. The non-transitory computer-readable medium of claim 16 , wherein the instructions cause the control circuitry to analyze audio in the media content by:

retrieve a first set of vocal characteristics corresponding to the retrieved portion of the speech;

retrieve a second set of vocal characteristics corresponding to the identified speech template; and

compare a first vocal characteristics among the first set of vocal characteristics to a corresponding second vocal characteristics among the second set of vocal characteristics.

20. The non-transitory computer-readable medium of claim 19 , wherein the vocal characteristics is one of a pitch, pauses, rate, and rhythm or combinations thereof.

Assignments (3)
CHANGE OF NAME Recorded Oct 2, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069086/0199 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2020
From: SANCHEZ, MARIO; MILLER, ASHLEIGH; STATHACOPOULOS, PAUL T.
To: ROVI GUIDES, INC.
Reel/Frame 053477/0075 →
Continuity (2)
Continuation 16610225
Related Publication 20200349961A1 · Nov 5, 2020
Cited By (1)
US 12,614,540