IP Library Granted Patent US 10,755,724
Granted Patent B2
US 10,755,724 · App. 16/610,225 · Granted Aug 25, 2020

Systems and methods for adjusting dubbed speech based on context of a scene

Inventors: Mario Sanchez (San Jose, CA); Ashleigh Miller (Denver, CO); Paul T. Stathacopoulos (San Carlos, CA)
Assignee: Rovi Guides, Inc.
G10L21/0202G06K9/00302G10L15/1815G10L15/22G10L25/51H04N9/802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,755,724
App. No.
16/610,225
Granted
Aug 25, 2020
Kind
B2
Abstract

Systems and methods are disclosed herein for detecting dubbed speech in a media asset and receiving metadata corresponding to the media asset. The systems and methods may determine a plurality of scenes in the media asset based on the metadata, retrieve a portion of the dubbed speech corresponding to the first scene, and process the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene. Further, the systems and methods may determine whether the speech characteristic of the character featured in the first scene matches the context of the first scene, and if the match fails, perform a function to adjust the portion of the Dubbed Speech of the dubbed speech so that the speech characteristic of the character featured in the first scene matches the context of the first scene.

Claims (128)

1. A method comprising:

detecting dubbed speech in a media asset;

receiving metadata corresponding to the media asset;

determining a plurality of scenes in the media asset based on the metadata;

determining context of a first scene from the plurality of scenes based on the metadata;

retrieving a portion of the dubbed speech corresponding to the first scene;

processing the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene;

determining whether the speech characteristic of the character featured in the first scene matches the context of the first scene; and

in response to determining that the speech characteristic of the character featured in the first scene fails to match the context of the first scene, performing a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context of the first scene.

2. A method for detecting and correcting a mismatch between a speech characteristic of a portion of a dubbed speech of a character featured in a first scene and context of the first scene, comprising:

detecting dubbed speech in a media asset;

in response to detecting the dubbed speech in the media asset, receiving metadata corresponding to the media asset;

determining a plurality of scenes in the media asset based on the metadata;

receiving metadata corresponding to a first scene from the plurality of scenes;

determining context of the first scene based on the metadata corresponding to the first scene;

retrieving a context speech characteristic for the context of the first scene;

retrieving a portion of the dubbed speech corresponding to the first scene;

retrieving a set of speech templates corresponding to a character featured in the first scene, wherein each speech template from the set of speech templates corresponds to a different speech characteristic of the character featured in the first scene;

comparing the retrieved portion of the dubbed speech corresponding to the first scene to each speech template from the set of speech templates corresponding to the character featured in the first scene to identify a speech template that corresponds to the retrieved portion;

identifying a speech characteristic associated with the identified speech template;

determining whether the identified speech characteristic of the character featured in the first scene matches the context speech characteristic for the context of the first scene; and

in response to determining that the speech characteristic of the character featured in the first scene fails to match the context speech characteristic for the context of the first scene, performing a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context speech characteristic for the context of the first scene.

3. The method of claim 2 , wherein detecting the dubbed speech in the media asset, comprises:

receiving the media asset;

retrieving video information corresponding to the media asset;

retrieving audio information corresponding to the media asset;

retrieving speech information of the character corresponding to the audio information;

retrieving facial movements of the character corresponding to the video information;

determining whether the facial movements of the character correspond to the speech information; and

in response to determining that the facial movements of the character do not correspond to the speech information, detecting the dubbed speech in the media asset.

4. The method of claim 2 , wherein determining the context of the first scene based on the metadata corresponding to the first scene, comprises:

retrieving a plurality of keywords corresponding to the metadata corresponding to the first scene; and

determining context of the first scene based on a subset of the plurality of the keywords corresponding to the metadata corresponding to the first scene.

5. The method of claim 2 , wherein retrieving the context speech characteristic for the context of the first scene, comprises:

retrieving a personality metadata corresponding to the character, wherein the personality metadata comprises expected speech characteristic for each context for the character; and

identifying the context speech characteristic as the expected speech characteristic for the context of the first scene based on the personality metadata corresponding to the character.

6. The method of claim 2 , wherein retrieving the set of speech templates corresponding to the character featured in the first scene, comprises:

retrieving a language of the dubbed speech corresponding to the media asset; and

retrieving the set of speech templates corresponding to the character featured in the first scene and corresponding to the language of the dubbed speech.

7. The method of claim 2 , wherein retrieving the set of speech templates corresponding to the character featured in the first scene, comprises:

retrieving an original speech corresponding to the media asset; and

retrieving a portion of the original speech corresponding to the first scene; and

retrieving the set of speech templates corresponding to the character featured in the first scene based on the retrieved portion of the original speech corresponding to the first scene.

8. The method of claim 2 , wherein comparing the retrieved portion of the dubbed speech corresponding to the first scene to each speech template from the set of speech templates corresponding to the character featured in the first scene, comprises:

retrieving a first set of vocal characteristics corresponding to the retrieved portion;

retrieving a second set of vocal characteristics corresponding to a speech template from the set of speech templates corresponding to the character featured in the first scene; and

comparing a first vocal characteristic from the first set of vocal characteristics to a corresponding second vocal characteristic from the second set of vocal characteristics.

9. The method of claim 2 , wherein performing the function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context speech characteristic for the context of the first scene, comprises:

retrieving a first set of vocal characteristics corresponding to the portion of the dubbed speech;

identifying a speech template from the set of speech templates corresponding to the character featured in the first scene, that has a speech characteristic that matches the context speech characteristic for the context of the first scene;

retrieving a second set of vocal characteristics corresponding to the speech template that has the speech characteristic that matched the context speech characteristic for the context of the first scene;

identifying a first vocal characteristic from the first set of vocal characteristics that does not match a corresponding second vocal characteristic from the second set of vocal characteristics; and

adjusting the first vocal characteristic from the first set of vocal characteristics to match the corresponding second vocal characteristic from the second set of vocal characteristics.

10. The method of claim 2 , wherein performing the function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context speech characteristic for the context of the first scene, comprises:

receiving metadata corresponding to a second scene from the plurality of scenes;

determining context of the second scene based on the metadata corresponding to the second scene;

retrieving a context speech characteristic for the context of the second scene;

determining that the context speech characteristic for the context of the first scene matches the context speech characteristic for the context of the second scene;

retrieving a portion of the dubbed speech corresponding to the second scene;

retrieving a first set of vocal characteristics corresponding to the portion of the dubbed speech corresponding to the first scene;

retrieving a second set of vocal characteristics corresponding to the portion of the dubbed speech corresponding to the second scene;

identifying a first vocal characteristic from the first set of vocal characteristics that does not match a corresponding second vocal characteristic from the second set of vocal characteristics; and

adjusting the first vocal characteristic from the first set of vocal characteristics to match the corresponding second vocal characteristic from the second set of vocal characteristics.

11. The method of claim 10 , further comprising:

retrieving a portion of an adjusted dubbed speech corresponding to the first scene;

comparing the retrieved portion of the adjusted dubbed speech corresponding to the first scene to each speech template from the set of speech templates corresponding to the character featured in the first scene to identify a speech template that corresponds to the retrieved portion of the adjusted dubbed speech;

identifying a speech characteristic associated with the identified speech template that corresponds to the retrieved portion of the adjusted dubbed speech; and

determining that the identified speech characteristic of the character featured in the first scene that corresponds to the retrieved portion of the adjusted dubbed speech matches the context speech characteristic for the context of the first scene.

12. A system for detecting and correcting a mismatch between a speech characteristic of a portion of a dubbed speech of a character featured in a first scene and context of the first scene, the system comprising:

control circuitry configured to:

detect dubbed speech in a media asset;

in response to detecting the dubbed speech in the media asset, receive metadata corresponding to the media asset;

determine a plurality of scenes in the media asset based on the metadata;

receive metadata corresponding to a first scene from the plurality of scenes;

determine context of the first scene based on the metadata corresponding to the first scene;

retrieve a context speech characteristic for the context of the first scene;

retrieve a portion of the dubbed speech corresponding to the first scene;

retrieve a set of speech templates corresponding to a character featured in the first scene, wherein each speech template from the set of speech templates corresponds to a different speech characteristic of the character featured in the first scene;

compare the retrieved portion of the dubbed speech corresponding to the first scene to each speech template from the set of speech templates corresponding to the character featured in the first scene to identify a speech template that corresponds to the retrieved portion;

identify a speech characteristic associated with the identified speech template;

determine whether the identified speech characteristic of the character featured in the first scene matches the context speech characteristic for the context of the first scene; and

in response to determining that the speech characteristic of the character featured in the first scene fails to match the context speech characteristic for the context of the first scene, perform a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context speech characteristic for the context of the first scene.

13. The system of claim 12 , wherein the control circuitry is further configured, when detecting the dubbed speech in the media asset, to:

receive the media asset;

retrieve video information corresponding to the media asset;

retrieve audio information corresponding to the media asset;

retrieve speech information of the character corresponding to the audio information;

retrieve facial movements of the character corresponding to the video information;

determine whether the facial movements of the character correspond to the speech information; and

in response to determining that the facial movements of the character do not correspond to the speech information, detect the dubbed speech in the media asset.

14. The system of claim 12 , wherein the control circuitry is further configured, when determining the context of the first scene based on the metadata corresponding to the first scene, to:

retrieve a plurality of keywords corresponding to the metadata corresponding to the first scene; and

determine context of the first scene based on a subset of the plurality of the keywords corresponding to the metadata corresponding to the first scene.

15. The system of claim 12 , wherein the control circuitry is further configured, when retrieving the context speech characteristic for the context of the first scene, to:

retrieve a personality metadata corresponding to the character, wherein the personality metadata comprises expected speech characteristic for each context for the character; and

identify the context speech characteristic as the expected speech characteristic for the context of the first scene based on the personality metadata corresponding to the character.

16. The system of claim 12 , wherein the control circuitry is further configured, when retrieving the set of speech templates corresponding to the character featured in the first scene, to:

retrieve a language of the dubbed speech corresponding to the media asset; and

retrieve the set of speech templates corresponding to the character featured in the first scene and corresponding to the language of the dubbed speech.

17. The system of claim 12 , wherein the control circuitry is further configured, when retrieving the set of speech templates corresponding to the character featured in the first scene, to:

retrieve an original speech corresponding to the media asset; and

retrieve a portion of the original speech corresponding to the first scene; and

retrieve the set of speech templates corresponding to the character featured in the first scene based on the retrieved portion of the original speech corresponding to the first scene.

18. The system of claim 12 , wherein the control circuitry is further configured, when comparing the retrieved portion of the dubbed speech corresponding to the first scene to each speech template from the set of speech templates corresponding to the character featured in the first scene, to:

retrieve a first set of vocal characteristics corresponding to the retrieved portion;

retrieve a second set of vocal characteristics corresponding to a speech template from the set of speech templates corresponding to the character featured in the first scene; and

compare a first vocal characteristic from the first set of vocal characteristics to a corresponding second vocal characteristic from the second set of vocal characteristics.

19. The system of claim 12 , wherein the control circuitry is further configured, when performing the function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context speech characteristic for the context of the first scene, to:

retrieve a first set of vocal characteristics corresponding to the portion of the dubbed speech;

identify a speech template from the set of speech templates corresponding to the character featured in the first scene, that has a speech characteristic that matches the context speech characteristic for the context of the first scene;

retrieve a second set of vocal characteristics corresponding to the speech template that has the speech characteristic that matched the context speech characteristic for the context of the first scene;

identify a first vocal characteristic from the first set of vocal characteristics that does not match a corresponding second vocal characteristic from the second set of vocal characteristics; and

adjust the first vocal characteristic from the first set of vocal characteristics to match the corresponding second vocal characteristic from the second set of vocal characteristics.

20. The system of claim 12 , the control circuitry is further configured, when performing the function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene match the context speech characteristic for the context of the first scene, to:

receive metadata corresponding to a second scene from the plurality of scenes;

determine context of the second scene based on the metadata corresponding to the second scene;

retrieve a context speech characteristic for the context of the second scene;

determine that the context speech characteristic for the context of the first scene matches the context speech characteristic for the context of the second scene;

retrieve a portion of the dubbed speech corresponding to the second scene;

retrieve a first set of vocal characteristics corresponding to the portion of the dubbed speech corresponding to the first scene;

retrieve a second set of vocal characteristics corresponding to the portion of the dubbed speech corresponding to the second scene;

identify a first vocal characteristic from the first set of vocal characteristics that does not match a corresponding second vocal characteristic from the second set of vocal characteristics; and

adjust the first vocal characteristic from the first set of vocal characteristics to match the corresponding second vocal characteristic from the second set of vocal characteristics.

21. The system of claim 20 , the control circuitry is further configured to:

retrieve a portion of an adjusted dubbed speech corresponding to the first scene;

compare the retrieved portion of the adjusted dubbed speech corresponding to the first scene to each speech template from the set of speech templates corresponding to the character featured in the first scene to identify a speech template that corresponds to the retrieved portion of the adjusted dubbed speech;

identify a speech characteristic associated with the identified speech template that corresponds to the retrieved portion of the adjusted dubbed speech; and

determine that the identified speech characteristic of the character featured in the first scene that corresponds to the retrieved portion of the adjusted dubbed speech matches the context speech characteristic for the context of the first scene.

Assignments (3)
CHANGE OF NAME Recorded Oct 2, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069086/0199 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: SANCHEZ, MARIO; MILLER, ASHLEIGH; STATHACOPOULOS, PAUL T.
To: ROVI GUIDES, INC.
Reel/Frame 051585/0125 →
Continuity (1)
Related Publication 20200066293A1 · Feb 27, 2020
Cited By (1)
US 12,614,540