IP Library Granted Patent US 11,736,773
Granted Patent B2
US 11,736,773 · App. 17/502,205 · Granted Aug 22, 2023

Interactive pronunciation learning system

Inventor: Serhad Doken (Bryn Mawr, PA)
Assignee: Rovi Guides, Inc.
H04N21/4884G10L15/187G10L15/26H04N21/4856
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,736,773
App. No.
17/502,205
Granted
Aug 22, 2023
Kind
B2
Abstract

Systems and methods for generating audible pronunciation of a closed captioning word in a content item. For example, a system generates for output on a first device a content item comprising dialogue. The system generates for display on the first device a closed captioning word corresponding to the dialogue where the closed captioning word is selectable via a user interface of the first device. The system receives a selection of the closed captioning word via the user interface of the first device. In response to receiving the selection of the closed captioning word, the system generates for playback on the first device at least a portion of the dialogue corresponding to the selected closed captioning word.

Claims (76)

1. A method comprising:

generating for output on a first device a content item comprising dialogue;

generating for display on the first device a closed captioning word corresponding to the dialogue, the closed captioning word being selectable via a user interface of the first device;

receiving a selection of the closed captioning word via the user interface of the first device;

identifying a plurality of pronunciation styles for the selected closed captioning word in a first language stored in a database;

generating for display a list of the plurality of pronunciation styles on the first device;

receiving a selection of a pronunciation style of the plurality of pronunciation styles;

retrieving an audio file containing audible pronunciation of the selected closed captioning word in the selected style; and

generating for output the audio file containing the audible pronunciation of the selected closed captioning word in the selected style.

2. The method of claim 1 , further comprising:

retrieving metadata of the content item, wherein the metadata of the content item comprises the dialogue and a respective timestamp corresponding to a word in the dialogue; and

retrieving the closed captioning word corresponding to the dialogue from a database of a content item.

3. The method of claim 2 , further comprising:

comparing the metadata of the content item to the closed captioning word corresponding to the dialogue; and

based on the comparison, determining that the at least the portion of the dialogue corresponds to the selected closed captioning word.

4. The method of claim 1 , further comprising:

determining that temporal proximity of a first set of words in the dialogue is less than a threshold; and

in response to determining that the temporal proximity of the first set of words in the dialogue is less than the threshold, categorizing the first set of words as a first phrase.

5. The method of claim 4 , further comprising:

retrieving an audio file containing audible pronunciation of the first phrase in the first language;

receiving a selection of at least one word of the first set of words via the user interface of the first device; and

generating for output the audible pronunciation of the first phrase.

6. The method of claim 1 , further comprising:

receiving a vocal input corresponding to the selected closed captioning word; and

comparing the vocal input to the audio file containing the audible pronunciation of the selected closed captioning word to calculate a similarity score.

7. The method of claim 6 , further comprising:

transmitting the vocal input to a server to enable rendering of the vocal input on a second device that is different from the first device.

8. The method of claim 1 , wherein the plurality of pronunciation styles includes at least one of a standard accent, a non-standard accent, a dialect, or a slang.

9. The method of claim 1 , further comprising:

in response to receiving the selection of the closed captioning word, pausing generating for output a video of the content item.

10. A method comprising:

generating for output on a first device a content item comprising dialogue;

generating for display on the first device a closed captioning word corresponding to the dialogue, the closed captioning word being selectable via a user interface of the first device;

receiving a selection of the closed captioning word via the user interface of the first device;

identifying that the selected closed captioning word is spoken by one or more characters of the content item;

generating for display a list of one or more characters of the content item;

receiving a selection of a character of the one or more characters;

retrieving an audio file containing audible pronunciation of the selected closed captioning word spoken by the selected character; and

generating for output the retrieved audio file containing audible pronunciation of the selected closed captioning word spoken by the selected character.

11. A system comprising:

control circuitry configured to:

generate for output on a first device a content item comprising dialogue;

generate for display on the first device a closed captioning word corresponding to the dialogue, the closed captioning word being selectable via a user interface of the first device;

receive a selection of the closed captioning word via the user interface of the first device;

identify a plurality of pronunciation styles for the selected closed captioning word in a first language stored in a database;

generate for display a list of the plurality of pronunciation styles on the first device;

receive a selection of a pronunciation style of the plurality of pronunciation styles;

retrieve an audio file containing audible pronunciation of the selected closed captioning word in the selected style; and

generate for output the audio file containing the audible pronunciation of the selected closed captioning word in the selected style.

12. The system of claim 11 , wherein the control circuitry is further configured to:

retrieve metadata of the content item, wherein the metadata of the content item comprises the dialogue and a respective timestamp corresponding to a word in the dialogue; and

retrieve the closed captioning word corresponding to the dialogue from a database of a content item.

13. The system of claim 12 , wherein the control circuitry is further configured to:

compare the metadata of the content item to the closed captioning word corresponding to the dialogue; and

based on the comparison, determine that the at least the portion of the dialogue corresponds to the selected closed captioning word.

14. The system of claim 11 , wherein the control circuitry is further configured to:

determine that temporal proximity of a first set of words in the dialogue is less than a threshold; and

in response to determining that the temporal proximity of the first set of words in the dialogue is less than the threshold, categorize the first set of words as a first phrase.

15. The system of claim 14 , wherein the control circuitry is further configured to:

retrieve an audio file containing audible pronunciation of the first phrase in the first language;

receive a selection of at least one word of the first set of words via the user interface of the first device; and

generate for output the audible pronunciation of the first phrase.

16. The system of claim 11 , wherein the control circuitry is further configured to:

receive a vocal input corresponding to the selected closed captioning word; and

compare the vocal input to the audio file containing the audible proninciation of the selected closed captioning word to calculate a similarity score.

17. The system of claim 16 , wherein the control circuitry is further configured to:

transmit the vocal input to a server to enable rendering of the vocal input on a second device that is different from the first device.

18. The system of claim 11 , wherein the plurality of pronunciation styles includes at least one of a standard accent, a non-standard accent, a dialect, or a slang.

19. The system of claim 11 , wherein the control circuitry is further configured to:

in response to receiving the selection of the closed captioning word, pause generating for output a video of the content item.

20. The system of claim 11 , wherein the control circuitry is further configured to:

identify that the selected closed captioning word is spoken by one or more characters of the content item;

generate for display a list of one or more characters of the content item;

receive a selection of a character of the one or more characters;

retrieve an audio file containing audible pronunciation of the selected closed captioning word spoken by the selected character; and

generate for output the retrieved audio file containing the audible pronunciation of the selected closed captioning word spoken by the selected character.

Assignments (3)
CHANGE OF NAME Recorded Oct 4, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069113/0399 →
SECURITY INTEREST Recorded May 19, 2023
From: ADEIA GUIDES INC.; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063707/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: DOKEN, SERHAD
To: ROVI GUIDES, INC.
Reel/Frame 057992/0657 →
Continuity (1)
Related Publication 20230124847A1 · Apr 20, 2023