IP Library › Granted Patent US 12,198,700
Granted Patent B2
US 12,198,700 · App. 18/328,358 · Granted Jan 14, 2025

Media system with closed-captioning data and/or subtitle data generation features

Inventors: Snehal Karia (Fremont, CA); Greg Garner (San Jose, CA); Sunil Ramesh (Cupertino, CA)
Assignee: Roku, Inc.
G10L15/26G10L15/25G10L25/57G10L25/78H04N21/4884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,700
App. No.
18/328,358
Granted
Jan 14, 2025
Kind
B2
Abstract

In one aspect, an example method includes (i) obtaining media, wherein the obtained media includes (a) audio representing speech and (b) video; (ii) using at least the audio representing speech as a basis to generate speech text; (iii) using at least the audio representing speech to determine starting and ending time points of the speech; and (iv) using at least the generated speech text and the determined starting and ending time points of the speech to (a) generate closed-captioning or subtitle data that includes closed-captioning or subtitle text based on the generated speech text and (b) associating the generated closed-captioning or subtitle data with the obtained media, such that the closed-captioning or subtitle text is time-aligned with the video based on the determined starting and ending time points of the speech.

Claims (53)

1. A method comprising:

obtaining media, wherein the obtained media includes (i) audio representing speech, (ii) video, and (iii) metadata associated with the obtained media;

using at least the audio representing speech as a basis to generate speech text, wherein using at least the audio representing speech as the basis to generate the speech text comprises:

(i) providing to a trained model, at least audio data for the audio representing speech and the metadata associated with the obtained media, wherein the metadata associated with the obtained media includes a rating of the obtained media; and

(ii) responsive to the providing, receiving from the trained model, generated speech text generated by the trained model;

using at least the audio representing speech as a basis to determine starting and ending time points of the speech; and

using at least the generated speech text and the determined starting and ending time points of the speech to (i) generate closed-captioning or subtitle data that includes closed-captioning or subtitle text based on the generated speech text and (ii) associating the generated closed-captioning or subtitle data with the obtained media, such that the closed-captioning or subtitle text is time-aligned with the video based on the determined starting and ending time points of the speech.

2. The method of claim 1 , further comprising:

extracting from the obtained media, the audio representing speech, wherein using the audio representing speech as the basis to generate the speech text comprises using the extracted audio representing speech as the basis to generate the speech text.

3. The method of claim 1 , wherein using at least the audio representing speech as the basis to generate the speech text comprises:

using at least (i) the audio representing speech and (ii) mouth movement depictions of the video, as the basis to generate the speech text.

4. The method of claim 1 , wherein using at least the audio representing speech to determine starting and ending time points of the speech comprises:

providing to a trained model, at least audio data for the audio representing speech; and

responsive to the providing, receiving from the trained model, starting and ending time points determined by the trained model.

5. The method of claim 1 , wherein the obtained media further includes audio representing a sound effect, and wherein the method further comprises:

using at least the audio representing the sound effect as a basis to generate sound effect description text.

6. The method of claim 5 , wherein using at least the audio representing the sound effect as the basis to generate the sound effect description text comprises:

providing to a trained model, at least audio data for the audio representing the sound effect; and

responsive to the providing, receiving from the trained model, sound effect description text generated by the trained model.

7. The method of claim 1 , wherein the generated closed-captioning or subtitle data is generated closed-captioning data, and wherein associating the generated closed-captioning data with the obtained media comprises:

storing the generated closed-captioning data as metadata associated with the obtained media.

8. The method of claim 7 , further comprising:

transmitting to a media-presentation device, the obtained media and the generated closed-captioning data as metadata of the media, wherein the media-presentation device is configured to (i) receive the transmitted media and closed-captioning data as metadata of the media, and (ii) present the received media with closed-captioning text overlaid thereon in accordance with the received closed-captioning data.

9. The method of claim 1 , wherein the generated closed-captioning or subtitle data is generated subtitle data, and wherein associating the generated subtitle data with the obtained media comprises:

modifying the obtained media by overlaying on it subtitle text in accordance with the subtitle data.

10. The method of claim 9 , further comprising:

transmitting to a media-presentation device, the modified media, wherein the media-presentation device is configured to receive and output for presentation the modified media.

11. The method of claim 1 , wherein the generated closed-captioning or subtitle data is generated closed-captioning data, and wherein the method further comprises:

outputting for presentation, by a media presentation device, media with closed-captioning text overlaid thereon in accordance with the closed-captioning data.

12. The method of claim 1 , wherein the generated closed-captioning or subtitle data is generated subtitle data, and wherein the method further comprises:

outputting for presentation, by a media presentation device, media modified to include subtitle text in accordance with the subtitle data.

13. The method of claim 1 , further comprising:

determining that the speech was spoken by a character associated with a given region within the video; and

based on the determining, outputting the speech text in or near that given region.

14. The method of claim 1 , further comprising:

determining that the speech was spoken by a given character; and

based on the determining, outputting the speech text in a font color associated with the given character.

15. A computing system comprising a processor and a non-transitory computer-readable storage medium having stored thereon program instructions that upon execution by the processor, cause the computing system to perform a set of acts comprising:

obtaining media, wherein the obtained media includes (i) audio representing speech, (ii) video, and (iii) metadata associated with the obtained media;

using at least the audio representing speech as a basis to generate speech text, wherein using at least the audio representing speech as the basis to generate the speech text comprises:

(i) providing to a trained model, at least audio data for the audio representing speech and the metadata associated with the obtained media, wherein the metadata associated with the obtained media includes a rating of the obtained media; and

(ii) responsive to the providing, receiving from the trained model, generated speech text generated by the trained model;

using at least the audio representing speech to determine starting and ending time points of the speech; and

using at least the generated speech text and the determined starting and ending time points of the speech to (i) generate closed-captioning or subtitle data that includes closed-captioning or subtitle text based on the generated speech text and (ii) associating the generated closed-captioning or subtitle data with the obtained media, such that the closed-captioning or subtitle text is time-aligned with the video based on the determined starting and ending time points of the speech.

16. The computing system of claim 15 , wherein the generated closed-captioning or subtitle data is generated closed-captioning data, and wherein associating the generated closed-captioning data with the obtained media comprises:

storing the generated closed-captioning data as metadata associated with the obtained media.

17. A non-transitory computer-readable storage medium having stored thereon program instructions that upon execution by a processor, cause a computing system to perform a set of acts comprising:

obtaining media, wherein the obtained media includes (i) audio representing speech, (ii) video, and (iii) metadata associated with the obtained media;

using at least the audio representing speech as a basis to generate speech text, wherein using at least the audio representing speech as the basis to generate the speech text comprises:

(i) providing to a trained model, at least audio data for the audio representing speech and metadata associated with the obtained media, wherein the metadata associated with the obtained media includes a rating of the obtained media; and

(ii) responsive to the providing, receiving from the trained model, generated speech text generated by the trained model;

using at least the audio representing speech to determine starting and ending time points of the speech; and

using at least the generated speech text and the determined starting and ending time points of the speech to (i) generate closed-captioning or subtitle data that includes closed-captioning or subtitle text based on the generated speech text and (ii) associating the generated closed-captioning or subtitle data with the obtained media, such that the closed-captioning or subtitle text is time-aligned with the video based on the determined starting and ending time points of the speech.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2023
From: KARIA, SNEHAL; GARNER, GREG; RAMESH, SUNIL
To: ROKU, INC.
Reel/Frame 063890/0655 →
Continuity (1)
Related Publication 20240404525A1 · Dec 5, 2024
References Cited (7)
US 8924210B2 · Basson · 2014 [cited by applicant]
US 9418152B2 · Nissan · 2016 [cited by applicant]
US 20050267894A1 · Camahan · 2005 [cited by examiner]
US 20070011012A1 · Yurick · 2007 [cited by examiner]
US 20080002291A1 · Balamane · 2008 [cited by applicant]
US 20200084521A1 · Adler · 2020 [cited by applicant]
US 20230252980A1 · Kumar · 2023 [cited by examiner]
Cited By (1)
US 12,513,374