IP Library › Granted Patent US 12,541,948
Granted Patent B2
US 12,541,948 · App. 18/485,891 · Granted Feb 3, 2026

Frame classification to generate target media content

Inventors: Bruce Patrick Robert Williams (Bristol, GB); Joseph William Bignell (Cardiff, GB); Russell Stuart Love (Cardiff, GB)
Assignee: Roku, Inc.
G06V10/764G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,948
App. No.
18/485,891
Granted
Feb 3, 2026
Kind
B2
Abstract

Aspects of the disclosed technology provide solutions for processing media content to generate customized media content of a target duration. An example method can include receiving media content of a first duration. The media content may include a plurality of video frames. The method can include steps for receiving one or more parameters, which may include a target duration, classifying each of the plurality of video frames of the media content based on a relevance level of each frame, and generating a target media content of the target duration based on the classification of the plurality of video frames of the media content of the first duration. Systems and machine-readable media are also provided.

Claims (47)

1 . A system, comprising:

one or more memories; and

at least one processor coupled to at least one of the one or more memories and configured to perform operations comprising:

receiving media content of a first duration, the media content comprising a plurality of video frames;

receiving one or more parameters associated with a target media content, wherein the one or more parameters include a parameter indicating a predetermined runtime duration;

classifying each of the plurality of video frames of the media content based on a respective relevance level associated with each of the plurality of video frames with respect to the target media content, wherein each of the plurality of video frames are classified as a critical frame or an uncritical frame;

selecting, from the plurality of video frames, a set of video frames classified as the critical frame to include in the target media content, wherein a runtime duration of the set of video frames meets the predetermined runtime duration;

generating the target media content including the set of video frames of the plurality of video frames; and

presenting the target media content via a display device.

2 . The system of claim 1 , wherein the respective relevance level is determined based on a presence or absence of text displayed on each of the plurality of video frames.

3 . The system of claim 1 , wherein the respective relevance level is determined based on a location of each frame within the media content.

4 . The system of claim 1 , wherein the respective relevance level is determined based on the one or more parameters associated with the target media content, which include at least one of a provider of the media content, one or more characteristics of the target media content, a geographic region for streaming the target media content, and target audience demographics.

5 . The system of claim 1 , wherein the at least one processor is configured to perform operations comprising:

generating, using a generative machine learning (ML) model, a transcript of the media content based on audio signals of the media content; and

regenerating audio signals corresponding to the predetermined runtime duration of the target media content based on the transcript and the predetermined runtime duration.

6 . The system of claim 1 , wherein the at least one processor is configured to perform operations comprising:

generating a voiceover narrative for the target media content in a new language that is different from an original language associated with the media content.

7 . The system of claim 1 , wherein the at least one processor is configured to perform operations comprising:

generating audio signals of the target media content based on a depicted text on one or more frames of the target media content.

8 . The system of claim 1 , wherein the respective relevance level includes a relevance score based on the one or more parameters associated with the target media content.

9 . The system of claim 8 , wherein generating the target media content of the predetermined runtime duration comprises:

selecting one or more video frames of the plurality of video frames based on a corresponding relevance score.

10 . A computer-implemented method for processing media content, the computer-implemented method comprising:

receiving media content of a first duration, the media content comprising a plurality of video frames;

receiving one or more parameters associated with a target media content, wherein the one or more parameters include a parameter indicating a predetermined runtime duration;

classifying each of the plurality of video frames of the media content based on a respective relevance level associated with each of the plurality of video frames with respect to the target media content, wherein each of the plurality of video frames are classified as a critical frame or an uncritical frame;

selecting, from the plurality of video frames, a set of video frames classified as the critical frame to include in the target media content, wherein a runtime duration of the set of video frames meets the predetermined runtime duration;

generating the target media content including the set of video frames of the plurality of video frames; and

presenting the target media content via a display device.

11 . The computer-implemented method of claim 10 , wherein the respective relevance level is determined based on a presence or absence of text displayed on each of the plurality of video frames.

12 . The computer-implemented method of claim 10 , wherein the respective relevance level is determined based on a location of each frame within the media content.

13 . The computer-implemented method of claim 10 , wherein the respective relevance level is determined based on the one or more parameters associated with the target media content, which include at least one of a provider of the media content, one or more characteristics of the target media content, a geographic region for streaming the target media content, and target audience demographics.

14 . The computer-implemented method of claim 10 , further comprising:

generating, using a generative machine learning (ML) model, a transcript of the media content based on audio signals of the media content; and

regenerating audio signals corresponding to the predetermined runtime duration of the target media content based on the transcript and the predetermined runtime duration.

15 . The computer-implemented method of claim 10 , further comprising:

generating a voiceover narrative for the target media content in a new language that is different from an original language associated with the media content.

16 . The computer-implemented method of claim 10 , further comprising:

generating audio signals of the target media content based on a depicted text on one or more frames of the target media content.

17 . The computer-implemented method of claim 10 , wherein the respective relevance level includes a relevance score based on the one or more parameters associated with the target media content.

18 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

receiving media content of a first duration, the media content comprising a plurality of video frames;

receiving one or more parameters associated with a target media content, wherein the one or more parameters include a parameter indicating a predetermined runtime duration;

classifying each of the plurality of video frames of the media content based on a respective relevance level associated with each of the plurality of video frames with respect to the target media content, wherein each of the plurality of video frames are classified as a critical frame or an uncritical frame;

selecting, from the plurality of video frames, a set of video frames classified as the critical frame to include in the target media content, wherein a runtime duration of the set of video frames meets the predetermined runtime duration;

generating the target media content including the set of video frames of the plurality of video frames; and

presenting the target media content via a display device.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2023
From: WILLIAMS, BRUCE PATRICK ROBERT; BIGNELL, JOSEPH WILLIAM; LOVE, RUSSELL STUART
To: ROKU, INC.
Reel/Frame 065203/0615 →
Continuity (1)
Related Publication 20250124689A1 · Apr 17, 2025
References Cited (13)
US 6535639B1 · Uchihachi · 2003 [cited by examiner]
US 11120293B1 · Rosenzweig · 2021 [cited by examiner]
US 20060165379A1 · Agnihotri · 2006 [cited by examiner]
US 20140344702A1 · Edge · 2014 [cited by examiner]
US 20160014482A1 · Chen · 2016 [cited by examiner]
US 20180061459A1 · Song · 2018 [cited by examiner]
US 20190251360A1 · Cricri · 2019 [cited by examiner]
US 20190377955A1 · Swaminathan et al. · 2019 [cited by applicant]
US 20200334468A1 · Agarwal et al. · 2020 [cited by applicant]
US 20210224571A1 · Vartakavi · 2021 [cited by examiner]
US 20230306539A1 · Frei · 2023 [cited by examiner]
Extended European Search Report in corresponding European application No. 24205555.6 mailed Mar. 3, 2025 in 12 pages. [cited by applicant]
Mayu Otani, et al: “Video Summarization Overview”, Arxiv.org, Cornell University Library, 202 Olin Library Cornell University Ithaca, NY 14853, Oct. 21, 2022 (Oct. 21, 2022), XP091350362, DOI: 10.1561/0600000099 in 55 p… [cited by applicant]