IP Library Granted Patent US 12,621,506
Granted Patent B2
US 12,621,506 · App. 18/474,065 · Granted May 5, 2026

Content adaptive micro encoding optimization for video

Inventors: Yuanyi Xue (Alameda, CA); Roberto Gerson De Albuquerque Azevedo (Zürich, CH); Christopher Richard Schroers (Uster, CH); Scott Labrozzi (Carey, NC); Wenhao Zhang (Beijing, CN)
Assignees: Disney Enterprises, Inc.; Beijing YoJaJa Software Technology Development Co., Ltd.
H04N21/23418G06V10/25G06V10/762H04N19/142H04N19/154H04N19/172H04N19/176H04N19/179H04N19/192H04N19/46H04N21/23424H04N21/812H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,506
App. No.
18/474,065
Granted
May 5, 2026
Kind
B2
Abstract

In some embodiments, a method analyzes flagged locations from a plurality of locations in an encoding of a video to form a cluster of locations. Draft micro-chunk boundaries for the cluster are determined based on searching for a first start location and a first end location in the encoding. The method searches in a first search range before the first start location and a second search range after the first end location for a second start location in the first search range and a second end location in the second search range. The second start location and the second end location form a micro-chunk. An encoding parameter set is determined for the micro-chunk formed by the second start location and the second end location based on content characteristics of the micro-chunk. The method uses the encoding parameter set to encode the micro-chunk for insertion in the encoding of the video.

Claims (60)

1 . A method comprising:

analyzing flagged frames from a plurality of frames in an encoding of a video to form a cluster of frames, wherein the cluster includes a plurality of flagged frames;

determining draft micro-chunk boundaries for the cluster of a first start frame and a first end frame in the encoding that includes the cluster;

searching in a first search range before the first start frame and a second search range after the first end frame for a second start frame in the first search range and a second end frame in the second search range, wherein the second start frame and the second end frame form a micro-chunk;

analyzing multiple frames based on the second start frame and the second end frame to determine content characteristics for the micro-chunk;

determining an encoding parameter set for the micro-chunk formed by the second start frame and the second end frame based on mapping the content characteristics of the micro-chunk to the encoding parameter set; and

using the encoding parameter set to encode the plurality of frames in the micro-chunk for insertion in the encoding of the video.

2 . The method of claim 1 , wherein the flagged frames are flagged based on a quality metric of the flagged frames meeting a threshold.

3 . The method of claim 1 , wherein a quality control process analyzes the encoding of the video based on a quality metric and outputs frames for the flagged frames in the encoding of the video.

4 . The method of claim 1 , wherein analyzing the flagged frames comprises:

grouping a portion of flagged frames in the cluster when distance between frames of the portion of flagged frames in the encoding of the video meet a threshold.

5 . The method of claim 1 , further comprising:

generating the first search range and the second search range based on a pre-set distance from the first start frame or the first end frame, or based on a length of the cluster.

6 . The method of claim 1 , wherein determining the draft micro-chunk boundaries comprises:

determining the draft micro-chunk boundaries based on a minimum micro-chunk duration.

7 . The method of claim 1 , wherein searching in the first search range and the second search range comprises:

determining a first usage area of video buffer verifier usage that meets a threshold in the first search range, wherein the video buffer verifier usage is determined during the encoding of the video; and

determining a second usage area of video buffer verifier usage that meets a threshold in the second search range.

8 . The method of claim 7 , wherein the first usage area and the second usage area have video buffer verifier usage that is lower than another area of the first search range or the second search range.

9 . The method of claim 1 , wherein searching in the first search range and the second search range comprises:

analyzing scene changes to select the second start frame or the second end frame.

10 . The method of claim 9 , wherein analyzing scene changes comprises:

calculating, for a scene change, an intra-scene distance for a scene based on the scene change and an inter-scene distance between scenes, wherein the intra-scene distance is a length of forward time in the video between the scene change and a next scene change, and the inter-scene distance is an average of distances to two scene change frames; and

selecting the scene change as the second start frame or the second end frame based on the intra-scene distance and inter-scene distance.

11 . The method of claim 9 , wherein analyzing scene changes comprises:

selecting the scene change as the second start frame or the second end frame based on a distance from multiple scene changes that are not within a threshold distance from the scene change.

12 . The method of claim 1 , wherein determining the encoding parameter set comprises:

analyzing content of the micro-chunk to classify the micro-chunk in a pre-defined category, wherein the encoding parameter set that is associated with the pre-defined category is used for the micro-chunk.

13 . The method of claim 1 , wherein determining the encoding parameter set comprises:

analyzing content of the micro-chunk to classify the micro-chunk in a pre-defined category that is based on an aspect of a content characteristic from a plurality of aspects of the content characteristic for the micro-chunk, wherein the encoding parameter set that is associated with the pre-defined category is used for the micro-chunk.

14 . The method of claim 1 , wherein determining the encoding parameter set comprises:

analyzing content of the micro-chunk to classify the micro-chunk in a first pre-defined category; and

analyzing content of the micro-chunk to classify the micro-chunk in a second pre-defined category that is based on an aspect of a content characteristic from a plurality of aspects of the content characteristic for the micro-chunk, wherein the encoding parameter set that is associated with the first pre-defined category or the second pre-defined category is used for the micro-chunk.

15 . The method of claim 1 , wherein determining the encoding parameter set comprises:

analyzing frames of the micro-chunk to classify the micro-chunk in a plurality of pre-defined categories, wherein each of the plurality of pre-defined categories is associated with a pre-defined encoding parameter set;

selecting a pre-defined category from the plurality of pre-defined categories; and

using the pre-defined encoding parameter set as the encoding parameter set for the micro-chunk.

16 . The method of claim 1 , wherein determining the encoding parameter set comprises:

using a continuous learning process to select the encoding parameter set, wherein the continuous learning process continually learns the encoding parameter set to use when encoding micro-chunks.

17 . The method of claim 16 , wherein the continuous learning process comprises:

receiving a state of an encoder that is encoding micro-chunks; and

generating a reward that is used to adjust the encoding parameter set.

18 . A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for:

analyzing flagged frames from a plurality of frames in an encoding of a video to form a cluster of frames, wherein the cluster includes a plurality of flagged frames;

determining draft micro-chunk boundaries for the cluster of a first start frame and a first end frame in the encoding that includes the cluster;

searching in a first search range before the first start frame and a second search range after the first end frame for a second start frame in the first search range and a second end frame in the second search range, wherein the second start frame and the second end frame form a micro-chunk;

analyzing multiple frames based on the second start frame and the second end frame to determine content characteristics for the micro-chunk;

determining an encoding parameter set for the micro-chunk formed by the second start frame and the second end frame based on mapping the content characteristics of the micro-chunk to the encoding parameter set; and

using the encoding parameter set to encode the plurality of frames in the micro-chunk for insertion in the encoding of the video.

19 . An apparatus comprising:

one or more computer processors; and

a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for:

analyzing flagged frames from a plurality of frames in an encoding of a video to form a cluster of frames, wherein the cluster includes a plurality of flagged frames;

determining draft micro-chunk boundaries for the cluster of a first start frame and a first end frame in the encoding that includes the cluster;

searching in a first search range before the first start frame and a second search range after the first end frame for a second start frame in the first search range and a second end frame in the second search range, wherein the second start frame and the second end frame form a micro-chunk;

analyzing multiple frames based on the second start frame and the second end frame to determine content characteristics for the micro-chunk;

determining an encoding parameter set for the micro-chunk formed by the second start frame and the second end frame based on mapping the content characteristics of the micro-chunk to the encoding parameter set; and

using the encoding parameter set to encode the plurality of frames in the micro-chunk for insertion in the encoding of the video.

20 . The non-transitory computer-readable storage medium of claim 18 , wherein determining the encoding parameter set comprises:

analyzing content of the micro-chunk to classify the micro-chunk in a pre-defined category, wherein the encoding parameter set that is associated with the pre-defined category is used for the micro-chunk.

Assignments (5)
CHANGE OF NAME Recorded Sep 24, 2024
From: BEIJING HULU SOFTWARE TECHNOLOGY DEVELOPMENT CO., LTD.
To: BEIJING YOJAJA SOFTWARE TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 068684/0455 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: LABROZZI, SCOTT; XUE, YUANYI
To: DISNEY ENTERPRISES, INC.
Reel/Frame 065015/0362 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: ZHANG, WENHAO
To: BEIJING HULU SOFTWARE TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 065015/0390 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: DE ALBUQUERQUE AZEVEDO, ROBERTO GERSON; SCHROERS, CHRISTOPHER RICHARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 065015/0392 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 065015/0395 →
Continuity (1)
Related Publication 20250106408A1 · Mar 27, 2025
References Cited (24)
US 7864840B2 · Labrozzi et al. · 2011 [cited by applicant]
US 20060233236A1 · Labrozzi et al. · 2006 [cited by applicant]
US 20100272173A1 · Puri · 2010 [cited by examiner]
US 20110255535A1 · Tinsman · 2011 [cited by examiner]
US 20120259894A1 · Varley · 2012 [cited by examiner]
US 20120275512A1 · Bouton et al. · 2012 [cited by applicant]
US 20150189222A1 · John · 2015 [cited by examiner]
US 20160073106A1 · Su · 2016 [cited by examiner]
US 20180176616A1 · Green · 2018 [cited by examiner]
US 20190075299A1 · Sethuraman · 2019 [cited by examiner]
US 20210006782A1 · Ramchandran et al. · 2021 [cited by applicant]
US 20210076045A1 · Xue et al. · 2021 [cited by applicant]
US 20210120061A1 · Labrozzi · 2021 [cited by examiner]
US 20220067386A1 · Rotman · 2022 [cited by examiner]
EP 1177691B1 · 2002 [cited by applicant]
JP 2009500951A · 2009 [cited by applicant]
WO 2007005750A2 · 2007 [cited by applicant]
Wu et al. “Scene Consistency Representation Learning for Video Scene Segmentation” (Year: 2022). [cited by examiner]
Mandhane et al., “MuZero with Self-competition for Rate Control inVP9 Video Compression,” arXiv:2202.06626v1 [eess.IV] Feb. 14, 2022, 20 pages. [cited by applicant]
Mnih, et al., “Asynchronous Methods for Deep Reinforcement Learning,” arXiv:1602.01783v2 [cs.LG] Jun. 16, 2016, 19 pages. [cited by applicant]
Scott Labrozzi, et al., “Surgical Micro-Encoding of Content,” U.S. Appl. No. 17/861,063, filed Jul. 8, 2022, 39 pages. [cited by applicant]
Extended European Search Report, European Patent Application 24188285.1, mailed Mar. 19, 2025, 13 pages. [cited by applicant]
First Office Action, Japan Application No. 2024-107092, Jul. 8, 2025, 9 pages. [cited by applicant]
Decision of Rejection, Japanese Patent Application No. 2024-107092, mailed Jan. 6, 2026, 10 pages. [cited by applicant]