IP Library Granted Patent US 12,700,428
Granted Patent B2
US 12,700,428 · App. 18/865,266 · Granted Aug 4, 2026

Trim pass metadata prediction in video sequences using neural networks

Inventors: Sri Harsha Musunuri (Santa Clara, CA); Shruthi Suresh Rotti (Mountain House, CA); Anustup Kumar Atanu Choudhury (Campbell, CA)
Assignee: DOLBY LABORATORIES LICENSING CORPORATION
G11B27/031G06V10/42G06V10/7715G06V10/774G06V10/82G11B27/34
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,428
App. No.
18/865,266
Filed
Nov 12, 2024
Granted
Aug 4, 2026
Kind
B2
Art Unit
2481
USPC
386/239
Abstract

Methods and systems for generating trim-pass metadata for high dynamic range (HDR) video are described. The trim-pass prediction pipeline includes a feature extraction network followed by a fully connected network which maps extracted features to trim-pass values. In a first architecture, the feature extraction network is based on four cascaded convolutional networks. In a second architecture, the feature extraction network is based on a modified MobileNetV3 neural network. In both architectures, the fully connected network is formed by a set of three linear networks, each set customized to best match its corresponding feature extraction network.

Claims (21)

1 . A method for generating trim-pass metadata of pictures in a video sequence, wherein the trim-pass metadata are configured to perform adjustments of a tone-mapping curve that is applied to an input picture when being displayed on a target display, the method comprising:

receiving the input picture;

providing a feature extraction network, the feature extraction network comprising a convolutional neural network for feature extraction that is trained to identify high-level image features of the input picture;

applying the feature extraction network to the input picture to generate the image features;

providing a fully connected network, the fully connected network comprising a plurality of cascaded linear neural networks that are trained to map the image features to output trim-pass metadata values for the input picture; and

applying the fully connected network to the image features to map the image features to the output trim-pass metadata values for the input picture, wherein training the networks comprises:

receiving input training trim-pass parameters corresponding to the input picture;

applying an error loss unit to generate an error metric based on the input training trim-pass parameters and the output trim-pass metadata; and

training the feature extraction network and the fully connected network by minimizing the error metric.

2 . The method of claim 1 , wherein the input picture is a high-dynamic range (HDR) picture coded using PQ encoding in the ICtCp color space.

3 . The method of claim 1 , wherein the feature extraction network comprises four cascaded convolutional networks.

4 . The method of claim 3 , wherein the fully connected network comprises three cascaded linear networks.

5 . The method of claim 1 , wherein the feature extraction network comprises a modified MobileNetV3 neural network accepting inputs with non square aspect ratios.

6 . The method of claim 5 , wherein the fully connected network comprises three cascaded linear networks.

7 . The method of claim 1 , wherein computing the error metric comprises computing a minimum absolute error or a mean square error between the input training trim-pass parameters and the output trim-pass metadata.

8 . The method of claim 1 , wherein computing the error metric comprises:

generating a first tone-mapping function based at least on the input training trim-pass parameters;

generating a second tone-mapping function based at least on the output trim-pass metadata; and

computing the minimum absolute error or the mean square error between values of the first tone-mapping function and the second tone-mapping function.

9 . An apparatus comprising a processor and configured to perform the method recited in claim 1 .

10 . A non-transitory computer-readable storage medium having stored thereon computer-executable instruction for executing a method with one or more processors in accordance with claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2025
From: MUSUNURI, SRI HARSHA; ROTTI, SHRUTHI SURESH; CHOUDHURY, ANUSTUP KUMAR ATANU
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 071820/0183 →
Priority Claims (1)
EP 22182506 · Jul 1, 2022 · regional
Continuity (2)
Provisional Application 63342306 · May 16, 2022
Related Publication 20250316291A1 · Oct 9, 2025
References Cited (30)
US 8593480B1 · Ballestad · 2013 [cited by applicant]
US 9098906B2 · Bruls · 2015 [cited by applicant]
US 9715642B2 · Szegedy · 2017 [cited by examiner]
US 9819974B2 · Kunkel · 2017 [cited by applicant]
US 10244244B2 · Piramanayagam · 2019 [cited by applicant]
US 10264287B2 · Wen · 2019 [cited by examiner]
US 10600166B2 · Pytlarz · 2020 [cited by applicant]
US 11158286B2 · Yaacob · 2021 [cited by applicant]
US 11310509B2 · Topiwala · 2022 [cited by applicant]
US 11361506B2 · Su · 2022 [cited by applicant]
US 20190080440A1 · Eriksson · 2019 [cited by applicant]
US 20210076042A1 · Choudhury · 2021 [cited by examiner]
US 20210142068A1 · Aliamiri · 2021 [cited by examiner]
US 20210150812A1 · Su · 2021 [cited by examiner]
US 20210350512A1 · Kadu · 2021 [cited by examiner]
US 20210400286A1 · Kale · 2021 [cited by applicant]
US 20220078386A1 · Zink · 2022 [cited by applicant]
US 20220084170A1 · Aydin · 2022 [cited by applicant]
JP 2020017079A · 2020 [cited by applicant]
JP 2020119462A · 2020 [cited by applicant]
JP 2020533841A · 2020 [cited by applicant]
JP 2021135619A · 2021 [cited by applicant]
JP 2022524651A · 2022 [cited by applicant]
WO 2021030506A1 · 2021 [cited by applicant]
A Color-Volume Mapping System for Perception-Accurate Reproduction of HDR Imagery in SDR production workflows. Pascal Kutschbach. SMPTE 2020 Annual Technical Conference and Exhibition, Date of Conference: Nov. 10-12, 20… [cited by applicant]
A. Howard et al., “Searching for MobileNetV3,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), Nov. 20, 2019., pp. 1314-1324, 11 Pages. [cited by applicant]
Batch normalization: accelerating deep network training by reducing internal covariate shift, In Proceedings of the 32nd International Conference on International Conference on Machine Learning—vol. 37 (ICML'15). JMLR.o… [cited by applicant]
Dolby Professional Support Learning. (2022). Module 2.8—The Dolby Vision Metadata Trim Pass. Retrieved from https://learning.dolby.com/hc/en-us/articles/360056574431-Module-2-8-The-Dolby-Vision-Metadata-Trim-Pass. 9 pag… [cited by applicant]
International Telecommunication Union. (2020). High dynamic range television for production and international programme exchange (Report ITU-R BT.2390-8). Geneva, Switzerland: ITU. 59 Pages. [cited by applicant]
J. Hu, L. Shen and G. Sun, “Squeeze-and-Excitation Networks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 7132-7141, doi: 10.1109/CVPR.2018.00745. [cited by applicant]