IP Library › Granted Patent US 12,272,039
Granted Patent B2
US 12,272,039 · App. 17/635,684 · Granted Apr 8, 2025

Efficient user-defined SDR-to-HDR conversion with model templates

Inventors: Guan-Ming Su (Fremont, CA); Harshad Kadu (Santa Clara, CA)
Assignee: DOLBY LABORATORIES LICENSING CORPORATION
G06T5/92G06T2207/10024G06T2207/20081G06T2207/20208
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,039
App. No.
17/635,684
Granted
Apr 8, 2025
Kind
B2
Abstract

Backward reshaping metadata prediction models are trained with training SDR images and corresponding training HDR images. Content creation user input to define user adjusted HDR appearances for the corresponding training HDR images is received. Content-creation-user-specific modified backward reshaping metadata prediction models are generated based on the trained prediction models and the content creation user input. The content-creation-user-specific modified prediction models are used to predict operational parameter values of content-creation-user-specific backward reshaping mappings for backward reshaping SDR images into mapped HDR images of at least one content-creation-user-adjusted HDR appearance.

Claims (27)

1. A method comprising:

accessing a model template comprising backward reshaping metadata prediction models, wherein the backward reshaping metadata prediction models are trained with a plurality of training image feature vectors from a plurality of training standard dynamic range (SDR) images in a plurality of training image pairs and ground truth derived with a plurality of corresponding training high dynamic range (HDR) images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs comprises a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, wherein the training SDR image and the corresponding training HDR image in each such training image pair depict same visual content but with different luminance dynamic ranges;

receiving content creation user input that defines one or more content-creation-user-adjusted HDR appearances for the plurality of corresponding training HDR images;

generating, based on the model template and the content creation user input, content-creation-user-specific modified backward reshaping metadata prediction models; and

using the content-creation-user-specific modified backward reshaping metadata prediction models to predict operational parameter values of content-creation-user-specific backward reshaping mappings for backward reshaping SDR images into mapped HDR images of at least one of the one or more content-creation-user-adjusted HDR appearances, wherein the backward reshaping metadata prediction models comprise a plurality of Gaussian process regression (GPR) models for predicting a luminance backward reshaping mapping to backward reshape input luminance SDR codewords into mapped luminance HDR codewords, and the content creation user input modifies a plurality of sample points of the luminance backward reshaping mapping.

2. The method of claim 1 , wherein the plurality of sample points as modified by the content creation user input is constrained to maintain the luminance backward reshaping mapping as a monotonically increasing function.

3. The method of claim 1 , wherein the plurality of image pairs is classified into a plurality of image categories; wherein the content creation user input modifies the luminance backward reshaping mapping differently for at least two image categories in the plurality of image categories.

4. The method of claim 1 , wherein the content creation user input modifies the luminance backward reshaping mapping that applies to all image pairs in the plurality of image pairs.

5. The method of claim 1 , wherein the one or more backward reshaping metadata prediction models comprise a set of multivariate multiple regression (MMR) mapping matrixes for generating MMR coefficients to generate mapped chrominance HDR codewords from input SDR codewords; wherein the content creation user input modifies a proper subset of MMR mapping matrixes in the set of MMR mapping matrixes with multiplicative operations; wherein remaining MMR mapping matrixes in the set of MMR mapping matrixes are freed from being modified by the content creation user input.

6. The method of claim 5 , wherein the plurality of image pairs is classified into a plurality of image categories; wherein the content creation user input modifies the proper subset of MMR mapping matrixes differently for at least two image categories in the plurality of image categories.

7. The method of claim 6 , wherein the plurality image categories is classified based on a plurality of regions each of which comprises a different group of mean predicted Cb values and mean predicted Cr values.

8. The method of claim 6 , wherein the plurality image categories is classified based on a plurality of different angle sub-ranges formed by different combinations of mean predicted Cb values and mean predicted Cr values.

9. The method of claim 5 , wherein the content creation user input modifies the proper set of MMR mapping matrixes that applies to all image pairs in the plurality of image pairs.

10. The method of claim 1 , further comprising: encoding one or more of the operational parameter values of backward reshaping mappings used to backward reshape SDR images into mapped HDR images into a video signal, along with the SDR images, as image metadata, wherein the video signal causes one or more recipient devices to render display images derived from the mapped HDR images with one or more display devices.

11. The method of claim 1 , wherein the backward reshaping metadata prediction models in the model template comprise a set of hyperparameter values and a set of weight factor values; wherein the content-creation-user-specific modified backward reshaping metadata prediction models are derived from the backward reshaping metadata prediction models in the model template by altering the set of weight factor values while maintaining the set of hyperparameter values unchanged.

12. An apparatus comprising a processor and configured to perform the method recited in claim 1 .

13. A non-transitory computer-readable storage medium having stored thereon computer-executable instruction for executing a method with one or more processors in accordance with the method recited in claim 1 .

14. A method comprising:

decoding, from a video signal, a standard dynamic range (SDR) image to be backward reshaped into a corresponding mapped high dynamic range (HDR) image;

decoding, from the video signal, composer metadata that is used to derive one or more operational parameter values of content-user-specific backward reshaping mappings;

wherein the one or more operational parameter values of content-user-specific backward reshaping mappings are predicted by one or more content-creation-user-specific modified backward reshaping metadata prediction models;

wherein the one or more content-creation-user-specific modified backward reshaping metadata prediction models are generated based on a model template and content creation user input;

wherein the model template includes backward reshaping metadata prediction models, wherein the backward reshaping metadata prediction models are trained with a plurality of training image feature vectors from a plurality of training SDR images in a plurality of training image pairs and ground truth derived with a plurality of corresponding training HDR images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs comprises a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, wherein the training SDR image and the corresponding training HDR image in each such training image pair depict same visual content but with different luminance dynamic ranges,

wherein the backward reshaping metadata prediction models comprise a plurality of Gaussian process regression (GPR) models for predicting a luminance backward reshaping mapping to backward reshape input luminance SDR codewords into mapped luminance HDR codewords,

wherein content creation user input modifies the plurality of corresponding training HDR images into one or more content-creation-user-adjusted HDR appearances;

using the one or more operational parameter values of the content-user-specific backward reshaping mappings to backward reshape the SDR image into the mapped HDR image of at least one of the one or more content-creation-user-adjusted HDR appearances; and

causing a display image derived from the mapped HDR image to be rendered with a display device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2022
From: SU, GUAN-MING; KADU, HARSHAD
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 059118/0024 →
Priority Claims (1)
EP 19191921 · Aug 15, 2019 · regional
Continuity (2)
Provisional Application 62887123 · Aug 15, 2019
Related Publication 20220301124A1 · Sep 22, 2022
References Cited (18)
US 8811490B2 · Su · 2014 [cited by applicant]
US 20170330529A1 · Van Mourik et al. · 2017 [cited by applicant]
US 20180020224A1 · Su · 2018 [cited by applicant]
US 20180115777A1 · Piramanayagam · 2018 [cited by examiner]
US 20180350047A1 · Baar · 2018 [cited by applicant]
US 20210195221A1 · Song · 2021 [cited by applicant]
US 20220058783A1 · Kadu · 2022 [cited by applicant]
CN 105405107A · 2016 [cited by applicant]
CN 108431886A · 2018 [cited by applicant]
CN 107105223B · 2018 [cited by applicant]
WO 2018231968A1 · 2018 [cited by applicant]
WO 2020131731A1 · 2020 [cited by applicant]
ITU REC.ITU-R BT. 1886 “Reference Electro-Optical Transfer Function for Flat Panel Displays Used in HDTV Studio Production” Mar. 2011. [cited by applicant]
ITU-R BT.2100-0, “Image Parameter Values for High Dynamic Range Television for Use in Production and International Programme Exchange” Jul. 2016. [cited by applicant]
Luzardo Gonzalo et al: “Fully-Automatic Inverse Tone Mapping Preserving the Content Creator's Artistic Intentions”, 2018 Picture Coding Symposium (PCS), IEEE, Jun. 24, 2018 (Jun. 24, 2018), pp. 199-203. [cited by applicant]
SMPTE ST 2084:2014 “High Dynamic Range EOTF of Mastering Reference Displays”. [cited by applicant]
Du Junpei, Interpolation, Enhamcement and Reconstruction of Transcale Motion Images, Pub No. 225, Apr. 30, 2019, 3 pages, Beijing, Beijing university publisher. [cited by applicant]
Huo Guan Ying et al., Side-Scan Sonal Image Target Segmentation, Apr. 30, 2017, pp. 24-25, 2 pages, Engineering University Press, Harbin, CN. [cited by applicant]