IP Library › Granted Patent US 12,506,894
Granted Patent B2
US 12,506,894 · App. 18/687,768 · Granted Dec 23, 2025

Reshaper for learning based image/video coding

Inventors: Peng Yin (Ithaca, NY); Fangjun Pu (Sunnyvale, CA); Taoran Lu (Santa Clara, CA); Arjun Arora (Sunnyvale, CA); Guan-Ming Su (Fremont, CA); Tao Chen (Palo Alto, CA); Sean Thomas McCarthy (San Francisco, CA); Walter J. Husak (Simi Valley, CA)
Assignee: DOLBY LABORATORIES LICENSING CORPORATION
H04N19/517H04N19/117H04N19/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,506,894
App. No.
18/687,768
Granted
Dec 23, 2025
Kind
B2
Abstract

An input image represented in an input domain is received from an input video signal. Forward reshaping is performed on the input image to generate a forward reshaped image represented in a reshaped image domain. Non-reshaping encoding operations are performed to encode the reshaped image into an encoded video signal. At least one of the non-reshaping encoding operations is implemented with an ML model that has been previously trained with training images in one or more training datasets in a preceding training stage. A recipient device of the encoded video signal is caused to generate a reconstructed image from the forward reshaped image.

Claims (24)

1 . A method comprising:

receiving, from an input video signal, an input image represented in an input domain;

performing forward reshaping on the input image to generate a forward reshaped image represented in a reshaped image domain;

wherein the forward reshaped image is generated by performing the forward reshaping with a neural network that includes a non-linear mapping of input codewords in the input image to forward reshaped codewords in N channels, where N represents an integer no less than three;

performing non-reshaping encoding operations to encode the reshaped image into an encoded video signal, wherein at least one of the non-reshaping encoding operations is implemented with a machine learning (ML) model that has been previously trained with training images in one or more training datasets in a preceding training stage;

causing a recipient device of the encoded video signal to generate a reconstructed image from the forward reshaped image, wherein the reconstructed image is used to derive a display image to be rendered on an image display operating with the recipient device.

2 . The method of claim 1 , wherein the neural network represents a first convolutional neural network that uses a convolutional filter of spatial kernel size of 1 pixel×1 pixel to forward reshape each input codeword in the input image in three color channels to a respective forward reshaped codeword in N channels, where N represents an integer no less than three; wherein the reconstructed image is generated by inverse reshaping performed with a second convolutional neural network that uses a second convolutional filter of spatial kernel size of 1 pixel×1 pixel to inverse reshape each forward reshaped codeword in the input image in the N channels to a respective reconstructed codeword in the three color channels.

3 . The method of claim 1 , wherein the non-reshaping encoding operations include one or more of: optical flow analysis, motion vector encoding, motion vector decoding, motion vector quantization, motion compensation, residual encoding, residual decoding, or residual quantization.

4 . The method of claim 1 , wherein the forward reshaping is performed as out-of-loop image processing operations performed before the non-reshaping encoding operations.

5 . The method of claim 1 , wherein the forward reshaping is performed as a part of overall in-loop image processing operations that include the non-reshaping encoding operations.

6 . The method of claim 5 , wherein the overall in-loop image processing operations are encoding operations.

7 . The method of claim 1 , wherein an image metadata portion for the forward reshaped image is a part of image metadata carried by the encoded video signal; wherein the image metadata portion includes one or more of: forward reshaping parameters for the forward reshaping, or backward reshaping parameters for inverse reshaping.

8 . The method of claim 7 , wherein the image metadata portion includes reshaping parameters that explicitly specifies a reshaping mapping for one of the forward reshaping or the inverse reshaping.

9 . The method of claim 8 , wherein the reshaping parameters that explicitly specifies are reshaping mapping are generated by one of: a ML-based reshaping mapping prediction method, or a non-ML-based reshaping mapping generation method.

10 . The method of claim 1 , wherein the image metadata portion includes a reshaping parameter that identifies the forward reshaping as one of: global mapping or image adaptive mapping.

11 . The method of claim 1 , wherein the forward reshaping is performed with an implicit reshaping mapping embodied with weights and biases of a neural network that have been previously trained with training images in one or more training datasets.

12 . A method comprising:

decoding, from an encoded video signal, a forward reshaped image represented in a reshaped image domain, wherein the forward reshaped image was generated by an upstream device by forward reshaping an input image represented in an input image domain;

wherein the forward reshaped image was generated by the upstream device performing the forward reshaping with a neural network that includes a non-linear mapping of input codewords in the input image to forward reshaped codewords in N channels, where N represents an integer no less than three;

performing inverse reshaping on, as well as non-reshaping decoding operations in connection with, the forward reshaped image to generate a reconstructed image represented in a reconstructed image domain, wherein the inverse reshaping and forward reshaping form a reshaping operation pair, wherein at least one of the non-reshaping decoding operations is implemented with a machine learning (ML) model that has been previously trained with training images in one or more training datasets in a preceding training stage;

causing a display image derived from the reconstructed image to be rendered on an image display.

13 . The method of claim 12 , wherein the inverse reshaping is performed with an implicit reshaping mapping embodied with weights and biases of a neural network that have been previously trained with training images in one or more training datasets.

14 . The method of claim 12 , wherein the inverse reshaping is performed with a reshaping mapping signaled in an image metadata portion for the forward reshaped image carried in the encoded video signal as a part of image metadata.

15 . An apparatus comprising a processor and configured to perform the method recited in claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2024
From: ARORA, ARJUN; CHEN, TAO; HUSAK, WALTER J.; LU, TAORAN; MCCARTHY, SEAN THOMAS; PU, FANGJUN; SU, GUAN-MING; YIN, PENG
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 067543/0596 →
Priority Claims (1)
EP 21193790 · Aug 30, 2021 · regional
Continuity (2)
Provisional Application 63238529 · Aug 30, 2021
Related Publication 20240422345A1 · Dec 19, 2024
References Cited (29)
US 8811490B2 · Su · 2014 [cited by applicant]
US 10080026B2 · Su · 2018 [cited by applicant]
US 10136162B2 · Qu · 2018 [cited by applicant]
US 11277627B2 · Song · 2022 [cited by applicant]
US 11962760B2 · Su · 2024 [cited by applicant]
US 12149753B2 · Su · 2024 [cited by applicant]
US 20180020224A1 · Su · 2018 [cited by examiner]
US 20180124399A1 · Su · 2018 [cited by examiner]
US 20200090506A1 · Chen · 2020 [cited by examiner]
US 20200280831A1 · Booij · 2020 [cited by examiner]
US 20210076079A1 · Lu · 2021 [cited by applicant]
US 20210150812A1 · Su · 2021 [cited by applicant]
US 20230300381A1 · Su · 2023 [cited by applicant]
WO 2018064591A1 · 2018 [cited by applicant]
WO 2019199701A1 · 2019 [cited by applicant]
WO 2021168001A1 · 2021 [cited by applicant]
Alshina E et al, “JVET AHG report: Neural network-based video coding (AHGII)”, 23rd JVET Meeting; Jul. 7, 2021-Jul. 16, 2021; Teleconference; (The Joint Video Exploration Team of ISO/IEC JTC1/SC29/WG11AND ITU-T SG.16 ),… [cited by applicant]
Carl de Boor, “A Practical Guide to Splines”, 1978, 4 Pages. [cited by applicant]
G. Lu, et al., “DVC: An End-To-End Deep Video Compression Framework,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, Jan. 9, 2020, pp. 10998-11007, 10 Pages. [cited by applicant]
Hartmut Prautzsch, et al. “B'ezier- and B-spline technique”, Mathematics and Visualization, Mar. 26, 2002, 58 Pages. [cited by applicant]
ITU Rec. ITU-R BT. 1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production BT”, BT Series Broadcasting service, 2017, 7 Pages. [cited by applicant]
Jean Bégaint, et al., “CompressAI: a PyTorch library and evaluation platform for end-to-end compression research”, Computer Vision and Pattern Recognition, Nov. 5, 2020, 19 Pages. [cited by applicant]
E. Francois, et al., JVET-V0108, “AHG9: AHG9: Colour Transform Information SEI message”, Joint Video Experts Team (JVET) of ITU-T SG16WP 3 and ISO/IEC JTC 1/SC 29 22nd Meeting, by teleconference, Apr. 20-28, 2021, 6 Pag… [cited by applicant]
Lu T., et al: “CE12: Mapping functions (test CE12-1 and CE12-2)”, 125. MPEG Meeting; Jan. 14, 2019-Jan. 18, 2019; Marrakech; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m45700, JVET-M0427, Jan. 15, 2019… [cited by applicant]
Rec. ITU-R BT.2100, “Image parameter values for high dynamic range television for use in production and international programme exchange”, BT Series Broadcasting service, 2017, 16 Pages. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—“Coding of moving video”, ITU-T H.266, Aug. 2020, 516 Pages. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, “Advanced video coding for generic audiovisual services”, ITU-T H.264, Jun. 2019, 836 Pages. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, “High efficiency video coding”, ITU-T H.265 , Nov. 2019, 712 Pages. [cited by applicant]
SMPTE ST 2084:2014 “High Dynamic Range EOTF of Mastering Reference Displays”, Aug. 16, 2014, 15 Pages. [cited by applicant]