IP Library Granted Patent US 12,536,711
Granted Patent B2
US 12,536,711 · App. 18/053,253 · Granted Jan 27, 2026

Decoder-side fine-tuning of neural networks for video coding for machines

Inventors: Francesco Cricrì (Tampere, FI); Honglei Zhang (Tampere, FI); Miska Matias Hannuksela (Tampere, FI); Hamed Rezazadegan Tavakoli (Espoo, FI); Nam Hai Le (Tampere, FI); Ramin Ghaznavi Youvalari (Tampere, FI); Jukka Ahonen (Tampere, FI); Emre Baris Aksu (Tampere, FI)
Assignee: Nokia Technologies Oy
G06T9/002G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,711
App. No.
18/053,253
Granted
Jan 27, 2026
Kind
B2
Abstract

Various embodiments provide an apparatus, a method, and a computer program product. An example apparatus includes at least one processor; and at least one non-transitory memory comprising computer program code; wherein the at least one memory and the computer program code are configured to, with the at least one processor, cause the apparatus at least to iteratively perform following until a stopping criterion is met: provide a finetuning driving content (FDC) or a content derived from FDC to a decoder side neural network (DSNN); compute an output of the DSNN as a processed FDC; compute a loss based on the processed FDC and an approximated ground truth data (AGT) associated with the FDC; compute an update to the DSNN; and apply the computed update to the DSNN.

Claims (64)

1 . An apparatus comprising:

at least one processor; and

at least one memory comprising computer program code; wherein the at least one memory and the computer program code are configured to, with the at least one processor, cause the apparatus at least to iteratively perform following until a stopping criterion is met:

provide a finetuning driving content (FDC) or a content derived from FDC to a decoder side neural network (DSNN);

compute an output of the DSNN as a processed FDC;

compute a loss based on the processed FDC and an approximated ground truth data (AGT) associated with the FDC;

compute an update to the DSNN;

apply the computed update to the DSNN;

run a loss proxy neural network for one or more input frames; and

encode an output of the loss proxy neural network, representing the AGT, as metadata.

2 . The apparatus of claim 1 , wherein to compute the update to the DSNN, the apparatus is further caused to:

use a backpropagation algorithm to compute gradients of the loss with respect to one or more parameters of the DSNN; and

apply an optimizer routine to compute the update.

3 . The apparatus of claim 1 , wherein the apparatus is further caused to:

encode one or more parts of a video in higher quality (HQ) content and remaining parts of the video in lower quality (LQ) content; and

use the HQ content to obtain the AGT.

4 . The apparatus of claim 1 , wherein the apparatus is further caused to lower a quality of the HQ content, and wherein to lower the quality of the HQ content the apparatus is further caused to perform at least one of the following:

encode the HQ content by using an encoding configuration that causes a decoded content to be of lower quality with respect to the HQ content and decode the encoded HQ content;

downsample the HQ content to a lower resolution; or

input the HQ content to a NN, and treat the output of the NN as a lower quality (LQ) content, wherein the NN is trained to output a content that has similar quality as the LQ content.

5 . The apparatus of claim 3 , wherein the apparatus is further caused to obtain the FDC at a decoder side by having an encoder encode the FDC as a low quality version of the HQ content, or as LQ content which is related or similar to the HQ content.

6 . The apparatus of claim 5 , wherein the apparatus is further caused to:

encode the LQ and HQ contents; and

compute the AGT as an output of a loss proxy neural network when an input to the loss proxy neural network comprises the HQ content or a content derived from the HQ content.

7 . The apparatus of claim 5 , wherein the apparatus is further caused to:

encode the LQ and HQ version of an input content; and

compute the AGT as an output of a loss proxy neural network when an input to the loss proxy neural network comprises the HQ content or a content derived from the HQ content.

8 . The apparatus of claim 5 , wherein the apparatus is further caused to:

extract a high resolution patch from the HQ content; and

extract a low resolution patch from the LQ content.

9 . The apparatus of claim 1 , wherein the apparatus is caused to:

store the finetuned DSNN in a buffer; and

signal the stored DSNN when the stored DSNN is to be used to process subsequent regions or frames or when the stored DSNN needs to be updated.

10 . A method comprising:

providing a finetuning driving content (FDC) or a content derived from FDC to a decoder side neural network (DSNN);

computing an output of the DSNN as a processed FDC;

computing a loss based on the processed FDC and an approximated ground truth data (AGT) associated with the FDC;

computing an update to the DSNN;

applying the computed update to the DSNN;

running a loss proxy neural network for one or more input frames; and

encoding an output of the loss proxy neural network, representing the AGT, as metadata.

11 . The method of claim 10 , wherein computing the update to the DSNN comprises:

using a backpropagation algorithm to compute gradients of the loss with respect to one or more parameters of the DSNN; and

applying an optimizer routine to compute the update.

12 . The method claim 10 , further comprising:

encoding one or more parts of a video in a higher quality (HQ) content and remaining parts of the video in a lower quality (LQ) content; and

using the HQ content to obtain the AGT.

13 . The method of claim 10 , further comprising lowering a quality of the HQ content, wherein lowering the quality of the HQ content comprises:

encoding the HQ content by using an encoding configuration that causes a decoded content to be of lower quality with respect to the HQ content and decode the encoded HQ content;

downsampling the HQ content to a lower resolution; or

inputting the HQ content to a NN, and treat the output of the NN as a lower quality (LQ) content, wherein the NN is trained to output a content that has similar quality as the LQ content.

14 . The method of claim 13 , further comprising obtaining the FDC at a decoder side by having an encoder encode the FDC as a low quality version of the HQ content, or as LQ content which is related or similar to the HQ content.

15 . The method of claim 14 , further comprising:

encoding the LQ and HQ contents; and

computing the AGT as an output of a loss proxy neural network when an input to the loss proxy neural network comprises the HQ content or a content derived from the HQ content.

16 . The method of claim 14 , further comprising:

encoding the LQ and HQ version of an input content; and

computing the AGT as an output of a loss proxy neural network when an input to the loss proxy neural network comprises the HQ content or a content derived from the HQ content.

17 . The method of claim 14 , further comprising:

extracting a high resolution patch from the HQ content; and

extracting a low resolution patch from the LQ content.

18 . The method of claim 10 , further comprising:

storing the finetuned DSNN in a buffer; and

signaling the stored DSNN when the stored DSNN is to be used to process subsequent regions or frames or when the stored DSNN needs to be updated.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2022
From: CRICRÌ, FRANCESCO; ZHANG, HONGLEI; MATIAS HANNUKSELA, MISKA; REZAZADEGAN TAVAKOLI, HAMED; HAI LE, NAM; GHAZNAVI YOUVALARI, RAMIN; BARIS AKSU, EMRE
To: NOKIA TECHNOLOGIES OY
Reel/Frame 061853/0068 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2022
From: AHONEN, JUKKA
To: TAMPERE UNIVERSITY FOUNDATION SR
Reel/Frame 061853/0129 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2022
From: TAMPERE UNIVERSITY FOUNDATION SR
To: NOKIA TECHNOLOGIES OY
Reel/Frame 061853/0241 →
Continuity (2)
Provisional Application 63279260 · Nov 15, 2021
Related Publication 20230154054A1 · May 18, 2023
References Cited (36)
US 10176388B1 · Ghafarianzadeh · 2019 [cited by examiner]
US 11375204B2 · Zhang · 2022 [cited by examiner]
US 11782942B2 · Liang · 2023 [cited by examiner]
US 20190251360A1 · Cricri · 2019 [cited by examiner]
US 20200034707A1 · Kivatinos et al. · 2020 [cited by applicant]
US 20200160593A1 · Gu et al. · 2020 [cited by applicant]
US 20200311551A1 · Aytekin · 2020 [cited by examiner]
US 20210104076A1 · Zare · 2021 [cited by examiner]
US 20210127140A1 · Hannuksela · 2021 [cited by examiner]
US 20210192323A1 · Whatmough · 2021 [cited by examiner]
US 20210357739A1 · Ramesh · 2021 [cited by examiner]
US 20220148242A1 · Russell · 2022 [cited by examiner]
US 20230154054A1 · Cricrì · 2023 [cited by examiner]
US 20240146938A1 · Zou · 2024 [cited by examiner]
EP 3672241A1 · 2020 [cited by examiner]
WO WO2019197715A1 · 2019 [cited by examiner]
WO WO2020008104A1 · 2020 [cited by examiner]
WO WO2020183059A1 · 2020 [cited by examiner]
WO WO2021165569A1 · 2021 [cited by examiner]
WO WO2021170901A1 · 2021 [cited by examiner]
WO WO2021205066A1 · 2021 [cited by examiner]
“Video Coding For Low Bit Rate Communication”, Series H: Audiovisual And Multimedia Systems, Infrastructure of audiovisual services—Coding of moving Video, ITU-T Recommendation H.263, Jan. 2005, 226 pages. [cited by applicant]
“Advanced Video Coding For Generic Audiovisual services”, Series H: Audiovisual And Multimedia Systems, Infrastructure of audiovisual services—Coding of moving Video, Recommendation ITU-T H.264, Apr. 2017, 812 pages. [cited by applicant]
“High Efficiency Video Coding”, Series H: Audiovisual And Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Recommendation ITU-T H.265, Feb. 2018, 692 pages. [cited by applicant]
“Versatile Video Coding”, Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video, Recommendation ITU-T H.266, Aug. 2020, 516 pages. [cited by applicant]
“Versatile supplemental enhancement information messages for coded video bitstreams”, Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video, Recommendation ITU-T H.27… [cited by applicant]
Wang et al., “Foreground Detection with Deeply Learned Multi-Scale Spatial-Temporal Features”, Sensors, vol. 18, No. 12, 2018, pp. 1-16. [cited by applicant]
Song et al., “VR-DANN: Real-Time Video Recognition via Decoder-Assisted Neural Network Acceleration”, 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), Oct. 17-21, 2020, pp. 698-710. [cited by applicant]
“Information Technology—Generic Coding of Moving Pictures and Associated Audio Information: Systems”, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Transmission Multiplexing and Sy… [cited by applicant]
“Information technology—Generic coding of moving pictures and associated audio information: Video”, Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video, ITU-T Recom… [cited by applicant]
“Information technology—Universal coded character set (UCS)”, ISO/IEC 10646, Sixth edition, Dec. 2020, 9 pages. [cited by applicant]
“IEEE 802.11”, Wikipedia, Retrieved on Nov. 28, 2022, Webpage available at : https://en.wikipedia.org/wiki/IEEE_802.11. [cited by applicant]
“Information Technology—Coding Of Audio-Visual Objects—Part 12: ISO Base Media File Format”, ISO/IEC 14496-12, Fifth edition, Dec. 15, 2015, 248 pages. [cited by applicant]
“Information Technology—Coding Of Audio-Visual Objects—Part 15: Advanced Video Coding (AVC) File Format”, ISO/IEC 14496-15, First edition, Apr. 15, 2004, 29 pages. [cited by applicant]
Extended European Search Report received for corresponding European Patent Application No. 22205523.8, dated Jun. 23, 2023, 11 pages. [cited by applicant]
Partial European Search Report received for corresponding European Patent Application No. 22205523.8, dated Mar. 1, 23, 2023, 11 pages. [cited by applicant]