IP Library Granted Patent US 12,323,607
Granted Patent B2
US 12,323,607 · App. 18/001,987 · Granted Jun 3, 2025

Apparatus, method and computer program product for optimizing parameters of a compressed representation of a neural network

Inventors: Francesco Cricrì (Tampere, FI); Hamed Rezazadegan Tavakoli (Espoo, FI); Honglei Zhang (Tampere, FI); Nannan Zou (Tampere, FI)
Assignee: Nokia Technologies Oy
H04N19/42G06N3/0455G06N3/0985H04N19/124H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,607
App. No.
18/001,987
Granted
Jun 3, 2025
Kind
B2
Abstract

In example embodiments, an apparatus, a method, and a computer program product are provided. An example apparatus include processing circuitry; and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the processing circuitry, cause the apparatus at least to: overfit a neural network on each media item, from a batch of media items, for a number of iterations to obtain an overfitted neural network model for the each media item; evaluate the overfitted neural network model on the each media item to obtain evaluation errors; and update parameters of the neural network to be based on the evaluation errors.

Claims (53)

1. An apparatus comprising:

at least one processor; and

at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:

overfit a neural network on each media item, from a batch of media items, for a number of iterations to obtain an individual overfitted neural network model for each media item in the batch of media items, wherein to overfit the neural network, the following are performed:

selecting overfitting parameters;

performing a forward operation on the neural network;

computing loss based on the forward operation;

computing derivatives of a loss based on the selected overfitting parameters; and

updating the selected overfitting parameters based on the computed derivatives;

evaluate individual overfitted neural network models on a corresponding media item to obtain evaluation errors; and

update parameters of the neural network based on the evaluation errors.

2. The apparatus of claim 1 , wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus at least to perform a training iteration on the neural network based on the evaluation errors.

3. The apparatus of claim 2 , wherein the training iteration comprises two stages, when overfitting parameters comprise a latent tensor output by an encoder neural network, or the latent tensor output by a quantizer, and wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus at least to perform the two stages.

4. The apparatus of claim 1 , wherein the overfitting parameters comprise at least one of the following: encoder neural network parameters, output of an encoder neural network, output of a quantizer, or at least one decoder neural network parameter.

5. The apparatus of claim 4 , wherein the encoder neural network parameters comprise parameters of convolution layers, parameters of fully-connected layers, or parameters of a batch-normalization layer.

6. The apparatus of claim 4 , wherein the at least one decoder neural network parameter comprises at least one of one or more bias terms of convolution layers, bias terms of fully-connected layers in a decoder neural network, attention-multipliers in an instance an attention mechanism is used in an architecture of the decoder neural network, or a vector for deriving one or more bias terms by multiplying the vector by a weight matrix at a decoder side.

7. The apparatus of claim 4 , wherein the at least one decoder neural network parameter comprises at least one of one or more of a set of scalar values that act as gates or attention-multipliers in specifically designed blocks used a decoder neural network.

8. The apparatus of claim 4 , wherein to perform the overfitting when the overfitting parameters comprise the encoder neural network parameters, the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus at least to:

input the media item to an encoding pipeline to obtain a latent tensor;

input the latent tensor to a probability neural network and to a decoding pipeline;

obtain a loss comprising a rate loss and a reconstruction loss;

compute a derivative of the loss based on the selected overfitting parameters in the encoder neural network;

update the selected overfitting parameters based on derivatives of the loss; and

repeat preceding for updated overfitting parameters for a predefined number of iterations.

9. The apparatus of claim 8 , wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus at least to:

compute an evaluation loss comprising rate losses and reconstruction losses based on the updated overfitting parameters;

compute gradients of the evaluation loss based on neural network parameters; and

use the computed gradients to update the neural network parameters.

10. The apparatus of claim 1 , wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus at least to select the batch of media items as a subset of training media items.

11. The apparatus of claim 1 , wherein a loss comprises a rate loss and reconstruction loss.

12. An apparatus comprising:

at least one processor; and

at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:

overfit a neural network on each media item, from a batch of media items, for a number of iterations to obtain an individual overfitted neural network model for each media item in the batch of media items;

evaluate individual overfitted neural network models on a corresponding media item to obtain evaluation errors;

update parameters of the neural network based on the evaluation errors;

perform a training iteration on the neural network based on the evaluation errors;

wherein the training iteration comprises two stages, when overfitting parameters comprise a latent tensor output by an encoder neural network, or the latent tensor output by a quantizer, and wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to perform the two stages at least by:

running the encoder neural network or the quantizer to obtain an initial tensor; and

performing a meta-learning by overfitting an initial latent tensor and update the neural network based on performance of latent tensor overfitting.

13. A method comprising:

overfitting a neural network on each media item, from a batch of media items, for a number of iterations to obtain an individual overfitted neural network model for each media item in the batch of media items, wherein to overfit the neural network, the following are performed:

selecting overfitting parameters;

performing a forward operation on the neural network;

computing loss based on the forward operation;

computing derivatives of a loss based on the selected overfitting parameters; and

updating the selected overfitting parameters based on the computed derivatives;

evaluating individual overfitted neural network models on a corresponding media item to obtain evaluation errors; and

updating parameters of the neural network based on the evaluation errors.

14. The method of claim 13 , further comprising performing a training iteration on the neural network based on the evaluation errors.

15. The method of claim 13 , wherein the overfitting parameters comprise at least one of the following: encoder neural network parameters, output of an encoder neural network, output of a quantizer, or at least one decoder neural network parameters.

16. The method of claim 15 , wherein the encoder neural network parameters comprise parameters of convolution layers, parameters of fully-connected layers, or parameters of a batch-normalization layer.

17. The method of claim 15 , wherein the at least one decoder neural network parameters comprise at least one of one or more bias terms of convolution layers, bias terms of fully-connected layers in a decoder neural network, attention-multipliers in an instance an attention mechanism is used in an architecture of the decoder neural network, or a vector for deriving one or more bias terms by multiplying the vector by a weight matrix at a decoder side.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 4, 2024
From: REZAZADEGAN TAVAKOLI, HAMED; CRICRÌ, FRANCESCO; ZHANG, HONGLEI
To: NOKIA TECHNOLOGIES OY
Reel/Frame 067912/0690 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 4, 2024
From: ZOU, NANNAN
To: TAMPERE UNIVERSITY FOUNDATION SR
Reel/Frame 067912/0695 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 4, 2024
From: TAMPERE UNIVERSITY FOUNDATION SR
To: NOKIA TECHNOLOGIES OY
Reel/Frame 067912/0704 →
Continuity (2)
Provisional Application 63041286 · Jun 19, 2020
Related Publication 20230269387A1 · Aug 24, 2023
References Cited (25)
US 20190155973A1 · Morczinek · 2019 [cited by examiner]
US 20190311259A1 · Cricri et al. · 2019 [cited by applicant]
US 20190354858A1 · Chrzanowski et al. · 2019 [cited by applicant]
US 20200134461A1 · Chai · 2020 [cited by examiner]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
US 20200210824A1 · Poornaki · 2020 [cited by examiner]
US 20200251183A1 · Kashefhaghighi · 2020 [cited by examiner]
US 20210056412A1 · Jung · 2021 [cited by examiner]
US 20210073612A1 · Vahdat · 2021 [cited by examiner]
US 20220164655A1 · Gomez · 2022 [cited by examiner]
WO 2019197712A1 · 2019 [cited by applicant]
WO 2020008104A1 · 2020 [cited by applicant]
“Video Coding For Low Bit Rate Communication”, Series H: Audiovisual And Multimedia Systems, Infrastructure of audiovisual services—Coding of moving Video, ITU-T Recommendation H.263, Jan. 2005, 226 pages. [cited by applicant]
“Advanced Video Coding For Generic Audiovisual services”, Series H: Audiovisual And Multimedia Systems, Infrastructure of audiovisual services—Coding of moving Video, Recommendation ITU-T H.264, Apr. 2017, 812 pages. [cited by applicant]
“Versatile Video Coding”, Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video, Recommendation ITU-T H.266, Aug. 2020, 516 pages. [cited by applicant]
Finn et al., “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks”, arXiv, Jul. 18, 2017, 13 pages. [cited by applicant]
Nichol et al., “On First-Order Meta-Learning Algorithms”, arXiv, Oct. 22, 2018, pp. 1-15. [cited by applicant]
“Information Technology—Generic Coding of Moving Pictures and Associated Audio Information: Systems”, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Transmission Multiplexing and Sy… [cited by applicant]
“Information technology—Generic coding of moving pictures and associated audio information: Video”, Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video, ITU-T Recom… [cited by applicant]
“Information technology—Universal coded character set (UCS)”, ISO/IEC 10646, Sixth edition, Dec. 2020, 9 pages. [cited by applicant]
“IEEE 802.11”, Wikipedia, Retrieved on Dec. 27, 2022, Webpage available at : https://en.wikipedia.org/wiki/ IEEE 802.11. [cited by applicant]
“Information Technology—Coding Of Audio-Visual Objects—Part 12: ISO Base Media File Format”, ISO/IEC 14496-12, Fifth edition, Dec. 15, 2015, 248 pages. [cited by applicant]
“Information Technology—Coding Of Audio-Visual Objects—Part 15: Advanced Video Coding (AVC) File Format”, ISO/IEC 14496-15, First edition, Apr. 15, 2004, 29 pages. [cited by applicant]
Invitation to Pay Additional Fees received for corresponding Patent Cooperation Treaty Application No. PCT/IB2021/055178, dated Sep. 9, 2021, 14 pages. [cited by applicant]
International Search Report and Written Opinion received for corresponding Patent Cooperation Treaty Application No. PCT/IB2021/055178, dated Nov. 2, 2021, 19 pages. [cited by applicant]