IP Library Granted Patent US 12,400,113
Granted Patent B2
US 12,400,113 · App. 17/438,908 · Granted Aug 26, 2025

Method and apparatus for updating a neural network

Inventors: Christof Fersch (Neumarkt, DE); Arijit Biswas (Schwaig bei Nuernberg, DE)
Assignee: DOLBY INTERNATIONAL AB
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,113
App. No.
17/438,908
Granted
Aug 26, 2025
Kind
B2
Abstract

Described herein is a method of generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, wherein the method includes the steps of: (a) determining at least one set of parameters for updating the neural network; (b) encoding the at least one set of parameters and media data to generate the media bitstream; and (c) transmitting the media bitstream to the decoder for updating the neural network with the at least one set of parameters. Described herein are further a method for updating a neural network implemented in a decoder, an apparatus for generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, an apparatus for updating a neural network implemented in a decoder and computer program products comprising a computer-readable storage medium with instructions adapted to cause the device to carry out said methods when executed by a device having processing capability.

Claims (27)

1. A method of generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, the neural network having a plurality of layers, with a media data facing layer as a first layer of the plurality of layers and an output layer as a last layer of the plurality of layers, wherein the method includes the steps of:

(a) determining at least one set of parameters for updating weights of the plurality of layers of the neural network, including parameters for updating weights of the media data facing layer and/or the output layer;

(b)generating the media bitstream by encoding, of the at least one set of parameters for updating weights of the plurality of layers of the neural network, only the parameters for updating weights of the media data facing layer and/or the output layer, and media data, the media data including one or more of audio data and/or video data; and

(c) transmitting the media bitstream to the decoder for updating the neural network with the parameters that were included in the media bitstream for updating weights of the media data facing layer and/or the output layer.

2. The method according to claim 1 , wherein the at least one set of parameters is encoded based on a set of syntax elements.

3. The method according to claim 2 , wherein in step (a) two or more sets of parameters for updating the neural network are determined, and wherein the set of syntax elements includes one or more syntax elements identifying a respective set of parameters for a respective update of the neural network to be performed.

4. The method according to claim 1 , wherein the neural network implemented in the decoder is used for processing of media data, and wherein, in the media bitstream, the at least one set of parameters for updating the neural network is time-aligned with the media data which are processed by the neural network.

5. The method according to claim 4 , wherein the at least one set of parameters is determined based on one or more of codec modes, a content of the media data and encoding constraints.

6. The method according to claim 5 , wherein the codec modes include one or more of a bitrate, a video and/or audio framerate and a used core codec.

7. The method according to claim 5 , wherein the content of the media data includes one or more of speech, music and applause.

8. The method according to claim 5 , wherein the encoding constraints include one or more of constraints for performance scalability and constraints for adaptive processing.

9. The method according to claim 5 , wherein the at least one set of parameters are included in the media bitstream prior to the media data to be processed by the respective updated neural network.

10. The method according to claim 1 , wherein the media data is of MPEG-H Audio or MPEG-I Audio format and the media bitstream is a packetized media bitstream of MHAS format.

11. The method according to claim 10 , wherein the at least one set of parameters is encoded by encapsulating the at least one set of parameters in one or more MHAS packets of a new MHAS packet type.

12. The method according to claim 1 , wherein the media data is in AC-4, AC-3, EAC-3 format, MPEG-4 or MPEG-D USAC format.

13. The method according to claim 12 , wherein the at least one set of parameters is encoded in the media bitstream as one or more payload elements.

14. The method according to claim 13 , wherein the at least one set of parameters is encoded in the media bitstream as one or more payload elements or one or more data stream elements.

15. The method according to claim 1 , wherein the at least one set of parameters include an identifier identifying whether the parameters for updating weights represent relative values or absolute values.

16. A computer program product comprising a computer-readable storage medium with instructions adapted to cause the device to carry out the method according to claim 1 when executed by a device having processing capability.

17. A method of updating a neural network implemented in a decoder, the neural network having a plurality of layers, with a media data facing layer as a first layer of the plurality of layers and an output layer as a last layer of the plurality of layers, the method including the steps of:

(a) receiving a coded media bitstream including media data and parameters for updating weights of the media data facing layer and/or the output layer of the neural network;

(b) decoding the received media bitstream to obtain the decoded media data and the parameters for updating weights of the media data facing layer and/or the output layer of the neural network; and

(c) updating, by the decoder, the media data facing layer and/or the output layer with the received parameters that were included in the media bitstream for updating weights of the media data facing layer and/or the output layer of the neural network.

18. An apparatus for generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, the neural network having a plurality of layers, with a media data facing layer as a first layer of the plurality of layers and an output layer as a last layer of the plurality of layers, wherein the apparatus includes a processor configured to perform a method including the steps of:

(a) determining at least one set of parameters for updating weights of the plurality of layers of the neural network, including parameters for updating weights of the media data facing layer and/or the output layer;

(b) generating the media bitstream by encoding, of the at least one set of parameters for updating weights of the plurality of layers of the neural network, only the parameters for updating weights of the media data facing layer and/or the output layer, and media data, the media data including one or more of audio data and/or video data; and

(c) transmitting the media bitstream to the decoder for updating the neural network with the parameters that were included in the media bitstream for updating weights of the media data facing layer and/or the output layer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2021
From: FERSCH, CHRISTOF; BISWAS, ARIJIT
To: DOLBY INTERNATIONAL AB
Reel/Frame 058091/0879 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2021
From: FERSCH, CHRISTOF; BISWAS, ARIJIT
To: DOLBY INTERNATIONAL AB
Reel/Frame 057933/0262 →
Priority Claims (1)
EP 19174542 · May 15, 2019 · regional
Continuity (2)
Provisional Application 62818879 · Mar 15, 2019
Related Publication 20220156584A1 · May 19, 2022
References Cited (48)
US 5907822A · Prieto, Jr. · 1999 [cited by applicant]
US 7272556B1 · Aguilar · 2007 [cited by applicant]
US 10089383B1 · MacKay · 2018 [cited by applicant]
US 10210861B1 · Arel · 2019 [cited by applicant]
US 20040203559A1 · Stojanovic · 2004 [cited by applicant]
US 20050025053A1 · Izzat · 2005 [cited by applicant]
US 20110274162A1 · Zhou · 2011 [cited by applicant]
US 20140119551A1 · Bharitkar · 2014 [cited by applicant]
US 20150169774A1 · Budzienski · 2015 [cited by applicant]
US 20170177993A1 · Draelos · 2017 [cited by examiner]
US 20180122403A1 · Koretzky · 2018 [cited by examiner]
US 20180151177A1 · Gemmeke · 2018 [cited by applicant]
US 20180314716A1 · Kim · 2018 [cited by applicant]
US 20180336471A1 · Rezagholizadeh · 2018 [cited by applicant]
US 20180341839A1 · Malak · 2018 [cited by applicant]
US 20180357514A1 · Zisimopoulos · 2018 [cited by applicant]
US 20190045034A1 · Alam · 2019 [cited by applicant]
US 20190050728A1 · Sim · 2019 [cited by applicant]
US 20190065935A1 · Ozcan · 2019 [cited by applicant]
US 20210327445A1 · Biswas · 2021 [cited by applicant]
CN 1625880B · 2010 [cited by applicant]
CN 105142096B · 2018 [cited by applicant]
CN 105868829A · 2021 [cited by applicant]
GB 2553351A · 2016 [cited by applicant]
JP 0232679 · 1990 [cited by applicant]
WO WO2015180866A1 · 2015 [cited by examiner]
WO 2016199330A1 · 2016 [cited by applicant]
WO 2018078213A1 · 2018 [cited by applicant]
WO 2018150083A1 · 2018 [cited by applicant]
WO 2018163011A1 · 2018 [cited by applicant]
A. Bhattacharya; A.G. Parlos; A.F. Atiya, “Prediction of MPEG-coded video source traffic using recurrent neural networks”, 2003, IEEE (Year: 2003). [cited by examiner]
Feng Jiang; Wen Tao; Shaohui Liu; Jie Ren; Xun Guo; Debin Zhao, “An End-to-End Compression Framework Based on Convolutional Neural Networks”, Aug. 2017 (Year: 2017). [cited by examiner]
Yat Hong Lam, Alireza Zare, Caglar Aytekin, Francesco Cricri, “Compressing Weight-updates for Image Artifacts Removal Neural Networks”, 2019 (Year: 2019). [cited by examiner]
Li, Geng et al; DDP: Distributed Network Updates in SDN; 2018; IEEE 38th International Conference on Distributed Computing Systems; pp. 1468-1473. 6 pages. [cited by applicant]
Signal and Information Processing; Zhou Dashan. Professor Li Hua; The Design and Optimization of AVS-M Decoder; Tianjin University School of Electronic Information Engineering; Jan. 2006. p. 1-53. 60 pages. [cited by applicant]
Zeng Tao. Research of Adaptive Media Playout Based on Neural Network Control and Realization of Streaming Player. Aug. 15, 2006. pp. 1-96. 108 pages. [cited by applicant]
Zhu, Yinhao et al.; Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification; Center for Informatics and Computational Science; Journal of Computational Physics; Jan. 23… [cited by applicant]
Choi, K. et al “A Tutorial on Deep Learning for Music Information Retrieval” Sep. 13, 2017, pp. 1-16. [cited by applicant]
Schmidt, W. et al “Feedforward Neural Networks with Random Weights” Pattern Recognition vol. II, 11th IAPR International Conference on the Hague, Netherlands, Aug. 30-Sep. 3, 1992. [cited by applicant]
Wang, Y. et al “Transferring GANs: Generating Images from Limited Data” May 2018, Computer Vision and Pattern Recognition, pp. 1-22. [cited by applicant]
Yang, Li-Chia, et al “MidiNet: A Convolutional Generative Adversarial Network for Symbolic-Domain Music Generation” Mar. 2017, 8 pages, accepted to ISMIR (International Society of Music Information Retrieval). [cited by applicant]
Biswas, A. et al “Audio Codec Enhancement with Generative Adversarial Networks” IEEE International Conference on Acoustics, Speech and Signal Processing, May 2020. [cited by applicant]
ISO/IEC JTC 1/SC 29/WG 6, “Information technology—Coded representation of immersive media—Part 4: MPEG-I immersive audio”, ISO 23090-4:202, 2025, 6 Pages. [cited by applicant]
ETSI TS 103 190-2 V1.2.1, “Digital Audio Compression (AC-4) Standard; Part 2: Immersive and personalized audio”, Reference RTS/JTC-043-2, Feb. 2018, 250 Pages. [cited by applicant]
ETSI TS 103 190-1 V1.3.1, “Digital Audio Compression (AC-4) Standard; Part 1: Channel based coding”, Reference RTS/JTC-043-1, Feb. 2018, 293 Pages. [cited by applicant]
ISO/IEC23003-3, “Information technology—MPEG audio technologies—Part 3: Unified speech and audio coding”, First edition Apr. 1, 2012, Reference No. ISO/IEC 23003-3:2012(E), 286 Pages. [cited by applicant]
ISO/IEC 14496-3, “Information technology—Coding of audio-visual objects—Part 3: Audio”, Fourth edition Sep. 1, 2009, Reference No. ISO/IEC 14496-3:2009(E), 1416 Pages. [cited by applicant]
ISO/IEC 23008-3, “Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio”, Second edition Feb. 2019, Reference No. ISO/IEC 23008-3:2019(E) , 812 Pages. [cited by applicant]