IP Library › Granted Patent US 11,818,399
Granted Patent B2
US 11,818,399 · App. 17/476,824 · Granted Nov 14, 2023

Techniques for signaling neural network topology and parameters in the coded video stream

Inventors: Byeongdoo Choi (Palo Alto, CA); Zeqiang Li (Palo Alto, CA); Wei Jiang (Sunnyvale, CA); Wei Wang (Palo Alto, CA); Xiaozhong Xu (State College, PA); Stephan Wenger (Hillsborough, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
H04N19/70G06N3/04H04N19/184H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,818,399
App. No.
17/476,824
Granted
Nov 14, 2023
Kind
B2
Abstract

There is included a method and apparatus comprising computer code configured to cause a processor or processors to perform obtaining a video bitstream, coding the video bitstream at least partly by a neural network, determining topology information and parameters of the neural network, signaling the determined topology information and the parameters of the neural network in a plurality of syntax elements associated with the coded video bitstream.

Claims (87)

1. A method for video coding performed by at least one processor, the method comprising:

obtaining a video bitstream;

coding the video bitstream at least partly by a neural network;

determining topology information and parameters of the neural network; and

signaling the determined topology information and the parameters of the neural network in a plurality of syntax elements associated with the coded video bitstream,

wherein at least one of the plurality of syntax elements signals whether the topology information and parameters are all included in a single supplemental enhancement information (SEI) message or are partitioned into multiple SEI messages.

2. The method according to claim 1 ,

wherein the plurality of syntax elements are signaled via one or more of a parameter set, a metadata container box, and at least one of the single SEI message and the multiple SEI messages.

3. The method according to claim 2 ,

wherein the neural network comprises a plurality of operation nodes,

wherein coding the video bitstream by the neural network comprises:

feeding input tensor data of the video bitstream into a first operation node of the operation nodes;

processing the input tensor data with any of pre-trained constants and variables; and

outputting intermediate tensor data,

wherein the intermediate tensor data comprises a weighted summation of the input tensor data and any of trained constants and updated variables.

4. The method according to claim 3 ,

wherein the topology information and parameters are based on said coding of the video bitstream by the neural network.

5. The method according to claim 1 ,

wherein signaling the determined topology information and the parameters comprises providing external linkage information at which the determined topology information and the parameters are stored.

6. The method according to claim 5 ,

wherein the neural network comprises a plurality of operation nodes,

wherein coding the video bitstream by the neural network comprises:

feeding input tensor data of the video bitstream into a first operation node of the operation nodes;

processing the input tensor data with any of pre-trained constants and variables; and

outputting intermediate tensor data, and

wherein the intermediate tensor data comprises a weighted summation of the input tensor data and any of trained constants and updated variables.

7. The method according to claim 6 ,

wherein the topology information and parameters are based on said coding of the video bitstream by the neural network.

8. The method according to claim 1 ,

wherein signaling the determined topology information and the parameters comprises explicitly signaling the determined topology information and the parameters by at least one of a neural network exchange format (NNEF), an open neural network exchange (ONNX) format, and an MPEG neural network compression standard (NNR) format.

9. The method according to claim 8 ,

wherein the neural network comprises a plurality of operation nodes,

wherein coding the video bitstream by the neural network comprises:

feeding input tensor data of the video bitstream into a first operation node of the operation nodes;

processing the input tensor data with any of pre-trained constants and variables; and

outputting intermediate tensor data,

wherein the intermediate tensor data comprises a weighted summation of the input tensor data and any of trained constants and updated variables, and

wherein the topology information and parameters are based on said coding of the video bitstream by the neural network.

10. The method according to claim 9 ,

wherein the at least one of the neural network exchange format (NNEF), the open neural network exchange (ONNX) format, and the MPEG neural network compression standard (NNR) format is the MPEG NNR format, and

wherein at least one of the parameters are compressed into a data file and at least one of the single SEI message and the multiple SEI messages.

11. An apparatus for video coding, the apparatus comprising:

at least one memory configured to store computer program code;

at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code including:

obtaining code configured to cause the at least one processor to obtain a video bitstream;

coding code configured to cause the at least one processor to code the video bitstream at least partly by a neural network;

determining code configured to cause the at least one processor to determine topology information and parameters of the neural network; and

signaling code configured to cause the at least one processor to signal the determined topology information and the parameters of the neural network in a plurality of syntax elements associated with the coded video bitstream,

wherein at least one of the plurality of syntax elements signals whether the topology information and parameters are all included in a single supplemental enhancement information (SEI) message or are partitioned into multiple SEI messages.

12. The apparatus according to claim 11 ,

wherein the plurality of syntax elements are signaled via one or more of a parameter set, a meta-data container box, and any of the SEI message and the multiple SEI messages.

13. The apparatus according to claim 12 ,

wherein the neural network comprises a plurality of operation nodes,

wherein coding the video bitstream by the neural network comprises:

feeding input tensor data of the video bitstream into a first operation node of the operation nodes;

processing the input tensor data with any of pre-trained constants and variables; and

outputting intermediate tensor data, and

wherein the intermediate tensor data comprises a weighted summation of the input tensor data and any of trained constants and updated variables.

14. The apparatus according to claim 13 ,

wherein the topology information and parameters are based on said coding of the video bitstream by the neural network.

15. The apparatus according to claim 11 ,

wherein signaling the determined topology information and the parameters comprises providing external linkage information at which the topology information and the parameters are stored.

16. The apparatus according to claim 15 ,

wherein the neural network comprises a plurality of operation nodes,

wherein coding the video bitstream by the neural network comprises:

feeding input tensor data of the video bitstream into a first operation node of the operation nodes;

processing the input tensor data with any of pre-trained constants and variables; and

outputting intermediate tensor data, and

wherein the intermediate tensor data comprises a weighted summation of the input tensor data and any of trained constants and updated variables.

17. The apparatus according to claim 16 ,

wherein the topology information and parameters are based on said coding of the video bitstream by the neural network.

18. The apparatus according to claim 11 ,

wherein signaling the determined topology information and the parameters comprises explicitly signaling the determined topology information and the parameters by at least one of a neural network exchange format (NNEF), an open neural network exchange (ONNX) format, and an MPEG neural network compression standard (NNR) format.

19. The apparatus according to claim 17 ,

wherein the neural network comprises a plurality of operation nodes,

wherein coding the video bitstream by the neural network comprises:

feeding input tensor data of the video bitstream into a first operation node of the operation nodes;

processing the input tensor data with any of pre-trained constants and variables; and

outputting intermediate tensor data,

wherein the intermediate tensor data comprises a weighted summation of the input tensor data and any of trained constants and updated variables, and

wherein the topology information and parameters are based on said coding of the video bitstream by the neural network.

20. A non-transitory computer readable medium storing a program causing a computer to execute a process, the process comprising:

obtaining a video bistream;

coding the video bitstream at least partly by a neural network;

determining topology information and parameters of the neural network; and

signaling the determined topology information and the parameters of the neural network in a plurality of syntax elements associated with the coded video bitstream,

wherein at least one of the plurality of syntax elements signals whether the topology information and parameters are all included in a single supplemental enhancement information (SEI) message or are partitioned into multiple SEI messages.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2021
From: CHOI, BYEONGDOO; LI, ZEQIANG; JIANG, WEI; WANG, WEI; XU, XIAOZHONG; WENGER, STEPHAN; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 057525/0481 →
Continuity (2)
Provisional Application 63133682 · Jan 4, 2021
Related Publication 20220217403A1 · Jul 7, 2022