IP Library Granted Patent US 12,445,633
Granted Patent B2
US 12,445,633 · App. 18/340,704 · Granted Oct 14, 2025

Method and apparatus for decoding with signaling of feature map data

Inventors: Sergey Yurievich Ikonin (Moscow, RU); Mikhail Vyacheslavovich Sosulnikov (Munich, DE); Alexander Alexandrovich Karabutov (Munich, DE); Timofey Mikhailovich Solovyev (Munich, DE); Biao Wang (Shenzhen, CN); Elena Alexandrovna Alshina (Munich, DE)
Assignee: Huawei Technologies Co., Ltd.
H04N19/33H04N19/139H04N19/59H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,633
App. No.
18/340,704
Granted
Oct 14, 2025
Kind
B2
Abstract

A method and apparatus for decoding data for still or video processing into a bitstream are provided. In particular, two or more sets of feature map elements are obtained from the bitstream. Each set of feature map elements relates to a feature map. Each of the two or more sets of feature map elements is then respectively inputted into two or more feature map processing layers out of a plurality of cascaded layers. The decoded data for picture or video processing is then obtained as a result of the processing by the plurality of cascaded layers. According to the present disclosure, the data may be decoded from the bitstream in an efficient manner in the layered structure.

Claims (75)

1. A method, applied to a decoding device, for decoding data for picture or video processing from a bitstream, the method comprising:

obtaining, from the bitstream, segmentation information relating two or more feature map processing layers of a plurality of cascaded layers and a part of a feature map processed by the two or more feature map processing layers;

obtaining, from the bitstream, two or more sets of feature map elements,

wherein the two or more sets of feature map elements relate to two or more feature maps, respectively, such that each set of feature map elements relates to one of the two or more feature maps, and

wherein obtaining the feature map elements from the bitstream is based on the segmentation information:

simultaneously inputting, based on the segmentation information, each of the two or more sets of feature map elements into a different one of two or more feature map processing layers out of the plurality of cascaded layers; and

obtaining decoded data for picture or video processing as a result of processing each of the feature maps in two or more feature map processing layers of the plurality of cascaded layers,

wherein each of the feature maps processed in two or more feature map processing layers of the plurality of cascaded layers has a resolution different from each of the other two or more feature maps processed in two or more feature map processing layers of the plurality of cascaded layers, and

wherein the processing of each of the feature maps in two or more feature map processing layers of the plurality of cascaded layers includes upsampling.

2. The method according to claim 1 , wherein the plurality of cascaded layers further comprises a plurality of segmentation information processing layers, and the method further comprises:

processing the segmentation information in the plurality of segmentation information processing layers.

3. The method according to claim 2 , wherein processing the segmentation information in at least one of the plurality of segmentation information processing layers includes upsampling.

4. The method according to claim 3 , wherein upsampling the segmentation information and/or upsampling the feature map comprises a nearest neighbor upsampling.

5. The method according to claim 3 , wherein upsampling the segmentation information and/or upsampling the feature map comprises a transposed convolution.

6. The method according to claim 2 , wherein obtaining the feature map elements from the bitstream is based on processed segmentation information processed by at least one of the plurality of segmentation information processing layers.

7. The method according to claim 2 , wherein inputting each of the two or more sets of feature map elements into the two or more feature map processing layers is based on processed segmentation information processed by at least one of the plurality of segmentation information processing layers.

8. The method according to claim 1 , wherein the obtained segmentation information is represented by a set of syntax elements,

wherein a position of an element in the set of syntax elements indicates a position of the feature map element to which the syntax element relates,

wherein processing the feature map for each of the syntax elements further comprises:

based on the syntax element having a first value, parsing from the bitstream an element of the feature map on the position indicated by the position of the syntax element within the bitstream, and

based on the syntax element not having the first value, bypassing parsing from the bitstream the element of the feature map on the position indicated by the position of the syntax element within the bitstream.

9. The method according to claim 8 , wherein based on the syntax element having a first value,

parsing from the bitstream an element of the feature map; and

bypassing parsing from the bitstream the element of the feature map, based on the syntax element having a second value or based on segmentation information processed by a preceding segmentation information processing layer having a first value.

10. The method according to claim 8 , wherein the syntax element parsed from the bitstream representing the segmentation information is a binary flag.

11. The method according to claim 10 , wherein the processed segmentation information is represented by a set of binary flags.

12. The method according to claim 2 , wherein processing the feature map by each layer j of the two or more N feature map processing layers further comprises:

parsing segmentation information elements for the j-th feature map processing layer from the bitstream;

obtaining the feature map processed by a preceding feature map processing layer; and

parsing, from the bitstream, a feature map element and associating the parsed feature map element with the obtained feature map,

wherein the position of the feature map element in the processed feature map is indicated by the parsed segmentation information element, and segmentation information processed by a preceding segmentation information processing layer, and

wherein N and j are positive integers and 1<j<N.

13. The method according to claim 12 , wherein upsampling the segmentation information in each segmentation information processing layer j further comprises:

for each p-th position in the obtained feature map that is indicated by the inputted segmentation information; and

determining as upsampled segmentation information, indications for feature map positions included in the same area in a reconstructed picture as the p-th position,

wherein p is a positive integer.

14. The method according to claim 1 , wherein the data for picture or video processing comprise picture data, prediction residual data and/or prediction information data.

15. The method according to claim 1 , wherein the data for picture or video processing comprises a motion vector field.

16. The method according to claim 1 , wherein a filter is used in the upsampling of the feature map, and the shape of the filter is any one of square, horizontal rectangular and vertical rectangular.

17. The method according to claim 1 , wherein a filter is used in the upsampling of the feature map, and obtaining the segmentation information from the bitstream further comprises:

obtaining information indicating the filter shape and/or filter coefficients from the bitstream.

18. The method according to claim 17 , wherein the obtained information indicating the filter shape indicates a mask comprised of flags, and the mask represents the filter shape in that a flag having a third value indicates a non-zero filter coefficient and the flag having a fourth value different from the third value indicates a zero filter coefficient.

19. The method according to claim 1 , wherein the plurality of cascaded layers comprises convolutional layers without upsampling between layers with different resolutions.

20. A device for decoding an image or video including a processing circuitry which is configured to perform the method according to claim 1 .

21. A non-transitory computer readable medium, having processor-executable instructions stored thereon, which when executed on one or more processors, cause the one or more processors to perform a method including:

obtaining, from the bitstream, segmentation information relating two or more feature map processing layers of a plurality of cascaded layers and a part of a feature map processed by the two or more feature map processing layers;

obtaining, from the bitstream, two or more sets of feature map elements,

wherein the two or more sets of feature map elements relate to two or more feature maps, respectively, such that each set of feature map elements relates to one of the two or more feature maps, and

wherein obtaining the feature map elements from the bitstream is based on the segmentation information:

simultaneously inputting, based on the segmentation information, each of the two or more sets of feature map elements into a different one of two or more feature map processing layers out of the plurality of cascaded layers; and

obtaining decoded data for picture or video processing as a result of processing each of the feature maps in two or more feature map processing layers of the plurality of cascaded layers,

wherein each of the feature maps processed in two or more feature map processing layers of the plurality of cascaded layers has a resolution different from each of the other two or more feature maps processed in two or more feature map processing layers of the plurality of cascaded layers, and

wherein the processing of each of the feature maps in two or more feature map processing layers of the plurality of cascaded layers includes upsampling.

22. The non-transitory computer readable medium of claim 21 , wherein the plurality of cascaded layers further comprises a plurality of segmentation information processing layers, and the method further comprises processing the segmentation information in the plurality of segmentation information processing layers.

23. The non-transitory computer readable medium of claim 21 , wherein the obtained segmentation information is represented by a set of syntax elements,

wherein a position of an element in the set of syntax elements indicates a position of the feature map element to which the syntax element relates,

wherein processing the feature map for each of the syntax elements further comprises:

based on the syntax element having a first value, parsing from the bitstream an element of the feature map on the position indicated by the position of the syntax element within the bitstream, and

based on the syntax element not having the first value, bypassing parsing from the bitstream the element of the feature map on the position indicated by the position of the syntax element within the bitstream.

24. A device for decoding data for picture or video processing from a bitstream, the device comprising:

a memory having processor-executable instructions stored thereon; and

a processor connected to the memory, wherein the processor is configured to execute the processor-executable instructions to facilitate the following:

obtaining, from the bitstream, segmentation information relating two or more feature map processing layers of a plurality of cascaded layers and a part of a feature map processed by the two or more feature map processing layers;

obtaining from the bitstream two or more sets of feature map elements,

wherein the two or more sets of feature map elements relate to two or more feature maps, respectively, such that each set of feature map elements relates to one of the two or more feature maps, and

wherein obtaining the feature map elements from the bitstream is based on the segmentation information;

simultaneously inputting, based on the segmentation information, each of the two or more sets of feature map elements into a different one of two or more feature map processing layers out of the plurality of cascaded layers; and

obtaining decoded data for picture or video processing as a result of processing each of the feature maps in two or more feature map processing layers of the plurality of cascaded layers,

wherein each of the feature maps processed in two or more feature map processing layers of the plurality of cascaded layers has a resolution different from each of the other two or more feature maps processed in two or more feature map processing layers of the plurality of cascaded layers, and

wherein the processing of each of the feature maps in two or more feature map processing layers of the plurality of cascaded layers includes upsampling.

25. The device for decoding data for picture or video processing from the bitstream according to claim 24 , wherein the obtained segmentation information is represented by a set of syntax elements,

wherein a position of an element in the set of syntax elements indicates a position of the feature map element to which the syntax element relates,

wherein processing the feature map for each of the syntax elements further comprises:

based on the syntax element having a first value, parsing from the bitstream an element of the feature map on the position indicated by the position of the syntax element within the bitstream, and

based on the syntax element not having the first value, bypassing parsing from the bitstream the element of the feature map on the position indicated by the position of the syntax element within the bitstream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2025
From: IKONIN, SERGEY YURIEVICH; SOSULNIKOV, MIKHAIL VYACHESLAVOVICH; KARABUTOV, ALEXANDER ALEXANDROVICH; SOLOVYEV, TIMOFEY MIKHAILOVICH; ALSHINA, ELENA ALEXANDROVNA; WANG, BIAO
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 071825/0255 →
Continuity (2)
Continuation PCTRU2020000748 · Dec 24, 2020
Related Publication 20230353764A1 · Nov 2, 2023
References Cited (62)
US 20170324962A1 · Karczewicz · 2017 [cited by examiner]
US 20180089834A1 · Spizhevoy · 2018 [cited by examiner]
US 20180350110A1 · Cho · 2018 [cited by examiner]
US 20190042923A1 · Janedula · 2019 [cited by examiner]
US 20190373293A1 · Bortman · 2019 [cited by examiner]
US 20200092552A1 · Coelho et al. · 2020 [cited by applicant]
US 20200143457A1 · Manmatha · 2020 [cited by examiner]
US 20200210844A1 · Gao et al. · 2020 [cited by applicant]
US 20210366123A1 · Wang · 2021 [cited by examiner]
US 20210409783A1 · Wan · 2021 [cited by examiner]
US 20230396801A1 · Racape · 2023 [cited by examiner]
CN 110300301A · 2019 [cited by applicant]
CN 111368699A · 2020 [cited by applicant]
EP 3624016A1 · 2020 [cited by applicant]
WO 2020106871A1 · 2020 [cited by applicant]
WO 2020113355A1 · 2020 [cited by applicant]
Hu et al., “Improving Deep Video Compression by Resolution-Adaptive Flow Coding,” Topics in Cryptology—CT-RSA 2020, The Cryptographers Track at the RSA Conference 2020, San Francisco, CA, USA, Cornell University Library… [cited by applicant]
C. E. Shannon, “A Mathematical Theory of Communication,” Reprinted with corrections from The Bell System Technical Journal, vol. 27, pp. 379-423, 623-656, Total 55 pages (Jul., Oct. 1948). [cited by applicant]
Akbari et al., “DSSLIC: Deep Semantic Segmentation-based Layered Image Compression,” Arxiv.org, arXiv: 1806.03348v [cs.CV], Cornell University Library, 201 Online Library Cornell University Ithaca, NY 14853, XP081200536… [cited by applicant]
Gersho et al., “Vector Quantization and Signal Compression,” The Springer International Series in Engineering and Computer Science, Springer Link, Total 7 pages (Nov. 1991). With the English Abstract. [cited by applicant]
Wintz, “Transform Picture Coding,” Proceedings of the IEEE, vol. 60, No. 7, Total 18 pages, Institute of Electrical Electronics Engineers, New York, New York (Jul. 1972). [cited by applicant]
Netravali et al., “Picture coding: a Review,” Proceedings of the IEEE vol. 68, No. 3, Total 48 pages, Institute of Electrical Electronics Engineers, New York, New York (Mar. 1980). [cited by applicant]
Balle et al., “End-to-End Optimized Image Compression,” Published as a conference paper at ICLR 2017, Presented at: Int'l Conf on Learning Representations, Toulon, France, Total 27 pages (Apr. 2017). [cited by applicant]
Balle et al., “Density Modeling of Images using a Generalized Normalization Transformation,” Published as a conference paper at ICLR 2016, arXiv:1511.06281v4 [cs.LG], Total 15 pages (Feb. 2016). [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” ICLR 2014 conference submission, arXiv:1312.6114v11 [stat.ML], Total 14 pages (Dec. 2013). [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” arXiv:1401.4082v3 [stat.ML], Total 8 pages (May 2014). [cited by applicant]
Wiegand et al., “Overview of the H.264/AVC Video Coding Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, Total 17 pages, Institute of Electrical Electronics Engineers, New York,… [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Total 20 pages, Institute of Electrical Electronics Engin… [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 7),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WVG 11, 16th Meeting: Geneva, CH, Document: JVET-P2001-vE, Total 492 pages (Oct. 1-11, 2019). [cited by applicant]
Choi et al., “Near-Lossless Deep Feature Compression for Collaborative Intelligence,” 2018 IEEE 20th International Workshop on Multimedia Signal Processing (MMSP), Total 6 pages, Institute of Electrical Electronics Engi… [cited by applicant]
Kang et al., “Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,” ASPLOS '17, Xi'an, China, Total 15 pages (Apr. 2017). [cited by applicant]
Eshratifar et al., “JointDNN: an Efficient Training and Inference Engine for Intelligent Mobile Cloud Computing Services,” arXiv:1801.08618v1 [cs.DC], Total 13 pages (Jan. 2018). [cited by applicant]
Choi et al., “Deep Feature Compression for Collaborative Object Detection,” ICIP 2018, Total 5 pages (Feb. 2018). [cited by applicant]
Redmon et al., “YOLO9000: Better, Faster, Stronger,” 2017 IEEE Conference on Computer Vision and Pattern Recognition, Total 9 pages, Institute of Electrical Electronics Engineers, New York, New York (Jul. 2017). [cited by applicant]
Bossen et al., “HEVC Complexity and Implementation Analysis,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Total 12 pages, Institute of Electrical Electronics Engineers, New York, New… [cited by applicant]
Luo et al., “DeepSIC: Deep Semantic Image Compression,” arXiv:1801.09468v1 [cs.CV], Total 8 pages (Jan. 2018). [cited by applicant]
Marpe et al., “Context-Based Adaptive Binary Arithmetic Coding in the H.264/AVC Video Compression Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, Total 17 pages, Institute of E… [cited by applicant]
“Cisco Global Cloud Index: Forecast and Methodology, 2016-2021,” White Paper, CISCO, Total 46 pages (2018). [cited by applicant]
Wu et al., “Compressed Video Action Recognition,” arXiv:1712.00636v2 [cs.CV], Total 14 pages (Mar. 2018). [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding,” Under review as a conference paper at ICLR 2016, arXiv:1510.00149v3 [cs.CV], Total 13 pages (Nov. 2… [cited by applicant]
Toderici et al., “Variable Rate Image Compression With Recurrent Neural Networks,” Under review as a conference paper at ICLR 2016, arXiv:1511.066085v2 [cs.CV], Total 9 pages (Nov. 2015). [cited by applicant]
Toderici et al., “Full Resolution Image Compression with Recurrent Neural Networks,” CVPR 2017, Computer Vision Foundation, Total 9 pages (Aug. 2016). [cited by applicant]
Agustsson et al., “Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, Total 11 pages (Apr. … [cited by applicant]
Balle et al., “Variational Image Compression With a Scale Hyperprior,” Published as a conference paper at ICLR 2018, arXiv:1802.01436v2 [eess..IV], Total 23 pages (May 2018). [cited by applicant]
Johnston et al., “Improved Lossy Image Compression with Priming and Spatially Adaptive Bit Rates for Recurrent Networks,” CVPR 2018, Computer Vision Foundation, Total 9 pages (Mar. 2017). [cited by applicant]
Theis et al., “Lossy Image Compression With Compressive Autoencoders,” Published as a conference paper at ICLR 2017, arXiv:1703.00395v1 [stat.ML], Total 19 pages (Mar. 2017). [cited by applicant]
Li et al., “Learning Convolutional Networks for Content-weighted Image Compression,” CVPR 2018, Computer Vision Foundation, Total 10 pages (Mar. 2017). [cited by applicant]
Rippel et al., “Real-Time Adaptive Image Compression,” Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, Total 9 pages (May 2017). [cited by applicant]
Agustsson et al., “Generative Adversarial Networks for Extreme Learned Image Compression,” arXiv:1804.02958v2 [cs.CV], Total 27 pages (Oct. 2018). [cited by applicant]
Wallace, “The JPEG Still Picture Compression Standard,” IEEE Transactions on Consumer Electronics, vol. 38, No. 1, Total 17 pages (Feb. 1992). [cited by applicant]
Skodras et al., “The JPEG 2000 Still Image Compression Standard,” IEEE Signal Processing Magazine, Total 23 pages (Sep. 2001). [cited by applicant]
“BPG Image format,” Bellard, org, Release 0.9.8 is available, Total 2 pages (Apr. 21, 2018). [cited by applicant]
Xue et al., “Video Enhancement with Task-Oriented Flow,” Intervational Journal of Computer Vision, arXiv:1711.09078v3 [cs.CV], Total 20 pages (Nov. 2019). [cited by applicant]
Lu et al., “DVC: an End-to-end Deep Video Compression Framework,” CVPR 2019, Computer Vision Foundation, Total 10 pages (Nov. 2018). [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 10),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 19th Meeting: by teleconference, Document: JVET-S2001-vH, Total 548 pages, Internatio… [cited by applicant]
Choi et al., “Text of ISO/IEC CD 23094-1, Essential Video Coding,” ISO/IEC JTC1/SC29/WG11 N18568, Gothenburg, Sweden, Total 292 pages (Jul. 2019). [cited by applicant]
He et al., “Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition,” arXiv:1406.4729v4 [cs.CV], Total 14 pages (Apr. 2015). [cited by applicant]
“Autoencoder,” Wikipedia, The Free Encyclopedia, Total 15 pages (Sep. 2006). [cited by applicant]
“Convolutional neural network,” Wikipedia, The Free Encyclopedia, Total 39 pages (Aug. 31, 2013). [cited by applicant]
Wu et al., “End-to-end Optimized Video Compression with MV-Residual Prediction,” CVPR May 2020, Computer Vision Foundation, Total 4 pages (May 2020). [cited by applicant]
“Quadtree,” Wikipedia, The Free Encyclopedia, Total 11 pages (Apr. 5, 2004). [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, High efficiency video coding,” Telecommunication Standardization Sector of ITU, ITU-T H.265, Total 712 pages,… [cited by applicant]