IP Library › Granted Patent US 12,506,891
Granted Patent B2
US 12,506,891 · App. 18/339,783 · Granted Dec 23, 2025

Decoding with signaling of segmentation information

Inventors: Sergey Yurievich Ikonin (Moscow, RU); Mikhail Vyacheslavovich Sosulnikov (Munich, DE); Alexander Alexandrovich Karabutov (Munich, DE); Timofey Mikhailovich Solovyev (Munich, DE); Biao Wang (Shenzhen, CN); Elena Alexandrovna Alshina (Munich, DE)
Assignee: Huawei Technologies Co., Ltd.
H04N19/46G06N3/045G06N3/0455G06N3/088G06T9/002H04N19/105H04N19/107H04N19/117H04N19/119H04N19/124H04N19/132H04N19/33H04N19/82G06N3/047G06N3/048H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,506,891
App. No.
18/339,783
Granted
Dec 23, 2025
Kind
B2
Abstract

The present disclosure relates to methods and apparatuses for decoding data for (still or video processing into a bitstream). Two or more sets of segmentation information elements are obtained from the bitstream. Then, each of the two or more sets of segmentation information elements are inputted respectively into two or more segmentation information processing layers out of a plurality of cascaded layers. In each of the two or more segmentation information processing layers, the respective sets of segmentation information are processed. The decoded data for picture or video processing are obtained based on the segmentation information processed by the plurality of cascaded layers. Accordingly, the data may be decoded from the bitstream in an efficient manner in the layered structure.

Claims (39)

1 . A method for decoding data for picture or video processing from a bitstream, the method comprising:

obtaining, from the bitstream, two or more sets of segmentation information elements, wherein at least one set of segmentation information elements of the two or more sets of segmentation information elements are represented by a set of binary flags;

inputting each of the two or more sets of segmentation information elements respectively into two or more segmentation information processing layers out of a plurality of cascaded layers;

processing, in each of the two or more segmentation information processing layers, the respective sets of segmentation information, wherein the segmentation information processed respectively in the two or more segmentation information processing layers differ in resolution,

wherein the processing of the segmentation information in the two or more segmentation information processing layers includes upsampling,

wherein obtaining the decoded data for picture or video processing is based on the segmentation information processed by the plurality of cascaded layers for motion estimation, and

wherein for each segmentation information processing layer j of the plurality of N segmentation information processing layers out of the plurality of cascaded layers:

the inputting comprises, inputting initial segmentation information from the bitstream if j=1, and otherwise inputting segmentation information processed by the (j−1)-th segmentation information processing layer; and

outputting the processed segmentation information, wherein the processed segmentation information comprises an upsampled set of binary segmentation flags corresponding to the at least one set of segmentation information elements.

2 . The method according to claim 1 , wherein the obtaining of the sets of segmentation information elements is based on segmentation information processed by at least one segmentation information processing layer out of the plurality of cascaded layers.

3 . The method according to claim 1 , wherein the inputting of the sets of segmentation information elements is based on the processed segmentation information outputted by at least one of the plurality of cascaded layers.

4 . The method according to claim 1 , wherein said upsampling of the segmentation information comprises a nearest neighbor upsampling.

5 . The method according to claim 1 , wherein said upsampling of the segmentation information comprises a transposed convolution.

6 . The method according to claim 1 , wherein the processing of the inputted segmentation information by each layer j<N of the plurality of N segmentation information processing layers further comprises:

parsing, from the bitstream, a segmentation information element and associating the parsed segmentation information element with the segmentation information outputted by a preceding layer, wherein the position of the parsed segmentation information element in the associated segmentation information is determined based on the segmentation information outputted by the preceding layer.

7 . The method according to claim 6 , wherein the amount of segmentation information elements parsed from the bitstream is determined based on segmentation information outputted by the preceding layer.

8 . The method according to claim 1 , wherein obtaining decoded data for picture or video processing comprises determining of at least one of:

intra- or inter-picture prediction mode;

picture reference index;

single-reference or multiple-reference prediction (including bi-prediction);

presence or absence prediction residual information;

quantization step size;

motion information prediction type;

length of the motion vector

motion vector resolution;

motion vector prediction index

motion vector difference size

motion vector difference resolution

motion interpolation filter

in-loop filter parameters

post-filter parameters;

based on segmentation information.

9 . The method according to claim 1 , further comprising:

obtaining, from the bitstream, sets of feature map elements and inputting the sets of feature map elements respectively into a feature map processing layer out of the plurality of layers based on the segmentation information processed by a segmentation information processing layer; and

obtaining the decoded data for picture or video processing based on a feature map processed by the plurality of cascaded layers.

10 . The method according to claim 9 , wherein at least one out of the plurality of cascaded layers is a segmentation information processing layer and a feature map processing layer.

11 . The method according to claim 9 , wherein, each layer out of the plurality of layers is either a segmentation information processing layer or a feature map processing layer.

12 . A computer program product stored on a non-transitory medium, which when executed on one or more processors performs the method according to claim 1 .

13 . A device for decoding an Image or video including a processing circuitry which is configured to perform the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 4, 2025
From: IKONIN, SERGEY YURIEVICH; SOSULNIKOV, MIKHAIL VYACHESLAVOVICH; KARABUTOV, ALEXANDER ALEXANDROVICH; SOLOVYEV, TIMOFEY MIKHAILOVICH; ALSHINA, ELENA ALEXANDROVNA; WANG, BIAO
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 072162/0794 →
Continuity (2)
Continuation PCTRU2020000750 · Dec 24, 2020
Related Publication 20230336759A1 · Oct 19, 2023
References Cited (61)
US 20170124409A1 · Choi · 2017 [cited by examiner]
US 20170316312A1 · Goyal · 2017 [cited by examiner]
US 20180350110A1 · Cho · 2018 [cited by examiner]
US 20190042923A1 · Janedula · 2019 [cited by examiner]
US 20190251360A1 · Cricri et al. · 2019 [cited by applicant]
US 20190373293A1 · Bortman · 2019 [cited by examiner]
US 20200143457A1 · Manmatha · 2020 [cited by examiner]
US 20200242774A1 · Park et al. · 2020 [cited by applicant]
US 20210279519A1 · Krim · 2021 [cited by examiner]
US 20210366123A1 · Wang · 2021 [cited by examiner]
CN 108986124A · 2018 [cited by applicant]
CN 111801945A · 2020 [cited by applicant]
JP 2020191630A · 2020 [cited by applicant]
WO 2005043882A2 · 2005 [cited by applicant]
Shannon, “A Mathematical Theory of Communication,” Reprinted with corrections from The Bell System Technical Journal, vol. 27, Total 55 pages (Jul., Oct. 1948). [cited by applicant]
Balle et al., “End-to-end Optimized Image Compression,” Published as a conference paper ICLR 2017, Total 27 pages (Mar. 2017). [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” arXiv e-prints, arXiv:1312.6114, Total 14 pages (Dec. 2022). [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” Proceedings of the 31st International Conference on Machine Learning, arXiv e-prints, Total 9 pages (Jan. 2014). [cited by applicant]
Wiegand et al., “Overview of the H.264/AVC video coding standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, Total 17 pages, Institute of Electrical and Electronics Engineers, New Y… [cited by applicant]
Sullivan et al., “Overview of the high efficiency video coding (HEVC) standard,” TCSVT, IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Total 20 pages (Dec. 2012). [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 7),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-P2001vE, Total 494 pages (Oct. 1-11, 2019). [cited by applicant]
Choi et al., “Near-Lossless Deep Feature Compression for Collaborative Intelligence,” 2018 IEEE 20th International Workshop on Multimedia Signal Processing (MMSP), Vancouver, BC, Total 6 pages (Jun. 2018). [cited by applicant]
Kang et al., “Neurosurgeon: Collaborative intelligence between the cloud and mobile edge,” Proc. 22nd ACM Int. Conf. Arch. Support Programming Languages and Operating Syst., Total 15 pages (Apr. 2017). [cited by applicant]
Eshratifar et al., “JointDNN: An efficient training and inference engine for intelligent mobile cloud computing services,” arXiv preprint, arXiv:1801.08618, Submitted 2018, Total 12 pages (Last Revised Feb. 2020). [cited by applicant]
Choi et al., “Deep feature compression for collaborative object detection,” arXiv:1802.03931v1 [cs.CV], Total 5 pages (Feb. 2018). [cited by applicant]
Redmon et al., “YOLO9000: Better, Faster, Stronger,” arXiv:1612.08242v1 [cs.CV], Total 9 pages (Dec. 2016). [cited by applicant]
Bossen et al., “HEVC Complexity and Implementation Analysis,” IEEE Transactions on Circuits Systems for Video Technology, vol. 22, No. 12, Total 12 pages (Dec. 2012). [cited by applicant]
Luo et al., “DeepSIC: Deep semantic image compression,” arXiv preprint, arXiv:1801.09468, Total 8 pages (Jan. 2018). [cited by applicant]
Marpe et al., “Context-based adaptive binary arithmetic coding in the H.264/AVC video compression standard,” IEEE Transactions on Circuits Systems for Video Technology, vol. 13, No. 7, Total 17 pages (Jul. 2003). [cited by applicant]
“Cisco Visual Networking Index: Forecast and Methodology,” White paper, CISCO, San Jose, CA, USA, Total 17 pages (Jun. 6, 2017). [cited by applicant]
Wu et al., “Compressed video action recognition,” Computer Vision and Pattern Recognition, CVPR, Total 14 pages (Mar. 2018). [cited by applicant]
Han et al., “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” arXiv preprint, arXiv:1510.00149v5, Submitted 2015, Total 14 pages (Last Revised Feb. 2016). [cited by applicant]
Toderici et al., “Variable rate image compression with recurrent neural networks,” arXiv preprint, arXiv:1511.06085v5, Total 12 pages (Mar. 2016). [cited by applicant]
Toderici et al., “Full resolution image compression with recurrent neural networks,” CVPR, Total 9 pages (Jul. 2017). [cited by applicant]
Agustsson et al., “Soft-to-hard vector quantization for end-to-end learning compressible representations,” NIPS, Total 16 pages (Jun. 2017). [cited by applicant]
Balle et al., “Variational image compression with a scale hyperprior,” Published as a conference paper at ICLR 2015, arXiv preprint, arXiv: 1802.01436v2, Total 23 pages (May 2018). [cited by applicant]
Johnston et al., “Improved lossy image compression with priming and spatially adaptive bit rates for recurrent networks,” Computer Vision and Pattern Recognition, CVPR, Jun. 2018, Total 9 pages (May 2017). [cited by applicant]
Theis et al., “Lossy image compression with compressive autoencoders,” arXiv preprint, arXiv:1703.00395v1, Total 19 pages (Mar. 2017). [cited by applicant]
Li et al., “Learning convolutional networks for content-weighted image compression,” Computer Vision and Pattern Recognition, CVPR, arXiv: 1703.10553v2 [cs.CV], Total 11 pages (Sep. 2017). [cited by applicant]
Rippel et al., “Real-time adaptive image compression,” ICML, arXiv:1705.05823v1 [stat.ML], Total 16 pages (May 2017). [cited by applicant]
Agustsson et al., “Generative adversarial networks for extreme learned image compression,” arXiv preprint, arXiv:1804.02958, Total 26 pages (Aug. 2019). [cited by applicant]
Wallace et al., “The JPEG still picture compression standard,” IEEE Transactions on Consumer Electronics, vol. 38, Issue 1, Total 17 pages (Feb. 1992). [cited by applicant]
Skodras et al., “The JPEG 2000 still image compression standard,” IEEE Signal Processing Magazine, vol. 18, Issue 5, Total 23 pages (Sep. 2001). [cited by applicant]
Bellarde et al., “BPG Image format,” http://bellard. org/bpg/, Accessed: Oct. 30, 2018, Total 2 pages (Apr. 2018). [cited by applicant]
Xue et al., “Video enhancement with task-oriented flow,” arXiv preprint, arXiv:1711.09078v2, Total 19 pages (Mar. 2019). [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 10),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-S2001-v, Total 548 pages (Jun. 22-Jul. 1, 2020). [cited by applicant]
PyTorch Master Documentation, “Torch.Masked_Select,” URL:https://pytorch.org/docs/master/generated/torch.masked_select.html, Total 1 pages (2023). [cited by applicant]
PyTorch Master Documentation, “Torch.Tensor,” URL:https://pytorch.org/docs/master/tensors.html, Total 23 pages (2023). [cited by applicant]
PyTorch 2.4 Documentation, “torch.gather,” URL:https://pytorch.org/docs/stable/generated/torch.gather.html, Total 2 pages (2023). [cited by applicant]
Garg et al., “Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue,” arXiv: 1603.04992v2 [cs.CV] Total 16 pages (Jul. 29, 2016). [cited by applicant]
Wu et al., “End-to-end Optimized Video Compression with MV-Residual Prediction,” mage and Video Processing, arXiv:2005.12945v1 [eess.IV], Total 4 pages (May 2020). [cited by applicant]
Wintz, “Transform Picture Coding,” Proceedings of the IEEE vol. 60, No. 7, Total 18 pages (Jul. 1972). [cited by applicant]
“Autoencoder,” WIKIPEDIA, The Free Encyclopedia, URL:https://en.wikipedia.org/wiki/Autoencoder, Total 15 pages, Sep. 2006 (Last edited on Sep. 19, 2024). [cited by applicant]
WIKIPEDIA The Free Encyclopedia, “Residual neural network”, URL: https://en.wikipedia.org/wiki/Convolutional_neural_network, total 7 pages, Nov. 2017 (Last edited Sep. 22, 2024). [cited by applicant]
Netravali et al., “Picture Coding: A Review,” Proceedings of the IEEE vol. 68, No. 3, Total 48 pages (Mar. 1980). [cited by applicant]
Choi et al., “Text of ISO/IEC CD 23094-1, Essential Video Coding,” ISO/IEC CD 23094-1, JTC1/SC29/WG11 N18568, Gothenburg, Sweden, Total 292 pages (Jul. 2019). [cited by applicant]
Gersho et al., “Vector Quantization and Signal Compression,” Springer Link, Kluwer Academic Publishers, Total 3 pages (Nov. 1991). [cited by applicant]
Akbari et al., “DSSLIC: Deep Semantic Segmentation-based Layered Image Compression,” Cornell University Library, 201 Online Library, Cornell University, Ithaca, NY, arXiv:1806.03348v3 [cs.CV] , Total 12 pages (Apr. 18, … [cited by applicant]
Lu et al., “DVC: An End-to-end Deep Video Compression Framework,” arXiv:1812.00101v3, total 14 pages (Apr. 7, 2019). [cited by applicant]
Balléet al., “Density Modeling of Images using a Generalized Normalization Transformation,” Published as a conference paper at ICLR 2016, arXiv:1511.06281v4, total 14 pages (Feb. 29, 2016). [cited by applicant]
Alshina et al., JPEG AI CfP response by Huawei: “Device agnostic learnable image coding using primary component extraction and conditional coding,” ISO/IEC JTC 1/SC29/WG1 M96016 (ITU-T SG16), 96th JPEG Meeting, Online, … [cited by applicant]
Cited By (1)
US 12,647,611