IP Library Granted Patent US 12,499,581
Granted Patent B2
US 12,499,581 · App. 17/981,163 · Granted Dec 16, 2025

Scalable coding of video and associated features

Inventors: Alexander Alexandrovich Karabutov (Moscow, RU); Hyomin Choi (Burnaby, CA); Ivan Bajic (Burnaby, CA); Robert A. Cohen (Burnaby, CA); Saeed Ranjbar Alvar (Burnaby, CA); Sergey Yurievich Ikonin (Moscow, RU); Elena Alexandrovna Alshina (Munich, DE); Yin Zhao (Hangzhou, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06T9/00G06V10/44G06V10/761G06V10/82H04N19/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,581
App. No.
17/981,163
Granted
Dec 16, 2025
Kind
B2
Abstract

The present disclosure relates to scalable encoding and decoding of pictures. In particular, a picture is processed by one or more network layers of a trained module to obtain base layer features. Then, enhancement layer features are obtained, e.g. by a trained network processing in sample domain. The base layer features are for use in computer vision processing. The base layer features together with enhancement layer features are for use in picture reconstruction, e.g. for human vision. The base layer features and the enhancement layer features are coded in a respective base layer bitstream and an enhancement layer bitstream. Accordingly, a scalable coding is provided which supports computer vision processing and/or picture reconstruction.

Claims (47)

1 . An apparatus for processing a bitstream, the apparatus comprising:

a memory comprising instructions and a processing circuitry configured to execute the instructions to cause the apparatus to:

obtain a base layer bitstream including base layer features of a latent space and an enhancement layer bitstream including enhancement layer features;

extract from the base layer bitstream the base layer features; and

perform:

computer vision processing based on the base layer features; and

extracting the enhancement layer features from the enhancement layer bitstream and reconstructing a picture based on the base layer features and the enhancement layer features, and the reconstructing of the picture comprises:

combining the base layer features and the enhancement layer features; and

reconstructing the picture based on the combined features.

2 . The apparatus according to claim 1 , wherein the processing circuitry is further configured to execute the instructions to cause the apparatus to perform computer-vision processing based on the base layer features, and the computer-vision processing includes processing of the base layer features by one or more network layers of a first trained subnetwork.

3 . The apparatus according to claim 2 , wherein the reconstructing of the picture includes processing the combined features by one or more network layers of a second trained subnetwork different from the first trained subnetwork.

4 . The apparatus according to claim 1 , wherein the reconstructing of the picture includes:

reconstructing a base layer picture based on the base layer features; and

adding the enhancement layer features to the base layer picture.

5 . The apparatus according to claim 4 , wherein the enhancement layer features are based on differences between an encoder-side input picture and the base layer picture.

6 . The apparatus according to claim 1 , wherein the processing circuitry is configured to execute the instructions to cause the apparatus to perform the extracting the enhancement layer features from the enhancement layer bitstream and reconstructing the picture based on the base layer features and the enhancement layer features, and the reconstructed picture is a frame of a video, and the base layer features and the enhancement layer features are for a plurality of frames of the video.

7 . The apparatus according to claim 1 , wherein the processing circuitry is further configured to execute the instructions to further cause the apparatus to de-multiplex the base layer features and the enhancement layer features from a bitstream per frame.

8 . The apparatus according to claim 1 , wherein the processing circuitry is further configured to execute the instructions to further cause the apparatus to decrypt a portion of a bitstream including the enhancement layer features.

9 . An apparatus for processing a bitstream, the apparatus comprising:

a memory comprising instructions and a processing circuitry configured to execute the instructions to cause the apparatus to:

obtain a base layer bitstream including base layer features of a latent space and an enhancement layer bitstream including enhancement layer features;

extract from the base layer bitstream the base layer features; and

perform:

computer-vision processing based on the base layer features, including processing of the base layer features by one or more network layers of a first trained subnetwork.

10 . The apparatus according to claim 9 , wherein the processing circuitry is further configured to execute the instructions to cause the apparatus to perform extracting the enhancement layer features from the enhancement layer bitstream and reconstructing a picture based on the base layer features and the enhancement layer features, and the reconstructing of the picture includes:

combining the base layer features and the enhancement layer features; and

reconstructing the picture based on the combined features.

11 . The apparatus according to claim 10 , wherein the processing circuitry is further configured to execute the instructions to cause the apparatus to perform extracting the enhancement layer features from the enhancement layer bitstream and reconstructing the picture based on the base layer features and the enhancement layer features, and the reconstructing of the picture includes processing the combined features by one or more network layers of a second trained subnetwork different from the first trained subnetwork.

12 . The apparatus according to claim 9 , wherein the processing circuitry is further configured to execute the instructions to cause the apparatus to perform extracting the enhancement layer features from the enhancement layer bitstream and reconstructing a picture based on the base layer features and the enhancement layer features, and the reconstructing of the picture includes:

reconstructing a base layer picture based on the base layer features; and

adding the enhancement layer features to the base layer picture.

13 . The apparatus according to claim 12 , wherein the enhancement layer features are based on differences between an encoder-side input picture and the base layer picture.

14 . The apparatus according to claim 9 , wherein the processing circuitry is configured to execute the instructions to further cause the apparatus to perform extracting the enhancement layer features from the enhancement layer bitstream and reconstructing a picture based on the base layer features and the enhancement layer features, and the reconstructed picture is a frame of a video, and the base layer features and the enhancement layer features are for a plurality of frames of the video.

15 . The apparatus according to claim 9 , wherein the processing circuitry is further configured to execute the instructions to further cause the apparatus to de-multiplex the base layer features and the enhancement layer features from a bitstream per frame.

16 . The apparatus according to claim 9 , wherein the processing circuitry is further configured to execute the instructions to further cause the apparatus to decrypt a portion of a bitstream including the enhancement layer features.

17 . An apparatus for processing a bitstream, the apparatus comprising:

a memory comprising instructions and a processing circuitry configured to execute the instructions to cause the apparatus to:

obtain a base layer bitstream including base layer features of a latent space and an enhancement layer bitstream including enhancement layer features;

extract from the base layer bitstream the base layer features; and

perform:

extracting the enhancement layer features from the enhancement layer bitstream and reconstructing a picture based on the base layer features and the enhancement layer features, and the reconstructing of the picture comprises:

reconstructing a base layer picture based on the base layer features; and

adding the enhancement layer features to the base layer picture.

18 . The apparatus according to claim 17 , wherein the enhancement layer features are based on differences between an encoder-side input picture and the base layer picture.

19 . The apparatus according to claim 17 , wherein the processing circuitry is configured to execute the instructions to cause the apparatus to perform the extracting the enhancement layer features from the enhancement layer bitstream and reconstructing the picture based on the base layer features and the enhancement layer features, and the reconstructed picture is a frame of a video, and the base layer features and the enhancement layer features are for a plurality of frames of the video.

20 . The apparatus according to claim 17 , wherein the processing circuitry is further configured to execute the instructions to further cause the apparatus to de-multiplex the base layer features and the enhancement layer features from a bitstream per frame.

21 . The apparatus according to claim 17 , wherein the processing circuitry is further configured to execute the instructions to further cause the apparatus to decrypt a portion of a bitstream including the enhancement layer features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2025
From: KARABUTOV, ALEXANDER ALEXANDROVICH; CHOI, HYOMIN; BAJIC, IVAN; COHEN, ROBERT A.; ALVAR, SAEED RANJBAR; IKONIN, SERGEY YURIEVICH; ALSHINA, ELENA ALEXANDROVNA; ZHAO, YIN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 072064/0666 →
Continuity (2)
Continuation PCTRU2021000013 · Jan 13, 2021
Related Publication 20230065862A1 · Mar 2, 2023
References Cited (29)
US 7933456B2 · Han · 2011 [cited by examiner]
US 8126054B2 · Hsiang · 2012 [cited by examiner]
US 8160158B2 · Choi · 2012 [cited by examiner]
US 8432968B2 · Ye · 2013 [cited by examiner]
US 9137539B2 · Kashiwagi · 2015 [cited by examiner]
US 9219913B2 · Tu · 2015 [cited by examiner]
US 9571840B2 · Regunathan · 2017 [cited by examiner]
US 10284864B2 · Kim · 2019 [cited by examiner]
US 10455242B2 · Wang · 2019 [cited by examiner]
US 10542286B2 · Yu · 2020 [cited by examiner]
US 11677967B2 · Minoo · 2023 [cited by examiner]
US 20180174047A1 · Bourdev et al. · 2018 [cited by applicant]
US 20210042964A1 · Yeung · 2021 [cited by examiner]
WO 2020055279A1 · 2020 [cited by applicant]
WO 2020188273A1 · 2020 [cited by applicant]
Nokia, [VCM] Uses Cases for Video Coding for Machines, International Organisation for Standardisation Organisation Internationale De Normalisation ISO/IEC JTC1/SC29/WG 11 Coding of Moving Pictures and Audio, ISO/IEC JTC… [cited by applicant]
Karen Simonyan et al, Very Deep Convolutional Networks for Large-Scale Image Recognition, arXiv:1409.1556v6 [cs.CV] Apr. 10, 2015, 14 pages. [cited by applicant]
Zhejiang University, Potential Chances of Standarization on Video Coding for Machines (VCM), International Organisation for Standardisation Organisation Intern a Tio Nale De Normalisation ISO/IEC JTC1/SC29/WG11 Coding o… [cited by applicant]
Olaf Ronneberger et al, U-Net: Convolutional Networks for Biomedical Image Segmentation, arXiv:1505.04597v1 [cs.CV] May 18, 2015, 8 pages. [cited by applicant]
Joseph Redmon et al, You Only Look Once:Unified, Real-Time Object Detection, arXiv:1506.02640v5 [cs.CV] May 9, 2016, 10 pages. [cited by applicant]
Johannes Ball et al, Density Modeling of Images Using a Generalized Normalization Transformation, arXiv:1511.06281v4 [cs.LG] Feb. 29, 2016, 14 pages. [cited by applicant]
Wei Liu et al, SSD: Single Shot MultiBox Detector, arXiv:1512.02325v5 [cs.CV] Dec. 29, 2016, 17 pages. [cited by applicant]
Kaiming He et al, Deep Residual Learning for Image Recognition, arXiv:1512.03385v1 [cs.CV] Dec. 10, 2015, 12 pages. [cited by applicant]
Saeed Ranjbar Alvar et al, Multi-Task Learning With Compressible Features for Collaborative Intelligence, arXiv:1902.05179v2 [cs.MM] May 15, 2019, 5 pages. [cited by applicant]
Yueyu Hu et al, Towards Coding for Human and Machine Vision: a Scalable Image Coding Approach, arXiv:2001.02915v2 [cs.CV] Jan. 10, 2020, 6 pages. [cited by applicant]
Xiang Zhang et al, A Joint Compression Scheme of Video Feature Descriptors and Visual Content, IEEE Transactions on Image Processing, vol. 26, No. 2, Feb. 2017, 15 pages. [cited by applicant]
Ling-Yu Duan et al, Compact Descriptors for Video Analysis: The Emerging MPEG Standard, Visual Descriptor; Video Analytics; Standard, 2018 IEEE, 11 pages. [cited by applicant]
Hyomin Choi et al, Deep Feature Compression for Collaborative Object Detection, arXiv:1802.03931v1 [cs.CV] Feb. 12, 2018, 6 pages. [cited by applicant]
MPEG Requirements, Draft Call for Evidence for Video Coding for Machines, International Organization for Standardization Organisation Internationale De Normalisation ISO/IEC JTC1/SC29/WG11CODING of Moving Pictures and A… [cited by applicant]