IP Library › Granted Patent US 11,516,478
Granted Patent B2
US 11,516,478 · App. 17/565,545 · Granted Nov 29, 2022

Method and apparatus for coding machine vision data using prediction

Inventors: Je Won Kang (Seoul, KR); Chae Hwa Yoo (Seoul, KR); Seung Wook Park (Gyeonggi-do, KR)
Assignees: Hyundai Motor Company; Kia Corporation; Ewha University-Industry Collaboration Foundation
H04N19/147G06V10/7715H04N19/105H04N19/176H04N19/196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,516,478
App. No.
17/565,545
Granted
Nov 29, 2022
Kind
B2
Abstract

The present disclosure relates to an apparatus for and a method of coding machine vision data by using prediction, and for improving the efficiency of encoding the data used for machine vision, provides an apparatus for Video Coding for Machines (VCM) which sets reference data according to a correlation between the data, generates, based on the reference data, prediction data for original data having a high correlation with the reference data, and generates residual data between the prediction data and the original data, and provides a coding method performed by the apparatus for VCM.

Claims (60)

1. A coding method performed by a coding apparatus of a machine vision system for coding feature maps of video frames, the coding method comprising:

extracting, from a key frame, a reference feature map that is a feature map of the key frame by using a machine task model that is based on deep learning, the key frame being selected from the video frames in terms of bit rate distortion optimization

extracting, from remaining frames other than the key frame, an original feature map of each of the remaining frames by using the machine task model;

generating a predicted feature map of each of the remaining frames based on the reference feature map;

generating a residual feature map by subtracting the predicted feature map from the original feature map of each of the remaining frames;

encoding the reference feature map; and

encoding the residual feature map of each of the remaining frames.

2. The coding method of claim 1 , wherein the generating of the predicted feature map comprises:

performing an inter prediction based on the reference feature map to generate the predicted feature map.

3. The coding method of claim 1 , wherein the generating of the predicted feature map comprises:

using a predictive model that is based on deep learning to generate the predicted feature map from the reference feature map.

4. The coding method of claim 3 , wherein the predictive model is configured to be pre-trained based on a loss function that includes a loss for promoting the predicted feature map to predict the original feature map from the reference feature map, a loss for reducing a bit number of the residual feature map, and a loss for reducing a difference between the original feature map and a reconstructed feature map that is generated by a decoding apparatus in the machine vision system.

5. The coding method of claim 3 , wherein the predictive model is configured to be pre-trained end-to-end along with the machine task model.

6. The coding method of claim 1 , wherein the encoding of the reference feature map comprises:

setting a neighboring block of a transport block in the key frame as a reference block;

generating a feature map of a prediction block by performing prediction based on a feature map of the reference block;

generating a residual block by subtracting the feature map of the prediction block from a feature map of the transport block; and

encoding the residual block.

7. The coding method of claim 6 , wherein the generating of the feature map of the prediction block comprises:

performing intra prediction based on the feature map of the reference block to generate the feature map of the prediction block.

8. The coding method of claim 6 , wherein the generating of the feature map of the prediction block comprises:

using a deep learning-based block prediction model to generate the feature map of the prediction block from the feature map of the reference block.

9. A coding method performed by a coding apparatus of a machine vision system for coding a feature map of a main task and feature maps of subtasks, the coding method comprising:

extracting a reference feature map that is the feature map of the main task set among target tasks by using a machine task model that is based on deep learning;

extracting, from the subtasks, an original feature map of each of the subtasks by using the machine task model;

generating a predicted feature map of each of the subtasks based on the reference feature map;

generating a residual feature map by subtracting the predicted feature map from the original feature map of each of the subtasks;

encoding the reference feature map; and

encoding a residual feature map of each of the subtasks.

10. The coding method of claim 9 , wherein the generating of the predicted feature map comprises:

using a predictive model that is based on deep learning to generate the predicted feature map from the reference feature map.

11. The coding method of claim 10 , wherein the predictive model is configured to be pre-trained based on a loss function that includes a loss for promoting the predicted feature map to predict the original feature map from the reference feature map, a loss for reducing a bit number of the residual feature map, and a loss for reducing a difference between the original feature map and a reconstructed feature map that is generated by a decoding apparatus in the machine vision system.

12. The coding method of claim 9 , wherein the encoding of the reference feature map comprises:

setting a neighboring block of a transport block in a frame representing the main task as a reference block;

generating a feature map of a prediction block by performing prediction based on a feature map of the reference block;

generating a residual block by subtracting the feature map of the prediction block from a feature map of the transport block; and

encoding the residual block.

13. The coding method of claim 12 , wherein the generating of the feature map of the prediction block comprises:

performing intra prediction based on the feature map of the reference block to generate the feature map of the prediction block.

14. The coding method of claim 12 , wherein the generating of the feature map of the prediction block comprises:

using a deep learning-based block prediction model to generate the feature map of the prediction block from the feature map of the reference block.

15. A coding method performed by a coding apparatus of a machine vision system for coding a feature map of a machine task model including a plurality of layers, the coding method comprising:

extracting, by using the machine task model and from an input image, a reference feature map that is an output feature map of a first layer;

extracting, by using the machine task model and from the input image, an original feature map that is an output feature map of a second layer that is a layer deeper than the first layer in the machine task model;

generating a predicted feature map based on the reference feature map;

generating a residual feature map of the second layer by subtracting the predicted feature map from the original feature map;

encoding the reference feature map; and

encoding the residual feature map of the second layer.

16. The coding method of claim 15 , wherein the generating of the predicted feature map comprises:

using a predictive model that is based on deep learning to generate the predicted feature map from the reference feature map.

17. The coding method of claim 16 , wherein the predictive model is configured to be pre-trained based on a loss function that includes a loss for promoting the predicted feature map to predict the original feature map from the reference feature map, a loss for reducing a bit number of the residual feature map, and a loss for reducing a difference between the original feature map and a reconstructed feature map that is generated by a decoding apparatus in the machine vision system.

18. The coding method of claim 15 , wherein the encoding of the reference feature map comprises:

setting a neighboring block of a transport block in the input image as a reference block;

generating a feature map of a prediction block by performing prediction based on a feature map of the reference block;

generating a residual block by subtracting the feature map of the prediction block from a feature map of the transport block; and

encoding the residual block.

19. The coding method of claim 18 , wherein the generating of the feature map of the prediction block comprises:

performing intra prediction based on the feature map of the reference block to generate the feature map of the prediction block.

20. The coding method of claim 18 , wherein the generating of the feature map of the prediction block comprises:

using a deep learning-based block prediction model to generate the feature map of the prediction block from the feature map of the reference block.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2021
From: KANG, JE WON; YOO, CHAE HWA; PARK, SEUNG WOOK
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION; EWHA UNIVERSITY - INDUSTRY COLLABORATION FOUNDATION
Reel/Frame 058505/0544 →
Priority Claims (2)
KR 10-2020-0187062 · Dec 30, 2020 · national
KR 10-2021-0182334 · Dec 20, 2021 · national
Continuity (1)
Related Publication 20220210435A1 · Jun 30, 2022