IP Library Granted Patent US 12,136,254
Granted Patent B2
US 12,136,254 · App. 17/159,653 · Granted Nov 5, 2024

Method and apparatus with image processing

Inventors: Donghyuk Kwon (Seoul, KR); Minkyoung Cho (Icheon, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V10/7715G06N3/04G06N3/08G06T3/40G06T7/33G06V10/955G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,136,254
App. No.
17/159,653
Granted
Nov 5, 2024
Kind
B2
Abstract

An Image processing method and apparatus are provided. The Image processing method includes obtaining a kernel that is deterministically defined based on an extension ratio of a first feature map, up-sampling the first feature map to a second feature map by performing a transposed convolution operation between the first feature map and the kernel, and outputting the second feature map.

Claims (78)

1. An image processing method comprising:

obtaining kernel values of a kernel, where distributed positions of same kernel values in the kernel are based on an extension ratio;

up-sampling a first feature map to a second feature map; and

outputting the second feature map,

wherein the extension ratio is a ratio between a size of the second feature map and a size of the first feature map, and

wherein the up-sampling of the first feature map includes selectively, based on whether the size of the second feature map matches a predetermined formulation that is dependent on at least on of the size of the first feature map or the extension ratio, performing one between an interpolation operation and a transposed convolution operation to up-sample the first feature map, where the transposed convolution is between the first feature map and the kernel.

2. The method of claim 1 , wherein the outputting of the second feature map includes inputting the second feature map to a neural network layer configured to perform another transposed convolution operation between the second feature map and a select kernel that is generated based on an extension ratio of the second feature map.

3. The method of claim 1 , wherein the image processing method is a method of an image processing device that includes a neuro processing unit that performs the up-sampling of the first feature map to the second feature map without performing the interpolation operation using an external information provider.

4. The method of claim 1 , wherein the kernel comprises weights that vary dependent on a distance between a central pixel of the kernel and pixels neighboring the central pixel.

5. The method of claim 4 , wherein values of the weights are inversely proportional to the distance between the central pixel and the neighboring pixels.

6. The method of claim 1 , wherein the kernel has a size that is greater than the size of the first feature map.

7. The method of claim 1 , wherein a size of the kernel, a stride parameter of the transposed convolution operation, and a padding parameter of the transposed convolution operation are determined based on the extension ratio.

8. The method of claim 7 ,

wherein the size of the kernel, the stride parameter of the transposed convolution operation, and the padding parameter are kernel parameters, and

wherein the performing of the transposed convolution operation includes performing the transposed convolution operation using a neural network layer configured according to the kernel parameters and input the first feature map.

9. The method of claim 7 , wherein the size of the kernel, the stride parameter of the transposed convolution operation, and the padding parameter of the transposed convolution operation are determined based on an interpolation scheme in addition to the extension ratio.

10. The method of claim 9 , wherein the method further comprises, at a different time than a time of the selective performing of the one between the interpolation operation and the transposed convolution operation, another selectively performing between one of:

a first generating of another second feature map by performing another transposed convolution operation based on another kernel to generate the other second feature map from another first feature map, where distributed positions of same kernel values in the other kernel are based on another extension ratio; and

a second generating of the other second feature map by interpolating, according to the interpolation scheme, the other second feature map from the other first feature map using an external information provider.

11. The method of claim 1 , wherein pixel alignment information of another second feature map is predetermined,

wherein the method further comprises determining whether there will be pixel alignment between another first feature map and the other second feature map based on the pixel alignment information of the other second feature map, and

the method further comprises, dependent on a result of the determining of whether there will be pixel alignment, another selectively performing, at a different time than a time than a time of the selective performing of the one between the interpolation operation and the transposed convolution operation, between one of:

a first generating of the other second feature map by performing another transposed convolution to generate the other second feature map from the other first feature map based on another kernel, where distributed positions of same kernel values in the other kernel are based on another extension ratio; and

a second generating of the other second feature map by interpolating, according to an interpolation scheme, the other second feature map from the other first feature map using an external information provider.

12. The method of claim 1 , wherein the first feature map comprises a plurality of channels, and the transposed convolution operation is a depth-wise transposed convolution operation between the channels and the kernel.

13. The method of claim 1 , wherein

the first feature map comprises quantized activation values, and

the kernel comprises quantized weights.

14. The method of claim 1 , wherein the kernel comprises a transposed convolution kernel matrix obtained by converting a weight multiplied by a first value of a first pixel of the first feature map into a parameter through a quantization, to convert the first value into a second value of a second pixel of the second feature map corresponding to a position of the first pixel.

15. The method of claim 1 , wherein the extension ratio is determined so that the size of the first feature map and the size of the second feature map satisfy a predetermined integer ratio.

16. The method of claim 1 , wherein the first feature map is an output feature map of a previous convolutional neural network layer in a neural network that is input an original image, or an output up-sampled feature map of a transposed convolution neural network layer in the previous layer in the neural network.

17. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

18. An image processing apparatus comprising:

one or more processors configured to:

obtain kernel values of a kernel, where distributed positions of same kernel values in the kernel are based on an extension ratio; and

up-sample a first feature map to a second feature map,

wherein the extension ratio is a ratio between a size of the second feature map and a size of the first feature map, and

wherein the up-sampling of the first feature map includes a selective, based on whether the size of the second feature map matches a predetermined formulation that is dependent on at least one of the size of the first feature map or the extension ratio, performance of one between an interpolation operation and a transposed convolution operation to up-sample the first feature map, where the transposed convolution operation is between the first feature map and the kernel.

19. The apparatus of claim 18 , wherein the one or more processors are configured to perform the transposed convolution operation using a neural network layer, configured according the kernel and kernel parameters, input the first feature map.

20. The apparatus of claim 19 , wherein the kernel parameters include a size of the kernel, a stride parameter of the transposed convolution operation, and a padding parameter.

21. The apparatus of claim 18 , wherein the one or more processors include a neural processing unit configured to perform the up-sampling of the first feature map to the second feature map without performing the interpolation operation using an external information provider.

22. The apparatus of claim 18 , wherein the kernel comprises weights that vary dependent on a distance between a central pixel of the kernel and pixels neighboring the central pixel.

23. The apparatus of claim 22 , wherein the weights are inversely proportional to the distance between the central pixel and the neighboring pixels.

24. The apparatus of claim 18 , wherein the kernel has a size greater than the first feature map.

25. The apparatus of claim 18 , wherein a size of the kernel, a stride parameter of the transposed convolution operation, and a padding parameter of the transposed convolution operation are determined based on the extension ratio.

26. The apparatus of claim 25 , wherein the size of the kernel, the stride parameter of the transposed convolution operation, and the padding parameter are kernel parameters, and

wherein the one or more processors are configured to perform the transposed convolution operation using a neural network layer configured according to the kernel parameters and input the first feature map.

27. The apparatus of claim 25 , wherein the size of the kernel, the stride parameter of the transposed convolution operation, and the padding parameter of the transposed convolution operation are determined based on an interpolation scheme in addition to the extension ratio.

28. The apparatus of claim 18 , wherein pixel alignment information of the second feature map is predetermined,

the one or more processors are further configured to determine whether there will be pixel alignment between the first feature map and the second feature map based on the pixel alignment information of the second feature map, and

dependent on a result of the determining of whether there will be pixel alignment, the one or more processors perform the transposed convolution using a neural network layer, or perform interpolation according to the interpolation scheme to generate the second feature map.

29. The apparatus of claim 18 , wherein the first feature map comprises a plurality of channels, and the transposed convolution operation is a depth-wise transposed convolution operation between the channels and the kernel.

30. The apparatus of claim 18 , wherein

the first feature map comprises quantized activation values, and

the kernel comprises quantized weights.

31. The apparatus of claim 18 , wherein the kernel comprises a transposed convolution kernel matrix obtained by converting a weight multiplied by a first value of a first pixel of the first feature map into a parameter through a quantization, to convert the first value into a second value of a second pixel of the second feature map corresponding to a position of the first pixel.

32. The apparatus of claim 18 , wherein the extension ratio is determined so that a size of the first feature map and a size of the second feature map satisfy a predetermined integer ratio.

33. The apparatus of claim 18 , wherein the first feature map is an output feature map of a previous convolutional neural network layer in a neural network that is input an original image, or an output up-sampled feature map of a transposed convolution neural network layer in the previous layer in the neural network.

34. The apparatus of claim 18 , further comprising:

a hardware accelerator configured to dynamically generate the kernel based on the extension ratio.

35. The apparatus of claim 18 , wherein the image processing apparatus is a head-up display (HUD) device, a three-dimensional (3D) digital information display (DID), a navigation device, a 3D mobile device, a smartphone, a smart television (TV), or a smart vehicle.

36. An apparatus comprising:

one or more processors, including a CPU, a GPU, and neuro processor,

wherein the neural processor is configured to:

implement a neural network through respective performances of plural convolutional layers;

obtain kernel values of a kernel, where distributed positions of same kernel values in the kernel are based on an extension ratio of a first feature map generated by one of the plural convolutional layers; and

up-sample a first feature map to a second feature map,

wherein the extension ratio is a ratio between a size of the second feature map and a size of the first feature map, and

wherein the up-sampling of the first feature map includes a selective, based on whether the size of the second feature map matches a predetermined formulation that is dependent on at least one of the size of the first feature map or the extension ratio, performance of one between an interpolation operation and a transposed convolution operation to up-sample the first feature map, where the transposed convolution operation is between the first feature map and the kernel.

37. The apparatus of claim 36 , further comprising a hardware accelerator configured to dynamically generate the kernel, based on the extension ratio, and provide the kernel to the neural processor.

38. The apparatus of claim 36 , wherein the neural network includes a neural network layer that is input the second feature map.

39. The method of claim 1 , wherein the kernel is dynamically generated by an accelerator based on a size of the first feature map.

40. The apparatus of claim 36 , wherein the generating of the kernel values is performed during an inference operation or after training of the neural network.

41. The apparatus of claim 36 , wherein the predetermined formulation is dependent on the size of the first feature map and the extension ratio, and

wherein, for the selective performance of the one between the interpolation operation and the transposed convolution operation, the neural processor is configured to:

perform the transposed convolution operation when the predetermined formulation matches the size of the second feature map; and

perform the interpolation operation when the predetermined formulation does not match the size of the second feature map.

42. The apparatus of claim 41 , wherein the interpolation operation is a bilinear interpolation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2021
From: KWON, DONGHYUK; CHO, MINKYOUNG
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 055051/0870 →
Priority Claims (1)
KR 10-2020-0111842 · Sep 2, 2020 · national
Continuity (1)
Related Publication 20220067429A1 · Mar 3, 2022
Cited By (1)
US 12,651,321