IP Library Granted Patent US 12,670,554
Granted Patent B2
US 12,670,554 · App. 17/483,074 · Granted Jun 30, 2026

Conditional kernel prediction network and adaptive depth prediction for image and video processing

Inventors: Anbang Yao (Beijing, CN); Ming Lu (Beijing, CN); Yikai Wang (Beijing, CN); Yurong Chen (Beijing, CN); Attila Tamas Afra (Satu Mare, RO); Sungye Kim (Folsom, CA); Karthik Vaidyanathan (San Francisco, CA)
Assignee: INTEL CORPORATION
G06T5/70G06N3/08G06N20/10G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,554
App. No.
17/483,074
Filed
Sep 23, 2021
Granted
Jun 30, 2026
Kind
B2
Examiner
LU, ZHIYU
Art Unit
2665
USPC
382/156
Abstract

Embodiments are generally directed to a Conditional Kernel Prediction Network (CKPN) for image and video de-noising and other related image and video processing applications. Disclosed is an embodiment of a method for de-noising an image or video frame by a convolutional neural network implemented on a compute engine, the image including a plurality of pixels, the method comprising: for each of the plurality of pixels of the image, generating a convolutional kernel having a plurality of kernel weights for the pixel, the plurality of kernel weights respectively corresponding to pixels within a region surrounding the pixel; adjusting the plurality of kernel weights of the convolutional kernel for the pixel based on convolutional kernels generated respectively for the corresponding pixels within the region surrounding the pixel; and filtering the pixel with the adjusted plurality of kernel weights and pixel values of the corresponding pixels within the region surrounding the pixel to obtain a de-noised pixel.

Claims (25)

1 . An apparatus comprising:

processor circuitry coupled to a memory, the processor circuitry to:

generate convolutional kernels having kernel weights corresponding to pixels associated with an image, wherein the pixels are within a region surrounding a pixel associated with the image;

adjust a kernel weight of a convolutional kernel corresponding to the pixel based on the convolutional kernels generated respectively for the pixels, wherein the kernel weight is further adjusted based on one or more of a position of the pixel relative to the pixels within the region surrounding the pixel associated with the image or a correlation between the convolutional kernel corresponding to the pixel and the convolutional kernels generated for the pixels within the region surrounding the pixel;

obtain a de-noised pixel by filtering the pixel based on the adjusted kernel weight and a pixel value associated with the pixel and further based on the kernel weights and pixel values associated with the pixels within the region surrounding the pixel; and

infer the convolutional kernel corresponding to the pixel based on a target depth of a convolutional neural network, wherein the target depth is individually determined for the pixels associated with the image by the convolutional neural network during inference, wherein the convolutional neural network is trained with a loss calculated based on a randomly sampled target depth and predicted target depths during iterations associated with training, wherein the target depths are predicted based on adaptive depth prediction to dynamically select depth values for the pixels during training based on image characteristics, and wherein the adaptive depth prediction adjusts the target depth for the pixels to optimize kernel generation for local neighborhood characteristics associated with the pixels.

2 . The apparatus of claim 1 , wherein the processor circuitry is further to:

determine the target depth of the convolutional neural network for generating the convolutional kernels for the pixels, wherein the target depth being less than or equal to a full depth associated with the convolutional neural network.

3 . The apparatus of claim 1 , wherein the processor circuitry comprises one or more of graphics processor circuitry or application processor circuitry.

4 . A method comprising:

generating, by a processor of a computing device, convolutional kernels having kernel weights corresponding to pixels associated with an image, wherein the pixels are within a region surrounding a pixel associated with the image;

adjusting a kernel weight of a convolutional kernel corresponding to the pixel based on the convolutional kernels generated respectively for the pixels, wherein the kernel weight is further adjusted based on one or more of a position of the pixel relative to the pixels within the region surrounding the pixel associated with the image or a correlation between the convolutional kernel corresponding to the pixel and the convolutional kernels generated for the pixels within the region surrounding the pixel;

obtaining a de-noised pixel by filtering the pixel based on the adjusted kernel weight and a pixel value associated with the pixel and further based on the kernel weights and pixel values associated with the pixels within the region surrounding the pixel; and

inferring the convolutional kernel corresponding to the pixel based on a target depth of a convolutional neural network, wherein the target depth is individually determined for the pixels associated with the image by the convolutional neural network during inference, wherein the convolutional neural network is trained with a loss calculated based on a randomly sampled target depth and predicted target depths during iterations associated with training, wherein the target depths are predicted based on adaptive depth prediction to dynamically select depth values for the pixels during training based on image characteristics, and wherein the adaptive depth prediction adjusts the target depth for the pixels to optimize kernel generation for local neighborhood characteristics associated with the pixels.

5 . The method of claim 4 , further comprising:

determine a target depth of a convolutional neural network for generating the convolutional kernels for the pixels, wherein the target depth being less than or equal to a full depth associated with the convolutional neural network.

6 . The method of claim 4 , wherein the processor is coupled to a memory, the processor comprises one or more of a graphics processor or an application processor.

7 . At least one non-transitory computer-readable medium having stored thereon instructions which, when executed, cause a computing device to perform operations comprising:

generating convolutional kernels having kernel weights corresponding to pixels associated with an image, wherein the pixels are within a region surrounding a pixel associated with the image;

adjusting a kernel weight of a convolutional kernel corresponding to the pixel based on the convolutional kernels generated respectively for the pixels, wherein the kernel weight is further adjusted based on one or more of a position of the pixel relative to the pixels within the region surrounding the pixel associated with the image or a correlation between the convolutional kernel corresponding to the pixel and the convolutional kernels generated for the pixels within the region surrounding the pixel;

obtaining a de-noised pixel by filtering the pixel based on the adjusted kernel weight and a pixel value associated with the pixel and further based on the kernel weights and pixel values associated with the pixels within the region surrounding the pixel; and

inferring the convolutional kernel corresponding to the pixel based on a target depth of a convolutional neural network, wherein the target depth is individually determined for the pixels associated with the image by the convolutional neural network during inference, wherein the convolutional neural network is trained with a loss calculated based on a randomly sampled target depth and predicted target depths during iterations associated with training, wherein the target depths are predicted based on adaptive depth prediction to dynamically select depth values for the pixels during training based on image characteristics, and wherein the adaptive depth prediction adjusts the target depth for the pixels to optimize kernel generation for local neighborhood characteristics associated with the pixels.

8 . The non-transitory computer-readable medium of claim 7 , wherein the operations further comprise:

determine a target depth of a convolutional neural network for generating the convolutional kernels for the pixels, wherein the target depth being less than or equal to a full depth associated with the convolutional neural network.

9 . The non-transitory computer-readable medium of claim 7 , wherein the computing device comprises processor circuitry coupled to a memory, the processor circuitry includes one or more of graphics processor circuitry or an application processor circuitry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2021
From: YAO, ANBANG; LU, MING; WANG, YIKAI; CHEN, YURONG; AFRA, ATTILA TAMAS; KIM, SUNGYE; VAIDYANATHAN, KARTHIK
To: INTEL CORPORATION
Reel/Frame 057577/0899 →
Priority Claims (1)
CN 202011565220.0 · Dec 25, 2020 · national
Continuity (1)
Related Publication 20220207656A1 · Jun 30, 2022
References Cited (33)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 8780971B1 · Bankoski · 2014 [cited by examiner]
US 9344729B1 · Grange · 2016 [cited by examiner]
US 9838690B1 · Grange · 2017 [cited by examiner]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20080240592A1 · Lee · 2008 [cited by examiner]
US 20150131885A1 · Kim · 2015 [cited by examiner]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180089806A1 · Bitterli · 2018 [cited by examiner]
US 20180314935A1 · Lewis · 2018 [cited by examiner]
US 20200051260A1 · Shen · 2020 [cited by examiner]
US 20220067429A1 · Kwon · 2022 [cited by examiner]
US 20220148135A1 · Isik · 2022 [cited by examiner]
US 20220291387A1 · Pacala · 2022 [cited by examiner]
CN 106612386 · 2017 [cited by examiner]
CN 107958286 · 2018 [cited by examiner]
CN 109740734 · 2019 [cited by examiner]
CN 112258565 · 2021 [cited by examiner]
CN 114693850A · 2022 [cited by applicant]
EP 4020377A1 · 2022 [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junking, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Extended European Search Report for EP Application No. 21199015.5 mailed Apr. 4, 2022, 10 pages. [cited by applicant]
Back Jonghee et al: “Deep combiner for independent and correlated pixel estimates” ACM Transactions on Graphics, ACM, NY, US, vol. 39, No. 6, Nov. 26, 2020 (Nov. 26, 2020), pp. 1-12, XP058644856, ISSN: 0730-0301, DOI: 1… [cited by applicant]
Ben Mildenhall et al: “Burst Denoising with Kernel Prediction Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Dec. 6, 2017 (Dec. 6, 2017), XP080845370, * the whole… [cited by applicant]
Marinc Talmaj et al: “Multi-Kernel Prediction Networks for Denoising of Burst Images”, 2019 IEEE International Conference on Image Processing (ICIP), IEEE, Sep. 22, 2019 (Sep. 22, 2019), pp. 2404-2408, XP033647290, DOI:… [cited by applicant]
Zhang Bin et al: “Attention Mechanism Enhanced Kernel Prediction Networks for Denoising of Burst Images”, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, … [cited by applicant]