IP Library Granted Patent US 12,561,763
Granted Patent B2
US 12,561,763 · App. 18/615,050 · Granted Feb 24, 2026

Apparatus and method of guided neural network model for image processing

Inventors: Anbang Yao (Beijing, CN); Ming Lu (Beijing, CN); Yikai Wang (Beijing, CN); Shandong Wang (Beijing, CN); Yurong Chen (Beijing, CN); Sungye Kim (Folsom, CA); Attila Tamas Afra (Satu Mare, RO)
Assignee: INTEL CORPORATION
G06T5/50G06N3/02G06T7/13G06V40/161G06V40/171G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,763
App. No.
18/615,050
Granted
Feb 24, 2026
Kind
B2
Abstract

The present disclosure provides an apparatus and method of guided neural network model for image processing. An apparatus may comprise a guidance map generator, a synthesis network and an accelerator. The guidance map generator may receive a first image as a content image and a second image as a style image, and generate a first plurality of guidance maps and a second plurality of guidance maps, respectively from the first image and the second image. The synthesis network may synthesize the first plurality of guidance maps and the second plurality of guidance maps to determine guidance information. The accelerator may generate an output image by applying the style of the second image to the first image based on the guidance information.

Claims (80)

1 . An apparatus comprising:

processing circuitry to:

receive a first image as a content image and a second image as a style image;

generate first guidance maps and second guidance maps, respectively, from the first image and the second image;

generate an output image by applying a style of the second image to content of the first image based on guidance information associated with the first and second guidance maps; and

generate a first semantic guidance map and a second semantic guidance map, respectively, based on the first image and the second image, wherein the first and second semantic guidance maps are further generated by combining component masks with boundaries of parts representing facial landmarks.

2 . The apparatus of claim 1 , wherein the processing circuitry is further to:

synthesize the first guidance maps and the second guidance maps to determine the guidance information;

generate a first neural guidance map and a second neural guidance map, respectively, from the first image and the second image; and

generate a first position guidance map and a second position guidance map, respectively, from the first image and the second image.

3 . The apparatus of claim 2 , wherein the processing circuitry is further to:

generate the first neural guidance map by removing a style from the first image;

generate the second neural guidance map by removing a style from the second image;

determine the boundaries of the parts of the first image and the second image through semantic parsing;

detect the facial landmarks within the boundaries of the parts;

fit boundary curves of the facial components based on the facial landmarks;

generate the component masks based on the fitted boundary curves; and

generate the first semantic guidance map and the second semantic guidance map by combining the component masks with the boundaries of the parts.

4 . The apparatus of claim 2 , wherein the processing circuitry is further to:

fit boundary curves of eyeballs;

generate the first semantic guidance map and the second semantic guidance map by combining the component masks with the boundary curves of eyeballs.

5 . The apparatus of claim 2 , wherein the processing circuitry is further to:

determine the boundaries of the parts of the first image and the second image through semantic parsing;

determine a score for pixels within a boundary of a part based on a minimal distance of the pixel from the boundary; and

generate the first position guidance map and the second position guidance map based on scores of pixels.

6 . The apparatus of claim 1 , wherein the processing circuitry is further to synthesize one or more of the first and second neural guidance maps, the first and second semantic guidance maps, or the first and second position guidance maps to determine the guidance information.

7 . The apparatus of claim 1 , wherein the processing circuitry is coupled to a memory, the processing circuitry comprising one or more of graphics processing circuitry or application processing circuitry.

8 . A method comprising:

receiving, by processing circuitry of a computing device, a first image as a content image and a second image as a style image;

generating first guidance maps and second guidance maps, respectively, from the first image and the second image;

generating an output image by applying a style of the second image to content of the first image based on guidance information associated with the first and second guidance maps; and

generating a first semantic guidance map and a second semantic guidance map, respectively, based on the first image and the second image, wherein the first and second semantic guidance maps are further generated by combining component masks with boundaries of parts representing facial landmarks.

9 . The method of claim 8 , further comprising:

synthesizing the first guidance maps and the second guidance maps to determine the guidance information;

generating a first neural guidance map and a second neural guidance map, respectively, from the first image and the second image; and

generating a first position guidance map and a second position guidance map, respectively, from the first image and the second image.

10 . The method of claim 9 , further comprising:

generating the first neural guidance map by removing a style from the first image; and

generating the second neural guidance map by removing a style from the second image.

11 . The method of claim 9 , further comprising:

determining the boundaries of the parts of the first image and the second image through semantic parsing;

detecting the facial landmarks within the boundaries of the parts;

fitting boundary curves of the facial components based on the facial landmarks;

generating component masks based on the fitted boundary curves;

generating the first semantic guidance map and the second semantic guidance map by combining the component masks with the boundaries of the parts;

fitting boundary curves of eyeballs; and

generating the first semantic guidance map and the second semantic guidance map by combining the component masks with the boundary curves of eyeballs.

12 . The method of claim 9 , further comprising:

determining the boundaries of the parts of the first image and the second image through semantic parsing;

determining a score for pixels within a boundary of a part based on a minimal distance of the pixel from the boundary; and

generating the first position guidance map and the second position guidance map based on scores of pixels,

wherein synthesizing one or more of the first guidance maps and the second guidance maps comprises synthesizing the first and second neural guidance maps, the first and second semantic guidance maps, or the first and second position guidance maps to determine the guidance information.

13 . The method of claim 8 , wherein one or more of the style of the first image or the style of the second image comprises a portrait style or a lighting style, wherein the parts comprise a face, a body, or hair.

14 . The method of claim 8 , wherein the processing circuitry comprises one or more of graphics processing circuitry or application processing circuitry.

15 . At least one non-transitory computer-readable medium comprising instructions which, when executed, cause a computing device to perform operations comprising:

receiving a first image as a content image and a second image as a style image;

generating first guidance maps and second guidance maps, respectively, from the first image and the second image;

generating an output image by applying a style of the second image to content of the first image based on guidance information associated with the first and second guidance maps; and

generating a first semantic guidance map and a second semantic guidance map, respectively, based on the first image and the second image, wherein the first and second semantic guidance maps are further generated by combining component masks with boundaries of parts representing facial landmarks.

16 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

synthesizing the first guidance maps and the second guidance maps to determine the guidance information;

generating a first neural guidance map and a second neural guidance map, respectively, from the first image and the second image; and

generating a first position guidance map and a second position guidance map, respectively, from the first image and the second image.

17 . The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

generating the first neural guidance map by removing a style from the first image; and

generating the second neural guidance map by removing a style from the second image.

18 . The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

determining the boundaries of the parts of the first image and the second image through semantic parsing;

detecting the facial landmarks within the boundaries of the parts;

fitting boundary curves of the facial components based on the facial landmarks;

generating component masks based on the fitted boundary curves;

generating the first semantic guidance map and the second semantic guidance map by combining the component masks with the boundaries of the parts;

fitting boundary curves of eyeballs; and

generating the first semantic guidance map and the second semantic guidance map by combining the component masks with the boundary curves of eyeballs.

19 . The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

determining the boundaries of the parts of the first image and the second image through semantic parsing;

determining a score for pixels within a boundary of a part based on a minimal distance of the pixel from the boundary; and

generating the first position guidance map and the second position guidance map based on scores of pixels,

wherein synthesizing the first guidance maps and the second guidance maps comprises synthesizing one or more of the first and second neural guidance maps, the first and second semantic guidance maps, or the first and second position guidance maps to determine the guidance information.

20 . The non-transitory computer-readable medium of claim 15 , wherein one or more of the style of the first image or the style of the second image comprises a portrait style or a lighting style, wherein the parts comprise a face, a body, or hair, wherein the computing device comprises processing circuitry having one or more of graphics processing circuitry or application processing circuitry.

Priority Claims (1)
CN 202011562902.6 · Dec 25, 2020 · national
Continuity (2)
Continuation 17482998 · Sep 23, 2021
Related Publication 20240257316A1 · Aug 1, 2024
References Cited (23)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10515456B2 · Aksit · 2019 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11972545B2 · Yao · 2024 [cited by examiner]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180350030A1 · Simons · 2018 [cited by examiner]
US 20220415011A1 · Yang · 2022 [cited by applicant]
CN 114693501A · 2022 [cited by applicant]
EP 4020369A1 · 2022 [cited by applicant]
Extended European Search Report for EP Application No. 21 19 7634.5, mailed Mar. 21, 2022, 13 pages. [cited by applicant]
Leon A. Gatys et al: “Image Style Transfer Using Convolutional Neural Networks”, IEEE Conference on Computer Vision and Pattern Recognition, Jun. 1, 2016, pp. 2414-2423, XP055571216. [cited by applicant]
Shiri Fatemeh et al: “Face Destylization”, International Conference on Digital Image Computing: Techniques and Applications, Nov. 29, 2017, pp. 1-8, XP033287369. [cited by applicant]
Zhao Hui-Huang et al.: “Automatic semantic style transfer using deep convolutional neural networks and soft masks”, Arxiv, Aug. 31, 2017, pp. 1-12, XP55898823, Retrieved from the Internet: URL:https://arxiv.org/pdf/1708… [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junking, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
EP Application No. 21 197 634.5 Communication Pursuant to Article 94(3) mailed Dec. 4, 2024, 6 pages. [cited by applicant]
Lu Ming, et al., “Exemplar-Based Portrait Style Transfer”, IEEE Access, vol. 6 (2018), pp. 58532-58542, XP011694009. [cited by applicant]