IP Library Granted Patent US 12,327,328
Granted Patent B2
US 12,327,328 · App. 17/379,622 · Granted Jun 10, 2025

Aesthetics-guided image enhancement

Inventors: Xiaohui Shen (San Jose, CA); Zhe Lin (Fremont, CA); Xin Lu (Mountain View, CA); Sarah Aye Kong (Cupertino, CA); I-Ming Pao (Palo Alto, CA); Yingcong Chen (Hong Kong, TW)
Assignee: Adobe Inc.
G06T5/00G06V10/454G06V10/774G06V10/82G06V20/10G06V30/19173G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,328
App. No.
17/379,622
Granted
Jun 10, 2025
Kind
B2
Abstract

Methods and systems are provided for generating enhanced image. A neural network system is trained where the training includes training a first neural network that generates enhanced images conditioned on content of an image undergoing enhancement and training a second neural network that designates realism of the enhanced images generated by the first neural network. The neural network system is trained by determine loss and accordingly adjusting the appropriate neural network(s). The trained neural network system is used to generate an enhanced aesthetic image from a selected image where the output enhanced aesthetic image has increased aesthetics when compared to the selected image.

Claims (35)

1. One or more non-transitory computer-readable media having a plurality of executable instructions embodied thereon, which, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising:

obtaining scored images, wherein a plurality of images are scored based on aesthetic attributes;

designating the scored images that have aesthetic scores within a predefined range as input images;

designating the scored images that have aesthetic scores above a predefined threshold as reference images;

training an aesthetic image enhancing neural network system using the input images with a first neural network and reference images with a second neural network, wherein the trained aesthetic image enhancing neural network system is used to generate an enhanced image from an image input;

analyzing portions of the scored images using a segmentation network, wherein analyzing portions of the scored images using the segmentation network causes adaptive adjustments by the first neural network; and

outputting the enhanced image, wherein the enhanced image has increased aesthetics when compared with the image input into the trained aesthetic image enhancing neural network system.

2. The non-transitory media of claim 1 ,

wherein the adaptive adjustments by the first neural network are dependent on the content of the scored images.

3. The non-transitory media of claim 1 , wherein the aesthetic scores are assigned based on one or more aesthetic attribute.

4. The non-transitory media of claim 3 , wherein the one or more aesthetic attributes include one or more of interesting content, object emphasis, lighting, color harmony, vivid color, shallow depth of field, motion blur, rule of thirds, balancing element, repetition, and symmetry.

5. The non-transitory media of claim 1 , wherein training the aesthetic image enhancing neural network system further comprises:

during a first iteration, determining loss using a comparison between a training enhanced image and an input image based in part on differences conditioned on a corresponding segmentation map; and

adjusting the aesthetic image enhancing neural network system based on the determined loss.

6. The non-transitory media of claim 5 , wherein the loss is one or more of content loss and adversarial loss.

7. The non-transitory media of claim 1 , wherein training the aesthetic image enhancing neural network system further comprises:

during a first iteration, determining loss using a comparison between a reference image and realism designation; and

adjusting the aesthetic image enhancing neural network system based on the determined loss.

8. The non-transitory media of claim 7 , wherein the loss is one or more of content loss and adversarial loss.

9. A computing system comprising:

one or more processors; and

one or more non-transitory computer-readable storage media, coupled with the one or more processors, having instructions stored thereon, which, when executed by the one or more processors, cause the computing system to:

obtain a set of images for training a first neural network of an aesthetic image enhancing neural network system;

determine a plurality of aesthetic scores for the set of images;

analyze portions of the scored images using a segmentation network, thereby causing adaptive adjustments by the first neural network;

designate a subset of the set of images, having aesthetic scores above a threshold level, as reference images for training a second neural network of the aesthetic image enhancing neural network system; and

train the aesthetic image enhancing neural network system using the set of images and the reference images, wherein the trained aesthetic image enhancing neural network system generates an enhanced image from a first input image.

10. The computing system of claim 9 , further comprising:

obtain a parsed image corresponding to the first input image; and

input the parsed image into the trained aesthetic image enhancing neural network system, wherein the first neural network converts the first input image into the enhanced aesthetic image using the parsed image.

11. The computing system of claim 9 , wherein the first neural network acts as a generator in the trained aesthetic image enhancing neural network system based on a generative adversarial type architecture.

12. The computing system of claim 9 , wherein the trained aesthetic image enhancing neural network system further includes a deactivated second neural network that acts as a discriminator during training.

13. The computing system of claim 9 , wherein the first input image is selected from a set of images stored on a user device, the set of images taken using a camera function of the user device.

14. The computing system of claim 9 , wherein the first neural network parses the set of images using a pyramid parsing network.

15. The computing system of claim 9 , wherein the first neural network parses the set of images using grayscale.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2021
From: SHEN, XIAOHUI; LIN, ZHE; LU, XIN; KONG, SARAH AYE; PAO, I-MING; CHEN, YINGCONG
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 057139/0276 →
CHANGE OF NAME Recorded Aug 10, 2021
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 057156/0421 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2021
From: SHEN, XIAOHUI; LIN, ZHE; PAO, I-MING; CHEN, YINGCONG; LU, XIN; KONG, SARAH AYE
To: ADOBE INC.
Reel/Frame 056905/0081 →
Continuity (2)
Division 15928706 · Mar 22, 2018
Related Publication 20210350504A1 · Nov 11, 2021
References Cited (18)
US 6816847B1 · Toyama · 2004 [cited by examiner]
US 10147216B1 · Miao · 2018 [cited by examiner]
US 10540757B1 · Bouhnik et al. · 2020 [cited by applicant]
US 20140354768A1 · Mei et al. · 2014 [cited by applicant]
US 20170294010A1 · Shen et al. · 2017 [cited by applicant]
US 20190108621A1 · Condorovici · 2019 [cited by examiner]
US 20190147582A1 · Lee et al. · 2019 [cited by applicant]
US 20190251674A1 · Chang et al. · 2019 [cited by applicant]
Choe, G.—“RANUS RGB and NIR Urban Scene Dataset for Deep Scene Parsing”—IEEE Robotics—Jan. 2018—pp. 1-8 (Year: 2018). [cited by examiner]
Agrawal, M., and Sawhney, K., “Exploring Convolutional Neural Networks for Automatic Image Colorization”, pp. 1-9 (2017). [cited by applicant]
Cao, Y., et al., “Unsupervised Diverse Colorization via Generative Adversarial Networks”, arXiv, pp. 1-16 (Jul. 2017). [cited by applicant]
Deng, Y., et al., “Aesthetic-Driven Image Enhancement by Adversarial Learning”, arXiv:1707.05251, pp. 1-16 (2017). [cited by applicant]
Fang, H., and Zhang, M., “Creatism: A deep-learning photographer capable of creating professional work”, arXiv:1707.03491, pp. 1-14 (2017). [cited by applicant]
Hensman, P., and Aizawa, K., “cGAN-based Manga Colorization Using a Single Training Image”, arXiv:1706.0691, pp. 1-8 (Jun. 2017). [cited by applicant]
Isola, P., et al., “Image-to-Image Translation with Conditional Adversarial Networks”, arXiv:1611.07004, pp. 1-17 (2017). [cited by applicant]
Johnson, J., et al., “Perceptual Losses for Real-Time Style Transfer and Super-Resolution”, In European Conference on Computer Vision, pp. 1-17 (Oct. 2016). [cited by applicant]
Kong, S., et al., “Photo Aesthetics Ranking Network with Attributes and Content Adaptation”, In European Conference on Computer Vision, pp. 1-16 (Oct. 2016). [cited by applicant]
Zhao, H., et al., “Pyramid scene parsing network”, In IEEE Conference on Computer Vision and Pattern Recognition 9 (CVPR), pp. 2881-2890 (Jul. 2017). [cited by applicant]