IP Library › Granted Patent US 12,417,558
Granted Patent B2
US 12,417,558 · App. 17/556,716 · Granted Sep 16, 2025

Generating stylized digital images via drawing stroke optimization utilizing a multi-stroke neural network

Inventors: Aaron Phillip Hertzmann (San Francisco, CA); Manuel Rodriguez Ladron de Guevara (Pittsburgh, PA); Matthew Fisher (San Francisco, CA)
Assignee: Adobe Inc.
G06T11/00G06N3/045G06N3/084G06V10/40G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,558
App. No.
17/556,716
Granted
Sep 16, 2025
Kind
B2
Abstract

Methods, systems, and non-transitory computer readable storage media are disclosed for utilizing a multi-stroke neural network for modifying a digital image via a plurality of generated stroke parameters in a single pass of the neural network. Specifically, the disclosed system utilizes an encoder neural network to generate an encoding of a digital image. The disclosed system then utilizes a decoder neural network that generates a sequence of stroke parameters for digital drawing strokes from the encoding in a single pass of the encoder neural network and decoder neural network. Additionally, the disclosed system utilizes a renderer neural network to render the digital drawing strokes on a digital canvas according to the sequence of stroke parameters. In additional embodiments, the disclosed system utilizes a balance of loss functions to learn parameters of the multi-stroke neural network to generate stroke parameters according to various rendering styles.

Claims (64)

1. A method comprising:

generating, utilizing an encoder neural network, an encoding comprising feature maps from a digital image;

determining, based on a user input via a graphical user interface of a client device, a selection of a particular rendering style for rendering digital drawing strokes;

generating, from the encoding comprising the feature maps utilizing a decoder neural network trained to generate stroke parameters according to the particular rendering style, a plurality of feature representations that define a sequence of stroke parameters for a plurality of digital drawing strokes; and

sequentially rendering the plurality of digital drawing strokes within a digital canvas according to the sequence of stroke parameters by utilizing a differentiable renderer neural network to sequentially convert the plurality of feature representations into the plurality of digital drawing strokes on a plurality of sequential instances of the digital canvas.

2. The method as recited in claim 1 , wherein generating the plurality of feature representations comprises generating the plurality of feature representations in a single pass of the encoding via the decoder neural network.

3. The method as recited in claim 1 , wherein generating the plurality of feature representations comprises generating a vector comprising the plurality of feature representations according to a number and an order of stroke parameters for the plurality of digital drawing strokes.

4. The method as recited in claim 3 , wherein sequentially rendering the plurality of digital drawing strokes comprises rendering the plurality of digital drawing strokes according to the number and the order of the stroke parameters from the vector comprising the plurality of feature representations.

5. The method as recited in claim 1 , wherein sequentially rendering the plurality of digital drawing strokes comprises:

converting a first feature representation of the plurality of feature representations into a first digital drawing stroke on a first instance of the digital canvas; and

converting a second feature representation of the plurality of feature representations into a second digital drawing stroke on a second instance of the digital canvas, the second instance of the digital canvas comprising the first digital drawing stroke and the second digital drawing stroke.

6. The method as recited in claim 1 , further comprising:

determining a loss based on a plurality of differences between the digital image and a plurality of instances of the digital canvas corresponding to rendering the plurality of digital drawing strokes within the digital canvas; and

modifying parameters of one or more of the encoder neural network or the decoder neural network based on the loss.

7. The method as recited in claim 1 , further comprising:

determining a loss based on a difference between the digital image and a final instance of the digital canvas after rendering the plurality of digital drawing strokes within the digital canvas; and

modifying parameters of one or more of the encoder neural network or the decoder neural network based on the difference.

8. The method as recited in claim 1 , further comprising:

determining, from a plurality of instances of the decoder neural network corresponding to a plurality of rendering styles, an instance of the decoder neural network for the particular rendering style of the selection from the user input.

9. The method as recited in claim 1 , further comprising:

determining a first loss based on a plurality of differences between the digital image and a plurality of instances of the digital canvas based on rendering the plurality of digital drawing strokes within the digital canvas;

determining a second loss based on a difference between the digital image and a final instance of the digital canvas after rendering the plurality of digital drawing strokes within the digital canvas;

determining a third loss based on a difference between the digital image and an intermediate instance of the digital canvas after rendering a subset of the plurality of digital drawing strokes within the digital canvas;

determining a plurality of weights for the first loss, the second loss, and the third loss in connection with a rendering style for rendering the plurality of digital drawing strokes within the digital canvas; and

modifying parameters of one or more of the encoder neural network or the decoder neural network based on the first loss, the second loss, and the third loss weighted according to the plurality of weights.

10. A system comprising:

one or more computer memory devices comprising a digital image, an encoder neural network, a decoder neural network having a stack of fully-connected neural network layers and trained to generate stroke parameters according to a particular rendering style, and a differentiable renderer neural network; and

one or more processors configured to cause the system to:

generate, utilizing the encoder neural network, an encoding comprising feature maps from the digital image;

determine, based on a user input via a graphical user interface of a client device, a selection of the particular rendering style for rendering digital drawing strokes;

generate, from the encoding comprising the feature maps via a single pass of the decoder neural network having the stack of fully-connected neural network layers, a plurality of feature representations that define a sequence of stroke parameters for a plurality of digital drawing strokes; and

sequentially render the plurality of digital drawing strokes within a digital canvas according to the sequence of stroke parameters by utilizing the differentiable renderer neural network to sequentially convert the plurality of feature representations into the plurality of digital drawing strokes on a plurality of sequential instances of the digital canvas.

11. The system as recited in claim 10 , wherein the one or more processors are further configured to:

generate a vector comprising the plurality of feature representations according to a number of digital drawing strokes for the sequence of stroke parameters for the plurality of digital drawing strokes; and

render the plurality of digital drawing strokes according to the vector comprising the plurality of feature representations.

12. The system as recited in claim 11 , wherein the one or more processors are further configured to:

generate the vector comprising a first feature representation for one or more first stroke parameters and a second feature representation for one or more second stroke parameters ordered after the first feature representation;

render a first digital drawing stroke within the digital canvas according to the one or more first stroke parameters; and

render a second digital drawing stroke within the digital canvas according to the one or more second stroke parameters after rendering the first digital drawing stroke.

13. The system as recited in claim 10 , wherein the one or more processors are further configured to:

determine a plurality of losses based on differences between the digital image and instances of the digital canvas corresponding to rendering a plurality of digital drawing strokes within the digital canvas;

determine weights of the plurality of losses according to a rendering style for rendering the plurality of digital drawing strokes within the digital canvas; and

modify parameters of the encoder neural network and the decoder neural network utilizing backpropagation of the plurality of losses according to the weights of the plurality of losses.

14. The system as recited in claim 13 , wherein the one or more processors are further configured to:

generate a plurality of instances of the decoder neural network corresponding to a plurality of rendering styles; and

generate the plurality of feature representations utilizing an instance of the decoder neural network corresponding to a selected rendering style.

15. The system as recited in claim 13 , wherein the one or more processors are further configured to determine the plurality of losses by:

determining a first loss based on a plurality of differences between the digital image and a first plurality of instances of the digital canvas corresponding to rendering a plurality of digital drawing strokes of the plurality of digital drawing strokes within the digital canvas;

determining a second loss based on a difference between the digital image and a final instance of the digital canvas after rendering the plurality of digital drawing strokes within the digital canvas; and

determining a third loss based on a difference between the digital image and an intermediate instance of the digital canvas after rendering a subset of the plurality of digital drawing strokes within the digital canvas.

16. A non-transitory computer readable storage medium comprising instructions that, when executed by at least one processor, cause a computing device to:

generate, utilizing an encoder neural network, an encoding comprising feature maps from a digital image;

determine, based on a user input via a graphical user interface of a client device, a selection of a particular rendering style for rendering digital drawing strokes;

generate, from the encoding comprising the feature maps via a single pass of a decoder neural network comprising a long short-term memory neural network layer and trained to generate stroke parameters according to the particular rendering style, a plurality of feature representations that define a sequence of stroke parameters for a plurality of digital drawing strokes; and

sequentially render the plurality of digital drawing strokes within a digital canvas according to the sequence of stroke parameters by sequentially converting the plurality of feature representations into the plurality of digital drawing strokes on a plurality of sequential instances of the digital canvas.

17. The non-transitory computer readable storage medium as recited in claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate, via the single pass of the decoder neural network from the encoding comprising the feature maps, the plurality of feature representations comprising information indicating stroke widths, stroke colors, and stroke positions of the plurality of digital drawing strokes.

18. The non-transitory computer readable storage medium as recited in claim 17 , further comprising instructions that, when executed by the at least one processor, cause the computing device to render the plurality of digital drawing strokes according to the stroke widths, the stroke colors, and the stroke positions of the plurality of digital drawing strokes from the plurality of feature representations.

19. The non-transitory computer readable storage medium as recited in claim 18 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

render, within a first instance of the digital canvas, a first digital drawing stroke of the plurality of digital drawing strokes according to a first stroke width, a first stroke color, and a first stroke position based on a feature representation corresponding to the first digital drawing stroke; and

render, within a second instance of the digital canvas, a second digital drawing stroke of the plurality of digital drawing strokes according to a second stroke width, a second stroke color, and a second stroke position based on a feature representation corresponding to the second digital drawing stroke, the second instance of the digital canvas comprising the first digital drawing stroke and the second digital drawing stroke.

20. The non-transitory computer readable storage medium as recited in claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

determine one or more losses based on differences between the digital image and one or more instances of the digital canvas corresponding to rendering one or more digital drawing strokes within the digital canvas based on the plurality of feature representations;

determine one or more weights of the one or more losses according to a rendering style for rendering the plurality of digital drawing strokes within the digital canvas; and

modify parameters of the encoder neural network or the decoder neural network according to the one or more weights of the one or more losses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2021
From: HERTZMANN, AARON PHILLIP; LADRON DE GUEVARA, MANUEL RODRIGUEZ; FISHER, MATTHEW
To: ADOBE INC.
Reel/Frame 058436/0829 →
Continuity (1)
Related Publication 20230196630A1 · Jun 22, 2023
References Cited (43)
US 10643130B2 · Fidler · 2020 [cited by examiner]
US 20220156987A1 · Chandran · 2022 [cited by examiner]
US 20230074420A1 · Yin · 2023 [cited by examiner]
US 20240095972A1 · Zhang · 2024 [cited by examiner]
Reimann, Max, et al. Interactive Multi-Level Stroke Control for Neural Style Transfer. arXiv:2106.13787, arXiv, Jun. 25, 2021. arXiv.org, https://doi.org/10.48550/arXiv.2106.13787. (Year: 2021). [cited by examiner]
Mihai, Daniela, and Jonathon Hare. Differentiable Drawing and Sketching. arXiv:2103.16194, arXiv, Jul. 19, 2021. arXiv.org, https://doi.org/10.48550/arXiv.2103.16194. (Year: 2021). [cited by examiner]
Kotovenko, Dmytro, et al. “Rethinking style transfer: From pixels to parameterized brushstrokes.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021. (Year: 2021). [cited by examiner]
Xie, Ning, et al. “Stroke-based stylization by learning sequential drawing examples.” Journal of Visual Communication and Image Representation 51 (2018): 29-39. (Year: 2018). [cited by examiner]
Singh, Jaskirat, et al. “Intelli-paint: Towards developing human-like painting agents.” arXiv preprint arXiv:2112.08930 (2021). (Year: 2021). [cited by examiner]
Kerdreux, Thomas, Louis Thiry, and Erwan Kerdreux. “Interactive neural style transfer with artists.” arXiv preprint arXiv:2003.06659 (2020). (Year: 2020). [cited by examiner]
Pierre Benard and Aaron Hertzmann. Line drawings from 3D models. Foundations and Trends in Computer Grapics and Vision, 11(1-2), 2019. [cited by applicant]
D. Berio, S. Calinon, and F. Fol Leymarie. Learning dynamic graffiti strokes with a compliant robot. In Proc. IEEE/RSJ Intl Conf. on Intelligent Robots and Systems (IROS), Dae-jeon, Korea, Oct. 2016. [cited by applicant]
Yang Chen, Yu-Kun Lai, and Yong-Jin Liu. Cartoongan: Generative adversarial networks for photo cartoonization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 9465-9474, 2018. [cited by applicant]
Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, SM Ali Eslami, and Oriol Vinyals. Synthesizing programs for images using reinforced adversarial learning. In International Conference on Machine Learning, pp. 1666-1675. … [cited by applicant]
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. In Proc. Neural Information Processing Systems, 2014. [cited by applicant]
Stephane Grabli, Emmanuel Turquin, Fredo Durand, and Francois X Sillion. Programmable rendering of line drawing from 3d scenes. ACM Transactions on Graphics (TOG), 29(2):1-20, 2010. [cited by applicant]
David Ha and Douglas Eck. A neural representation of sketch drawings. arXiv preprint arXiv:1704.03477, 2017. [cited by applicant]
Paul Haeberli. Paint By Numbers: Abstract Image Representations. In Proc. SIGGRAPH, 1990. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. CoRR, abs/1512.03385, 2015. [cited by applicant]
Aaron Hertzmann. Painterly Rendering with Curved Brush Strokes of Multiple Sizes. In Proc. SIGGRAPH, 1998. [cited by applicant]
Aaron Hertzmann. Paint by relaxation. In Proc. CGI, 2001. [cited by applicant]
Aaron Hertzmann. A Survey of Stroke-Based Rendering. IEEE Computer Graphics & Applications, 23(4), 2003. [cited by applicant]
Zhewei Huang, Wen Heng, and Shuchang Zhou. Learning to paint with model-based deep reinforcement learning, 2019. [cited by applicant]
Biao Jia, Jonathan Brandt, Radomir Mech, Byungmoon Kim, and Dinesh Manocha. Lpaintb: Learning to paint from self-supervision. In Proc. Pacific Graphics, 2019. [cited by applicant]
Junhwan Kim and Fabio Pellacini. Jigsaw image mosaics. ACM Transactions on Graphics, 21(3):657-664, 2002. [cited by applicant]
Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. [cited by applicant]
Tzu-Mao Li, Michal Lukac, Michael Gharbi, and Jonathan Ragan-Kelley. Differentiable vector graphics rasterization for editing and learning. ACM Trans. Graph., 39(6), 2020. [cited by applicant]
Peter Litwinowicz. Processing Images and Video for an Impressionist Effect. In Proc. SIGGRAPH, 1997. [cited by applicant]
Songhua Liu, Tianwei Lin, Dongliang He, Fu Li, Ruifeng Deng, Xin Li, Errui Ding, and Hao Wang. Paint transformer: Feed forward neural painting with stroke prediction. In Proceedings of the IEEE/CVF International Confere… [cited by applicant]
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), Dec. 2015. [cited by applicant]
John F. J. Mellor, Eunbyung Park, Yaroslav Ganin, Igor Babuschkin, Tejas Kulkarni, Dan Rosenbaum, Andy Ballard, Theophane Weber, Oriol Vinyals, and S. M. Ali Eslami. Unsupervised doodling and painting with improved spir… [cited by applicant]
Haoran Mo, Edgar Simo-Serra, Chengying Gao, Changqing Zou, and Ruomei Wang. General virtual sketching frame-work for vector line art. ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH 2021), 40(4):51:1-51:14, 20… [cited by applicant]
Paul Rosin and John Collomosse. Image and Video-Based Artistic Stylisation. Springer, 2013. [cited by applicant]
Peter Schaldenbrand and Jean Oh. Content masked loss: Human-like brush stroke planning in a reinforcement learning painting agent. InProc. AAAI, 2021. [cited by applicant]
Adrian Secord. Weighted voronoi stippling. In Proceedings of the 2nd international symposium on Non-photorealistic animation and rendering, pp. 37-43, 2002. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition.arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
Jaskirat Singh and Liang Zheng. Combining semantic guidance and deep reinforcement learning for generating human level paintings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.… [cited by applicant]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InProc. CVPR, 2018. [cited by applicant]
Ningyuan Zheng, Yifan Jiang, and Dingjiang Huang. Strokenet: A neural painting environment. In International Conference on Learning Representations, 2019. [cited by applicant]
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Computer Vision (ICCV), 2017 IEEE International Conference on, 2017. [cited by applicant]
Zhengxia Zou, Tianyang Shi, Shuang Qiu, Yi Yuan, and Zhenwei Shi. Stylized neural painting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 15689-15698, Jun. 2021. [cited by applicant]
Tao Zhou, Chen Fang, Zhaowen Wang, Jimei Yang, Byungmoon Kim, Zhili Chen, Jonathan Brandt, and Demetri Terzopoulos. Learning to sketch with deep q networks and demonstrated strokes. arXiv preprint arXiv:1810.05977, 2018. [cited by applicant]
Ning Xie, Hirotaka Hachiya, and Masashi Sugiyama. Artist agent: A reinforcement learning approach to automatic stroke generation in oriental ink painting. IEICE Transactions on Information and Systems, 96(5):1134-1144, … [cited by applicant]