IP Library › Granted Patent US 12,198,308
Granted Patent B2
US 12,198,308 · App. 18/369,810 · Granted Jan 14, 2025

Machine learning based image processing techniques

Inventors: Clarence Chui (Los Altos Hills, CA); Manu Parmar (Sunnyvale, CA)
Assignee: Outward, Inc.
G06T5/70G06F18/214G06F18/217G06N20/00G06T7/40G06T7/60G06T15/06G06T19/20G06V10/774G06V10/776G06V20/10G06V20/64G06N3/045G06T2207/20081G06T2219/2024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,308
App. No.
18/369,810
Granted
Jan 14, 2025
Kind
B2
Abstract

A machine learning based image processing architecture and associated applications are disclosed herein. In some embodiments, a machine learning framework is trained to learn low level image attributes such as object/scene types, geometries, placements, materials and textures, camera characteristics, lighting characteristics, contrast, noise statistics, etc. Thereafter, the machine learning framework may be employed to detect such attributes in other images and process the images at the attribute level.

Claims (33)

1. A method, comprising:

detecting a set of one or more attributes of an input image using a machine learning framework, wherein the machine learning framework is trained at least in part on a set of training images comprising a prescribed scene type to which the input image belongs and wherein training images comprising the set of training images are labeled with corresponding sets of labels from which attributes associated with the prescribed scene type are learned by the machine learning framework; and

generating an output image comprising a modified version of the input image by modifying at least a subset of the detected set of attributes.

2. The method of claim 1 , wherein the prescribed scene type comprises one or more of a constrained set of objects.

3. The method of claim 1 , wherein the set of training images includes images comprising different combinations of objects and object arrangements, camera configurations, lighting types and locations, and materials and textures.

4. The method of claim 1 , wherein a training image of the set of training images is rendered using one or more three-dimensional models, is captured by an imaging or a scanning device, or is generated from one or more other existing images.

5. The method of claim 1 , wherein at least the subset of the detected set of attributes is associated with an aesthetic.

6. The method of claim 5 , wherein the output image comprises a different aesthetic than the input image.

7. The method of claim 1 , wherein at least the subset of the detected set of attributes is associated with a style.

8. The method of claim 7 , wherein the output image comprises a restyled version of the input image.

9. The method of claim 1 , wherein at least the subset of the detected set of attributes is associated with an object in the input image.

10. The method of claim 9 , wherein the object in the input image is replaced by a different object in the output image.

11. The method of claim 1 , wherein at least the subset of the detected set of attributes is associated with lighting.

12. The method of claim 11 , wherein the output image comprises a relit version of the input image.

13. The method of claim 1 , wherein at least the subset of the detected set of attributes is associated with noise.

14. The method of claim 13 , wherein the output image comprises a denoised version of the input image.

15. The method of claim 1 , further comprising labeling or tagging the output image with a corresponding set of attributes.

16. The method of claim 1 , wherein the detected set of attributes comprises one or more attributes associated with object or scene types, geometries, placements, materials, textures, camera characteristics, lighting characteristics, noise statistics, and contrast.

17. The method of claim 1 , wherein a corresponding set of labels of a training image of the set of training images comprises one or more scene-based labels, image-based labels, or both.

18. The method of claim 1 , wherein a training image of the set of training images is labeled with ground truth data associated with generating the training image.

19. The method of claim 1 , wherein the output image comprises a photograph or a photorealistic rendering.

20. The method of claim 1 , wherein the output image comprises a video frame.

21. The method of claim 1 , wherein the prescribed scene type is associated with a retailer or a brand.

22. The method of claim 1 , wherein the prescribed scene type is associated with an animation or a video sequence.

23. The method of claim 1 , wherein the machine learning framework comprises a deep neural network, a convolutional neural network, or both.

24. A system, comprising:

a processor configured to:

detect a set of one or more attributes of an input image using a machine learning framework, wherein the machine learning framework is trained at least in part on a set of training images comprising a prescribed scene type to which the input image belongs and wherein training images comprising the set of training images are labeled with corresponding sets of labels from which attributes associated with the prescribed scene type are learned by the machine learning framework; and

generate an output image comprising a modified version of the input image by modifying at least a subset of the detected set of attributes; and

a memory coupled to the processor and configured to provide the processor with instructions.

25. A computer program product embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

detecting a set of one or more attributes of an input image using a machine learning framework, wherein the machine learning framework is trained at least in part on a set of training images comprising a prescribed scene type to which the input image belongs and wherein training images comprising the set of training images are labeled with corresponding sets of labels from which attributes associated with the prescribed scene type are learned by the machine learning framework; and

generating an output image comprising a modified version of the input image by modifying at least a subset of the detected set of attributes.

Continuity (5)
Continuation 17870830 · Jul 22, 2022
Continuation 17131586 · Dec 22, 2020
Continuation 16056110 · Aug 6, 2018
Provisional Application 62541603 · Aug 4, 2017
Related Publication 20240005456A1 · Jan 4, 2024
References Cited (87)
US 5526446A · Adelson · 1996 [cited by applicant]
US 6249610B1 · Matsumoto · 2001 [cited by applicant]
US 6353816B1 · Tsukimoto · 2002 [cited by applicant]
US 6816621B1 · Handley · 2004 [cited by applicant]
US 7573508B1 · Kondo · 2009 [cited by applicant]
US 8712184B1 · Liao · 2014 [cited by applicant]
US 9940753B1 · Grundhöfer · 2018 [cited by applicant]
US 10388059B2 · Luebke · 2019 [cited by applicant]
US 10451900B2 · Fonte · 2019 [cited by applicant]
US 10453165B1 · Kostov · 2019 [cited by applicant]
US 10496104B1 · Liu · 2019 [cited by applicant]
US 10567641B1 · Rueckner · 2020 [cited by applicant]
US 10783549B2 · Sinha · 2020 [cited by applicant]
US 20020037101A1 · Aihara · 2002 [cited by applicant]
US 20020146179A1 · Handley · 2002 [cited by applicant]
US 20030039401A1 · Gallagher · 2003 [cited by applicant]
US 20030187354A1 · Rust · 2003 [cited by applicant]
US 20050060142A1 · Visser · 2005 [cited by applicant]
US 20050083556A1 · Carlson · 2005 [cited by applicant]
US 20050147291A1 · Huang · 2005 [cited by applicant]
US 20050147292A1 · Huang · 2005 [cited by applicant]
US 20080158400A1 · Reyneri · 2008 [cited by applicant]
US 20100066868A1 · Shohara · 2010 [cited by applicant]
US 20100111436A1 · Jung · 2010 [cited by applicant]
US 20100189358A1 · Kaneda · 2010 [cited by applicant]
US 20100246991A1 · Naito · 2010 [cited by applicant]
US 20110137837A1 · Kobayashi · 2011 [cited by applicant]
US 20120195378A1 · Zheng · 2012 [cited by applicant]
US 20130051519A1 · Yang · 2013 [cited by applicant]
US 20130120354A1 · Falco, Jr. · 2013 [cited by applicant]
US 20130128056A1 · Chuang · 2013 [cited by applicant]
US 20130156297A1 · Shotton · 2013 [cited by applicant]
US 20130182779A1 · Lim · 2013 [cited by applicant]
US 20130235067A1 · Cherna · 2013 [cited by applicant]
US 20140043436A1 · Bell · 2014 [cited by applicant]
US 20140104450A1 · Cox · 2014 [cited by applicant]
US 20140133775A1 · Wang · 2014 [cited by applicant]
US 20140136137A1 · Tarshish-Shapir · 2014 [cited by applicant]
US 20150055085A1 · Fonte · 2015 [cited by applicant]
US 20150172677A1 · Norkin · 2015 [cited by applicant]
US 20150189186A1 · Fahn · 2015 [cited by applicant]
US 20150228110A1 · Hecht · 2015 [cited by applicant]
US 20150287172A1 · Choudhury · 2015 [cited by applicant]
US 20150294193A1 · Tate · 2015 [cited by applicant]
US 20150296152A1 · Fanello · 2015 [cited by applicant]
US 20150310602A1 · Lee · 2015 [cited by applicant]
US 20150338722A1 · Bonnier · 2015 [cited by applicant]
US 20150378325A1 · El Zur · 2015 [cited by applicant]
US 20160171753A1 · Park · 2016 [cited by applicant]
US 20160321523A1 · Sen · 2016 [cited by applicant]
US 20160373743A1 · Zhao · 2016 [cited by applicant]
US 20170097948A1 · Kerr · 2017 [cited by examiner]
US 20170098152A1 · Kerr · 2017 [cited by applicant]
US 20170103512A1 · Mailhe · 2017 [cited by applicant]
US 20170108569A1 · Harvey · 2017 [cited by applicant]
US 20170148158A1 · Najarian · 2017 [cited by applicant]
US 20170150180A1 · Lin · 2017 [cited by applicant]
US 20170193332A1 · Kim · 2017 [cited by applicant]
US 20170294010A1 · Shen · 2017 [cited by applicant]
US 20170300785A1 · Merhav · 2017 [cited by applicant]
US 20170337657A1 · Cornell · 2017 [cited by applicant]
US 20170344807A1 · Jillela · 2017 [cited by applicant]
US 20180018970A1 · Heyl · 2018 [cited by applicant]
US 20180089583A1 · Iyer · 2018 [cited by applicant]
US 20180114096A1 · Sen · 2018 [cited by applicant]
US 20180121762A1 · Han · 2018 [cited by applicant]
US 20180147015A1 · She · 2018 [cited by applicant]
US 20180286037A1 · Zaharchuk · 2018 [cited by applicant]
US 20180293496A1 · Vogels · 2018 [cited by applicant]
US 20180293710A1 · Meyer · 2018 [cited by applicant]
US 20180300865A1 · Weiss · 2018 [cited by applicant]
US 20180315172A1 · Smirnov · 2018 [cited by applicant]
US 20180316918A1 · Drugeon · 2018 [cited by applicant]
US 20180365856A1 · Parasnis · 2018 [cited by applicant]
US 20190026884A1 · Huang · 2019 [cited by applicant]
US 20190026956A1 · Gausebeck · 2019 [cited by applicant]
US 20190205946A1 · Mellina · 2019 [cited by applicant]
US 20190244011A1 · Liu · 2019 [cited by applicant]
US 20190261945A1 · Funka-Lea · 2019 [cited by applicant]
US 20190311199A1 · Mukherjee · 2019 [cited by applicant]
Baharul et al., A Survey of Aesthetics-Driven Image Recomposition, Multimedia Tools and Applications, vol. 76, No. 7, pp. 9517-9542, May 13, 2016. [cited by applicant]
Cheng et al., “ImageSpirit”, ACM Transactions on Graphics, ACM, NY, US, vol. 34, No. 1, Dec. 29, 2014, pp. 1-11. [cited by applicant]
Laffont et al., Transient Attributes for High-Level Understanding and Editing of Outdoor Scenes, ACM Transactions on Graphics, vol. 33, No. 4, Article 149, Jul. 27, 2014. [cited by applicant]
Sarkar et al., “Trained 3D Models for CNN based Object Recognition”, Proceedings of the 12th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Mar. 1, 2017, pp. 13… [cited by applicant]
Yu et al., “Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images”, arxiv.org, Cornell University Library, Ithaca, NY, Dec. 19, 2016. [cited by applicant]
Zhou et al., GeneGAN: Learning Object Transfiguration and Attribute Subspace from Unpaired Data, Student, Prof, Collaborator: BMVC Author Guidelines, pp. 1-13, May 14, 2017. [cited by applicant]
Innamorati et al., Decomposing Single Images for Layered Photo Retouching, Computer Graphics Forum, Journal of the European Association for Computer Graphics, vol. 36, No. 4, Jul. 5, 2017, pp. 15-25. [cited by applicant]