IP Library › Granted Patent US 12,524,955
Granted Patent B2
US 12,524,955 · App. 18/639,771 · Granted Jan 13, 2026

Arbitrary view generation

Inventors: Clarence Chui (Los Altos Hills, CA); Manu Parmar (Sunnyvale, CA); Amogh Subbakrishna Adishesha (State College, PA); Harshul Gupta (La Jolla, CA); Avinash Venkata Uppuluri (Sunnyvale, CA)
Assignee: Outward, Inc.
G06T15/205G06F16/58G06T5/50G06T7/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,955
App. No.
18/639,771
Granted
Jan 13, 2026
Kind
B2
Abstract

A machine learning based image processing and generation framework is disclosed. In some embodiments, depth values of an object or asset in a received input image are at least in part determined using a machine learning based framework that is constrained to a known prescribed environment. Determined depth values facilitate generation of other views of the object or asset.

Claims (65)

1 . A method, comprising:

receiving an input image of an object or asset, wherein the input image comprises a stereo pair comprising corresponding left and right images;

determining depth values of the object or asset in the input image based at least in part on the left and right stereo pair comprising the input image; and

generating an output image of the object or asset comprising a prescribed perspective that is different than the input image perspective at least in part by performing a perspective transformation of the input image using the determined depth values.

2 . The method of claim 1 , wherein the input image of the object or asset comprises a photograph of the object or asset captured in a known physical environment.

3 . The method of claim 1 , wherein the input image of the object or asset comprises a photograph of the object or asset captured in a prescribed imaging apparatus.

4 . The method of claim 1 , wherein the left and right images comprising the stereo pair are captured by corresponding left and right cameras.

5 . The method of claim 1 , wherein determining depth values comprises one or both of determining a depth estimate and refining the determined depth estimate.

6 . The method of claim 1 , wherein depth values are determined on a per pixel basis.

7 . The method of claim 1 , wherein depth values are determined at least in part using a neural network.

8 . The method of claim 1 , further comprising removing a background of the input image.

9 . The method of claim 1 , wherein the prescribed perspective comprises any desired or requested camera view of the object or asset.

10 . The method of claim 1 , wherein the prescribed perspective comprises an orthographic view.

11 . The method of claim 1 , wherein performing the perspective transformation comprises one or both of determining a perspective transformation estimate and refining the determined perspective transformation estimate.

12 . The method of claim 1 , wherein the perspective transformation is at least in part determined from a mathematical transformation.

13 . The method of claim 1 , wherein the perspective transformation is at least in part determined from a neural network.

14 . The method of claim 1 , wherein determining depth values of the object or asset in the input image is based at least in part on using a machine learning based framework.

15 . The method of claim 14 , wherein the machine learning based framework is constrained to one or more textures.

16 . The method of claim 14 , wherein the machine learning based framework is constrained to a prescribed scene type to which the input image belongs.

17 . The method of claim 14 , wherein the machine learning based framework is constrained to a prescribed environment.

18 . The method of claim 17 , wherein the prescribed environment comprises a physical environment in which the input image is photographed and a model environment that simulates the physical environment for training datasets of the machine learning based framework.

19 . A system, comprising:

a processor configured to:

receive an input image of an object or asset, wherein the input image comprises a stereo pair comprising corresponding left and right images;

determine depth values of the object or asset in the input image based at least in part on the left and right stereo pair comprising the input image; and

generate an output image of the object or asset comprising a prescribed perspective that is different than the input image perspective at least in part by performing a perspective transformation of the input image using the determined depth values; and

a memory coupled to the processor and configured to provide the processor with instructions.

20 . The system of claim 19 , wherein the input image of the object or asset comprises a photograph of the object or asset captured in a known physical environment.

21 . The system of claim 19 , wherein the input image of the object or asset comprises a photograph of the object or asset captured in a prescribed imaging apparatus.

22 . The system of claim 19 , wherein the left and right images comprising the stereo pair are captured by corresponding left and right cameras.

23 . The system of claim 19 , wherein to determine depth values comprises one or both of to determine a depth estimate and to refine the determined depth estimate.

24 . The system of claim 19 , wherein depth values are determined on a per pixel basis.

25 . The system of claim 19 , wherein depth values are determined at least in part using a neural network.

26 . The system of claim 19 , wherein the processor is further configured to remove a background of the input image.

27 . The system of claim 19 , wherein the prescribed perspective comprises any desired or requested camera view of the object or asset.

28 . The system of claim 19 , wherein the prescribed perspective comprises an orthographic view.

29 . The system of claim 19 , wherein performing the perspective transformation comprises one or both of determining a perspective transformation estimate and refining the determined perspective transformation estimate.

30 . The system of claim 19 , wherein the perspective transformation is at least in part determined from a mathematical transformation.

31 . The system of claim 19 , wherein the perspective transformation is at least in part determined from a neural network.

32 . The system of claim 19 , wherein to determine depth values of the object or asset in the input image is based at least in part on using a machine learning based framework.

33 . The system of claim 32 , wherein the machine learning based framework is constrained to one or more textures.

34 . The system of claim 32 , wherein the machine learning based framework is constrained to a prescribed scene type to which the input image belongs.

35 . The system of claim 32 , wherein the machine learning based framework is constrained to a prescribed environment.

36 . The system of claim 35 , wherein the prescribed environment comprises a physical environment in which the input image is photographed and a model environment that simulates the physical environment for training datasets of the machine learning based framework.

37 . A computer program product embodied in a non-transitory computer readable storage medium comprising computer instructions which when executed cause a computer to:

receive an input image of an object or asset, wherein the input image comprises a stereo pair comprising corresponding left and right images;

determine depth values of the object or asset in the input image based at least in part on the left and right stereo pair comprising the input image; and

generate an output image of the object or asset comprising a prescribed perspective that is different than the input image perspective at least in part by performing a perspective transformation of the input image using the determined depth values.

38 . The computer program product of claim 37 , wherein the input image of the object or asset comprises a photograph of the object or asset captured in a known physical environment.

39 . The computer program product of claim 37 , wherein the input image of the object or asset comprises a photograph of the object or asset captured in a prescribed imaging apparatus.

40 . The computer program product of claim 37 , wherein the left and right images comprising the stereo pair are captured by corresponding left and right cameras.

41 . The computer program product of claim 37 , wherein determining depth values comprises one or both of determining a depth estimate and refining the determined depth estimate.

42 . The computer program product of claim 37 , wherein depth values are determined on a per pixel basis.

43 . The computer program product of claim 37 , wherein depth values are determined at least in part using a neural network.

44 . The computer program product of claim 37 , further comprising computer instructions for removing a background of the input image.

45 . The computer program product of claim 37 , wherein the prescribed perspective comprises any desired or requested camera view of the object or asset.

46 . The computer program product of claim 37 , wherein the prescribed perspective comprises an orthographic view.

47 . The computer program product of claim 37 , wherein performing the perspective transformation comprises one or both of determining a perspective transformation estimate and refining the determined perspective transformation estimate.

48 . The computer program product of claim 20 , wherein the perspective transformation is at least in part determined from a mathematical transformation.

49 . The computer program product of claim 20 , wherein the perspective transformation is at least in part determined from a neural network.

50 . The computer program product of claim 37 , wherein determining depth values of the object or asset in the input image is based at least in part on using a machine learning based framework.

51 . The computer program product of claim 50 , wherein the machine learning based framework is constrained to one or more textures.

52 . The computer program product of claim 50 , wherein the machine learning based framework is constrained to a prescribed scene type to which the input image belongs.

53 . The computer program product of claim 50 , wherein the machine learning based framework is constrained to a prescribed environment.

54 . The computer program product of claim 53 , wherein the prescribed environment comprises a physical environment in which the input image is photographed and a model environment that simulates the physical environment for training datasets of the machine learning based framework.

Continuity (9)
Continuation 17090794 · Nov 5, 2020
Continuation In Part 16523888 · Jul 26, 2019
Continuation In Part 16181607 · Nov 6, 2018
Continuation 15721426 · Sep 29, 2017
Continuation In Part 15081553 · Mar 25, 2016
Provisional Application 62933261 · Nov 8, 2019
Provisional Application 62933258 · Nov 8, 2019
Provisional Application 62541607 · Aug 4, 2017
Related Publication 20240346746A1 · Oct 17, 2024
References Cited (68)
US 6222947B1 · Koba · 2001 [cited by applicant]
US 6362822B1 · Randel · 2002 [cited by applicant]
US 6377257B1 · Borrel · 2002 [cited by applicant]
US 8655052B2 · Spooner · 2014 [cited by applicant]
US 9407904B2 · Sandrew · 2016 [cited by applicant]
US 9996914B2 · Chui · 2018 [cited by applicant]
US 10163249B2 · Chui · 2018 [cited by applicant]
US 10163250B2 · Chui · 2018 [cited by applicant]
US 10163251B2 · Chui · 2018 [cited by applicant]
US 10909749B2 · Chui · 2021 [cited by applicant]
US 20050018045A1 · Thomas · 2005 [cited by applicant]
US 20060280368A1 · Petrich · 2006 [cited by applicant]
US 20070242284A1 · Schalkwijk · 2007 [cited by applicant]
US 20080143715A1 · Moden · 2008 [cited by applicant]
US 20090028403A1 · Bar-Aviv · 2009 [cited by applicant]
US 20090276105A1 · Lacaze · 2009 [cited by applicant]
US 20110001826A1 · Hongo · 2011 [cited by applicant]
US 20120120240A1 · Muramatsu · 2012 [cited by applicant]
US 20120140027A1 · Curtis · 2012 [cited by applicant]
US 20120163672A1 · Mckinnon · 2012 [cited by applicant]
US 20120314937A1 · Kim · 2012 [cited by applicant]
US 20130100290A1 · Sato · 2013 [cited by applicant]
US 20130222369A1 · Huston · 2013 [cited by applicant]
US 20130259448A1 · Stankiewicz · 2013 [cited by applicant]
US 20140063061A1 · Reitan · 2014 [cited by applicant]
US 20140198182A1 · Ward · 2014 [cited by applicant]
US 20140254908A1 · Strommer · 2014 [cited by applicant]
US 20140267343A1 · Arcas · 2014 [cited by applicant]
US 20150015581A1 · Lininger · 2015 [cited by applicant]
US 20150169982A1 · Perry · 2015 [cited by applicant]
US 20150317822A1 · Haimovitch-Yogev · 2015 [cited by applicant]
US 20170103512A1 · Mailhe · 2017 [cited by applicant]
US 20170277979A1 · Allen · 2017 [cited by applicant]
US 20170278251A1 · Peeper · 2017 [cited by applicant]
US 20170304732A1 · Velic · 2017 [cited by applicant]
US 20170334066A1 · Levine · 2017 [cited by applicant]
US 20170372193A1 · Mailhe · 2017 [cited by applicant]
US 20180012330A1 · Holzer · 2018 [cited by applicant]
US 20190080506A1 · Chui · 2019 [cited by applicant]
US 20190325621A1 · Wang · 2019 [cited by applicant]
CN 101246599 · 2008 [cited by applicant]
CN 101283375 · 2008 [cited by applicant]
CN 101281640 · 2012 [cited by applicant]
CN 103152518 · 2013 [cited by applicant]
CN 103179339 · 2013 [cited by applicant]
CN 103828359 · 2014 [cited by applicant]
CN 203870604 · 2014 [cited by applicant]
JP 2000137815 · 2000 [cited by applicant]
JP 2003187261 · 2003 [cited by applicant]
JP 2004287517 · 2004 [cited by applicant]
JP 2009211335 · 2009 [cited by applicant]
JP 2010140097 · 2010 [cited by applicant]
JP 2017212593 · 2017 [cited by applicant]
JP 2018081672 · 2018 [cited by applicant]
JP 2019516202 · 2019 [cited by applicant]
KR 20060029140 · 2006 [cited by applicant]
KR 20120137295 · 2012 [cited by applicant]
KR 20140021766 · 2014 [cited by applicant]
KR 20190094254 · 2019 [cited by applicant]
WO 2018197984 · 2018 [cited by applicant]
WO 2019167453 · 2019 [cited by applicant]
Bouwmans et al.: “Deep neural network concepts for background subtraction: A systematic review and comparative evaluation”, Neural Networks, Elsevier Science Publishers, Barking, GB, vol. 117, May 15, 2019 (May 15, 2019… [cited by applicant]
Daniel Scharstein. “A Survey of Image-Based Rendering and Stereo”. In: “View Synthesis Using Stereo Vision”, Lecture Notes in Computer Science, vol. 1583, Jan. 1, 1999, pp. 23-39. [cited by applicant]
Inamoto et al. “Virtual Viewpoint Replay for a Soccer Match by View Interpolation from Multiple Cameras”. IEEE Transactions on Multimedia, vol. 9 No. 6, Oct. 1, 2007, pp. 1155-1166. [cited by applicant]
Ledig et al., “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”, Sep. 15, 2016 (Sep. 15, 2016), Retrieved from the Internet: URL:https://arxiv.org/pdf/1609.04802.pdf. [cited by applicant]
Stamatios Lefkimmiatis: “Non-local Color Image Denoising with Convolutional Neural Networks”, IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 1, 2017 (Jul. 1, 2017), pp. 5882-5891, ISBN: 978… [cited by applicant]
Sun et al. “An overview of free viewpoint Depth-Image-Based Rendering (DIBR).” Proceedings of the Second APSIPA Annual Summit and Conference. Dec. 14, 2010, pp. 1-8. [cited by applicant]
Sheng et al., “Virtue Plane Mapping: A Method of Rendering into Depth Images”, Journal of Software, vol. 19, No. 7, Jul. 2008, pp. 1806-1816. [cited by applicant]