IP Library › Granted Patent US 12,731,309
Granted Patent B2
US 12,731,309 · App. 18/304,134 · Granted Sep 8, 2026

Generating scale fields indicating pixel-to-metric distances relationships in digital images via neural networks

Inventors: Yannick Hold-Geoffroy (Quebec City, CA); Jianming Zhang (Campbell, CA); Byeonguk Lee (Daejeon, KR)
Assignee: Adobe Inc.
G06T11/60G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,309
App. No.
18/304,134
Granted
Sep 8, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify two-dimensional images via scene-based editing using three-dimensional representations of the two-dimensional images. For instance, in one or more embodiments, the disclosed systems utilize three-dimensional representations of two-dimensional images to generate and modify shadows in the two-dimensional images according to various shadow maps. Additionally, the disclosed systems utilize three-dimensional representations of two-dimensional images to modify humans in the two-dimensional images. The disclosed systems also utilize three-dimensional representations of two-dimensional images to provide scene scale estimation via scale fields of the two-dimensional images. In some embodiments, the disclosed systems utilizes three-dimensional representations of two-dimensional images to generate and visualize 3D planar surfaces for modifying objects in two-dimensional images. The disclosed systems further use three-dimensional representations of two-dimensional images to customize focal points for the two-dimensional images.

Claims (64)

1 . A computer-implemented method comprising:

generating, by at least one processor utilizing one or more neural networks, a feature representation of a two-dimensional image;

generating, by the at least one processor utilizing the one or more neural networks and based on the feature representation, a scale field for the two-dimensional image comprising pixel-to-metric ratios between:

pixel distances from pixels to a horizon line in the two-dimensional image; and

metric distances from corresponding three-dimensional points to a three-dimensional horizon line in a three-dimensional space corresponding to the two-dimensional image; and

performing at least one of:

generating, by the at least one processor, a metric distance of content portrayed in the two-dimensional image utilizing the pixel-to-metric ratios of the scale field of the two-dimensional image; or

modifying, by the at least one processor, the two-dimensional image utilizing the pixel-to-metric ratios of the scale field of the two-dimensional image.

2 . The computer-implemented method of claim 1 , further comprising generating, utilizing the one or more neural networks and based on the feature representation, a plurality of ground-to-horizon vectors in the three-dimensional space according to the horizon line of the two-dimensional image projected into the three-dimensional space.

3 . The computer-implemented method of claim 2 , wherein generating the plurality of ground-to-horizon vectors comprises generating a ground-to-horizon vector indicating a distance and a direction from a three-dimensional point corresponding to a pixel of the two-dimensional image to the three-dimensional horizon line in the three-dimensional space.

4 . The computer-implemented method of claim 1 , wherein generating the scale field for the two-dimensional image comprises generating, for a pixel of the two-dimensional image, a pixel-to-metric ratio between a pixel distance in the two-dimensional image and a corresponding three-dimensional distance in the three-dimensional space relative to a camera height of the two-dimensional image.

5 . The computer-implemented method of claim 1 , wherein generating the metric distance of the content portrayed in the two-dimensional image comprises:

determining a pixel distance between a first pixel corresponding to the content and a second pixel corresponding to the content; and

generating the metric distance based on the pixel distance and the pixel-to-metric ratios between the pixel distances in the two-dimensional image and the metric distances in the three-dimensional space.

6 . The computer-implemented method of claim 5 , wherein generating the metric distance of the content portrayed in the two-dimensional image comprises:

determining a value of the scale field corresponding to the first pixel; and

converting the value of the scale field corresponding to the first pixel to the metric distance based on the pixel distance between the first pixel and the second pixel.

7 . The computer-implemented method of claim 1 , wherein modifying the two-dimensional image comprises:

determining a pixel position of an object placed within the two-dimensional image; and

determining a scale of the object based on the pixel position and the scale field.

8 . The computer-implemented method of claim 7 , wherein determining the scale of the object comprises:

determining an initial size of the object; and

inserting the object at the pixel position with a modified size based on a ratio indicated by a value from the scale field at the pixel position of the object.

9 . The computer-implemented method of claim 1 , further comprising learning parameters of the one or more neural networks by:

generating, for an additional two-dimensional image, estimated depth values for a plurality of pixels of the additional two-dimensional image projected to a corresponding three-dimensional space;

determining, for the additional two-dimensional image, an additional horizon line according to an estimated camera height of the additional two-dimensional image;

generating, for the additional two-dimensional image, a ground-truth scale field based on a plurality of ground-to-horizon vectors in the corresponding three-dimensional space according to the estimated depth values for the plurality of pixels and the additional horizon line; and

modifying parameters of the one or more neural networks based on the ground-truth scale field of the additional two-dimensional image.

10 . A non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating, utilizing one or more neural networks comprising parameters learned from a plurality of digital images with annotated horizon lines and ground-to-horizon vectors, a feature representation of a two-dimensional image;

generating, utilizing the one or more neural networks and based on the feature representation, a scale field for the two-dimensional image comprising a plurality of values indicating ratios of pixel distances relative to a camera height of the two-dimensional image; and

performing at least one of:

generating a metric distance of an object portrayed in the two-dimensional image according to the scale field of the two-dimensional image; or

modifying the two-dimensional image according to the scale field of the two-dimensional image.

11 . The non-transitory computer readable medium of claim 10 , wherein generating the scale field comprises generating, for a pixel of the two-dimensional image, a value representing a ratio between a pixel distance from the pixel to a horizon line of the two-dimensional image and a camera height of the two-dimensional image.

12 . The non-transitory computer readable medium of claim 10 , wherein modifying the two-dimensional image comprises inserting an object at a location of the two-dimensional image by:

determining a pixel corresponding to the location of the two-dimensional image;

determining a scaled size of the object based on a value from the scale field for the pixel corresponding to the location of the two-dimensional image; and

inserting the object at the location of the two-dimensional image according to the scaled size of the object.

13 . A system comprising:

one or more memory devices comprising a two-dimensional image; and

one or more processors configured to cause the system to:

generate, utilizing one or more neural networks, a feature representation of the two-dimensional image;

generate, utilizing the one or more neural networks and based on the feature representation, a scale field for the two-dimensional image comprising pixel-to-metric ratios between:

pixel distances from pixels to a horizon line in the two-dimensional image; and

metric distances from corresponding three-dimensional points to a three-dimensional horizon line in a three-dimensional space corresponding to the two-dimensional image; and

perform at least one of:

generating a metric distance of content portrayed in the two-dimensional image utilizing the pixel-to-metric ratios of the scale field of the two-dimensional image; or

modifying the two-dimensional image utilizing the pixel-to-metric ratios of the scale field of the two-dimensional image.

14 . The system of claim 13 , wherein the one or more processors are further configured to cause the system to generate, utilizing the one or more neural networks and based on the feature representation, a plurality of ground-to-horizon vectors in the three-dimensional space according to the horizon line of the two-dimensional image projected into the three-dimensional space.

15 . The system of claim 14 , wherein the one or more processors are further configured to cause the system to generate the plurality of ground-to-horizon vectors by generating a ground-to-horizon vector indicating a distance and a direction from a three-dimensional point corresponding to a pixel of the two-dimensional image to the three-dimensional horizon line in the three-dimensional space.

16 . The system of claim 13 , wherein the one or more processors are further configured to cause the system to generate the scale field for the two-dimensional image by generating, for a pixel of the two-dimensional image, a pixel-to-metric ratio between a pixel distance in the two-dimensional image and a corresponding three-dimensional distance in the three-dimensional space relative to a camera height of the two-dimensional image.

17 . The system of claim 13 , wherein the one or more processors are further configured to cause the system to generate the metric distance of the content portrayed in the two-dimensional image by:

determining a pixel distance between a first pixel corresponding to the content and a second pixel corresponding to the content; and

generating the metric distance based on the pixel distance and the pixel-to-metric ratios between the pixel distances in the two-dimensional image and the metric distances in the three-dimensional space.

18 . The system of claim 17 , wherein the one or more processors are further configured to cause the system to generate the metric distance of the content portrayed in the two-dimensional image by:

determining a value of the scale field corresponding to the first pixel; and

converting the value of the scale field corresponding to the first pixel to the metric distance based on the pixel distance between the first pixel and the second pixel.

19 . The system of claim 13 , wherein the one or more processors are further configured to cause the system to modify the two-dimensional image by:

determining a pixel position of an object placed within the two-dimensional image; and

determining a scale of the object based on the pixel position and the scale field.

20 . The system of claim 19 , wherein the one or more processors are further configured to cause the system to determine the scale of the object by:

determining an initial size of the object; and

inserting the object at the pixel position with a modified size based on a ratio indicated by a value from the scale field at the pixel position of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2023
From: ZHANG, JIANMING; LEE, BYEONGUK; HOLD-GEOFFROY, YANNICK
To: ADOBE INC.
Reel/Frame 063393/0031 →
Continuity (26)
Continuation In Part 18190544 · Mar 27, 2023
Continuation In Part 18190500 · Mar 27, 2023
Continuation In Part 18190556 · Mar 27, 2023
Continuation In Part 18190513 · Mar 27, 2023
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058630 · Nov 23, 2022
Continuation In Part 18058575 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058630 · Nov 23, 2022
Continuation In Part 18058575 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058622 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058622 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Provisional Application 63378616 · Oct 6, 2022
Provisional Application 63378212 · Oct 3, 2022
Related Publication 20240127509A1 · Apr 18, 2024
References Cited (143)
US 8289318B1 · Hadap et al. · 2012 [cited by applicant]
US 8830237B2 · Zimmermann · 2014 [cited by applicant]
US 9153209B2 · Dmitriev · 2015 [cited by applicant]
US 10043279B1 · Eshet · 2018 [cited by applicant]
US 10460214B2 · Lu et al. · 2019 [cited by applicant]
US 10679046B1 · Black et al. · 2020 [cited by applicant]
US 10691286B2 · King et al. · 2020 [cited by applicant]
US 10930075B2 · Costa et al. · 2021 [cited by applicant]
US 11094083B2 · Eisenmann et al. · 2021 [cited by applicant]
US 11217035B2 · El Hajjar · 2022 [cited by applicant]
US 11263823B2 · Gauseback et al. · 2022 [cited by applicant]
US 11494995B2 · Berkebile · 2022 [cited by applicant]
US 11514638B2 · Lafter et al. · 2022 [cited by applicant]
US 11694450B1 · Spinelli · 2023 [cited by examiner]
US 11741668B2 · Jones et al. · 2023 [cited by applicant]
US 11869152B2 · Koh et al. · 2024 [cited by applicant]
US 11881049B1 · Soltz · 2024 [cited by applicant]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 12141916B2 · Brown et al. · 2024 [cited by applicant]
US 20110018873A1 · Chang et al. · 2011 [cited by applicant]
US 20110298799A1 · Mariani et al. · 2011 [cited by applicant]
US 20120081357A1 · Habbecke et al. · 2012 [cited by applicant]
US 20130135305A1 · Bystrov et al. · 2013 [cited by applicant]
US 20140359536A1 · Cheng et al. · 2014 [cited by applicant]
US 20160035068A1 · Wilensky et al. · 2016 [cited by applicant]
US 20160063669A1 · Wilensky et al. · 2016 [cited by applicant]
US 20160119670A1 · Izutsu et al. · 2016 [cited by applicant]
US 20180158230A1 · Yan et al. · 2018 [cited by applicant]
US 20180218535A1 · Ceylan et al. · 2018 [cited by applicant]
US 20180329485A1 · Carothers et al. · 2018 [cited by applicant]
US 20190026958A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190065026A1 · Kiemele et al. · 2019 [cited by applicant]
US 20190354699A1 · Pekelny et al. · 2019 [cited by applicant]
US 20200020173A1 · Sharif et al. · 2020 [cited by applicant]
US 20200175756A1 · Crowe et al. · 2020 [cited by applicant]
US 20200234498A1 · Price · 2020 [cited by examiner]
US 20210005026A1 · Lesbordes · 2021 [cited by applicant]
US 20210065440A1 · Sunkavalli et al. · 2021 [cited by applicant]
US 20210074062A1 · Madonna et al. · 2021 [cited by applicant]
US 20210097776A1 · Faulkner et al. · 2021 [cited by applicant]
US 20210335039A1 · Jones et al. · 2021 [cited by applicant]
US 20210343080A1 · Kim et al. · 2021 [cited by applicant]
US 20210398351A1 · Papandreou et al. · 2021 [cited by applicant]
US 20220068007A1 · Lafer et al. · 2022 [cited by applicant]
US 20220068037A1 · Pardeshi · 2022 [cited by applicant]
US 20220222887A1 · Hundal et al. · 2022 [cited by applicant]
US 20220284613A1 · Yin et al. · 2022 [cited by applicant]
US 20220292352A1 · Jourdan et al. · 2022 [cited by applicant]
US 20220375025A1 · Ardö · 2022 [cited by examiner]
US 20220375113A1 · Sosnovik · 2022 [cited by examiner]
US 20220414834A1 · Du et al. · 2022 [cited by applicant]
US 20230033956A1 · Yong et al. · 2023 [cited by applicant]
US 20230080584A1 · Zohar et al. · 2023 [cited by applicant]
US 20230117686A1 · Jain et al. · 2023 [cited by applicant]
US 20230140460A1 · Munkberg et al. · 2023 [cited by applicant]
US 20230141494A1 · Brown et al. · 2023 [cited by applicant]
US 20230143034A1 · Wu et al. · 2023 [cited by applicant]
US 20230230321A1 · Schreckenbert et al. · 2023 [cited by applicant]
US 20230239458A1 · Clares et al. · 2023 [cited by applicant]
US 20230281925A1 · Aigerman et al. · 2023 [cited by applicant]
US 20230290090A1 · Su · 2023 [cited by applicant]
US 20230298269A1 · Genova et al. · 2023 [cited by applicant]
US 20230306686A1 · Zangenehpour et al. · 2023 [cited by applicant]
US 20230326028A1 · Zhang et al. · 2023 [cited by applicant]
US 20230368339A1 · Zheng et al. · 2023 [cited by applicant]
US 20240013462A1 · Seol et al. · 2024 [cited by applicant]
US 20240037717A1 · Amirghodsi et al. · 2024 [cited by applicant]
US 20240095898A1 · Loyd · 2024 [cited by examiner]
US 20240112396A1 · Spencer · 2024 [cited by applicant]
US 20240135572A1 · Singh et al. · 2024 [cited by applicant]
US 20240144520A1 · Gori et al. · 2024 [cited by applicant]
US 20240144586A1 · Hold-Geoffroy et al. · 2024 [cited by applicant]
US 20240144623A1 · Gori et al. · 2024 [cited by applicant]
US 20240161320A1 · Gadelha et al. · 2024 [cited by applicant]
US 20240161366A1 · Mech et al. · 2024 [cited by applicant]
US 20240161405A1 · Mech et al. · 2024 [cited by applicant]
US 20240161406A1 · Mech et al. · 2024 [cited by applicant]
CN 103578116A · 2014 [cited by examiner]
CN 115639718A · 2023 [cited by examiner]
U.S. Appl. No. 18/055,590, Apr. 8, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,162, Apr. 25, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,584, Jun. 21, 2024, Office Action. [cited by applicant]
Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105, 2017. [cited by applicant]
Olga Barinova, Victor Lempitsky, Elena Tretiak, and Push-meet Kohli. Geometric image parsing in man-made environments. In European conference on computer vision, pp. 57-70. Springer, 2010. [cited by applicant]
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Adabins: Depth estimation using adaptive bins. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4009-4018, 2021. [cited by applicant]
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conferen… [cited by applicant]
Jia-Ren Chang and Yong-Sheng Chen. Pyramid stereo matching network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5410-5418, 2018. [cited by applicant]
Antonio Criminisi, Ian Reid, and Andrew Zisserman. Single view metrology. International Journal of Computer Vision, 40(2):123-148, 2000. [cited by applicant]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. Ieee, 2009. [cited by applicant]
Patrick Denis, James H Elder, and Francisco J Estrada. Efficient edge-based methods for estimating manhattan frames in urban imagery. In European conference on computer vision, pp. 197-210. Springer, 2008. [cited by applicant]
Jonathan Deutscher, Michael Isard, and John MacCormick. Automatic camera calibration from a single manhattan image. In European Conference on Computer Vision, pp. 175-188. Springer, 2002. [cited by applicant]
Derek Hoiem, Alexei A Efros, and Martial Hebert. Putting objects in perspective. International Journal of Computer Vision, 80(1):3-15, 2008. [cited by applicant]
Yannick Hold-Geoffroy, Kalyan Sunkavalli, Jonathan Eisenmann, Matthew Fisher, Emiliano Gambaretto, Sunil Hadap, and Jean-Francois Lalonde. A perceptual measure for deep single image camera calibration. In Proceedings of… [cited by applicant]
Abhishek Kar, Shubham Tulsiani, Joao Carreira, and Jitendra Malik. Amodal completion and size constancy in natural scenes. In Proceedings of the IEEE international conference on computer vision, pp. 127-135, 2015. [cited by applicant]
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1611-1621, 2021. [cited by applicant]
Byeong-Uk Lee, Kyunghyun Lee, and In So Kweon. Depth completion using plane-residual representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13916-13925, 2021. [cited by applicant]
Hyunjoon Lee, Eli Shechtman, Jue Wang, and Seungyong Lee. Automatic upright adjustment of photographs with robust camera calibration. IEEE transactions on pattern analysis and machine intelligence, 36(5):833-844, 2013. [cited by applicant]
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, Il Hong Suh. From big to small: Multi-scale local planar guidance for monocular depth estimation. arXiv preprint arXiv:1907.10326, 2019. [cited by applicant]
Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker, Noah Snavely, Ce Liu, and William T Freeman. Learning the depths of moving people by watching frozen people. In Proceedings of the IEEE/CVF conference on computer … [cited by applicant]
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. Consistent video depth estimation. ACM Transactions on Graphics (ToG), 39(4):71-1, 2020. [cited by applicant]
Fangchang Ma and Sertac Karaman. Sparse-to-dense: Depth prediction from sparse depth samples and a single image. In 2018 IEEE international conference on robotics and automation (ICRA), pp. 4796-4803. IEEE, 2018. [cited by applicant]
Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEE… [cited by applicant]
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, 942 Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation.… [cited by applicant]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99-106, 2021. [cited by applicant]
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. arXiv preprint arXiv:2201.05989, 2022. [cited by applicant]
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on compu… [cited by applicant]
Rene' Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis a… [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pp. 234-241… [cited by applicant]
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. Demon: Depth and motion network for learning monocular stereo. In Proceedings of the IEEE conference on com… [cited by applicant]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415-424, 2022. [cited by applicant]
Scott Workman, Connor Greenwell, Menghua Zhai, Ryan Baltenberger, and Nathan Jacobs. Deepfocal: A method for direct focal length estimation. In 2015 IEEE International Conference on Image Processing (ICIP), pp. 1369-137… [cited by applicant]
Scott Workman, Menghua Zhai, and Nathan Jacobs. Horizon lines in the wild. arXiv preprint arXiv:1604.02129, 2016. [cited by applicant]
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vision (ECCV), pp. 767-783, 2018. [cited by applicant]
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Menghua Zhai, Scott Workman, and Nathan Jacobs. Detecting vanishing points using global image context in a non-manhattan world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5657-… [cited by applicant]
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1851-1858, 201… [cited by applicant]
Rui Zhu, Xingyi Yang, Yannick Hold-Geoffroy, Federico Perazzi, Jonathan Eisenmann, Kalyan Sunkavalli, and Manmohan Chandraker. Single view metrology in the wild. In European Conference on Computer Vision, pp. 316-333. S… [cited by applicant]
Enric Corona, Albert Pumarola, Guillem Alenyà, Gerard Pons-Moll, Francesc Moreno-Noguer—SMPLicit: Topology-aware Generative Model for Clothed People, Corona et al., 2021. [cited by applicant]
Kevin Lin, Lijuan Wang, Zicheng Liu—End-to-End Human Pose and Mesh Reconstruction with Transformers, CVPR (2021). [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, Christian Theobalt—Style and Pose Control for Image Synthesis of Humans from a Single Monocular View, Sarkar et al., 2021. [cited by applicant]
Xie et al., SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, 2021. [cited by applicant]
Zhe Cao, Gines Hildago, Tomas Simon, Shih-En Wei, Yaser Sheikh—OpenPose—Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields, CVPS (2019). [cited by applicant]
Office Action as received in CN Application No. 2023112860680.7 dated Nov. 2, 2023. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2312456.3 dated Feb. 15, 2024. [cited by applicant]
Lili Wang et al., “Bidirectional Shadow Rendering for Interactive Mixed 360° Videos”, 2021 IEEE Virtual Reality and 3D User Interfaces (VR), p. 170-178, 2021. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2402407.7 dated Jul. 9, 2024. [cited by applicant]
U.S. Appl. No. 18/055,585, Aug. 26, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/055,590, Sep. 17, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,162, Sep. 11, 2024, Notice of Allowance. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2403090.0 dated Jan. 28, 2025. [cited by applicant]
Chung-Yi Weng et al. “Photo Wake-Up: 3D Character Animation from a Single Photo” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5908-5917 (2019). [cited by applicant]
Federica Bogo et al. “Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image”, ECCV, pp. 561-578 (2016). [cited by applicant]
Imry Kissos et al. “Beyond Weak Perspective for Monocular 3D Human Pose Estimation”, Springer Nature, pp. 541-554 (2021). [cited by applicant]
S. Lv, X. Yang, L. Gu, X. Xing, L. Pan and M. Fang, “Delaunay Mesh Reconstruction from 3D Medical Images Based on Centroidal Voronoi Tessellations,” 2009 International Conference on Computational Intelligence and Softwa… [cited by applicant]
Shih-En Wei et al. “Convolutional Pose Machines”, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4724-4732 (2016). [cited by applicant]
Shrivastava, S. (2020). Stereo Vision Based Object Detection Using V-Disparity and 3D Density-Based Clustering. In: Arai, K., Kapoor, S. (eds) Advances in Computer Vision. CVC 2019. Advances in Intelligent Systems and C… [cited by applicant]
Watson, Jamie, Oisin Mac Aodha, Daniyar Turmukhambetov, Gabriel J. Brostow and Michael Firman. University of Edinburgh. “Learning Stereo from Single Images.” European Conference on Computer Vision (2020). Aug. 2020. (Ye… [cited by applicant]
U.S. Appl. No. 18/055,584, Dec. 16, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/055,585, Jan. 13, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,590, Feb. 3, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,594, Mar. 5, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/304,144, Mar. 10, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,147, Jan. 21, 2025, Office Action. [cited by applicant]