IP Library Granted Patent US 12,277,652
Granted Patent B2
US 12,277,652 · App. 18/055,585 · Granted Apr 15, 2025

Modifying two-dimensional images utilizing segmented three-dimensional object meshes of the two-dimensional images

Inventors: Radomir Mech (Mountain View, CA); Nathan Carr (San Jose, CA); Matheus Gadelha (San Jose, CA)
Assignee: Adobe Inc.
G06T17/205G06T7/11G06T7/50G06T7/70G06T11/60G06V20/70G06T2200/08G06T2200/24G06T2207/20021G06T2207/20084G06T2207/20228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,652
App. No.
18/055,585
Granted
Apr 15, 2025
Kind
B2
Abstract

Methods, systems, and non-transitory computer readable storage media are disclosed for generating three-dimensional meshes representing two-dimensional images for editing the two-dimensional images. The disclosed system utilizes a first neural network to determine density values of pixels of a two-dimensional image based on estimated disparity. The disclosed system samples points in the two-dimensional image according to the density values and generates a tessellation based on the sampled points. The disclosed system utilizes a second neural network to estimate camera parameters and modify the three-dimensional mesh based on the estimated camera parameters of the pixels of the two-dimensional image. In one or more additional embodiments, the disclosed system generates a three-dimensional mesh to modify a two-dimensional image according to a displacement input. Specifically, the disclosed system maps the three-dimensional mesh to the two-dimensional image, modifies the three-dimensional mesh in response to a displacement input, and updates the two-dimensional image.

Claims (69)

1. A method comprising:

generating, by at least one processor utilizing one or more neural networks, a three-dimensional mesh by determining displacement of vertices of a tessellation of a two-dimensional image based on pixel depth values of the two-dimensional image;

segmenting, by the at least one processor utilizing the one or more neural networks, the three-dimensional mesh into a plurality of three-dimensional object meshes corresponding to a plurality of distinct objects of the two-dimensional image;

modifying, by the at least one processor in response to a displacement input to the two-dimensional image within a graphical user interface displaying the two-dimensional image, a selected three-dimensional object mesh of the plurality of three-dimensional object meshes based on a displaced portion of the selected three-dimensional object mesh corresponding to the displacement input to the two-dimensional image; and

generating, by the at least one processor, a modified two-dimensional image comprising at least one modified portion according to the displaced portion of the selected three-dimensional object mesh.

2. The method of claim 1 , wherein generating the three-dimensional mesh comprises:

generating the three-dimensional mesh by determining the displacement of the vertices of the tessellation of the two-dimensional image based on the pixel depth values and estimated camera parameters; or

generating the three-dimensional mesh based on a plurality of points sampled in the two-dimensional image according to density values determined from the pixel depth values of the two-dimensional image.

3. The method of claim 1 , wherein segmenting the three-dimensional mesh comprises:

detecting one or more objects in the two-dimensional image or in the three-dimensional mesh; and

separating a first portion of the three-dimensional mesh from a second portion of the three-dimensional mesh based on the one or more objects detected in the two-dimensional image or in the three-dimensional mesh.

4. The method of claim 3 , wherein segmenting the three-dimensional mesh comprises:

determining a semantic map comprising labels indicating object classifications of pixels in the two-dimensional image; and

detecting the one or more objects in the two-dimensional image based on the labels of the semantic map.

5. The method of claim 3 , wherein segmenting the three-dimensional mesh comprises:

determining a depth discontinuity at a portion of the three-dimensional mesh based on corresponding pixel depth values of the two-dimensional image; and

detecting the one or more objects in the three-dimensional mesh based on the depth discontinuity at the portion of the three-dimensional mesh.

6. The method of claim 1 , wherein modifying the selected three-dimensional object mesh comprises:

determining a projection from the two-dimensional image onto the plurality of three-dimensional object meshes in a three-dimensional environment; and

determining the selected three-dimensional object mesh based on a two-dimensional position of the displacement input relative to the two-dimensional image and the projection from the two-dimensional image onto the plurality of three-dimensional object meshes.

7. The method of claim 1 , wherein modifying the selected three-dimensional object mesh comprises:

determining that the displacement input indicates a displacement direction for a portion of the selected three-dimensional object mesh; and

modifying a portion of the selected three-dimensional object mesh by displacing the portion of the selected three-dimensional object mesh according to the displacement direction.

8. The method of claim 1 , wherein generating the modified two-dimensional image comprises:

determining a two-dimensional position of the two-dimensional image corresponding to a three-dimensional position of the displaced portion of the selected three-dimensional object mesh based on a mapping between the plurality of three-dimensional object meshes and the two-dimensional image; and

generating the modified two-dimensional image comprising the at least one modified portion at the two-dimensional position based on the three-dimensional position of the displaced portion of the selected three-dimensional object mesh.

9. The method of claim 1 , wherein modifying the selected three-dimensional object mesh comprises modifying the selected three-dimensional object mesh according to the displacement input without modifying one or more additional three-dimensional object meshes adjacent to the selected three-dimensional object mesh within a three-dimensional environment.

10. A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

generating, utilizing one or more neural networks, a three-dimensional mesh by determining displacement of vertices of a tessellation of a two-dimensional image based on pixel depth values of the two-dimensional image;

detecting, utilizing one or more object detection models, a plurality of distinct objects of the two-dimensional image;

segmenting, in response to detecting the plurality of distinct objects, the three-dimensional mesh into a plurality of three-dimensional object meshes corresponding to the plurality of distinct objects of the two-dimensional image;

modifying, in response to a displacement input to the two-dimensional image within a graphical user interface displaying the two-dimensional image, a selected three-dimensional object mesh of the plurality of three-dimensional object meshes based on a displaced portion of the selected three-dimensional object mesh corresponding to the displacement input to the two-dimensional image; and

generating a modified two-dimensional image comprising at least one modified portion according to the displaced portion of the selected three-dimensional object mesh.

11. The system of claim 10 , wherein detecting the plurality of distinct objects comprises:

generating, utilizing the one or more object detection models, a semantic map comprising labels indicating object classifications of pixels in the two-dimensional image; and

detecting the plurality of distinct objects of the two-dimensional image based on the labels of the semantic map.

12. The system of claim 10 , wherein detecting the plurality of distinct objects comprises:

determining, based on the pixel depth values of the two-dimensional image, a portion of the two-dimensional image comprising a depth discontinuity between adjacent regions of the two-dimensional image; and

determining that a first region of the adjacent regions corresponds to a first object and a second region of the adjacent regions corresponds to a second object.

13. The system of claim 10 , wherein modifying the selected three-dimensional object mesh comprises:

determining a two-dimensional position of the displacement input relative to the two-dimensional image; and

determining a three-dimensional position corresponding to a three-dimensional object mesh of the plurality of three-dimensional object meshes based on a mapping between the two-dimensional image and a three-dimensional environment comprising the plurality of three-dimensional object meshes.

14. The system of claim 10 , wherein modifying the selected three-dimensional object mesh comprises:

determining, based on an attribute of the displacement input, that the displacement input indicates one or more displacement directions for a portion of the selected three-dimensional object mesh; and

displacing the portion of the selected three-dimensional object mesh in the one or more displacement directions.

15. The system of claim 10 , wherein generating the modified two-dimensional image comprises:

determining that the displacement input indicates a displacement direction for the selected three-dimensional object mesh, the selected three-dimensional object mesh being adjacent an additional three-dimensional object mesh; and

displacing a portion of the selected three-dimensional object mesh according to the displacement direction without modifying the additional three-dimensional object mesh.

16. The system of claim 10 , wherein generating the modified two-dimensional image comprises:

determining, based on a mapping between the two-dimensional image and the three-dimensional mesh, a two-dimensional position of the two-dimensional image corresponding to a three-dimensional position of the displaced portion of the selected three-dimensional object mesh; and

generating the modified two-dimensional image comprising the at least one modified portion at the two-dimensional position of the two-dimensional image according to the displaced portion of the selected three-dimensional object mesh.

17. A non-transitory computer readable medium comprising instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

generating, by at least one processor utilizing one or more neural networks, a three-dimensional mesh by determining displacement of vertices of a tessellation of a two-dimensional image based on pixel depth values of the two-dimensional image;

segmenting, by the at least one processor utilizing the one or more neural networks, the three-dimensional mesh into a plurality of three-dimensional object meshes corresponding to a plurality of distinct objects of the two-dimensional image;

modifying, by the at least one processor in response to a displacement input to the two-dimensional image within a graphical user interface displaying the two-dimensional image, a selected three-dimensional object mesh of the plurality of three-dimensional object meshes based on a displaced portion of the selected three-dimensional object mesh corresponding to the displacement input to the two-dimensional image; and

generating, by the at least one processor, a modified two-dimensional image comprising at least one modified portion according to the displaced portion of the selected three-dimensional object mesh.

18. The non-transitory computer readable medium of claim 17 , wherein segmenting the three-dimensional mesh into the plurality of three-dimensional object meshes comprises:

detecting the plurality of distinct objects of the two-dimensional image according to: a semantic map corresponding to the two-dimensional image; or

depth discontinuities based on the pixel depth values of the two-dimensional image;

and

separating the three-dimensional mesh into the plurality of three-dimensional object meshes in response to detecting the plurality of distinct objects of the two-dimensional image.

19. The non-transitory computer readable medium of claim 17 , wherein modifying the selected three-dimensional object mesh comprises displacing a portion of the selected three-dimensional object mesh according to a displacement direction of the displacement input.

20. The non-transitory computer readable medium of claim 17 , wherein generating the modified two-dimensional image comprises:

determining a mapping between the two-dimensional image and the three-dimensional

mesh;

determining a three-dimensional position of the displaced portion of the selected three-dimensional object mesh; and

generating, based on the mapping between the two-dimensional image and the three-dimensional mesh, the modified two-dimensional image comprising the at least one modified portion at a two-dimensional position of the two-dimensional image corresponding to the three-dimensional position of the displaced portion of the selected three-dimensional object mesh.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: MECH, RADOMIR; CARR, NATHAN; GADELHA, MATHEUS
To: ADOBE INC.
Reel/Frame 061776/0246 →
Continuity (1)
Related Publication 20240161405A1 · May 16, 2024
References Cited (135)
US 8289318B1 · Hadap et al. · 2012 [cited by applicant]
US 8830237B2 · Zimmermann · 2014 [cited by examiner]
US 9153209B2 · Dmitriev · 2015 [cited by examiner]
US 10043279B1 · Eshet · 2018 [cited by applicant]
US 10460214B2 · Lu et al. · 2019 [cited by applicant]
US 10679046B1 · Black et al. · 2020 [cited by applicant]
US 10691286B2 · King et al. · 2020 [cited by applicant]
US 10930075B2 · Costa et al. · 2021 [cited by applicant]
US 11094083B2 · Eisenmann et al. · 2021 [cited by applicant]
US 11217035B2 · Hajjar · 2022 [cited by applicant]
US 11263823B2 · Gauseback et al. · 2022 [cited by applicant]
US 11494995B2 · Berkebile · 2022 [cited by applicant]
US 11514638B2 · Lafer · 2022 [cited by examiner]
US 11741668B2 · Jones et al. · 2023 [cited by applicant]
US 11869152B2 · Koh et al. · 2024 [cited by applicant]
US 11881049B1 · Soltz · 2024 [cited by examiner]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 12141916B2 · Brown et al. · 2024 [cited by applicant]
US 20110018873A1 · Chang et al. · 2011 [cited by applicant]
US 20110298799A1 · Mariani et al. · 2011 [cited by applicant]
US 20120081357A1 · Habbecke et al. · 2012 [cited by applicant]
US 20130135305A1 · Bystrov et al. · 2013 [cited by applicant]
US 20140359536A1 · Cheng et al. · 2014 [cited by applicant]
US 20160035068A1 · Wilensky et al. · 2016 [cited by applicant]
US 20160063669A1 · Wilensky et al. · 2016 [cited by applicant]
US 20160119670A1 · Izutsu et al. · 2016 [cited by applicant]
US 20180158230A1 · Yan et al. · 2018 [cited by applicant]
US 20180218535A1 · Ceylan et al. · 2018 [cited by applicant]
US 20180329485A1 · Carothers et al. · 2018 [cited by applicant]
US 20190026958A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190065026A1 · Kiemele et al. · 2019 [cited by applicant]
US 20190130229A1 · Lu et al. · 2019 [cited by applicant]
US 20190354699A1 · Pekelny · 2019 [cited by examiner]
US 20200020173A1 · Sharif et al. · 2020 [cited by applicant]
US 20200175756A1 · Crowe et al. · 2020 [cited by applicant]
US 20210005026A1 · Lesbordes · 2021 [cited by applicant]
US 20210065440A1 · Sunkavalli et al. · 2021 [cited by applicant]
US 20210074062A1 · Madonna et al. · 2021 [cited by applicant]
US 20210097776A1 · Faulkner et al. · 2021 [cited by applicant]
US 20210335039A1 · Jones · 2021 [cited by examiner]
US 20210343080A1 · Kim et al. · 2021 [cited by applicant]
US 20210398351A1 · Papandreou et al. · 2021 [cited by applicant]
US 20220068007A1 · Lafer et al. · 2022 [cited by applicant]
US 20220068037A1 · Pardeshi · 2022 [cited by applicant]
US 20220222887A1 · Hundal et al. · 2022 [cited by applicant]
US 20220284613A1 · Yin et al. · 2022 [cited by applicant]
US 20220292352A1 · Jourdan · 2022 [cited by examiner]
US 20220414834A1 · Du et al. · 2022 [cited by applicant]
US 20230033956A1 · Yong et al. · 2023 [cited by applicant]
US 20230080584A1 · Zohar et al. · 2023 [cited by applicant]
US 20230117686A1 · Jain · 2023 [cited by examiner]
US 20230140460A1 · Munkberg et al. · 2023 [cited by applicant]
US 20230141494A1 · Brown et al. · 2023 [cited by applicant]
US 20230143034A1 · Wu · 2023 [cited by examiner]
US 20230230321A1 · Schreckenbert et al. · 2023 [cited by applicant]
US 20230239458A1 · Clares et al. · 2023 [cited by applicant]
US 20230281925A1 · Aigerman et al. · 2023 [cited by applicant]
US 20230290090A1 · Su · 2023 [cited by examiner]
US 20230298269A1 · Genova et al. · 2023 [cited by applicant]
US 20230306686A1 · Zangenehpour et al. · 2023 [cited by applicant]
US 20230326028A1 · Zhang et al. · 2023 [cited by applicant]
US 20230368339A1 · Zheng et al. · 2023 [cited by applicant]
US 20240013462A1 · Seol et al. · 2024 [cited by applicant]
US 20240037717A1 · Amirghodsi et al. · 2024 [cited by applicant]
US 20240112396A1 · Spencer · 2024 [cited by applicant]
US 20240135572A1 · Singh et al. · 2024 [cited by applicant]
US 20240144520A1 · Gori et al. · 2024 [cited by applicant]
US 20240144586A1 · Hold-Geoffroy et al. · 2024 [cited by applicant]
US 20240144623A1 · Gori et al. · 2024 [cited by applicant]
US 20240161320A1 · Gadelha et al. · 2024 [cited by applicant]
US 20240161366A1 · Mech et al. · 2024 [cited by applicant]
US 20240161405A1 · Mech et al. · 2024 [cited by applicant]
US 20240161406A1 · Mech et al. · 2024 [cited by applicant]
U.S. Appl. No. 18/055,590, Apr. 8, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,162, Apr. 25, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,584, Jun. 21, 2024, Office Action. [cited by applicant]
Abhishek Kar, Shubham Tulsiani, Joao Carreira, and Jitendra Malik. Amodal completion and size constancy in natural scenes. In Proceedings of the IEEE international conference on computer vision, pp. 127-135, 2015. [cited by applicant]
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conferen… [cited by applicant]
Antonio Criminisi, Ian Reid, and Andrew Zisserman. Single view metrology. International Journal of Computer Vision, 40(2):123-148, 2000. [cited by applicant]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99-106, 2021. [cited by applicant]
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. Demon: Depth and motion network for learning monocular stereo. In Proceedings of the IEEE conference on com… [cited by applicant]
Byeong-Uk Lee, Kyunghyun Lee, and In So Kweon. Depth completion using plane-residual representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13916-13925, 2021. [cited by applicant]
Derek Hoiem, Alexei A Efros, and Martial Hebert. Putting objects in perspective. International Journal of Computer Vision, 80(1):3-15, 2008. [cited by applicant]
Enric Corona, Albert Pumarola, Guillem Alenyà, Gerard Pons-Moll, Francesc Moreno-Noguer—SMPLicit: Topology-aware Generative Model for Clothed People, Corona et al., 2021. [cited by applicant]
Fangchang Ma and Sertac Karaman. Sparse-to-dense: Depth prediction from sparse depth samples and a single image. In 2018 IEEE international conference on robotics and automation (ICRA), pp. 4796-4803. IEEE, 2018. [cited by applicant]
Hyunjoon Lee, Eli Shechtman, Jue Wang, and Seungyong Lee. Automatic upright adjustment of photographs with robust camera calibration. IEEE transactions on pattern analysis and machine intelligence, 36(5):833-844, 2013. [cited by applicant]
Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105, 2017. [cited by applicant]
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on compu… [cited by applicant]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. Ieee, 2009. [cited by applicant]
Jia-Ren Chang and Yong-Sheng Chen. Pyramid stereo matching network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5410-5418, 2018. [cited by applicant]
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, Il Hong Suh. From big to small: Multi-scale local planar guidance for monocular depth estimation. arXiv preprint arXiv:1907.10326, 2019. [cited by applicant]
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1611-1621, 2021. [cited by applicant]
Jonathan Deutscher, Michael Isard, and John MacCormick. Automatic camera calibration from a single manhattan image. In European Conference on Computer Vision, pp. 175-188. Springer, 2002. [cited by applicant]
Kevin Lin, Lijuan Wang, Zicheng Liu—End-to-End Human Pose and Mesh Reconstruction with Transformers, CVPR (2021). [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, Christian Theobalt—Style and Pose Control for Image Synthesis of Humans from a Single Monocular View, Sarkar et al., 2021. [cited by applicant]
Menghua Zhai, Scott Workman, and Nathan Jacobs. Detecting vanishing points using global image context in a non-manhattan world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5657-… [cited by applicant]
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, 942 Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation.… [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pp. 234-241… [cited by applicant]
Olga Barinova, Victor Lempitsky, Elena Tretiak, and Pushmeet Kohli. Geometric image parsing in man-made environments. In European conference on computer vision, pp. 57-70. Springer, 2010. [cited by applicant]
Patrick Denis, James H Elder, and Francisco J Estrada. Efficient edge-based methods for estimating manhattan frames in urban imagery. In European conference on computer vision, pp. 197-210. Springer, 2008. [cited by applicant]
Renè Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis an… [cited by applicant]
Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEE… [cited by applicant]
Rui Zhu, Xingyi Yang, Yannick Hold-Geoffroy, Federico Perazzi, Jonathan Eisenmann, Kalyan Sunkavalli, and Man-mohan Chandraker. Single view metrology in the wild. In European Conference on Computer Vision, pp. 316-333. … [cited by applicant]
Scott Workman, Connor Greenwell, Menghua Zhai, Ryan Baltenberger, and Nathan Jacobs. Deepfocal: A method for direct focal length estimation. In 2015 IEEE International Conference on Image Processing (ICIP), pp. 1369-137… [cited by applicant]
Scott Workman, Menghua Zhai, and Nathan Jacobs. Horizon lines in the wild. arXiv preprint arXiv:1604.02129, 2016. [cited by applicant]
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Adabins: Depth estimation using adaptive bins. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4009-4018, 2021. [cited by applicant]
Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. arXiv preprint arXiv:2201.05989, 2022. [cited by applicant]
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1851-1858, 201… [cited by applicant]
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415-424, 2022. [cited by applicant]
Xie et al., SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, 2021. [cited by applicant]
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. Consistent video depth estimation. ACM Transactions on Graphics (ToG), 39(4):71-1, 2020. [cited by applicant]
Yannick Hold-Geoffroy, Kalyan Sunkavalli, Jonathan Eisenmann, Matthew Fisher, Emiliano Gambaretto, Sunil Hadap, and Jean-Francois Lalonde. A perceptual measure for deep single image camera calibration. In Proceedings of… [cited by applicant]
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vision (ECCV), pp. 767-783, 2018. [cited by applicant]
Zhe Cao, Gines Hildago, Tomas Simon, Shih-En Wei, Yaser Sheikh—OpenPose—Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields, CVPS (2019). [cited by applicant]
Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker, Noah Snavely, Ce Liu, and William T Freeman. Learning the depths of moving people by watching frozen people. In Proceedings of the IEEE/CVF conference on computer … [cited by applicant]
Office Action as received in CN Application No. 2023112860680.7 dated Nov. 2, 2023. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2312456.3 dated Feb. 15, 2024. [cited by applicant]
Lili Wang et al., “Bidirectional Shadow Rendering for Interactive Mixed 360° Videos”, 2021 IEEE Virtual Reality and 3D User Interfaces (VR), p. 170-178, 2021. [cited by applicant]
U.S. Appl. No. 18/055,590 filed Sep. 17, 2024. [cited by applicant]
U.S. Appl. No. 18/304,162 filed Sep. 11, 2024. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2402407.7 dated Jul. 9, 2024. [cited by applicant]
S. Lv, X. Yang, L. Gu, X. Xing, L. Pan and M. Fang, “Delaunay Mesh Reconstruction from 3D Medical Images Based on Centroidal Voronoi Tessellations,” 2009 International Conference on Computational Intelligence and Softwa… [cited by applicant]
Shrivastava, S. (2020). Stereo Vision Based Object Detection Using V-Disparity and 3D Density-Based Clustering. In: Arai, K., Kapoor, S. (eds) Advances in Computer Vision. CVC 2019. Advances in Intelligent Systems and C… [cited by applicant]
Watson, Jamie, Oisin Mac Aodha, Daniyar Turmukhambetov, Gabriel J. Brostow and Michael Firman. University of Edinburgh. “Learning Stereo from Single Images.” European Conference on Computer Vision (2020). Aug. 2020. (Ye… [cited by applicant]
U.S. Appl. No. 18/055,584, Dec. 16. 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/055,590, Feb. 3, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,594, Mar. 5, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/304,144, Mar. 10, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,147, Jan. 21, 2025, Office Action. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2403090.0 dated Jan. 28, 2025. [cited by applicant]
Chung-Yi Weng et al. “Photo Wake-Up: 3D Character Animation from a Single Photo”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5908-5917 (2019). [cited by applicant]
Federica Bogo et al. “Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image”, ECCV, pp. 561-578 (2016). [cited by applicant]
Imry Kissos et al. “Beyond Weak Perspective for Monocular 3D Human Pose Estimation”, Springer Nature, pp. 541-554 (2020). [cited by applicant]
Shih-En Wei et al. “Convolutional Pose Machines”, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4724-4732 (2016). [cited by applicant]
Cited By (2)
US 12,518,170 US 12,569,761