IP Library Granted Patent US 12,499,574
Granted Patent B2
US 12,499,574 · App. 18/304,144 · Granted Dec 16, 2025

Generating three-dimensional human models representing two-dimensional humans in two-dimensional images

Inventors: Giorgio Gori (San Jose, CA); Yi Zhou (Los Angeles, CA); Yangtuanfeng Wang (London, GB); Yang Zhou (San Jose, CA); Krishna Kumar Singh (San Jose, CA); Jae Shin Yoon (San Jose, CA); Duygu Ceylan Aksit (Mountain View, CA)
Assignee: Adobe Inc.
G06T7/73G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,574
App. No.
18/304,144
Granted
Dec 16, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify two-dimensional images via scene-based editing using three-dimensional representations of the two-dimensional images. For instance, in one or more embodiments, the disclosed systems utilize three-dimensional representations of two-dimensional images to generate and modify shadows in the two-dimensional images according to various shadow maps. Additionally, the disclosed systems utilize three-dimensional representations of two-dimensional images to modify humans in the two-dimensional images. The disclosed systems also utilize three-dimensional representations of two-dimensional images to provide scene scale estimation via scale fields of the two-dimensional images. In some embodiments, the disclosed systems utilizes three-dimensional representations of two-dimensional images to generate and visualize 3D planar surfaces for modifying objects in two-dimensional images. The disclosed systems further use three-dimensional representations of two-dimensional images to customize focal points for the two-dimensional images.

Claims (52)

1 . A system comprising:

one or more memory devices comprising a two-dimensional image; and

one or more processors configured to cause the system to:

extract, utilizing one or more neural networks, two-dimensional pose data corresponding to a two-dimensional skeleton with two-dimensional bones for a two-dimensional human extracted from the two-dimensional image;

extract, utilizing the one or more neural networks, three-dimensional pose data and three-dimensional shape data corresponding to a three-dimensional skeleton for the two-dimensional human extracted from the two-dimensional image; and

generate, within a three-dimensional space corresponding to the two-dimensional image, a three-dimensional human model representing the two-dimensional human by refining the three-dimensional skeleton of the three-dimensional pose data according to the two-dimensional skeleton of the two-dimensional pose data and the three-dimensional shape data.

2 . The system of claim 1 , wherein the one or more processors are configured to cause the system to:

extract the two-dimensional pose data from the two-dimensional image utilizing a first neural network of the one or more neural networks; and

extract the three-dimensional pose data and the three-dimensional shape data utilizing a second neural network of the one or more neural networks.

3 . The system of claim 1 , wherein the one or more processors are configured to cause the system to extract the three-dimensional pose data by:

generating a body bounding box corresponding to a body portion of the two-dimensional human; and

extracting, utilizing a neural network, three-dimensional pose data corresponding to the body portion of the two-dimensional human according to the body bounding box.

4 . The system of claim 3 , wherein the one or more processors are configured to cause the system to extract the three-dimensional pose data by:

generating one or more hand bounding boxes corresponding to one or more hands of the two-dimensional human; and

extracting, utilizing an additional neural network, additional three-dimensional pose data corresponding to the one or more hands of the two-dimensional human according to the one or more hand bounding boxes.

5 . The system of claim 4 , wherein the one or more processors are configured to cause the system to generate the three-dimensional human model by combining the three-dimensional pose data corresponding to the body portion of the two-dimensional human with the additional three-dimensional pose data corresponding to the one or more hands of the two-dimensional human.

6 . The system of claim 1 , wherein the one or more processors are configured to cause the system to generate the three-dimensional human model by iteratively modifying positions of bones in the three-dimensional skeleton based on positions of bones in the two-dimensional skeleton.

7 . The system of claim 1 , wherein the one or more processors are configured to cause the system to generate a modified two-dimensional image by:

modifying a pose of the three-dimensional human model within the three-dimensional space;

generating a modified pose of the two-dimensional human within the two-dimensional image according to the pose of the three-dimensional human model in the three-dimensional space; and

generating, utilizing the one or more neural networks, the modified two-dimensional image comprising a modified two-dimensional human according to the modified pose of the two-dimensional human and a camera position associated with the two-dimensional image.

8 . A non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

extracting, utilizing one or more neural networks, two-dimensional pose data, comprising a two-dimensional skeleton with two-dimensional bones, from a two-dimensional human extracted from a two-dimensional image;

extracting, utilizing the one or more neural networks, three-dimensional pose data and three-dimensional shape data corresponding to the two-dimensional human extracted from the two-dimensional image; and

generating, within a three-dimensional space corresponding to the two-dimensional image, a three-dimensional human model representing the two-dimensional human by combining the two-dimensional pose data with the three-dimensional pose data and the three-dimensional shape data.

9 . The non-transitory computer readable medium of claim 8 , wherein:

extracting the two-dimensional pose data comprises extracting a two-dimensional skeleton from a cropped portion of the two-dimensional image utilizing a first neural network of the one or more neural networks; and

extracting the three-dimensional pose data comprises extracting a three-dimensional skeleton from the cropped portion of the two-dimensional image utilizing a second neural network of the one or more neural networks.

10 . The non-transitory computer readable medium of claim 8 , wherein extracting the three-dimensional pose data comprises:

extracting a first three-dimensional skeleton corresponding to a first portion of the two-dimensional human utilizing a first neural network; and

extracting a second three-dimensional skeleton corresponding to a second portion of the two-dimensional human comprising a hand utilizing a second neural network.

11 . The non-transitory computer readable medium of claim 10 , wherein generating the three-dimensional human model comprises:

iteratively modifying positions of bones of the second three-dimensional skeleton according to positions of bones of the first three-dimensional skeleton within the three-dimensional space to merge the first three-dimensional skeleton and the second three-dimensional skeleton; and

iteratively modifying positions of bones in the first three-dimensional skeleton according to positions of bones of a two-dimensional skeleton from the two-dimensional pose data.

12 . A computer-implemented method comprising:

extracting, by at least one processor utilizing one or more neural networks, two-dimensional pose data, comprising a two-dimensional skeleton with two-dimensional bones, from a two-dimensional human extracted from a two-dimensional image;

extracting, by the at least one processor utilizing the one or more neural networks, three-dimensional pose data and three-dimensional shape data corresponding to the two-dimensional human extracted from the two-dimensional image; and

generating, by the at least one processor and within a three-dimensional space corresponding to the two-dimensional image, a three-dimensional human model representing the two-dimensional human by combining the two-dimensional pose data with the three-dimensional pose data and the three-dimensional shape data.

13 . The computer-implemented method of claim 12 , wherein extracting the two-dimensional pose data comprises extracting, utilizing a first neural network, the two-dimensional pose data comprising a two-dimensional skeleton with two-dimensional bones and annotations indicating one or more portions of the two-dimensional skeleton.

14 . The computer-implemented method of claim 13 , wherein extracting the three-dimensional pose data and the three-dimensional shape data comprises extracting, utilizing a second neural network, the three-dimensional pose data comprising a three-dimensional skeleton with three-dimensional bones and the three-dimensional shape data comprising a three-dimensional mesh according to the two-dimensional human.

15 . The computer-implemented method of claim 14 , wherein extracting the three-dimensional pose data comprises extracting, utilizing a third neural network for hand-specific bounding boxes, three-dimensional hand pose data corresponding to one or more hands of the two-dimensional human.

16 . The computer-implemented method of claim 12 , wherein generating the three-dimensional human model comprises iteratively adjusting one or more bones in the three-dimensional pose data according to one or more corresponding bones in the two-dimensional pose data.

17 . The computer-implemented method of claim 12 , wherein generating the three-dimensional human model comprises iteratively connecting one or more hand skeletons with three-dimensional hand pose data to a body skeleton with three-dimensional body pose data.

18 . The computer-implemented method of claim 12 , further comprising:

generating, in response to an indication to modify a pose of the two-dimensional human within the two-dimensional image, a modified three-dimensional human model with modified three-dimensional pose data; and

generating a modified two-dimensional image comprising a modified two-dimensional human based on the modified three-dimensional human model.

19 . The computer-implemented method of claim 18 , further comprising:

determining, in response to the three-dimensional human model comprising the modified three-dimensional pose data, an interaction between the modified three-dimensional human model and an additional three-dimensional model within the three-dimensional space corresponding to the two-dimensional image; and

generating the modified two-dimensional image comprising the modified two-dimensional human according to the interaction between the modified three-dimensional human model and the additional three-dimensional model.

20 . The computer-implemented method of claim 12 , wherein extracting the two-dimensional pose data comprises:

generating a cropped image corresponding to a boundary of the two-dimensional human in the two-dimensional image; and

extracting the two-dimensional pose data from the cropped image utilizing the one or more neural networks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2023
From: AKSIT, DUYGU CEYLAN; GORI, GIORGIO; ZHOU, YI; ZHOU, YANG; SINGH, KRISHNA KUMAR; WANG, YANGTUANFENG; YOON, JAE SHIN
To: ADOBE INC.
Reel/Frame 063393/0057 →
Continuity (26)
Continuation In Part 18190544 · Mar 27, 2023
Continuation In Part 18190556 · Mar 27, 2023
Continuation In Part 18190636 · Mar 27, 2023
Continuation In Part 18190513 · Mar 27, 2023
Continuation In Part 18190654 · Mar 27, 2023
Continuation In Part 18190500 · Mar 27, 2023
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058630 · Nov 23, 2022
Continuation In Part 18058575 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058630 · Nov 23, 2022
Continuation In Part 18058575 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058622 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058622 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Provisional Application 63378616 · Oct 6, 2022
Related Publication 20240144520A1 · May 2, 2024
References Cited (135)
US 8289318B1 · Hadap et al. · 2012 [cited by applicant]
US 8830237B2 · Zimmermann · 2014 [cited by applicant]
US 9153209B2 · Dmitriev · 2015 [cited by applicant]
US 10043279B1 · Eshet · 2018 [cited by applicant]
US 10460214B2 · Lu et al. · 2019 [cited by applicant]
US 10679046B1 · Black et al. · 2020 [cited by applicant]
US 10691286B2 · King et al. · 2020 [cited by applicant]
US 10930075B2 · Costa et al. · 2021 [cited by applicant]
US 11094083B2 · Eisenmann et al. · 2021 [cited by applicant]
US 11217035B2 · Hajjar · 2022 [cited by applicant]
US 11263823B2 · Gauseback et al. · 2022 [cited by applicant]
US 11494995B2 · Berkebile · 2022 [cited by applicant]
US 11514638B2 · Lafer · 2022 [cited by examiner]
US 11741668B2 · Jones et al. · 2023 [cited by applicant]
US 11869152B2 · Koh et al. · 2024 [cited by applicant]
US 11881049B1 · Soltz · 2024 [cited by applicant]
US 12026845B2 · Pardeshi · 2024 [cited by examiner]
US 12141916B2 · Brown · 2024 [cited by examiner]
US 20110018873A1 · Chang et al. · 2011 [cited by applicant]
US 20110298799A1 · Mariani et al. · 2011 [cited by applicant]
US 20120081357A1 · Habbecke et al. · 2012 [cited by applicant]
US 20130135305A1 · Bystrov et al. · 2013 [cited by applicant]
US 20140359536A1 · Cheng et al. · 2014 [cited by applicant]
US 20160035068A1 · Wilensky et al. · 2016 [cited by applicant]
US 20160063669A1 · Wilensky et al. · 2016 [cited by applicant]
US 20160119670A1 · Izutsu et al. · 2016 [cited by applicant]
US 20180158230A1 · Yan et al. · 2018 [cited by applicant]
US 20180218535A1 · Ceylan et al. · 2018 [cited by applicant]
US 20180329485A1 · Carothers et al. · 2018 [cited by applicant]
US 20190026958A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190065026A1 · Kiemele et al. · 2019 [cited by applicant]
US 20190354699A1 · Pekelny et al. · 2019 [cited by applicant]
US 20200020173A1 · Sharif et al. · 2020 [cited by applicant]
US 20200175756A1 · Crowe et al. · 2020 [cited by applicant]
US 20210005026A1 · Lesbordes · 2021 [cited by applicant]
US 20210065440A1 · Sunkavalli et al. · 2021 [cited by applicant]
US 20210074062A1 · Madonna et al. · 2021 [cited by applicant]
US 20210097776A1 · Faulkner et al. · 2021 [cited by applicant]
US 20210335039A1 · Jones et al. · 2021 [cited by applicant]
US 20210343080A1 · Kim et al. · 2021 [cited by applicant]
US 20210398351A1 · Papandreou et al. · 2021 [cited by applicant]
US 20220068007A1 · Lafer · 2022 [cited by examiner]
US 20220068037A1 · Pardeshi · 2022 [cited by examiner]
US 20220222887A1 · Hundal et al. · 2022 [cited by applicant]
US 20220284613A1 · Yin et al. · 2022 [cited by applicant]
US 20220292352A1 · Jourdan et al. · 2022 [cited by applicant]
US 20220414834A1 · Du et al. · 2022 [cited by applicant]
US 20230033956A1 · Yong et al. · 2023 [cited by applicant]
US 20230080584A1 · Zohar et al. · 2023 [cited by applicant]
US 20230117686A1 · Jain et al. · 2023 [cited by applicant]
US 20230140460A1 · Munkberg et al. · 2023 [cited by applicant]
US 20230141494A1 · Brown · 2023 [cited by examiner]
US 20230143034A1 · Wu et al. · 2023 [cited by applicant]
US 20230230321A1 · Schreckenbert et al. · 2023 [cited by applicant]
US 20230239458A1 · Cantero Clares et al. · 2023 [cited by applicant]
US 20230281925A1 · Aigerman et al. · 2023 [cited by applicant]
US 20230290090A1 · Su · 2023 [cited by applicant]
US 20230298269A1 · Genova · 2023 [cited by examiner]
US 20230306686A1 · Zangenehpour et al. · 2023 [cited by applicant]
US 20230326028A1 · Zhang et al. · 2023 [cited by applicant]
US 20230368339A1 · Zheng et al. · 2023 [cited by applicant]
US 20240013462A1 · Seol et al. · 2024 [cited by applicant]
US 20240037717A1 · Amirghodsi et al. · 2024 [cited by applicant]
US 20240112396A1 · Spencer · 2024 [cited by applicant]
US 20240135572A1 · Singh et al. · 2024 [cited by applicant]
US 20240144520A1 · Gori et al. · 2024 [cited by applicant]
US 20240144586A1 · Hold-Geoffroy et al. · 2024 [cited by applicant]
US 20240144623A1 · Gori et al. · 2024 [cited by applicant]
US 20240161320A1 · Gadelha et al. · 2024 [cited by applicant]
US 20240161366A1 · Mech et al. · 2024 [cited by applicant]
US 20240161405A1 · Mech et al. · 2024 [cited by applicant]
US 20240161406A1 · Mech et al. · 2024 [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2402407.7 dated Jul. 9, 2024. [cited by applicant]
U.S. Appl. No. 18/055,585, filed Aug. 26, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/055,590, filed Sep. 17, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,162, filed Sep. 11, 2024, Notice of Allowance. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2312456.3 dated Feb. 15, 2024. [cited by applicant]
Lili Wang et al., “Bidirectional Shadow Rendering for Interactive Mixed 360° Videos”, 2021 IEEE Virtual Reality and 3D User Interfaces (VR), p. 170-178, 2021. [cited by applicant]
Zhe Cao, Gines Hildago, Tomas Simon, Shih-En Wei, Yaser Sheikh-OpenPose-Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields, CVPS (2019). [cited by applicant]
Kevin Lin, Lijuan Wang, Zicheng Liu—End-to-End Human Pose and Mesh Reconstruction with Transformers, CVPR (2021). [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, Christian Theobalt—Style and Pose Control for Image Synthesis of Humans from a Single Monocular View, Sarkar et al., 2021. [cited by applicant]
Enric Corona, Albert Pumarola, Guillem Alenyà, Gerard Pons-Moll, Francesc Moreno-Noguer—SMPLicit: Topology-aware Generative Model for Clothed People, Corona et al., 2021. [cited by applicant]
U.S. Appl. No. 18/055,590, filed Apr. 8, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/304,162, filed Apr. 25, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,584, filed Jun. 21, 2024, Office Action. [cited by applicant]
Abhishek Kar, Shubham Tulsiani, Joao Carreira, and Jitendra Malik. Amodal completion and size constancy in natural scenes. In Proceedings of the IEEE international conference on computer vision, pp. 127-135, 2015. [cited by applicant]
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conferen… [cited by applicant]
Antonio Criminisi, Ian Reid, and Andrew Zisserman. Single view metrology. International Journal of Computer Vision, 40(2):123-148, 2000. [cited by applicant]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99-106, 2021. [cited by applicant]
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. Demon: Depth and motion network for learning monocular stereo. In Proceedings of the IEEE conference on com… [cited by applicant]
Byeong-Uk Lee, Kyunghyun Lee, and In So Kweon. Depth completion using plane-residual representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13916-13925, 2021. [cited by applicant]
Derek Hoiem, Alexei A Efros, and Martial Hebert. Putting objects in perspective. International Journal of Computer Vision, 80(1):3-15, 2008. [cited by applicant]
Fangchang Ma and Sertac Karaman. Sparse-to-dense: Depth prediction from sparse depth samples and a single image. In 2018 IEEE international conference on robotics and automation (ICRA), pp. 4796-4803. IEEE, 2018. [cited by applicant]
Hyunjoon Lee, Eli Shechtman, Jue Wang, and Seungyong Lee. Automatic upright adjustment of photographs with robust camera calibration. IEEE transactions on pattern analysis and machine intelligence, 36(5):833-844, 2013. [cited by applicant]
Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105, 2017. [cited by applicant]
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on compu… [cited by applicant]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. Ieee, 2009. [cited by applicant]
Jia-Ren Chang and Yong-Sheng Chen. Pyramid stereo matching network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5410-5418, 2018. [cited by applicant]
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, Il Hong Suh. From big to small: Multi-scale local planar guidance for monocular depth estimation. arXiv preprint arXiv:1907.10326, 2019. [cited by applicant]
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1611-1621, 2021. [cited by applicant]
Jonathan Deutscher, Michael Isard, and John MacCormick. Automatic camera calibration from a single manhattan image. In European Conference on Computer Vision, pp. 175-188. Springer, 2002. [cited by applicant]
Menghua Zhai, Scott Workman, and Nathan Jacobs. Detecting vanishing points using global image context in a non-manhattan world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5657-… [cited by applicant]
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, 942 Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation.… [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pp. 234-241… [cited by applicant]
Olga Barinova, Victor Lempitsky, Elena Tretiak, and Pushmeet Kohli. Geometric image parsing in man-made environments. In European conference on computer vision, pp. 57-70. Springer, 2010. [cited by applicant]
Patrick Denis, James H Elder, and Francisco J Estrada. Efficient edge-based methods for estimating manhattan frames in urban imagery. In European conference on computer vision, pp. 197-210. Springer, 2008. [cited by applicant]
Rene' Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis a… [cited by applicant]
Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEE… [cited by applicant]
Rui Zhu, Xingyi Yang, Yannick Hold-Geoffroy, Federico Perazzi, Jonathan Eisenmann, Kalyan Sunkavalli, and Manmohan Chandraker. Single view metrology in the wild. In European Conference on Computer Vision, pp. 316-333. S… [cited by applicant]
Scott Workman, Connor Greenwell, Menghua Zhai, Ryan Baltenberger, and Nathan Jacobs. Deepfocal: A method for direct focal length estimation. In 2015 IEEE International Conference on Image Processing (ICIP), pp. 1369-137… [cited by applicant]
Scott Workman, Menghua Zhai, and Nathan Jacobs. Horizon lines in the wild. arXiv preprint arXiv:1604.02129, 2016. [cited by applicant]
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Adabins: Depth estimation using adaptive bins. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4009-4018, 2021. [cited by applicant]
Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. arXiv preprint arXiv:2201.05989, 2022. [cited by applicant]
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1851-1858, 201… [cited by applicant]
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415-424, 2022. [cited by applicant]
Xie et al., SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, 2021. [cited by applicant]
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. Consistent video depth estimation. ACM Transactions on Graphics (ToG), 39(4):71-1, 2020. [cited by applicant]
Yannick Hold-Geoffroy, Kalyan Sunkavalli, Jonathan Eisenmann, Matthew Fisher, Emiliano Gambaretto, Sunil Hadap, and Jean-Francois Lalonde. A perceptual measure for deep single image camera calibration. In Proceedings of… [cited by applicant]
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vision (ECCV), pp. 767-783, 2018. [cited by applicant]
Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker, Noah Snavely, Ce Liu, and William T Freeman. Learning the depths of moving people by watching frozen people. In Proceedings of the IEEE/CVF conference on computer … [cited by applicant]
Office Action as received in CN Application No. 2023112860680.7 dated Nov. 2, 2023. [cited by applicant]
Combined Search Report and Written Opinion as received in GB 2403090.0 dated Jan. 28, 2025. [cited by applicant]
Chung-Yi Weng et al. “Photo Wake-Up: 3D Character Animation from a Single Photo” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5908-5917 (2019). [cited by applicant]
Federica Bogo et al. “Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image”, ECCV, pp. 561-578 (2016). [cited by applicant]
Imry Kissos et al. “Beyond Weak Perspective for Monocular 3D Human Pose Estimation”, Springer Nature, pp. 541-554 (2021). [cited by applicant]
Shih-En Wei et al. “Convolutional Pose Machines”, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4724-4732 (2016). [cited by applicant]
S. Lv, X. Yang, L. Gu, X. Xing, L. Pan and M. Fang, “Delaunay Mesh Reconstruction from 3D Medical Images Based on Centroidal Voronoi Tessellations,” 2009 International Conference on Computational Intelligence and Softwa… [cited by applicant]
Shrivastava, S. (2020). Stereo Vision Based Object Detection Using V-Disparity and 3D Density-Based Clustering. In: Arai, K., Kapoor, S. (eds) Advances in Computer Vision. CVC 2019. Advances in Intelligent Systems and C… [cited by applicant]
Watson, Jamie, Oisin Mac Aodha, Daniyar Turmukhambetov, Gabriel J. Brostow and Michael Firman. University of Edinburgh. “Learning Stereo from Single Images.” European Conference on Computer Vision (2020). Aug. 2020. (Ye… [cited by applicant]
U.S. Appl. No. 18/055,584, filed Dec. 16, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/055,585, filed Jan. 13, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,590, filed Feb. 3, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/055,594, filed Mar. 5, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/304,147, filed Jan. 21, 2025, Office Action. [cited by applicant]