IP Library › Granted Patent US 12,586,306
Granted Patent B2
US 12,586,306 · App. 18/337,537 · Granted Mar 24, 2026

Method, electronic device, and computer program product for modeling an object utilizing a spatial distribution and differences between features of respective image pairs

Inventors: Zhisong Liu (Shenzhen, CN); Zijia Wang (Weifang, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T17/00G06T7/55G06T7/70G06T2200/08G06T2207/10024G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,306
App. No.
18/337,537
Granted
Mar 24, 2026
Kind
B2
Abstract

A method for modeling an object in one embodiment includes: acquiring first position spatial information of a camera at a first position and second position spatial information of the camera at a second position, wherein the first position is different from the second position; generating a spatial distribution based on the first position spatial information and the second position spatial information, wherein the spatial distribution represents a probability distribution of space occupied by the object captured from a pose of the camera; generating a third image based on a first image captured at the first position, the spatial distribution, and the first position spatial information; generating a fourth image based on a second image captured at the second position, the spatial distribution, and the second position spatial information; and adjusting a model of the object based on the first image, the second image, the third image, and the fourth image.

Claims (79)

1 . A method for modeling an object, comprising:

acquiring first position spatial information of a camera at a first position and second position spatial information of the camera at a second position, wherein the first position is different from the second position;

generating a spatial distribution based on the first position spatial information and the second position spatial information, wherein the spatial distribution represents a probability distribution of space occupied by the object captured from a pose of the camera;

generating a third image based on a first image captured at the first position, the spatial distribution, and the first position spatial information;

generating a fourth image based on a second image captured at the second position, the spatial distribution, and the second position spatial information; and

adjusting a model of the object based on the first image, the second image, the third image, and the fourth image;

wherein adjusting the model based on the first image, the second image, the third image, and the fourth image comprises:

calculating a first difference between features of the third image and features of the second image;

calculating a second difference between features of the fourth image and features of the first image; and

controlling adjustment of the model based on comparison of a function of at least the first difference and the second difference to at least one threshold.

2 . The method according to claim 1 , wherein generating the third image based on the first image captured at the first position, the spatial distribution, and the first position spatial information comprises:

generating a first model based on the first image and the spatial distribution;

generating a second model by performing a grid transformation of the first model based on the pose of the camera in the first position spatial information; and

generating the third image by performing a two-dimensional (2D) transformation of the second model.

3 . The method according to claim 2 , wherein generating the first model based on the first image and the spatial distribution comprises:

determining whether a probability value above a predetermined threshold of the spatial distribution exists at a position in the first image; and

acquiring one or more of color information, texture information, and depth information at the position in response to the probability value above the predetermined threshold of the spatial distribution existing at the position in the first image.

4 . The method according to claim 2 , wherein generating the second model by performing the grid transformation of the first model based on the pose of the camera in the first position spatial information further comprises:

predicting features in the second model by applying a bicubic interpolation method on features in the first model.

5 . The method according to claim 1 , wherein generating the fourth image based on the second image captured at the second position, the spatial distribution, and the second position spatial information comprises:

generating a third model based on the second image and the spatial distribution;

generating a fourth model by performing a grid transformation of the third model based on the pose of the camera in the second position spatial information; and

generating the fourth image by performing a 2D transformation of the fourth model.

6 . The method according to claim 1 , wherein features of the third image correspond to features of the second image, and features of the fourth image correspond to features of the first image.

7 . The method according to claim 1 , wherein adjusting controlling adjustment of the model comprises:

continuing to adjust the model of the object in response to a sum of the first difference and the second difference being greater than a predetermined threshold; and

stopping adjustment of the model of the object in response to the sum of the first difference and the second difference being smaller than the predetermined threshold.

8 . The method according to claim 1 , wherein adjusting the model based on the first image, the second image, the third image, and the fourth image further comprises:

adjusting the model of the object based on spatial density information, color information, texture information, and depth information in the first image, the second image, the third image, and the fourth image.

9 . The method according to claim 1 , wherein acquiring the first image based on the first position spatial information and acquiring the second image based on the second position spatial information comprises:

acquiring the first image and the second image by using the camera to shoot the same object at the first position and at the second position, respectively; and

wherein the camera is jittered during use.

10 . An electronic device, comprising:

at least one processor; and

memory coupled to the at least one processor and storing instructions, wherein the instructions, when executed by the at least one processor, cause the electronic device to perform actions comprising:

acquiring first position spatial information of a camera at a first position and second position spatial information of the camera at a second position, wherein the first position is different from the second position;

generating a spatial distribution based on the first position spatial information and the second position spatial information, wherein the spatial distribution represents a probability distribution of space occupied by an object captured from a pose of the camera;

generating a third image based on a first image captured at the first position, the spatial distribution, and the first position spatial information;

generating a fourth image based on a second image captured at the second position, the spatial distribution, and the second position spatial information; and

adjusting a model of the object based on the first image, the second image, the third image, and the fourth image;

wherein adjusting the model based on the first image, the second image, the third image, and the fourth image comprises:

calculating a first difference between features of the third image and features of the second image;

calculating a second difference between features of the fourth image and features of the first image; and

controlling adjustment of the model based on comparison of a function of at least the first difference and the second difference to at least one threshold.

11 . The electronic device according to claim 10 , wherein generating the third image based on the first image captured at the first position, the spatial distribution, and the first position spatial information comprises:

generating a first model based on the first image and the spatial distribution;

generating a second model by performing a grid transformation of the first model based on the pose of the camera in the first position spatial information; and

generating the third image by performing a two-dimensional (2D) transformation of the second model.

12 . The electronic device according to claim 11 , wherein generating the first model based on the first image and the spatial distribution comprises:

determining whether a probability value above a predetermined threshold of the spatial distribution exists at a position in the first image; and

acquiring one or more of color information, texture information, and depth information at the position in response to the probability value above the predetermined threshold of the spatial distribution existing at the position in the first image.

13 . The electronic device according to claim 11 , wherein generating the second model by performing the grid transformation of the first model based on the pose of the camera in the first position spatial information further comprises: predicting features in the second model by applying a bicubic interpolation method on features in the first model.

14 . The electronic device according to claim 10 , wherein generating the fourth image based on the second image captured at the second position, the spatial distribution, and the second position spatial information comprises:

generating a third model based on the second image and the spatial distribution;

generating a fourth model by performing a grid transformation of the third model based on the pose of the camera in the second position spatial information; and

generating the fourth image by performing a 2D transformation of the fourth model.

15 . The electronic device according to claim 10 , wherein features of the third image correspond to features of the second image, and features of the fourth image correspond to features of the first image.

16 . The electronic device according to claim 10 , wherein controlling adjustment of the model comprises:

continuing to adjust the model of the object in response to a sum of the first difference and the second difference being greater than a predetermined threshold; and

stopping adjustment of the model of the object in response to the sum of the first difference and the second difference being smaller than the predetermined threshold.

17 . The electronic device according to claim 10 , wherein adjusting the model based on the first image, the second image, the third image, and the fourth image further comprises:

adjusting the model of the object based on spatial density information, color information, texture information, and depth information in the first image, the second image, the third image, and the fourth image.

18 . The electronic device according to claim 10 , wherein acquiring the first image based on the first position spatial information and acquiring the second image based on the second position spatial information comprises:

acquiring the first image and the second image by using the camera to shoot the same object at the first position and at the second position, respectively; and

wherein the camera is jittered during use.

19 . A computer program product comprising a non-transitory computer-readable storage medium having machine-executable instructions stored therein, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising:

acquiring first position spatial information of a camera at a first position and second position spatial information of the camera at a second position, wherein the first position is different from the second position;

generating a spatial distribution based on the first position spatial information and the second position spatial information, wherein the spatial distribution represents a probability distribution of space occupied by an object captured from a pose of the camera;

generating a third image based on a first image captured at the first position, the spatial distribution, and the first position spatial information;

generating a fourth image based on a second image captured at the second position, the spatial distribution, and the second position spatial information; and

adjusting a model of the object based on the first image, the second image, the third image, and the fourth image;

wherein adjusting the model based on the first image, the second image, the third image, and the fourth image comprises:

calculating a first difference between features of the third image and features of the second image;

calculating a second difference between features of the fourth image and features of the first image; and

controlling adjustment of the model based on comparison of a function of at least the first difference and the second difference to at least one threshold.

20 . The computer program product according to claim 19 , wherein generating the third image based on the first image captured at the first position, the spatial distribution, and the first position spatial information comprises:

generating a first model based on the first image and the spatial distribution;

generating a second model by performing a grid transformation of the first model based on the pose of the camera in the first position spatial information; and

generating the third image by performing a two-dimensional (2D) transformation of the second model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: LIU, ZHISONG; WANG, ZIJIA; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 063993/0064 →
Priority Claims (1)
CN 202310645471.7 · Jun 1, 2023 · national
Continuity (1)
Related Publication 20240404191A1 · Dec 5, 2024
References Cited (29)
US 20110007072A1 · Khan · 2011 [cited by examiner]
US 20140025331A1 · Ma · 2014 [cited by examiner]
US 20200043122A1 · Xiao · 2020 [cited by examiner]
US 20200090575A1 · Martin · 2020 [cited by examiner]
US 20220301182A1 · Mahjourian · 2022 [cited by examiner]
US 20230104977A1 · Zaghetto · 2023 [cited by examiner]
US 20250124601A1 · Park · 2025 [cited by examiner]
B. Mildenhall et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” European Conference on Computer Vision, Aug. 2020, 17 pages. [cited by applicant]
T. Müller et al., “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding,” arXiv:2201.05989v2, May 4, 2022, 15 pages. [cited by applicant]
J. T. Barron et al., “Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields,” Conference on Computer Vision and Pattern Recognition, arXiv:2111.12077v3, Mar. 25, 2022, 18 pages. [cited by applicant]
G. Gafni et al., “Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 8649-8658. [cited by applicant]
S. Liu et al., “Editing Conditional Radiance Fields,” International Conference on Computer Vision, arXiv:2105.06466v2, Jun. 4, 2021, 24 pages. [cited by applicant]
C. Wang et al., “CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields,” arXiv:2112.05139v3, Mar. 2, 2022, 13 pages. [cited by applicant]
A. Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” International Conference on Machine Learning, arXiv:2103.00020v1, Feb. 26, 2021, 48 pages. [cited by applicant]
S. Zhi et al., “In-Place Scene Labelling and Understanding with Implicit Scene Representation,” International Conference on Computer Vision, arXiv:2103.15875v2, Aug. 21, 2021, 14 pages. [cited by applicant]
S. Kobayashi et al., “Decomposing NeRF for Editing via Feature Field Distillation,” arXiv:2205.15585v1, May 31, 2022, 23 pages. [cited by applicant]
I. Mehta et al., “Modulated Periodic Activations for Generalizable Local Functional Representations,” IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 14214-14223. [cited by applicant]
X. Zhang et al., “NeRFactor: Neural Factorization of Shape and Reflectance under an Unknown Illumination,” ACM Transactions on Graphics, vol. 40, No. 6, Dec. 2021, pp. 237:1-237:18. [cited by applicant]
V. Rudnev et al., “NeRF for Outdoor Scene Relighting,” 17th European Conference on Computer Vision, Oct. 2022, 17 pages. [cited by applicant]
M. Boss et al., “NeRD: Neural Reflectance Decomposition from Image Collections,” IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 12684-12694. [cited by applicant]
R. Basri et al., “Lambertian Reflectance and Linear Subspaces,” Weizmann Institute of Science, Technical Report MCS00-21, NEC Research Institute Technical Report 2000-172R, Mar. 2023, 27 pages. [cited by applicant]
B. Li et al., “Language-driven Semantic Segmentation,” International Conference on Learning Representations, arXiv:2201.03546v2, Apr. 3, 2022, 13 pages. [cited by applicant]
M. Caron et al., “Emerging Properties in Self-Supervised Vision Transformers,” IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jun. 2021, pp. 9650-9660. [cited by applicant]
J. L. Schonberger et al., “Structure-from-Motion Revisited,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 4104-4113. [cited by applicant]
W. Cheng et al., “Cascaded Parallel Filtering for Memory-Efficient Image-Based Localization,” IEEE/CVF International Conference on Computer Vision (ICCV), Nov. 2019, pp. 1032-1041. [cited by applicant]
D. Campbell et al., “Solving the Blind Perspective-n-Point Problem End-To-End With Robust Differentiable Geometric Optimization,” arXiv:2007.14628v2, Sep. 8, 2020, 18 pages. [cited by applicant]
U.S. Appl. No. 17/984,474 filed in the name of Zhisong Liu et al. on Nov. 10, 2022, and entitled “Method, Electronic Device, and Computer Program Product for Generating Three-Dimensional Scene.” [cited by applicant]
U.S. Appl. No. 18/144,289 filed in the name of Zhisong Liu et al. on May 8, 2023, and entitled “Method, Device, and Computer Program Product for Rendering.” [cited by applicant]
U.S. Appl. No. 18/319,577 filed in the name of Zhisong Liu et al. on May 18, 2023, and entitled “Method, Device, and Computer Program Product for Determining Camera Pose for an Image.” [cited by applicant]