IP Library Granted Patent US 12,614,306
Granted Patent B2
US 12,614,306 · App. 18/333,471 · Granted Apr 28, 2026

Inverting neural radiance fields for part and scene pose estimation

Inventors: Kin Gwn Lore (Belmont, MA); Ganesh Sundaramoorthi (Duluth, GA); Xinhua Zhang (Weatogue, CT)
Assignee: RTX CORPORATION
G06T7/75G06N3/06G06N3/084G06N20/00G06T7/11G06T19/006G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,306
App. No.
18/333,471
Granted
Apr 28, 2026
Kind
B2
Abstract

A method for generating an application specific pose estimation responsive to a camera pose, a scene pose and an object pose includes generating an observed image (I obs ) of an object and a scene using a camera and generating a current pose estimate having an object image render and a scene image render. The method further includes generating a final rendered output (I rndr ) by combining the object image render and the scene image render and processing the final rendered output (I rndr ) and the observed image (I obs ) to generate a final pose estimate, wherein the final pose estimate matches the observed image (I obs ).

Claims (45)

1 . A method that generates an application specific pose estimation responsive to a camera pose, a scene pose and an object pose, the method comprising:

generating an observed image (Iobs) of an object and a scene using a camera;

processing the observed image to generate a current pose estimate having a current object pose and a current scene pose, the processing including segmenting the observed image using a 2D segmentation technique to obtain from the observed image a first segmented component representing a scene image of the current scene pose and a second segmented component representing an object image representing the current object pose that is an isolated from the first segmented component;

processing the current pose estimate to,

generate an object render by applying the object image to a first Neural Radiance Field (NeRF), and

generate a scene render by applying, independently from the object image, the scene image to a second NeRF that is different from the first NeRF;

generating a final rendered output (Irndr) by combining the object render and the scene render;

generating a final pose estimate, wherein generating a final pose estimate includes,

comparing the observed image (Iobs) with the final rendered output (Irndr) to obtain a loss metric,

backpropagating the loss metric to update the current pose estimate, and

repeating the generating a final pose estimate until the loss metric is equal or greater than a loss metric threshold such that the final rendered output (Irndr) matches the observed image (Iobs),

wherein the object render and the scene render are generated independently based on the segmented components, and

wherein the object render and the scene render are composited to generate the final rendered output (Irndr).

2 . The method of claim 1 , wherein generating an object image includes operating a moving camera to generate an image of a stationary object from at least one predefined angle relative to a position of the object.

3 . The method of claim 1 , wherein generating an object image includes operating a stationary camera to generate an image of a moving object from at least one predefined angle relative to a position of the object.

4 . The method of claim 1 , wherein generating a final rendered output (Irndr) includes compositing the object render and the scene render together using a predetermined compositing technique.

5 . The method of claim 1 , wherein generating a final pose estimate includes,

comparing the observed image (Iobs) with the final rendered output (Irndr) to obtain a loss metric, wherein the loss metric is responsive to a difference between the observed image (Iobs) and the final rendered output (Irndr).

6 . The method of claim 5 , wherein in response to the loss metric being below the predetermined loss metric threshold, backpropagating the loss metric to update the current pose estimate until the loss metric is equal or greater than the loss metric threshold.

7 . The method of claim 6 , wherein in response to the loss metric being below the predetermined loss metric threshold, repeating the generating a final pose estimate until the loss metric is equal to or greater than the predetermined loss metric threshold.

8 . The method of claim 7 , wherein in response to the loss metric being equal to or greater than the predetermined loss metric threshold, the final rendered output (Irndr) matches the observed image (Iobs) and is determined to be the final pose estimate.

9 . The method of claim 1 , wherein

generating an object render includes applying the object image to a Bundle-Adjusting Neural Radiance Fields (BARF), and

generating a scene render by applying the scene image to a BARF.

10 . A method that generates an application specific pose estimation responsive to a camera pose, a scene pose and at least one object pose, the method comprising:

generating an observed image (Iobs) of at least one object and a scene using a camera;

processing the observed image to generate a current pose estimate having a current object pose and a current scene pose, the processing including segmenting the observed image using a 2D segmentation technique to obtain from the observed image a first segmented component representing a scene image of the current scene pose and a second segmented component representing an object image representing the current object pose that is an isolated from the first segmented component:

generating an object render by applying an-the object image to a first Neural Radiance Field (NeRF), and

generating a scene render by applying, independently from the object image, the scene image to a second NeRF that is different from the first NeRF;

generating a final rendered output (Irndr) by combining the a object render and the scene render;

comparing the observed image (Iobs) with the final rendered output (Irndr) to obtain a loss metric, wherein the loss metric is responsive to a difference between the observed image (Iobs) and the final rendered output (Irndr);

processing the final rendered output (Irndr) and the observed image (Iobs) to generate a final pose estimate, wherein the final pose estimate matches the observed image (Iobs) when the loss metric is equal or greater than a loss metric threshold,

wherein the object render and the scene render are generated independently based on the segmented components, and

wherein the object render and the scene render are composited to generate the final rendered output (Irndr).

11 . The method of claim 10 , wherein generating the object image includes at least one of,

operating a moving camera to generate an image of a stationary object from at least one predefined angle relative to a position of the at least one object, and

operating a stationary camera to generate an image of a moving object from at least one predefined angle relative to a position of the at least one object.

12 . The method of claim 10 , wherein processing the observed image includes segmenting the observed image using a 2D segmentation technique to isolate and obtain a current scene pose and at least one current object pose from the observed image to obtain the segmented components.

13 . The method of claim 10 , wherein generating a final rendered output (Irndr) includes compositing the at least one object render and the scene render together using a predetermined compositing technique.

14 . The method of claim 10 , wherein in response to the loss metric is below a predetermined loss metric threshold, backpropagating the loss metric to update the current pose estimate until the loss metric is equal or greater than the loss metric threshold.

15 . The method of claim 14 , wherein in response to the loss metric is below the predetermined loss metric threshold, repeating the generating a final pose estimate until the loss metric is equal to or greater than the predetermined loss metric threshold.

16 . The method of claim 15 , wherein in response to the loss metric is equal to or greater than the predetermined loss metric threshold, the final rendered output (Irndr) matches the observed image (Iobs) and is determined to be the final pose estimate.

17 . The method of claim 10 , wherein

generating the object render includes applying the at least one object image to a Bundle-Adjusting Neural Radiance Fields (BARF), and

generating a scene render by applying the scene image to a BARF.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2023
From: LORE, KIN GWN; SUNDARAMOORTHI, GANESH; ZHANG, XINHUA
To: RAYTHEON TECHNOLOGIES CORPORATION
Reel/Frame 064631/0606 →
CHANGE OF NAME Recorded Jul 27, 2023
From: RAYTHEON TECHNOLOGIES CORPORATION
To: RTX CORPORATION
Reel/Frame 064402/0837 →
Continuity (1)
Related Publication 20240412409A1 · Dec 12, 2024
References Cited (19)
US 12394166B2 · Gori · 2025 [cited by examiner]
US 20220398806A1 · Arksey et al. · 2022 [cited by applicant]
US 20230230275A1 · Lin · 2023 [cited by examiner]
US 20230281913A1 · Rematas · 2023 [cited by examiner]
US 20230410339A1 · Sawarkar · 2023 [cited by examiner]
US 20240005598A1 · Sucar · 2024 [cited by examiner]
US 20240273811A1 · Radwan · 2024 [cited by examiner]
US 20240312123A1 · Anwar · 2024 [cited by examiner]
WO 2022104178A1 · 2022 [cited by applicant]
WO WO2024123989A1 · 2024 [cited by examiner]
WO WO2024129646A1 · 2024 [cited by examiner]
WO WO2024163018A1 · 2024 [cited by examiner]
Extended European Search Report for EP Application No. 24181675.0, dated Oct. 25, 2024, pp. 1-11. [cited by applicant]
Lin et al., “BARF: Bundle-Adjusting Neural Radiance Fields”, arxiv.org, Cornell University Library, Aug. 19, 2021, XP091024204, pp. 5741-5751. [cited by applicant]
Stelzner et al., “Decomposing 3D Scenes into Objects via Unsupervised Volume Segmentation”, arxiv.org, Cornell University Library, Apr. 2, 2021, XP081931740, pp. 1-15. [cited by applicant]
Yen-Chen et al., “iNeRF: Inverting Neural Radiance Fields for Pose Estimation”, arxiv.org, Cornell University Ithaca, NY 14853, Dec. 10, 2020, XP081834080, pp. 1-12. [cited by applicant]
Yen-Chen, Lin, “iNeRF: Inverting Neural Radiance Fields for Pose Estimation”, Dec. 8, 2020, XP093214193, retrieved from the Internet: https://www.youtube.com/watch?v=eQuCZaQNOtl&ab_channel=Yen-ChenLin, pp. 1-2. [cited by applicant]
Schonberger et al., “Structure-from-Motion Revisited”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4104-4113. [cited by applicant]
Wei et al., “DeepSFM: Structure From Motion via Deep Bundle Adjustment”, European Conference on Computer Vision, 2020, pp. 1-29. [cited by applicant]