IP Library Granted Patent US 12,561,830
Granted Patent B2
US 12,561,830 · App. 18/539,157 · Granted Feb 24, 2026

Neural implicit scattering functions for inverse parameter estimation and dynamics modeling of multi-object interactions

Inventors: Stephen Tian (Stanford, CA); Yancheng Cai (Cambridge, GB); Hong-Xing Yu (Stanford, CA); Sergey Zakharov (Stanford, CA); Katherine Liu (Mountain View, CA); Adrien David Gaidon (Mountain View, CA); Yunzhu Li (Stanford, CA); Jiajun Wu (Stanford, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSSHA KABUSHIKI KAISHA; THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
G06T7/70B25J9/1697G06V10/60G06F16/95G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,830
App. No.
18/539,157
Granted
Feb 24, 2026
Kind
B2
Abstract

A method for dynamic modeling and manipulation of multi-object scenes is described. The method includes using object-centric neural implicit scattering functions (OSFs) as object representations in a model-predictive control (MPC) framework for the multi-object scenes. The method also includes modeling a per-object light transport to enable compositional scene re-rendering under object rearrangement and varying lighting conditions. The method further includes applying inverse parameter estimation and graph neural network (GNN) dynamics models to estimate initial object poses and a light position in the multi-object scene. The method also includes manipulating an object perceived in the multi-object scene according to the applying of the inverse parameter estimation and the GNN dynamics models.

Claims (46)

1 . A method for dynamic modeling and manipulation of multi-object scenes, the method comprising:

using object-centric neural implicit scattering functions (OSFs) as object representations in a model-predictive control (MPC) framework for the multi-object scenes;

modeling a per-object light transport to enable compositional scene re-rendering under object rearrangement and varying lighting conditions;

applying inverse parameter estimation and graph neural network (GNN) dynamics models to estimate initial object poses and a light position in the multi-object scene; and

manipulating an object perceived in the multi-object scene according to the applying of the inverse parameter estimation and the GNN dynamics models.

2 . The method of claim 1 , in which the multi-object scene exhibits multi-object interactions in extreme and/or harsh lighting conditions.

3 . The method of claim 1 , further comprising:

training a dynamics model using simulated data;

inputting 6D object poses of each object in the multi-object scene and a pusher's action; and

outputting future 6D poses for each object in the multi-object scene.

4 . The method of claim 3 , further comprising:

estimating the initial object poses in the multi-object scene and the light positions;

inferring the initial object poses in the multi-object scene and the light positions using the inverse parameter estimation using the OSFs;

performing model predictive control using the GNN dynamics model; and

updating object states at each step using the inverse parameter estimation.

5 . The method of claim 1 , further comprising training the GNN dynamics model using simulated data, including inputting 6D object poses of each object in the scene and a pusher's action, and to output future 6-D poses for each object.

6 . The method of claim 1 , further comprising using the inverse parameter estimation with OSF models to estimate the initial object poses and the light position at inference time.

7 . The method of claim 1 , further comprising using a learned, GNN dynamics model to perform model-predictive control while updating object states at each step using the inverse parameter estimation.

8 . The method of claim 1 , further comprising planning pushing and rearranging of an object by a robot according to predictions from video captured by the robot.

9 . A non-transitory computer-readable medium having program code recorded thereon for dynamic modeling and manipulation of multi-object scenes, the program code being executed by a processor and comprising:

program code to use object-centric neural implicit scattering functions (OSFs) as object representations in a model-predictive control (MPC) framework for the multi-object scenes;

program code to model a per-object light transport to enable compositional scene re-rendering under object rearrangement and varying lighting conditions;

program code to apply inverse parameter estimation and graph neural network (GNN) dynamics models to estimate initial object poses and a light position in the multi-object scene; and

program code to manipulate an object perceived in the multi-object scene according to the applying of the inverse parameter estimation and the GNN dynamics models.

10 . The non-transitory computer-readable medium of claim 9 , in which the multi-object scene exhibits multi-object interactions in extreme and/or harsh lighting conditions.

11 . The non-transitory computer-readable medium of claim 9 , further comprising:

program code to train a dynamics model using simulated data;

program code to input 6D object poses of each object in the multi-object scene and a pusher's action; and

program code to output future 6D poses for each object in the multi-object scene.

12 . The non-transitory computer-readable medium of claim 11 , further comprising:

program code to estimate the initial object poses in the multi-object scene and the light positions;

program code to infer the initial object poses in the multi-object scene and the light positions using the inverse parameter estimation using the OSFs;

program code to perform model predictive control using the GNN dynamics model; and

program code to update object states at each step using the inverse parameter estimation.

13 . The non-transitory computer-readable medium of claim 9 , further comprising program code to train the GNN dynamics model using simulated data, including inputting 6D object poses of each object in the scene and a pusher's action, and to output future 6-D poses for each object.

14 . The non-transitory computer-readable medium of claim 9 , further comprising program code to use inverse parameter estimation with OSF models to estimate the initial object poses and the light position at inference time.

15 . The non-transitory computer-readable medium of claim 9 , further comprising program code to use a learned, GNN dynamics model to perform model-predictive control while updating object states at each step using the inverse parameter estimation.

16 . The non-transitory computer-readable medium of claim 9 , further comprising planning pushing and rearranging of an object by a robot according to predictions from video captured by the robot.

17 . A system for dynamic modeling and manipulation of multi-object scenes, the system comprising:

a framework module to use object-centric neural implicit scattering functions (OSFs) as object representations in a model-predictive control (MPC) framework for the multi-object scenes;

a per object light transfer model to model a per-object light transport to enable compositional scene re-rendering under object rearrangement and varying lighting conditions;

a graph-based neural dynamics model to apply inverse parameter estimation and graph neural network (GNN) dynamics models to estimate initial object poses and a light position in the multi-object scene; and

an object manipulation module to manipulate an object perceived in the multi-object scene according to the applying of the inverse parameter estimation and the GNN dynamics models.

18 . The system of claim 17 , in which the multi-object scene exhibits multi-object interactions in extreme and/or harsh lighting conditions.

19 . The system of claim 17 , in which the graph-based neural dynamics model is further to use a learned, GNN dynamics model to perform model-predictive control while updating object states at each step using the inverse parameter estimation.

20 . The system of claim 17 , further comprising a planner module to push and/or rearrange an object by a robot according to predictions from video captured by the robot.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2026
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 074900/0380 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2023
From: ZAKHAROV, SERGEY; LIU, KATHERINE; GAIDON, ADRIEN DAVID
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 065884/0079 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2023
From: TIAN, STEPHEN; CAI, YANCHENG; YU, HONG-XING; LI, YUNZHU; WU, JIAJUN
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 065884/0427 →
Continuity (2)
Provisional Application 63440971 · Jan 25, 2023
Related Publication 20240249426A1 · Jul 25, 2024
References Cited (9)
US 20180253869A1 · Yumer · 2018 [cited by examiner]
US 20210133990A1 · Eckart · 2021 [cited by examiner]
US 20220076432A1 · Ramezani · 2022 [cited by examiner]
US 20220076447A1 · He · 2022 [cited by examiner]
US 20220126445A1 · Zhu · 2022 [cited by examiner]
US 20240153101A1 · Ye · 2024 [cited by examiner]
Guo, Michelle, et al. “Object-centric neural scene rendering.” arXiv preprint arXiv:2012.08503 (2020). (Year: 2020). [cited by examiner]
Driess, Danny, et al. “Learning multi-object dynamics with compositional neural radiance fields.” Conference on robot learning. PMLR, 2023. (Year: 2023). [cited by examiner]
Driess, Danny, et al., “Learning Multi-Object Dynamics with Compositional Neural Radiance Fields,” arXiv:2202.11855v3 [cs.CV] Jul. 27, 2022. [cited by applicant]