IP Library Granted Patent US 12,223,949
Granted Patent B2
US 12,223,949 · App. 17/930,349 · Granted Feb 11, 2025

Semantic rearrangement of unknown objects from natural language commands

Inventors: Christopher Jason Paxton (Pittsburgh, PA); Weiyu Liu (Atlanta, GA); Tucker Ryer Hermans (Salt Lake City, UT); Dieter Fox (Seattle, WA)
Assignee: NVIDIA Corporation
G10L15/1815B25J13/003G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,949
App. No.
17/930,349
Granted
Feb 11, 2025
Kind
B2
Abstract

A robotic system is provided for performing rearrangement tasks guided by a natural language instruction. The system can include a number of neural networks used to determine a selected rearrangement of the objects in accordance with the natural language instruction. A target object predictor network processes a point cloud of the scene and the natural language instruction to identify a set of query objects that are to-be-rearranged. A language conditioned prior network processes the point cloud, natural language instruction, and the set of query objects to sample a distribution of rearrangements to generate a number of sets of pose offsets for the set of query objects. A discriminator network then processes the samples to generate scores for the samples. The samples may be refined until a score for at least one of the sample generated by the discriminator network is above a threshold value.

Claims (40)

1. A method, comprising:

receiving a point cloud and a natural language instruction;

generating, via a first network, an instance mask based on the natural language instruction, wherein the instance mask identifies a set of query objects in the point cloud;

generating, via a second network, samples of pose offsets for the set of query objects in accordance with the natural language instruction; and

identifying, via a discriminator network, a selected rearrangement of the set of query objects in accordance with scores.

2. The method of claim 1 , wherein the point cloud represents a scene including a plurality of objects, and wherein the set of query objects is a subset of the objects in the scene that are identified to be moved in order to satisfy the natural language instruction.

3. The method of claim 1 , wherein at least one of the first network or the second network is configured to generate embeddings for the point cloud and the natural language instruction used as inputs to the first network or the second network.

4. The method of claim 1 , wherein at least one of the first network or the second network is a transformer-based neural network comprising an encoder and a decoder, and wherein each of the encoder and the decoder comprises a plurality of layers.

5. The method of claim 4 , wherein a layer of the encoder includes a first sub-layer that includes a multi-head attention mechanism and a second sub-layer that includes a fully-connected feed-forward layer.

6. The method of claim 5 , wherein a layer of the decoder includes a first sub-layer that includes a masked multi-head attention mechanism, a second sub-layer that includes a multi-head attention mechanism, and a third sub-layer that includes a fully-connected feed-forward layer.

7. The method of claim 1 , wherein each sample comprises a set of pose offsets corresponding to the set of query objects, and wherein each pose offset comprises a parameterization of a six degree of freedom (6-DOF) transformation.

8. The method of claim 1 , wherein identifying the selected rearrangement comprises:

comparing each of the scores to a threshold value; and

selecting a particular sample having a score that is greater than the threshold value as the selected rearrangement, or

responsive to determining that no samples have a score greater than the threshold value, refining the samples using a cross-entropy method.

9. The method of claim 1 , further comprising:

rearranging a number of objects using a robot based on the selected rearrangement.

10. The method of claim 1 , further comprising training the discriminator network using a set of scenes and a set of randomly perturbed scenes as positive and negative samples, respectively.

11. A system for rearranging objects in a scene, the system comprising:

a memory storing a point cloud and a natural language instruction; and

at least one processor configured to:

generate, via a first network, an instance mask based on the natural language instruction, wherein the instance mask identifies a set of query objects in the point cloud;

generate, via a second network, samples of pose offsets for the set of query objects in accordance with the natural language instruction; and

identify, via a discriminator network, a selected rearrangement of the set of query objects in accordance with scores.

12. The system of claim 11 , wherein the point cloud represents a scene including a plurality of objects, and wherein the set of query objects is a subset of the objects in the scene that are identified to be moved in order to satisfy the natural language instruction.

13. The system of claim 11 , wherein at least one of the first network or the second network is configured to generate embeddings for the point cloud and the natural language instruction used as inputs to the first network or the second network.

14. The system of claim 11 , wherein at least one of the first network or the second network is a transformer-based neural network comprising an encoder and a decoder, and wherein each of the encoder and the decoder comprises a plurality of layers.

15. The system of claim 14 , wherein a layer of the encoder includes a first sub-layer that includes a multi-head attention mechanism and a second sub-layer that includes a fully-connected feed-forward layer.

16. The system of claim 15 , wherein a layer of the decoder includes a first sub-layer that includes a masked multi-head attention mechanism, a second sub-layer that includes a multi-head attention mechanism, and a third sub-layer that includes a fully-connected feed-forward layer.

17. The system of claim 11 , wherein each sample comprises a set of pose offsets corresponding to the set of query objects, and wherein each pose offset comprises a parameterization of a six degree of freedom (6-DOF) transformation.

18. The system of claim 11 , wherein identifying the selected rearrangement comprises:

comparing each of the scores to a threshold value; and

selecting a particular sample having a score that is greater than the threshold value as the selected rearrangement, or

responsive to determining that no samples have a score greater than the threshold value, refining the samples using a cross-entropy method.

19. The system of claim 11 , further comprising a robot that rearranges a number of objects based on the selected rearrangement.

20. A non-transitory computer readable medium storing instructions that, responsive to being executed by at least one processor, cause a system to:

receive a point cloud and a natural language instruction;

generate, via a first network, an instance mask based on the natural language instruction, wherein the instance mask identifies a set of query objects in the point cloud;

generate, via a second network, samples of pose offsets for the set of query objects in accordance with the natural language instruction; and

identify, via a discriminator network, a selected rearrangement of the set of query objects in accordance with scores.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2022
From: PAXTON, CHRISTOPHER JASON; LIU, WEIYU; HERMANS, TUCKER RYER; FOX, DIETER
To: NVIDIA CORPORATION
Reel/Frame 061080/0687 →
Continuity (2)
Provisional Application 63241293 · Sep 7, 2021
Related Publication 20230073154A1 · Mar 9, 2023
References Cited (27)
US 20170372527A1 · Murali · 2017 [cited by examiner]
US 20180268065A1 · Parepally · 2018 [cited by examiner]
US 20190172261A1 · Alt · 2019 [cited by examiner]
US 20200368616A1 · Delamont · 2020 [cited by examiner]
US 20220335647A1 · Shrivastava · 2022 [cited by examiner]
Mees, O., et al., “Metric learning for generalizing spatial relations to new objects,” in 2017 IEEE/RSJ Int'l Conference on Intelligent Robots and Systems (IROS), IEEE, 2017, pp. 3175-3182. [cited by applicant]
Mees, O., et al., “Learning object placements for relational instructions by hallucinating scene representations,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2020, pp. 94-100. [cited by applicant]
Janner, M., et al., “Representation learning for grounded spatial reasoning,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 49-61, 2018. [cited by applicant]
Zhu, Y., et al., “Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs,” arXiv preprint arXiv:2012.07277, 2020. [cited by applicant]
Bisk, Y., et al., “Learning interpretable spatial operations in a rich 3D blocks world,” in Thirty-Second AAAI Conference no Artificial Intelligence, 2018. [cited by applicant]
Hristov, Y., et al., “Disentangled relational representations for explaining and learning from demonstration,” in Conference on Robot Learning, PMLR, 2020, pp. 870-884. [cited by applicant]
Johnson, J., et al., “CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2901-2910. [cited by applicant]
Zhang, C., et al., “RAVEN: a dataset for relational and analogical visual reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 5317-5327. [cited by applicant]
Barrett, D., et al., “Measuring abstract reasoning in neural networks,” in International Conference on Machine Learning, PMLR, 2018, pp. 511-520. [cited by applicant]
Suhr, A., et al., “A corpus of natural language for visual reasoning,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (vol. 2: short papers), 2018, pp. 217-223. [cited by applicant]
Achlioptas, P., et al., “Learning representations and generative models for 3D point clouds,” in International Conference on Machine Learning, PMLR, 2018, pp. 40-49. [cited by applicant]
Park, J.J., et al., “DeepSDF: Learning continuous signed distance functions for shape representation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 165-174. [cited by applicant]
Mo, K., et al., “StructureNet: Hierarchical graph networks for 3D shape generation,” arXiv preprint arXiv:1908.00575, 2019. [cited by applicant]
Wu, R., et al., “PQ-NET: A generative part seq2seq network for 3D shapes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 829-838. [cited by applicant]
Li, J., et al., “GRASS: Generative recursive autoencoders for shape structures,” ACM Transactions on Graphics (TOG), vol. 36, No. 4, pp. 1-14, 2017. [cited by applicant]
Li, M., et al., “GRAINS: Generative recursive autoencoders for INdoor scenes,” ACM Transactions on Graphics (TOG), vol. 38, No. 2, pp. 1-16, 2019. [cited by applicant]
Vaswani, A., et al., “Attention is all you need,” in Advances in neural information processing systems, 2017, pp. 5998-6008. [cited by applicant]
De Boer, P.T., et al., “A tutorial on the cross-entropy method,” Annals of operations research, vol. 134, No. 1, pp. 19-67, 2005. [cited by applicant]
Morrical, N., et al., “NViSII: A scriptable tool for photorealistic image generation,” arXiv preprint arXiv:2105.13962, 2021. [cited by applicant]
Wang, X., et al., “SceneFormer: Indoor scene generation with transformers,” arXiv preprint arXiv:2012.09793, 2020. [cited by applicant]
Radevski, G., et al., “Decoding language spatial relations to 2D arrangements,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, 2020, pp. 4549-4560. [cited by applicant]
Qureshi, A.H., et al., “NeRP: Neural rearrangement planning for unknown objects,” arXiv preprint arXiv:2106.01352, 2021. [cited by applicant]