IP Library Granted Patent US 12,304,081
Granted Patent B2
US 12,304,081 · App. 17/112,699 · Granted May 20, 2025

Deep compositional robotic planners that follow natural language commands

Inventors: Yen-Ling Kuo (Cambridge, MA); Boris Katz (Cambridge, MA); Andrei Barbu (Cambridge, MA)
Assignee: Massachusetts Institute of Technology
B25J9/1664G06N3/042G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,304,081
App. No.
17/112,699
Granted
May 20, 2025
Kind
B2
Abstract

The present approach similarly combines task and motion planning, but does so without symbolic representations and begins with simpler tasks than other models in such domains can handle. Unlike prior approaches, the present approach does so in continuous action and state spaces which require many precise steps in the configuration space to execute what otherwise is a single output token such as “pick up” for discrete problems.

Claims (48)

1. A method comprising:

determining, by a planner, a first point in a selected neighborhood of a configuration space;

receiving a natural language input, uttered by a human or generated by a computer;

parsing the received natural language input to determine (i) linguistic structure of the natural language input and (ii) a plurality of words in the natural language input;

modifying a first neural network (NN) to encode the determined linguistic structure of the natural language input by arranging a plurality of component neural networks of the first NN in an order corresponding to the determined linguistic structure of the natural language input, wherein each component neural network corresponds to one or more respective words of the determined plurality of words;

determining, by the modified first NN, a second point in the selected neighborhood of the configuration space;

choosing among the first point and second point to generate an additional node to add to a search tree; and

adding the additional node to the search tree by connecting the additional node to a node associated with the selected neighborhood;

wherein a second NN controls behavior of an agent corresponding to the search tree having the additional node.

2. The method of claim 1 , further comprising:

determining the selected neighborhood by:

determining a first neighborhood to add the additional node to the search tree by evaluating one or more selected nodes of the search tree with the planner, each selected node representing a coordinate in the configuration space;

determining a second neighborhood to add the additional node to the search tree by evaluating the one or more selected nodes of the search tree with the modified first NN; and

choosing, as the selected neighborhood, one neighborhood among the first neighborhood and second neighborhood based on at least one of: a respective level of confidence determined for the first neighborhood and second neighborhood, at least one extrinsic factor, and an impossibility factor indicative of whether a path to each node of the selected neighborhood is possible.

3. The method of claim 1 , wherein choosing includes producing hypotheses for the planner and for the modified first NN, the hypotheses including at least one of confidence weights and features.

4. The method of claim 1 , wherein choosing includes evaluating one or more nodes of the search tree with the planner or the modified first NN, said evaluating being based on observations.

5. The method of claim 1 , wherein the second NN includes component NNs.

6. The method of claim 1 , wherein the second NN outputs to the agent at least one of a path, destination, and stopping point.

7. A system comprising:

a processor; and

a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the processor to:

determine, by a planner, a first point in a selected neighborhood of a configuration space;

receive a natural language input, uttered by a human or generated by a computer;

parse the received natural language input to determine (i) linguistic structure of the natural language input and (ii) a plurality of words in the natural language input;

modify a first neural network (NN) to encode the determined linguistic structure of the natural language input by arranging a plurality of component neural networks of the first NN in an order corresponding to the determined linguistic structure of the natural language input, wherein each component neural network corresponds to one or more respective words of the determined plurality of words,

determine, by the modified first NN, a second point in the selected neighborhood of the configuration space;

choose among the first point and second point to generate an additional node to add to a search tree; and

add the additional node to the search tree by connecting the additional node to a node associated with the selected neighborhood;

wherein a second NN controls behavior of an agent corresponding to the search tree having the additional node.

8. The system of claim 7 , wherein choosing includes producing hypotheses for the planner and for the modified first NN, the hypotheses including at least one of confidence weights and features.

9. The system of claim 7 , wherein the instructions further configure the processor to:

determine the selected neighborhood by:

determining a first neighborhood to add the additional node to the search tree by evaluating one or more selected nodes of the search tree with the planner, each selected node representing a coordinate in the configuration space;

determining a second neighborhood to add the additional node to the search tree by evaluating the one or more selected nodes of the search tree with the modified first NN; and

choosing as the selected neighborhood, one neighborhood among the first neighborhood and second neighborhood based on at least one of: a respective level of confidence determined for the first neighborhood and second neighborhood, at least one extrinsic factor, and an impossibility factor indicative of whether a path to each node of the selected neighborhood is possible.

10. A method comprising:

determining, by a planner, a first neighborhood of a configuration space and a first point in the first neighborhood;

receiving a natural language input, uttered by a human or generated by a computer;

parsing the received natural language input to determine (i) linguistic structure of the natural language input and (ii) a plurality of words in the natural language input;

modifying a first neural network (NN) to encode the determined linguistic structure of the natural language input by arranging a plurality of component neural networks of the first NN in an order corresponding to the determined linguistic structure of the natural language input, wherein each component neural network corresponds to one or more respective words of the determined plurality of words;

determining, by the modified first NN, a second neighborhood of a configuration space and a second point in the second neighborhood;

choosing as a selected neighborhood, one neighborhood among the first neighborhood and second neighborhood based on at least one of: a respective level of confidence determined for the first neighborhood and second neighborhood, at least one extrinsic factor, and an impossibility factor indicative of whether a path to each node of the selected neighborhood is possible;

choosing among the first point and second point to generate an additional node to add to a search tree; and

adding the additional node to the search tree by connecting the additional node to a node associated with the selected neighborhood;

wherein a second NN controls behavior of an agent corresponding to the search tree having the additional node.

11. The method of claim 1 , wherein the first neural network and the second neural network are component neural networks of a given neural network.

12. The system of claim 7 , wherein the first neural network and the second neural network are component neural networks of a given neural network.

13. The method of claim 10 , wherein the first neural network and the second neural network are component neural networks of a given neural network.

Assignments (2)
CONFIRMATORY LICENSE Recorded Oct 4, 2021
From: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 057697/0416 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: KUO, YEN-LING; KATZ, BORIS; BARBU, ANDREI
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 054552/0525 →
Continuity (3)
Provisional Application 62944924 · Dec 6, 2019
Provisional Application 62944932 · Dec 6, 2019
Related Publication 20210170594A1 · Jun 10, 2021
References Cited (45)
US 10210860B1 · Ward · 2019 [cited by examiner]
US 10380236B1 · Ganu · 2019 [cited by examiner]
US 11170293B2 · Gao · 2021 [cited by examiner]
US 20080091628A1 · Srinivasa · 2008 [cited by examiner]
US 20180307779A1 · Tellex · 2018 [cited by examiner]
US 20210005182A1 · Han · 2021 [cited by examiner]
Y.-L. Kuo, A. Barbu and B. Katz, “Deep Sequential Models for Sampling-Based Planning,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 2018, pp. 6490-6497, doi: 10.1109/IR… [cited by examiner]
Zang, Xiaoxue, et al. “Translating navigation instructions in natural language to a high-level plan for behavioral robot navigation.” arXiv preprint arXiv:1810.00663 (2018). (Year: 2018). [cited by examiner]
Kuo, Yen-Ling, Andrei Barbu, and Boris Katz. “Deep sequential models for sampling-based planning” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (Year: 2018). [cited by examiner]
Mei, Hongyuan, Mohit Bansal, and Matthew Walter. “Listen, attend, and walk: Neural mapping of navigational instructions to action sequences.” Proceedings of the AAAI Conference on Artificial Intelligence. vol. 30. No. 1… [cited by examiner]
Paxton, Chris, et al. “Prospection: Interpretable plans from language by predicting the future.” 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019 (Year: 2019). [cited by examiner]
Karaman, Sertac, and Emilio Frazzoli. “Sampling-based algorithms for optimal motion planning.” The international journal of robotics research 30.7 (2011): 846-894. [cited by applicant]
Elbanhawi, Mohamed, and Milan Simic. “Sampling-based robot motion planning: A review.” Ieee access 2 (2014): 56-77. [cited by applicant]
Gammell, Jonathan D., et al. “Batch informed trees (BIT*): Sampling-based optimal planning via the heuristically guided search of implicit random geometric graphs.” 2015 IEEE international conference on robotics and aut… [cited by applicant]
Burget, Felix, et al. “BI 2 RRT*: An efficient sampling-based path planning framework for task-constrained mobile manipulation.” 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 201… [cited by applicant]
Adiyatov, Olzhas, and Huseyin Atakan Varol. “A novel RRT*—based algorithm for motion planning in Dynamic environments.” 2017 IEEE International Conference on Mechatronics and Automation (ICMA). IEEE, 2017. [cited by applicant]
Urmson, Chris, and Reid Simmons. “Approaches for heuristically biasing RRT growth.” Proceedings 2003 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2003)(Cat. No. 03CH37453). vol. 2. IEEE, 200… [cited by applicant]
Lindemann, Stephen R., and Steven M. LaValle. “Incrementally reducing dispersion by increasing Voronoi bias in RRTs.” IEEE International Conference on Robotics and Automation, 2004. Proceedings. ICRA'04. 2004. vol. 4. I… [cited by applicant]
L. E. Baum and T. Petrie, “Statistical inference for probabilistic functions of finite state markov chains,” The annals of mathematical statistics, vol. 37, No. 6, pp. 1554-1563, 1966. [cited by applicant]
Hochreiter, Sepp, and Jürgen Schmidhuber. “Long short-term memory.” Neural computation 9.8 (1997): 1735-1780. [cited by applicant]
Siddharth, Narayanaswamy, et al. “Seeing what you're told: Sentence-guided activity recognition in video.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2014. [cited by applicant]
Yu, Haonan, et al. “A compositional framework for grounding language inference, generation, and acquisition in video.” Journal of Artificial Intelligence Research 52 (2015): 601-713. [cited by applicant]
Donahue, Jeffrey, et al. “Long-term recurrent convolutional networks for visual recognition and description.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2015. [cited by applicant]
Fulgenzi, Chiara, et al. “Probabilistic navigation in dynamic environment using rapidly-exploring random trees and gaussian processes.” 2008 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 200… [cited by applicant]
Barbu, Andrei, et al. “Simultaneous object detection, tracking, and event recognition.” arXiv preprint arXiv:1204.2741 (2012). [cited by applicant]
Ramanathan, Vignesh, et al. “Detecting events and key actors in multi-person videos.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. [cited by applicant]
Aoude, Georges S., et al. “Probabilistically safe motion planning to avoid dynamic obstacles with uncertain motion patterns.” Autonomous Robots 35.1 (2013): 51-76. [cited by applicant]
Le, Tuan Anh, et al. “Inference compilation and universal probabilistic programming.” Artificial Intelligence and Statistics. PMLR, 2017. [cited by applicant]
Kulkarni, Tejas D., et al. “Picture: A probabilistic programming language for scene perception.” Proceedings of the ieee conference on computer vision and pattern recognition. 2015. [cited by applicant]
Narayanaswamy, Siddharth, et al. “Seeing unseeability to see the unseeable.” arXiv preprint arXiv:1204.2801 (2012). [cited by applicant]
Bowen, Chris, and Ron Alterovitz. “Closed-loop global motion planning for reactive execution of learned tasks.” 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2014. [cited by applicant]
Bowen, Chris, and Ron Alterovitz. “Asymptotically optimal motion planning for tasks using learned virtual landmarks.” IEEE Robotics and Automation Letters 1.2 (2016): 1036-1043. [cited by applicant]
Kim, Beomjoon, et al. “Guiding search in continuous state-action spaces by learning an action sampler from off-target search experience.” Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 32. No. 1. 20… [cited by applicant]
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NIPS, 2014. [cited by applicant]
Janson, Lucas, et al. “Monte Carlo motion planning for robot trajectory optimization under uncertainty.” Robotics Research. Springer, Cham, 2018. 343-361. [cited by applicant]
Arslan, Omur, et al. “Sensory steering for sampling-based motion planning.” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2017. [cited by applicant]
Schmitt, Philipp S., et al. “Optimal, sampling-based manipulation planning.” 2017 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2017. [cited by applicant]
Čáp, Michal, et al. “Multi-agent RRT*: Sampling-based cooperative pathfinding.” arXiv preprint arXiv:1302.2828 (2013). [cited by applicant]
Chen, Yufan, et al.. “Decoupled multiagent path planning via incremental sequential convex programming.” 2015 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2015. [cited by applicant]
Kiesel, Scott, et al. “An effort bias for sampling-based motion planning.” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2017. [cited by applicant]
Hopfield, John J. “Neural networks and physical systems with emergent collective computational abilities.” Proceedings of the national academy of sciences 79.8 (1982): 2554-2558. [cited by applicant]
Cho, Kyunghyun, et al. “On the properties of neural machine translation: Encoder-decoder approaches.” arXiv preprint arXiv:1409.1259 (2014). [cited by applicant]
Palangi, Hamid, et al. “Deep sentence embedding using long short-term memory networks: Analysis and application to Information retrieval.” IEEE/ACM Transactions on Audio, Speech, and Language Processing 24.4 (2016): 694… [cited by applicant]
Sucan, Ioan A., Mark Moll, and Lydia E. Kavraki. “The open motion planning library.” IEEE Robotics & Automation Magazine 19.4 (2012): 72-82. [cited by applicant]
Hsu, David, et al. “Path planning in expansive configuration spaces.” Proceedings of International Conference on Robotics and Automation. Vol. 3. IEEE, 1997. [cited by applicant]
Cited By (2)
US 12,443,881 US 12,443,885