Spatial mapping and planning with language models using knowledge graphs
In various examples, a technique for performing spatial mapping and planning with large language models using knowledge graphs may include querying a temporal knowledge graph to predict a next step that a robotic system is to traverse in an environment, wherein the temporal knowledge graph comprises a graph representation of entities and relationships in the environment, wherein the entities and relationships in the graph representation are produced based on sensor data captured from the environment. The technique also may include receiving additional sensor data captured from the environment. The technique further may include generating, via a machine learning model, spatial and temporal data associated with the additional sensor data. The technique still further may include updating, via the machine learning model, the temporal knowledge graph to include representations of the spatial and temporal data, based on a similarity between the spatial and temporal data and the graph representation.
1 . A method comprising:
querying a temporal knowledge graph to predict a next action that a robotic system is to execute in an environment, the temporal knowledge graph including a graph representation of entities and relationships between the entities in the environment that are produced based at least on sensor data captured from the environment;
receiving additional sensor data captured from the environment;
generating, using a machine learning model, spatial and temporal data associated with the additional sensor data; and
updating, using the machine learning model, the temporal knowledge graph to include representations of the spatial and temporal data based at least on a similarity between the spatial and temporal data and the graph representation.
2 . The method of claim 1 , wherein the temporal knowledge graph maintains a history of entities and relationships between the entities over a time period, and wherein the querying further comprises querying the history to predict the next action that the robotic system is to execute.
3 . The method of claim 1 , further comprising:
(i) querying the temporal knowledge graph to identify a subgraph that is similar to the additional sensor data;
(ii) comparing the similarity to a predetermined threshold; and
(iii) responsive to the similarity being above a predetermined threshold, adding entities to the temporal knowledge graph to update the underlying spatial representations for providing an updated prediction of the next action that the robotic system is to execute.
4 . The method of claim 3 , wherein (iii) further comprises adding one or more relations among the entities in the temporal knowledge graph.
5 . The method of claim 3 , further comprising:
(iv) responsive to the similarity being below the predetermined threshold, merging entities in the temporal knowledge graph to provide an updated prediction of the next action that the robotic system is to execute.
6 . The method of claim 5 , wherein (iv) further comprises altering the one or more relations in the temporal knowledge graph.
7 . The method of claim 1 , wherein the machine learning model comprises a graph neural network (GNN), a large language model (LLM), a vision language model (VLM), a multi-modal language model (MMLM), or a large action model (LAM).
8 . The method of claim 1 , wherein the robotic system comprises one or more of a robotic vehicle, a robotic arm, an autonomous or semi-autonomous vehicle, an aircraft, or a watercraft.
9 . The method of claim 1 , further comprising employing simultaneous localization and mapping (SLAM) to generate SLAM data for the machine learning model to provide the spatial data.
10 . The method of claim 9 , wherein the SLAM data is provided using one or more sensors of the robotic system, the one or more sensors comprising at least one of a light imaging detection and ranging (LiDAR) sensor, a global positioning system (GPS) sensor, a global navigation satellite system (GNSS), a visual camera, an infrared camera, a depth camera, a sonic sensor, or an ultrasonic sensor.
11 . One or more processors comprising processing circuitry to:
query a temporal knowledge graph to predict a next action that a robotic system is to execute in an environment, the temporal knowledge graph including a graph representation of entities and relationships between the entities in the environment that are produced based at least on sensor data captured from the environment;
receive additional sensor data captured from the environment;
generate, using a machine learning model, spatial and temporal data associated with the additional sensor data; and
update, using the machine learning model, the temporal knowledge graph to include representations of the spatial and temporal data based at least on a similarity between the spatial and temporal data and the graph representation.
12 . The one or more processors of claim 11 , wherein the temporal knowledge graph maintains a history of entities and relationships between the entities over a time period, and wherein the processing circuitry further queries the history to predict the next action that the robotic system is to execute.
13 . The one or more processors of claim 11 , wherein the processing circuitry further:
(i) queries the temporal knowledge graph to identify a subgraph that is similar to the additional sensor data;
(ii) compares the similarity to a predetermined threshold; and
(iii) responsive to the similarity being above a predetermined threshold, adds entities to the temporal knowledge graph to update the underlying spatial representations for providing an updated prediction of the next action that the robotic system is to execute.
14 . The one or more processors of claim 13 , wherein in (iii) the processing circuitry further adds one or more relations among the entities in the temporal knowledge graph.
15 . The one or more processors of claim 13 , wherein the processing circuitry further:
(iv) responsive to the similarity being below the predetermined threshold, merges entities in the temporal knowledge graph to provide an updated prediction of the next action that the robotic system is to execute.
16 . The one or more processors of claim 15 , wherein (iv) further comprises altering the one or more relations in the temporal knowledge graph.
17 . The one or more processors of claim 11 , wherein the machine learning model comprises a graph neural network (GNN), a large language model (LLM), a vision language model (VLM), a multi-modal large language model (MMLM), or a large action model (LAM).
18 . The one or more processors of claim 11 , wherein the one or more processors are comprised in at least one of:
a control system for the robotic system;
a perception system for the robotic system;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing one or more deep learning operations;
a system implemented using an edge device;
a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content;
a system implemented using a robot;
a system for performing one or more conversational AI operations;
a system implemented using one or more large language models (LLMs);
a system implementing one or more vision language models (VLMs);
a system implementing one or more multi-modal language models (MMLMs);
a system implementing one or more large action models (LAMs);
a system implementing one or more graph neural networks (GNNs);
a system for generating synthetic data;
a system for performing one or more generative AI operations;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.
19 . A system comprising one or more processors to:
query a temporal knowledge graph to predict a next action that a robotic system is to execute in an environment;
receive sensor data captured from the environment;
generate, via a machine learning model, spatial and temporal data associated with the additional sensor data; and
update, via the machine learning model, the temporal knowledge graph to include representations of the spatial and temporal data based at least on a similarity between the spatial and temporal data and the graph representation.
20 . The system of claim 19 , wherein the system is comprised in at least one of:
a control system for the robotic system;
a perception system for the robotic system;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing one or more deep learning operations;
a system implemented using an edge device;
a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content;
a system implemented using a robot;
a system for performing one or more conversational AI operations;
a system implemented using one or more large language models (LLMs);
a system implementing one or more vision language models (VLMs);
a system implementing one or more multi-modal language models (MMLMs);
a system implementing one or more large action models (LAMs);
a system implementing one or more graph neural networks (GNNs);
a system for generating synthetic data;
a system for performing one or more generative AI operations;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.