Category-level manipulation from visual demonstration
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for robotic control using demonstrations to learn category-level manipulation task. One of the methods includes obtaining a collection of object models for a plurality of different types of objects belonging to a same object category and training a category-level representation in a category-level space from the collection of object models. A category-level trajectory is generated the demonstration data of a demonstration object. For a new object in the object category, a trajectory projection is generated in the category-level space, which is used to cause a robot to perform the robotic manipulation task on the new object.
1 . A method performed by one or more computers, the method comprising:
obtaining demonstration data representing a trajectory of a demonstration object while a manipulation task is performed on the demonstration object;
generating a category-level trajectory from the demonstration data, including generating a sequence of poses of the demonstration object relative to a reference point in an environment of the demonstration object;
generating, for each of the sequence of poses, a respective corresponding attention heatmap and an anchor point defining an origin of a category-level coordinate system;
receiving data representing a new object belonging to the object category;
generating a trajectory projection of the category-level trajectory according to the representation of the new instance of the object category in the category-level space by training a neural network to learn a mapping between a partial point cloud representation of the new object and a category-level representation of object models of the object category in the category-level space, wherein the neural network generates a location of the new object in the category-level space from the partial point cloud representation and wherein generating the trajectory projection further comprises, after learning the mapping:
transferring the attention heatmap to the partial point cloud representation of the new object; and
determining, from the attention heatmap, an anchor point for the new object to align the coordinate frame of the new object with the category-level coordinate system for the trajectory projection;
using the trajectory projection to cause a robot to perform the manipulation task using the new object belonging to the object category.
2 . The method of claim 1 , further comprising:
obtaining a collection of object models for a plurality of different types of objects belonging to a same object category;
training a neural network to generate category-level representations in a category-level space from the collection of object models.
3 . The method of claim 2 , wherein training the network from the collection of object models comprises generating a non-uniform, normalized representation space and normalizing a point cloud representation of each object model along one or more dimensions.
4 . The method of claim 1 , wherein training the neural network comprises performing a training process entirely in simulation.
5 . The method of claim 4 , wherein training the neural network does not require gathering any real-world data or human annotated keypoints.
6 . The method of claim 1 , wherein the category-level trajectory is an object-centric trajectory representing a trajectory of an object.
7 . The method of claim 6 , wherein the category-level trajectory is agnostic to how the object is held by a robot.
8 . The method of claim 1 , wherein obtaining the demonstration data comprises obtaining video data of the demonstration object.
9 . The method of claim 8 , further comprising processing the video data to generate a sequence of partial point clouds of the demonstration object.
10 . The method of claim 9 , further comprising mapping the sequence of partial point clouds to the category-level space.
11 . The method of claim 1 , wherein the collection of object models comprises a plurality of CAD models of objects belonging to the same category.
12 . The method of claim 1 , wherein the category-level representation is generated from multiple different instances of objects belonging to the category.
13 . The method of claim 1 , wherein the new object is an object that has never been seen by the system.
14 . The method of claim 1 , wherein performing the robotic skill on the new object belonging to the object category does not require retraining a model or acquiring additional training data.
15 . A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
obtaining demonstration data representing a trajectory of a demonstration object while a manipulation task is performed on the demonstration object;
generating a category-level trajectory from the demonstration data, including generating a sequence of poses of the demonstration object relative to a reference point in the environment of the demonstration object;
generating, for each of the sequence of poses, a respective corresponding attention heatmap and an anchor point defining an origin of a category-level coordinate system;
receiving data representing a new object belonging to the object category;
generating a trajectory projection of the category-level trajectory according to the representation of the new instance of the object category in the category-level space by training a neural network to learn a mapping between a partial point cloud representation of the new object and a category-level representation of object models of the object category in the category-level space, wherein the neural network generates a location of the new object in the category-level space from the partial point cloud representation and wherein generating the trajectory projection further comprises, after learning the mapping:
transferring the attention heatmap to the partial point cloud representation of the new object; and
determining, from the attention heatmap, an anchor point for the new object to align the coordinate frame of the new object with the category-level coordinate system for the trajectory projection;
using the trajectory projection to cause a robot to perform the manipulation task using the new object belonging to the object category.
16 . The system of claim 15 , wherein the operations further comprise:
obtaining a collection of object models for a plurality of different types of objects belonging to a same object category;
training a neural network to generate category-level representations in a category-level space from the collection of object models.
17 . The system of claim 16 , wherein training the network from the collection of object models comprises generating a non-uniform, normalized representation space and normalizing a point cloud representation of each object model along one or more dimensions.
18 . The system of claim 16 , wherein the operations further comprise training the neural network to learn a mapping between partial point cloud representations of the object models and the category-level representation in a category-level space,
wherein the neural network generates a prediction of points in the category-level space.
19 . One or more non-transitory computer storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
obtaining demonstration data representing a trajectory of a demonstration object while a manipulation task is performed on the demonstration object;
generating a category-level trajectory from the demonstration data, including generating a sequence of poses of the demonstration object relative to a reference point in the environment of the demonstration object;
generating, for each of the sequence of poses, a respective corresponding attention heatmap and an anchor point defining an origin of a category-level coordinate system;
receiving data representing a new object belonging to the object category;
generating a trajectory projection of the category-level trajectory according to the representation of the new instance of the object category in the category-level space by training a neural network to learn a mapping between a partial point cloud representation of the new object and a category-level representation of object models of the object category in the category-level space, wherein the neural network generates a location of the new object in the category-level space from the partial point cloud representation and wherein generating the trajectory projection further comprises, after learning the mapping:
transferring the attention heatmap to the partial point cloud representation of the new object; and
determining, from the attention heatmap, an anchor point for the new object to align the coordinate frame of the new object with the category-level coordinate system for the trajectory projection;
using the trajectory projection to cause a robot to perform the manipulation task using the new object belonging to the object category.