IP Library › Granted Patent US 12,611,770
Granted Patent B2
US 12,611,770 · App. 18/103,256 · Granted Apr 28, 2026

Category-level manipulation from visual demonstration

Inventors: Wenzhao Lian (Fremont, CA); Bowen Wen (Bellevue, WA); Stefan Schaal (Mountain View, CA)
Assignee: Intrinsic Innovation LLC
B25J9/1664B25J9/163G05B19/4155G06N3/08G06V10/82G06V20/41G05B2219/50391
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,611,770
App. No.
18/103,256
Granted
Apr 28, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for robotic control using demonstrations to learn category-level manipulation task. One of the methods includes obtaining a collection of object models for a plurality of different types of objects belonging to a same object category and training a category-level representation in a category-level space from the collection of object models. A category-level trajectory is generated the demonstration data of a demonstration object. For a new object in the object category, a trajectory projection is generated in the category-level space, which is used to cause a robot to perform the robotic manipulation task on the new object.

Claims (49)

1 . A method performed by one or more computers, the method comprising:

obtaining demonstration data representing a trajectory of a demonstration object while a manipulation task is performed on the demonstration object;

generating a category-level trajectory from the demonstration data, including generating a sequence of poses of the demonstration object relative to a reference point in an environment of the demonstration object;

generating, for each of the sequence of poses, a respective corresponding attention heatmap and an anchor point defining an origin of a category-level coordinate system;

receiving data representing a new object belonging to the object category;

generating a trajectory projection of the category-level trajectory according to the representation of the new instance of the object category in the category-level space by training a neural network to learn a mapping between a partial point cloud representation of the new object and a category-level representation of object models of the object category in the category-level space, wherein the neural network generates a location of the new object in the category-level space from the partial point cloud representation and wherein generating the trajectory projection further comprises, after learning the mapping:

transferring the attention heatmap to the partial point cloud representation of the new object; and

determining, from the attention heatmap, an anchor point for the new object to align the coordinate frame of the new object with the category-level coordinate system for the trajectory projection;

using the trajectory projection to cause a robot to perform the manipulation task using the new object belonging to the object category.

2 . The method of claim 1 , further comprising:

obtaining a collection of object models for a plurality of different types of objects belonging to a same object category;

training a neural network to generate category-level representations in a category-level space from the collection of object models.

3 . The method of claim 2 , wherein training the network from the collection of object models comprises generating a non-uniform, normalized representation space and normalizing a point cloud representation of each object model along one or more dimensions.

4 . The method of claim 1 , wherein training the neural network comprises performing a training process entirely in simulation.

5 . The method of claim 4 , wherein training the neural network does not require gathering any real-world data or human annotated keypoints.

6 . The method of claim 1 , wherein the category-level trajectory is an object-centric trajectory representing a trajectory of an object.

7 . The method of claim 6 , wherein the category-level trajectory is agnostic to how the object is held by a robot.

8 . The method of claim 1 , wherein obtaining the demonstration data comprises obtaining video data of the demonstration object.

9 . The method of claim 8 , further comprising processing the video data to generate a sequence of partial point clouds of the demonstration object.

10 . The method of claim 9 , further comprising mapping the sequence of partial point clouds to the category-level space.

11 . The method of claim 1 , wherein the collection of object models comprises a plurality of CAD models of objects belonging to the same category.

12 . The method of claim 1 , wherein the category-level representation is generated from multiple different instances of objects belonging to the category.

13 . The method of claim 1 , wherein the new object is an object that has never been seen by the system.

14 . The method of claim 1 , wherein performing the robotic skill on the new object belonging to the object category does not require retraining a model or acquiring additional training data.

15 . A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining demonstration data representing a trajectory of a demonstration object while a manipulation task is performed on the demonstration object;

generating a category-level trajectory from the demonstration data, including generating a sequence of poses of the demonstration object relative to a reference point in the environment of the demonstration object;

generating, for each of the sequence of poses, a respective corresponding attention heatmap and an anchor point defining an origin of a category-level coordinate system;

receiving data representing a new object belonging to the object category;

generating a trajectory projection of the category-level trajectory according to the representation of the new instance of the object category in the category-level space by training a neural network to learn a mapping between a partial point cloud representation of the new object and a category-level representation of object models of the object category in the category-level space, wherein the neural network generates a location of the new object in the category-level space from the partial point cloud representation and wherein generating the trajectory projection further comprises, after learning the mapping:

transferring the attention heatmap to the partial point cloud representation of the new object; and

determining, from the attention heatmap, an anchor point for the new object to align the coordinate frame of the new object with the category-level coordinate system for the trajectory projection;

using the trajectory projection to cause a robot to perform the manipulation task using the new object belonging to the object category.

16 . The system of claim 15 , wherein the operations further comprise:

obtaining a collection of object models for a plurality of different types of objects belonging to a same object category;

training a neural network to generate category-level representations in a category-level space from the collection of object models.

17 . The system of claim 16 , wherein training the network from the collection of object models comprises generating a non-uniform, normalized representation space and normalizing a point cloud representation of each object model along one or more dimensions.

18 . The system of claim 16 , wherein the operations further comprise training the neural network to learn a mapping between partial point cloud representations of the object models and the category-level representation in a category-level space,

wherein the neural network generates a prediction of points in the category-level space.

19 . One or more non-transitory computer storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining demonstration data representing a trajectory of a demonstration object while a manipulation task is performed on the demonstration object;

generating a category-level trajectory from the demonstration data, including generating a sequence of poses of the demonstration object relative to a reference point in the environment of the demonstration object;

generating, for each of the sequence of poses, a respective corresponding attention heatmap and an anchor point defining an origin of a category-level coordinate system;

receiving data representing a new object belonging to the object category;

generating a trajectory projection of the category-level trajectory according to the representation of the new instance of the object category in the category-level space by training a neural network to learn a mapping between a partial point cloud representation of the new object and a category-level representation of object models of the object category in the category-level space, wherein the neural network generates a location of the new object in the category-level space from the partial point cloud representation and wherein generating the trajectory projection further comprises, after learning the mapping:

transferring the attention heatmap to the partial point cloud representation of the new object; and

determining, from the attention heatmap, an anchor point for the new object to align the coordinate frame of the new object with the category-level coordinate system for the trajectory projection;

using the trajectory projection to cause a robot to perform the manipulation task using the new object belonging to the object category.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2024
From: LIAN, WENZHAO; WEN, BOWEN; SCHAAL, STEFAN
To: INTRINSIC INNOVATION LLC
Reel/Frame 067266/0125 →
Continuity (2)
Provisional Application 63304533 · Jan 28, 2022
Related Publication 20230241773A1 · Aug 3, 2023
References Cited (44)
US 11507826B2 · Rohanimanesh · 2022 [cited by examiner]
US 12175703B2 · Birchfield · 2024 [cited by examiner]
US 20050107919A1 · Watanabe et al. · 2005 [cited by applicant]
US 20180250826A1 · Jiang · 2018 [cited by examiner]
US 20220277472A1 · Birchfield · 2022 [cited by examiner]
US 20220402128A1 · Lian · 2022 [cited by examiner]
US 20230130281A1 · Brown · 2023 [cited by examiner]
US 20230145208A1 · Bobu · 2023 [cited by examiner]
US 20240005547A1 · Lin · 2024 [cited by examiner]
US 20240371082A1 · Goyal · 2024 [cited by examiner]
Andrychowicz et al., “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research, 2020, 39 (1):3-20. [cited by applicant]
Bechtle et al., “Learning Extended Body Schemas from Visual Keypoints for Object Manipulation,” CoRR, Submitted on Nov. 8, 2020, arXiv:2011.03882, 7 pages. [cited by applicant]
Byravan et al., “Se3-pose-nets: Structured deep dynamics models for visuomotor control,” CoRR, Submitted on Oct. 2, 2017, arXiv:1710.00489v1, 8 pages. [cited by applicant]
Chai et al., “Multistep pick-and-place tasks using object-centric dense correspondences,” In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2019, pp. 4004-4011. [cited by applicant]
Coumans et al., “PyBullet, a Python module for physics simulation for games, robotics and machine learning,” 2016-2021. [cited by applicant]
Deng et al., “PoseRBPF: A Rao-Blackwellized Particle Filter for 6-D Object Pose Tracking,” CoRR, Submitted on May 22, 2019, arXiv:1905.09304v1, 10 pages. [cited by applicant]
Ebert et al., “Visual foresight: Model-based deep reinforcement learning for vision-based robotic control,” CoRR, Submitted on Dec. 3, 2018, arXiv:1812.00568v1, 14 pages. [cited by applicant]
Florence et al., “Dense object nets: Learning dense visual object descriptors by and for robotic manipulation,” CoRR, Submitted on Sep. 7, 2018, 12 pages. [cited by applicant]
Gao et al., “kPAM 2.0: Feedback Control for Category-Level Robotic Manipulation,” IEEE Robotics and Automation Letters, Apr. 2021, 6(2):2962-2969. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2023/011884, mailed on Aug. 8, 2024, 7 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2023/011884, mailed on May 8, 2023, 9 pages. [cited by applicant]
Kappler et al., “Real-time perception meets reactive motion generation,” IEEE Robotics and Automation Letters, Jul. 2018, 3(3):1864-1871. [cited by applicant]
Levine et al. “Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection,” The International Journal of Robotics Research, 2018, 37(45):421-436. [cited by applicant]
Levine et al., “End-to-end training of deep visuomotor policies,” CoRR, Submitted on Apr. 19, 2016, arXiv:1504.00702v5, 40 pages. [cited by applicant]
Mandlekar et al., “Learning to generalize across long-horizon tasks from human demonstrations,” CoRR, Submitted on Mar. 13, 2020, arXiv:2003.06085v1, 11 pages. [cited by applicant]
Manuelli et al., “Keypoint affordances for category-level robotic manipulation,” CoRR, Submitted on Oct. 29, 2019, arXiv:1903.06684v2, 26 pages. [cited by applicant]
Manuelli et al., “Keypoints into the Future: Self-Supervised Correspondence in Model-Based Reinforcement Learning,” CoRR, Submitted on Sep. 10, 2020, arXiv:2009.05085v1, 17 pages. [cited by applicant]
Mitash et al., “Task-Driven Perception and Manipulation for Constrained Placement of Unknown Objects,” CoRR, Submitted on Jun. 28, 2020, arXiv:2006.15503v1, 8 pages. [cited by applicant]
Morgan et al., “Vision-driven Compliant Manipulation for Reliable, High-Precision Assembly Tasks,” CoRR, Submitted on Jun. 26, 2021, arXiv:2106.14070v1, 13 pages. [cited by applicant]
Qi et al., “PointNet: Deep learning on point sets for 3D classification and segmentation,” CoRR, Submitted on Apr. 10, 2017, arXiv:1612.00593v2, 19 pages. [cited by applicant]
Qin et al., “KETO: Learning keypoint representations for tool manipulation,” 2020 IEEE International Conference on Robotics and Automation (ICRA), May 31, 2020, pp. 7278-7285. [cited by applicant]
Sieb et al., “Graph-structured visual imitation,” Proceedings of the Conference on Robot Learning, PMLR, 2020, 11 pages. [cited by applicant]
Simeonov et al., “Neural Descriptor Fields: SE (3)-Equivariant Object Representations for Manipulation,” CoRR, Submitted on Dec. 9, 2021, arXiv:2112.05124, 9 pages. [cited by applicant]
Tobin et al., “Domain randomization for transferring deep neural networks from simulation to the real world,” CoRR, Submitted on Mar. 20, 2017, arXiv:1703.06907v1, 8 pages. [cited by applicant]
Umeyama, “Least-squares estimation of transformation parameters between two point patterns,” IEEE Computer Architecture Letters, 1991, 13(04):376-380. [cited by applicant]
Vecerik et al., “S3K: Self-Supervised Semantic Keypoints for Robotic Manipulation via Multi-View Consistency,” CoRR, Submitted on Oct. 13, 2020, arXiv:2009.14711v2, 11 pages. [cited by applicant]
Viereck et al., “Learning a visuomotor controller for real world robotic grasping using simulated depth images,” Proceedings of the 1st Annual Conference on Robot Learning, PMLR, 2017, 10 pages. [cited by applicant]
Wang et al., “6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints,” 2020 IEEE International Conference on Robotics and Automation (ICRA), May 31, 2020, pp. 10059-10066. [cited by applicant]
Wang et al., “Normalized object coordinate space for category-level 6D object pose and size estimation,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, 2642-2651. [cited by applicant]
Wen et al., “BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models,” CoRR, Submitted on Aug. 1, 2021, arXiv:2108.00516v1, 8 pages. [cited by applicant]
Wen et al., “CaTGrasp: Learning Category-Level Task-Relevant Grasping in Clutter from Simulation,” CoRR, Submitted on Sep. 19, 2021, arXiv:2109.09163v1, 8 pages. [cited by applicant]
Wen et al., “Robust, occlusion-aware pose estimation for objects grasped by adaptive hands,” CoRR, Submitted on March, 7, 2020, arXiv:2003.03518v1, 8 pages. [cited by applicant]
Wen et al., “se(3)-TrackNet: Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains,” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 24, 2020, pp. 10367-1… [cited by applicant]
Yang et al., “Learning Multi-Object Dense Descriptor for Autonomous Goal-Conditioned Grasping,” IEEE Robotics and Automation Letters, 2021, 6(2):4109-4116. [cited by applicant]