IP Library › Granted Patent US 12,629,822
Granted Patent B2
US 12,629,822 · App. 17/935,879 · Granted May 19, 2026

Device and method for controlling a robot

Inventors: Andras Gabor Kupcsik (Boeblingen, DE); Meng Guo (Beijing, CN)
Assignee: ROBERT BOSCH GMBH
B25J9/163B25J9/161B25J9/1697B25J19/023
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,629,822
App. No.
17/935,879
Granted
May 19, 2026
Kind
B2
Abstract

A method for controlling a robot. The method includes performing demonstrations and descriptor images for the demonstrations from a point of view of the robot of the object; selecting a set of feature points, wherein the feature points are selected by searching an optimum of an objective function which rewards selected feature points being visible in the descriptor images; training a robot control model using the demonstrations and controlling the robot for a control scene with the object by determining a descriptor image of the object, locating the selected set of feature points in the descriptor image of the object; determining Euclidean coordinates of the located feature points; estimating a pose from the determined Euclidean coordinates; and controlling the robot to handle the object by means of the robot control model with the estimated pose.

Claims (48)

1 . A method for controlling a robot, comprising:

Performing demonstrations, whereineach demonstration of the demonstrations demonstrates a handling of an object;

Providing, for each demonstration, at least one descriptor image of the object from a point of view of the robot, wherein the descriptor image specifies feature points for locations on the object;

Selecting a set of feature points from the specified feature points, wherein the feature points are selected by searching an optimum of an objective function which rewards selected feature points being visible in the descriptor images;

Training a robot control modelusing the demonstrations, wherein the robot control model is configured to output control information depending on an input object pose; and

Controlling the robot for a control scene with the object by:

Determining a descriptor image of the object from the point of view of the robot;

Locating the selected set of feature points in the descriptor image of the object;

Determining Euclidean coordinates of the located feature points for the control scene;

Estimating a pose from the determined Euclidean coordinates; and

Controlling the robot to handle the object using the robot control model, wherein the estimated pose is supplied to the robot control model as input.

2 . The method of claim 1 , wherein the objective function further rewards one or more of the selected feature points being spaced apart in descriptor space, locations on the object corresponding to the selected feature points being spaced apart in Euclidean space, and a detection error for the selected feature points for the object being low.

3 . The method of claim 1 , further comprising:

Matching a plane to the object;

Selecting the feature points such that they define a coordinate frame on the plane; and

Estimating the pose from the determined Euclidean coordinates of the located feature points and information about a pose of the matched plane.

4 . The method of claim 3 , wherein estimating the pose from the determined Euclidean coordinates includes projecting the Euclidean coordinates of the located feature points to the matched plane.

5 . The method of claim 3 , wherein the feature points are selected such that they define a coordinate frame on the plane and the pose is estimated from the determined Euclidean coordinates of the located feature points and information about the pose of the plane when a variation of the object in a spatial direction is below a predetermined threshold.

6 . The method of claim 1 , further comprising:

Determininga derivation rule of a coordinate frame from Euclidean coordinates of the selected feature points;

Wherein estimating the pose from the determined Euclidean coordinates includes application of the derivation rule to the selected feature points, wherein the derivation rule is determined by searching a minimum of a dependency of the coordinate frame from noise in the Euclidean coordinates.

7 . The method of claim 1 , wherein the training of the robot control model using the demonstrations includes, for each demonstration, locating the selected set of feature points in the descriptor image of the object, determining Euclidean coordinates of the located feature points for the demonstration, and estimating a pose from the determined Euclidean coordinates for the demonstration.

8 . The method of claim 1 , further comprising determining the descriptor image of the object from a camera image of the object by a dense object net.

9 . The method of claim 1 , wherein the estimated pose is invariant with respect to a movement of the object such that when the object is moved the estimated pose is transformed in the same way.

10 . A robot controller, configured to control a robot, the robot controller configured to:

Perform demonstrations, wherein each demonstration of the demonstrations demonstrates a handling of an object;

Provide, for each demonstration, at least one descriptor image of the object from a point of view of the robot, wherein the descriptor image specifies feature points for locations on the object;

Select a set of feature points from the specified feature points, wherein the feature points are selected by searching an optimum of an objective function which rewards selected feature points being visible in the descriptor images;

Train a robot control model using the demonstrations, wherein the robot control model is configured to output control information depending on an input object pose; and

Control the robot for a control scene with the object by:

Determining a descriptor image of the object from the point of view of the robot;

Locating the selected set of feature points in the descriptor image of the object;

Determining Euclidean coordinates of the located feature points for the control scene;

Estimating a pose from the determined Euclidean coordinates; and

Controlling the robot to handle the object using the robot control model, wherein the estimated pose is supplied to the robot control model as input.

11 . The robot controller of claim 10 , wherein the estimated pose is invariant with respect to a movement of the object such that when the object is moved the estimated pose is transformed in the same way.

12 . A non-transitory computer-readable medium on which is stored a computer program for controlling a robot, the control program, when executed by a computer, causing the computer to perform:

Performing demonstrations, wherein each demonstration of the demonstrations demonstrates a handling of an object;

Providing, for each demonstration, at least one descriptor image of the object from a point of view of the robot, wherein the descriptor image specifies feature points for locations on the object;

Selecting a set of feature points from the specified feature points, wherein the feature points are selected by searching an optimum of an objective function which rewards selected feature points being visible in the descriptor images;

Training a robot control modelusing the demonstrations, wherein the robot control model is configured to output control information depending on an input object pose; and

Controlling the robot for a control scene with the object by:

Determining a descriptor image of the object from the point of view of the robot;

Locating the selected set of feature points in the descriptor image of the object;

Determining Euclidean coordinates of the located feature points for the control scene;

Estimating a pose from the determined Euclidean coordinates; and

Controlling the robot to handle the object using the robot control model, wherein the estimated pose is supplied to the robot control model as input.

13 . The non-transitory computer-readable medium of claim 12 , wherein the estimated pose is invariant with respect to a movement of the object such that when the object is moved the estimated pose is transformed in the same way.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2023
From: KUPCSIK, ANDRAS GABOR; GUO, MENG
To: ROBERT BOSCH GMBH
Reel/Frame 062395/0993 →
Priority Claims (1)
DE 10 2021 211 185.8 · Oct 5, 2021 · national
Continuity (1)
Related Publication 20230107993A1 · Apr 6, 2023
References Cited (15)
US 20150331415A1 · Feniello et al. · 2015 [cited by applicant]
US 20190126487A1 · Benaim · 2019 [cited by examiner]
US 20200276703A1 · Chebotar et al. · 2020 [cited by applicant]
US 20210170580A1 · Guo · 2021 [cited by examiner]
US 20210192784A1 · Taylor · 2021 [cited by examiner]
US 20210316449A1 · Wang · 2021 [cited by examiner]
DE 10355283A1 · 2004 [cited by applicant]
DE 102008020579B4 · 2014 [cited by applicant]
DE 102014106210A1 · 2015 [cited by applicant]
DE 102020110650A1 · 2020 [cited by applicant]
DE 112019001507T5 · 2020 [cited by applicant]
DE 102019216229A1 · 2021 [cited by applicant]
DE 102020127508A1 · 2021 [cited by applicant]
DE 102020128653A1 · 2021 [cited by applicant]
Pillai, Sudeep, Matthew R. Walter, and Seth Teller. “Learning articulated motions from visual demonstration.” arXiv preprint arXiv:1502.01659 (2015). [cited by examiner]