Techniques for adaptive robotic assembly
Techniques are disclosed for controlling robotic systems to perform assembly tasks. In some embodiments, a robot control application receives sensor data associated with one or more parts. The robot control application applies a grasp perception model to predict one or more grasp proposals indicating regions of the one or more parts that a robotic system can grasp. The robot control application causes the robotic system to grasp one of the parts based on a corresponding grasp proposal. If the pose of the grasped part needs to be changed in order to assemble the part with one or more other parts, the robot control application determines movements of the robotic system required to re-grasp the part in a different pose. In addition, the robot control application determines movements of the robot system for assembling the part with the one or more other parts based on results of a motion planning technique.
1 . A computer-implemented method for controlling a robotic system, the method comprising:
receiving sensor data associated with one or more parts;
executing, based on the sensor data, a first trained machine learning model that predicts one or more grasp proposals associated with the one or more parts, wherein each grasp proposal indicates a region of one of the one or more parts that the robotic system can grasp, wherein the robotic system comprises a first robot and a second robot;
causing the first robot to grasp a first part in a first pose based on the one or more grasp proposals, wherein the first part is included in the one or more parts;
receiving additional sensor data associated with the first part;
executing, based on the additional sensor data, a second trained machine learning model that predicts a pose proposal associated with a second pose of the first part;
in response to determining that the grasping of the first part in the first pose does not allow the first part to be assembled with the one or more parts, computing one or more movements of the robotic system to grasp the first part in the second pose by the first robot or the second robot, wherein the second pose of the first part is different than the first pose of the first part;
determining one or more additional movements of the robotic system to assemble the first part with the one or more other parts; and
causing the robotic system to perform the one or more movements and the one or more additional movements.
2 . The computer-implemented method of claim 1 , wherein computing the one or more movements comprises performing a search of a graph, and each node of the graph represents one of the first robot or the second robot grasping the first part in a different pose.
3 . The computer-implemented method of claim 1 , further comprising training the second trained machine learning model based on training data that includes (i) sensor data associated with a computer-aided design (CAD) model that represents the first part being grasped in a plurality of poses, and (ii) a plurality of pose proposals that each represents the first part being grasped in a different pose included in the plurality of pose proposals.
4 . The computer-implemented method of claim 1 , further comprising training the first trained machine learning model based on a training set of sensor data associated with one or more computer-aided design (CAD) models of a plurality of parts and one or more grasp proposals associated with the one or more CAD models.
5 . The computer-implemented method of claim 4 , further comprising generating each grasp proposal included in the one or more grasp proposals associated with the one or more CAD models based on one or more user-specified graspings of a CAD model included in the one or more CAD models.
6 . The computer-implemented method of claim 1 , wherein causing the robotic system to grasp the first part comprises:
computing a principal component analysis of a first grasp proposal, included in the one or more grasp proposals, that corresponds to the first part; and
determining how fingers of the robotic system can grasp the first part based on the principal component analysis.
7 . The computer-implemented method of claim 1 , wherein the one or more movements of the robotic system are computed based on results of one or more motion planning operations.
8 . The computer-implemented method of claim 1 , wherein the sensor data comprises depth data.
9 . The computer-implemented method of claim 1 , further comprising, prior to executing the first trained machine learning model, generating a height map associated with the one or more parts based on the sensor data, wherein the height map specifies a height of the one or more parts from a plane behind the one or more parts, wherein the first trained machine learning model is executed further based on the height map.
10 . One or more non-transitory computer-readable storage media including instructions that, when executed by at least one processor, cause the at least one processor to perform steps for controlling a robotic system, the steps comprising:
receiving sensor data associated with one or more parts;
executing, based on the sensor data, a first trained machine learning model that predicts one or more grasp proposals associated with the one or more parts, wherein each grasp proposal indicates a region of one of the one or more parts that the robotic system can grasp, wherein the robotic system comprises a first robot and a second robot;
causing the first robot to grasp a first part in a first pose based on the one or more grasp proposals, wherein the first part is included in the one or more parts;
receiving additional sensor data associated with the first part;
executing, based on the additional sensor data, a second trained machine learning model that predicts a pose proposal associated with a second pose of the first part;
in response to determining that the grasping of the first part in the first pose does not allow the first part to be assembled with the one or more parts, computing one or more movements of the robotic system to grasp the first part in the second pose by the first robot or the second robot, wherein the second pose of the first part is different than the first pose of the first part;
determining one or more additional movements of the robotic system to assemble the first part with the one or more parts; and
causing the robotic system to perform the one or more movements and the one or more additional movements.
11 . The one or more non-transitory computer-readable storage media of claim 10 , wherein computing the one or more movements comprises performing a search of a graph, and each node of the graph represents one of the first robot or the second robot grasping the first part in a different pose.
12 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the second trained machine learning model comprises a semantic segmentation neural network.
13 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training the first trained machine learning model based on a training set of sensor data associated with one or more computer-aided design (CAD) models of parts and one or more grasp proposals associated with the one or more CAD models.
14 . The one or more non-transitory computer-readable storage media of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of generating each grasp proposal included in the one or more grasp proposals associated with the one or more CAD models based on one or more user-specified graspings of a CAD model included in the one or more CAD models.
15 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the first part is closer in height to the robotic system than one or more other parts of the one or more parts.
16 . A robotic system comprising:
one or more memories storing instructions; and
one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to:
receive sensor data associated with one or more parts,
execute, based on the sensor data, a first trained machine learning model that predicts one or more grasp proposals associated with the one or more parts, wherein each grasp proposal indicates a region of one of the one or more parts that the robotic system can grasp, wherein the robotic system comprises a first robot and a second robot;
cause the first robot to grasp a first part in a first pose based on the one or more grasp proposals, wherein the first part is included in the one or more parts;
receive additional sensor data associated with the first part;
execute, based on the additional sensor data, a second trained machine learning model that predicts a pose proposal associated with a second pose of the first part;
in response to determining that the grasping of the first part in the first pose does not allow the first part to be assembled with the one or more parts, compute one or more movements of the robotic system to grasp the first part in the second pose by the first robot or the second robot, wherein the second pose of the first part is different than the first pose of the first part;
compute one or more additional movements of the robotic system to assemble the first part with the one or more parts; and
cause the robotic system to perform the one or more movements and the one or more additional movements.