Generating robotic control plans
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating a robotic control plan. One of the methods includes obtaining, from a user device, image data depicting an instruction manual for assembling a plurality of assembly components; processing the image data using a machine learning model to generate instruction data representing a sequence of instructions for assembling the plurality of assembly components, wherein the machine learning model has been configured through training to process images depicting instruction manuals and to generate instruction data characterizing sequences of instructions identified in the instruction manuals; processing the instruction data to generate a robotic control plan to be executed by one or more robotic components for assembling the plurality of assembly components; and providing the robotic control plan to a robotic control system for executing the robotic control plan using the one or more robotic components.
1 . A method comprising:
obtaining, from a user device, image data comprising a plurality of images of an instruction manual, wherein the instruction manual depicts assembly actions for assembling a plurality of assembly components into a product;
generating, for each image using a machine learning model, instruction data for the image, the instruction data representing a sequence of instructions for a robot to perform one or more assembly actions to assemble a respective plurality of assembly components represented in the image, wherein the machine learning model has been configured, through training on training examples that each include one or more illustrations of one or more assembly actions using a respective plurality of assembly components and ground-truth instruction data representing a sequence of instructions for a robot to perform the one or more assembly actions using the plurality of assembly components;
processing the instruction data to generate a sequence of stages, each stage having a plurality of subtasks to be executed by one or more robots for assembling the plurality of assembly components identified from the depictions of assembly actions for assembling the assembly components into the product in the one or more images of the image data; and
causing a robotic control system to execute the plurality of subtasks using the one or more robots.
2 . The method of claim 1 , wherein processing the instruction data to generate a sequence of stages comprises:
obtaining, from the user device, second image data depicting the plurality of assembly components;
obtaining, for one or more of the assembly components, assembly component data characterizing one or more properties of the assembly component; and
processing i) the instruction data and ii) the assembly component data to generate the sequence of stages.
3 . The method of claim 2 , wherein obtaining assembly component data for a particular assembly component comprises:
processing the second image data depicting the particular assembly component using a second machine learning model to generate the assembly component data for the particular assembly component, wherein the second machine learning model has been configured through training to process images depicting assembly components and to generate assembly component data characterizing one or more properties of the assembly components.
4 . The method of claim 2 , wherein obtaining assembly component data for a particular assembly component comprises:
identifying the particular assembly component in the second image data; and
obtaining, from a data store, predetermined assembly component data for the particular assembly component.
5 . The method of claim 2 , wherein the assembly component data comprises data identifying, for one or more of the plurality of assembly components, one or more of:
a material of the assembly component,
a weight of the assembly component,
a density of the assembly component,
a center of mass of the assembly component,
a strength of the assembly component,
a flexibility of the assembly component, or
one or more preferred or required touch points of the assembly component.
6 . The method of claim 2 , wherein processing i) the instruction data and ii) the assembly component data to generate the sequence of stages comprises:
identifying an order of assembly actions to be executed by the one or more robots from the instruction data;
identifying, for each assembly action, a skill type from the instruction data that represents a type of movement to be executed by the one or more robots on one or more assembly components of the plurality of assembly components;
identifying, for each assembly action, one or more standards from the instruction data that represent a measure of success for the assembly action; and
generating the sequence of stages from the order of assembly actions, and skill type and one or more standards for each assembly action.
7 . The method of claim 1 , wherein the plurality of assembly components have been manufactured by a particular manufacturer, and wherein the machine learning model has been trained using training examples corresponding to the particular manufacturer.
8 . The method of claim 7 , wherein the instruction data is represented using a computer language that can be used to represent instruction manuals produced by a plurality of different manufacturers.
9 . The method of claim 1 , wherein the one or more robots execute the plurality of subtasks in a temporary robotic operating environment.
10 . The method of claim 1 , wherein:
the image data comprises an image depicting a portion of the instruction manual that identifies each of the plurality of assembly components; and
the one or more images of the image data comprise one or more other images that correspond to respective other portions of the instruction manual.
11 . The method of claim 1 , wherein the instruction manual comprises one or more physical pages.
12 . The method of claim 1 , wherein the ground-truth instruction data of each training example represents instruction data for accomplishing one or more subtasks.
13 . The method of claim 1 , wherein each training example further comprises data identifying assembly components depicted in the respective instruction manual.
14 . A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
obtaining, from a user device, image data comprising a plurality of images of an instruction manual, wherein the instruction manual depicts assembly actions for assembling a plurality of assembly components into a product;
generating, for each image using a machine learning model, instruction data for the image, the instruction data representing a sequence of instructions for a robot to perform one or more assembly actions to assemble a respective plurality of assembly components represented in the image, wherein the machine learning model has been configured, through training on training examples that each include one or more illustrations of one or more assembly actions using a respective plurality of assembly components and ground-truth instruction data representing a sequence of instructions for a robot to perform the one or more assembly actions using the plurality of assembly components;
processing the instruction data to generate a sequence of stages, each stage having a plurality of subtasks to be executed by one or more robots for assembling the plurality of assembly components identified from the depictions of assembly actions for assembling the assembly components into the product in the one or more images of the image data; and
causing a robotic control system to execute the plurality of subtasks using the one or more robots.
15 . The system of claim 14 , wherein processing the instruction data to generate a sequence of stages comprises:
obtaining, from the user device, second image data depicting the plurality of assembly components;
obtaining, for one or more of the assembly components, assembly component data characterizing one or more properties of the assembly component; and
processing i) the instruction data and ii) the assembly component data to generate the sequence of stages.
16 . The system of claim 15 , wherein obtaining assembly component data for a particular assembly component comprises:
processing the second image data depicting the particular assembly component using a second machine learning model to generate the assembly component data for the particular assembly component, wherein the second machine learning model has been configured through training to process images depicting assembly components and to generate assembly component data characterizing one or more properties of the assembly components.
17 . The system of claim 15 , wherein obtaining assembly component data for a particular assembly component comprises:
identifying the particular assembly component in the second image data; and
obtaining, from a data store, predetermined assembly component data for the particular assembly component.
18 . The system of claim 14 , wherein the plurality of assembly components have been manufactured by a particular manufacturer, and wherein the machine learning model has been trained using training examples corresponding to the particular manufacturer.
19 . The system of claim 14 , wherein:
the image data comprises an image depicting a portion of the instruction manual that identifies each of the plurality of assembly components; and
the one or more images of the image data comprise one or more other images that correspond to respective other portions of the instruction manual.
20 . One or more non-transitory storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
obtaining, from a user device, image data comprising a plurality of images of an instruction manual, wherein the instruction manual depicts assembly actions for assembling a plurality of assembly components into a product;
generating, for each image using a machine learning model, instruction data for the image, the instruction data representing a sequence of instructions for a robot to perform one or more assembly actions to assemble a respective plurality of assembly components represented in the image, wherein the machine learning model has been configured, through training on training examples that each include one or more illustrations of one or more assembly actions using a respective plurality of assembly components and ground-truth instruction data representing a sequence of instructions for a robot to perform the one or more assembly actions using the plurality of assembly components;
processing the instruction data to generate a sequence of stages, each stage having a plurality of subtasks to be executed by one or more robots for assembling the plurality of assembly components identified from the depictions of assembly actions for assembling the assembly components into the product in the one or more images of the image data; and
causing a robotic control system to execute the plurality of subtasks using the one or more robots.
21 . The non-transitory storage media of claim 20 , wherein processing the instruction data to generate a sequence of stages comprises:
obtaining, from the user device, second image data depicting the plurality of assembly components;
obtaining, for one or more of the assembly components, assembly component data characterizing one or more properties of the assembly component; and
processing i) the instruction data and ii) the assembly component data to generate the sequence of stages.
22 . The non-transitory storage media of claim 21 , wherein obtaining assembly component data for a particular assembly component comprises:
processing the second image data depicting the particular assembly component using a second machine learning model to generate the assembly component data for the particular assembly component, wherein the second machine learning model has been configured through training to process images depicting assembly components and to generate assembly component data characterizing one or more properties of the assembly components.
23 . The non-transitory storage media of claim 21 , wherein obtaining assembly component data for a particular assembly component comprises:
identifying the particular assembly component in the second image data; and
obtaining, from a data store, predetermined assembly component data for the particular assembly component.
24 . The non-transitory storage media of claim 20 , wherein:
the image data comprises an image depicting a portion of the instruction manual that identifies each of the plurality of assembly components; and
the one or more images of the image data comprise one or more other images that correspond to respective other portions of the instruction manual.