Visual servoing with real-time diffusion-based image inpainting
Various aspects of visual servoing based on use of a generative diffusion model are described. One technique includes obtaining a current view image from a robot camera that captures a portion of an environment of the robot, and obtaining a spherical image from a previous view that depicts a surrounding area of the environment of the robot. The current view image is combined with the previous spherical image to produce a modified spherical image, and inpainting is performed in a mask area of the modified spherical image with a generative diffusion model. The inpainting blends the current view image with the previous spherical image, providing a view of the entire environment of the robot that can be used for visual servoing. The robot can perform path planning and visual servoing of an end effector based on the modified spherical image, independent of any use of an external positioning or localization system.
1 . At least one machine-readable medium, including instructions, which when executed by processing circuitry, cause the processing circuitry to perform operations to:
obtain a current view image from a camera of a robot, wherein the current view image captures a portion of an environment of the robot;
obtain a spherical image that depicts the environment of the robot surrounding the robot;
combine the current view image with the spherical image to produce a modified spherical image, based on inpainting a mask area of the modified spherical image with a generative diffusion model to blend the current view image with the spherical image; and
cause visual servoing of the robot based on the modified spherical image, the visual servoing to move the robot to a target located in the environment.
2 . The at least one machine-readable medium of claim 1 , wherein the generative diffusion model is pre-trained based on multiple images of the environment.
3 . The at least one machine-readable medium of claim 2 , wherein the instructions further cause the processing circuitry to perform operations to:
detect at least one occlusion in the current view image, wherein the at least one occlusion blocks a view of the portion of the environment of the robot; and
perform inpainting of the at least one occlusion in the current view image using the generative diffusion model, wherein the inpainting of the at least one occlusion is based on at least one previous view image that depicts the view of the portion of the environment.
4 . The at least one machine-readable medium of claim 1 , wherein the instructions further cause the processing circuitry to perform operations to:
measure performance of the visual servoing of the robot based on the inpainting of the mask area; and
modify the generative diffusion model based on the performance of the visual servoing of the robot.
5 . The at least one machine-readable medium of claim 1 , wherein the visual servoing of the robot is caused based on: feature matching between one or more features of the modified spherical image and a goal spherical image, and an error of the feature matching between the one or more features in the modified spherical image and the goal spherical image.
6 . The at least one machine-readable medium of claim 1 , wherein the spherical image is based on a 360-degree scene maintained for the environment around the robot.
7 . The at least one machine-readable medium of claim 6 , wherein the spherical image is constructed from multiple prior images of the environment, and wherein the prior images are combined by inpainting overlapping areas of the prior images using the generative diffusion model.
8 . The at least one machine-readable medium of claim 1 , wherein the mask area where inpainting is performed includes an outside portion of the current view image and a portion of the modified spherical image that surrounds the current view image.
9 . The at least one machine-readable medium of claim 1 , wherein the camera that provides the current view image is attached to an end effector, a joint, or a segment of the robot.
10 . The at least one machine-readable medium of claim 9 , wherein the instructions further cause the processing circuitry to perform operations to:
perform path planning of the end effector based on the modified spherical image.
11 . The at least one machine-readable medium of claim 10 , wherein the robot is configured to perform path planning and visual servoing of the end effector based on the modified spherical image independent of an external positioning system.
12 . A system comprising:
processing circuitry; and
memory including instructions, which when executed by the processing circuitry, cause the processing circuitry to:
access a current view image captured from a camera of a robot, wherein the current view image captures a portion of an environment of the robot;
access a spherical image that depicts the environment of the robot surrounding the robot;
combine the current view image with the spherical image to produce a modified spherical image, based on inpainting a mask area of the modified spherical image with a generative diffusion model to blend the current view image with the spherical image; and
control visual servoing of the robot based on the modified spherical image, the visual servoing to move the robot to a target located in the environment.
13 . The system of claim 12 , wherein the instructions further cause the processing circuitry to perform operations to:
detect at least one occlusion in the current view image, wherein the at least one occlusion blocks a view of the portion of the environment of the robot; and
perform inpainting of the at least one occlusion in the current view image using the generative diffusion model, wherein the inpainting of the at least one occlusion is based on at least one previous view image that depicts the view of the portion of the environment;
wherein the generative diffusion model is pre-trained based on multiple images of the environment.
14 . The system of claim 12 , wherein the instructions further cause the processing circuitry to perform operations to:
measure performance of the visual servoing of the robot based on the inpainting of the mask area; and
modify the generative diffusion model based on the performance of the visual servoing of the robot.
15 . The system of claim 12 , wherein the visual servoing of the robot is controlled based on: feature matching between one or more features of the modified spherical image and a goal spherical image, and an error of the feature matching between the one or more features in the modified spherical image and the goal spherical image.
16 . The system of claim 12 , wherein the spherical image provides a 360-degree view of the environment around the robot, and wherein the spherical image is constructed from multiple prior images of the environment, and wherein the prior images are combined by inpainting overlapping areas of the prior images using the generative diffusion model.
17 . The system of claim 12 , wherein the mask area where inpainting is performed includes an outside portion of the current view image and a portion of the modified spherical image that surrounds the current view image.
18 . The system of claim 12 , wherein the instructions further cause the processing circuitry to:
perform path planning of an end effector of the robot based on the modified spherical image;
wherein the robot is configured to perform path planning and visual servoing of the end effector based on the modified spherical image independent of an external positioning system.
19 . An apparatus comprising:
memory means for storing a current view image and a spherical image, wherein the current view image captures a portion of an environment of a robot, and wherein the spherical image depicts the environment of the robot surrounding the robot;
processing means for generating a modified spherical image, the processing means configured to combine the current view image with the spherical image to produce the modified spherical image and perform inpainting in a mask area of the modified spherical image using a generative diffusion model, wherein the inpainting blends the current view image with the spherical image; and
control means for servoing the robot based on the modified spherical image, wherein the servoing causes the robot to move to a target located in the environment.
20 . The apparatus of claim 19 , wherein the processing means is further configured to: detect at least one occlusion in the current view image that blocks a view of the portion of the environment of the robot, and perform inpainting of the at least one occlusion in the current view image using the generative diffusion model.