IP Library › Granted Patent US 12,205,188
Granted Patent B2
US 12,205,188 · App. 18/138,655 · Granted Jan 21, 2025

Multicamera image processing

Inventors: Kevin Jose Chavez (Redwood City, CA); Yuan Gao (Santa Clara, CA); Rohit Arka Pidaparthi (Mountain View, CA); Talbot Morris-Downing (Redwood City, CA); Harry Zhe Su (Union City, CA); Samir Menon (Atherton, CA)
Assignee: Dexterity, Inc.
G06T1/0014G06T1/20G06T7/174G06T7/593G06T7/90G06T2207/10024G06T2207/10028G06T2207/20221H04N13/204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,188
App. No.
18/138,655
Granted
Jan 21, 2025
Kind
B2
Abstract

A multicamera image processing system is disclosed. In various embodiments, image data is received from each of a plurality of sensors associated with a workspace, the image data comprising for each sensor in the plurality of sensors one or both of visual image information and depth information. Image data from the plurality of sensors is merged to generate a merged point cloud data. Segmentation is performed based on visual image data from at least a subset of the sensors in the plurality of sensors to generate a segmentation result. One or both of the merged point cloud data and the segmentation result is/are used to generate a merged three dimensional and segmented view of the workspace.

Claims (54)

1. A system, comprising:

a communication interface configured to receive image data from each of a plurality of cameras associated with a workspace; and

a processor coupled to the communication interface and configured to:

merge image data from the plurality of cameras to generate a merged point cloud data;

perform segmentation based on visual image data from a subset of the plurality of cameras to generate a segmentation result, wherein the segmentation result includes an indication of an object boundary for one or more objects;

use the merged point cloud data and the segmentation result to generate a merged three dimensional and segmented view of the workspace, including by:

for at least a subset of the plurality of cameras, generating a point cloud for each object and computing a corresponding centroid; and

segmenting objects within the workspace based at least in part on centroids of corresponding to each object; and

provide the merged three dimensional and segmented view of the workspace as an output to a module configured to determine a strategy to grasp an object present in the workspace using a robot.

2. The system of claim 1 , wherein:

the segmentation result is obtained based on performing segmentation using RGB data from a camera;

the segmentation result comprises a plurality of RGB pixels; and

a subset of the plurality of RGB pixels is identified based at least in part on determination that the corresponding RGB pixels are associated with an object boundary.

3. The system of claim 1 , wherein merging the image data from the plurality of cameras to generate the merged point cloud data comprises:

translating to a global three dimensional coordinate framework respective positions and orientations of objects and features of the workspace as captured in separately acquired views associated with image data from the plurality of cameras.

4. The system of claim 1 , where the image data from the plurality of cameras are dynamically merged to provide a continuously updated merged point cloud data.

5. The system of claim 1 , wherein the processor is configured to:

autonomously detect that at least one of the plurality of cameras requires recalibration; and

in response to detecting that the at least one of the plurality of cameras requires recalibration, recalibrate the at least one of the plurality of cameras.

6. The system of claim 1 , wherein recalibrating the at least one of the plurality of cameras comprises:

using a camera mounted to a robotic actuator to relocate a fiducial marker in the workspace;

re-estimating camera-to-workspace transformation using fiducial markers; and

recalibrating the at least one of the plurality of cameras to a marker on the robot.

7. The system of claim 1 , wherein generating the merged three dimensional and segmented view of the workspace comprises:

applying a workspace filter in connection with removing one or more of an image and point cloud data associated with portions of the workspace, features of the workspace, or items in the workspace.

8. The system of claim 7 , wherein the workspace filter removes image or point cloud data associated with portions of the workspace, features of the workspace, or items in the workspace that are not required for determining the strategy to grasp the object present in the workspace.

9. The system of claim 7 , wherein the workspace filter removes statistical outlier data.

10. The system of claim 1 , wherein the plurality of cameras includes one or more three dimensional (3D) cameras.

11. The system of claim 1 , wherein the image data includes RGB data.

12. The system of claim 1 , wherein the processor is configured to use one or both of the merged point cloud data and the segmentation result to perform a box fit with respect to the object in the workspace.

13. The system of claim 1 , wherein the processor is further configured to implement the strategy to grasp the object using the robot.

14. The system of claim 6 , wherein the processor is configured to grasp the object in connection with a robotic operation to pick the object from an origin location and place the object in a destination location in the workspace.

15. The system of claim 1 , wherein the processor is further configured to use the merged three dimensional and segmented view of the workspace to display a visualization of the workspace.

16. The system of claim 1 , wherein generating the merged three dimensional and segmented view of the workspace comprises de-projecting into the merged point cloud a set of points comprising the segmentation result.

17. The system of claim 1 , wherein the processor is further configured to subsample the merged point cloud data.

18. The system of claim 17 , wherein the processor is further configured to perform cluster processing on the subsampled point cloud data.

19. The system of claim 18 , wherein the processor is configured to use the subsampled and clustered point cloud data and the segmentation result to generate a box fit result with respect to the object in the workspace.

20. The system of claim 1 , wherein the merged point cloud data and the segmentation result is used to determine a trajectory via which a robotic arm is to move the object to a destination location.

21. A method, comprising:

receiving image data from each of a plurality of cameras associated with a workspace;

merging image data from the plurality of cameras to generate a merged point cloud data;

performing segmentation based on visual image data from a subset of the plurality of cameras to generate a segmentation result, wherein the segmentation result includes an indication of an object boundary for one or more objects;

using one or both of the merged point cloud data and the segmentation result to generate a merged three dimensional and segmented view of the workspace, including by:

for at least a subset of the plurality of cameras, generating a point cloud for each object and computing a corresponding centroid; and

segmenting objects within the workspace based at least in part on centroids of corresponding to each object; and

providing the merged three dimensional and segmented view of the workspace as an output to a module configured to determine a strategy to grasp an object present in the workspace using a robot.

22. A computer program product embodied in a non-transitory computer readable medium and comprising computer instructions for:

receiving image data from each of a plurality of cameras associated with a workspace;

merging image data from the plurality of cameras to generate a merged point cloud data;

performing segmentation based on visual image data from a subset of the plurality of cameras to generate a segmentation result, wherein the segmentation result includes an indication of an object boundary for one or more objects;

using one or both of the merged point cloud data and the segmentation result to generate a merged three dimensional and segmented view of the workspace, including by:

for at least a subset of the plurality of cameras, generating a point cloud for each object and computing a corresponding centroid; and

segmenting objects within the workspace based at least in part on centroids of corresponding to each object; and

providing the merged three dimensional and segmented view of the workspace as an output to a module configured to determine a strategy to grasp an object present in the workspace using a robotic arm.

Continuity (4)
Continuation 16667661 · Oct 29, 2019
Continuation In Part 16380859 · Apr 10, 2019
Provisional Application 62809389 · Feb 22, 2019
Related Publication 20230260071A1 · Aug 17, 2023
References Cited (63)
US 5501571A · Van Durrett · 1996 [cited by applicant]
US 5908283A · Huang · 1999 [cited by applicant]
US 8930019B2 · Allen · 2015 [cited by applicant]
US 9089969B1 · Theobald · 2015 [cited by applicant]
US 9315344B1 · Lehmann · 2016 [cited by applicant]
US 9327406B1 · Hinterstoisser · 2016 [cited by applicant]
US 9802317B1 · Watts · 2017 [cited by applicant]
US 9811892B1 · Silverstein · 2017 [cited by applicant]
US 10124489B2 · Chitta · 2018 [cited by applicant]
US 10207868B1 · Stubbs · 2019 [cited by applicant]
US 10549928B1 · Chavez · 2020 [cited by applicant]
US 10906188B1 · Sun · 2021 [cited by applicant]
US 11591169B2 · Chavez · 2023 [cited by applicant]
US 11741566B2 · Chavez · 2023 [cited by examiner]
US 20020106273A1 · Huang · 2002 [cited by applicant]
US 20020164067A1 · Askey · 2002 [cited by applicant]
US 20070280812A1 · Morency · 2007 [cited by applicant]
US 20090033655A1 · Boca · 2009 [cited by applicant]
US 20100324729A1 · Ruge · 2010 [cited by applicant]
US 20120259582A1 · Gloger · 2012 [cited by applicant]
US 20130315479A1 · Paris · 2013 [cited by applicant]
US 20150035272A1 · Johnson · 2015 [cited by applicant]
US 20150352721A1 · Wicks · 2015 [cited by applicant]
US 20160016311A1 · Konolige · 2016 [cited by applicant]
US 20160075031A1 · Gotou · 2016 [cited by applicant]
US 20160207195A1 · Eto · 2016 [cited by applicant]
US 20160229061A1 · Takizawa · 2016 [cited by applicant]
US 20160272354A1 · Nammoto · 2016 [cited by applicant]
US 20170246744A1 · Chitta · 2017 [cited by applicant]
US 20170267467A1 · Kimoto · 2017 [cited by applicant]
US 20180086572A1 · Kimoto · 2018 [cited by applicant]
US 20180144458A1 · Xu · 2018 [cited by applicant]
US 20180162660A1 · Saylor · 2018 [cited by applicant]
US 20180308254A1 · Fu · 2018 [cited by applicant]
US 20190000564A1 · Navab · 2019 [cited by applicant]
US 20190016543A1 · Turpin · 2019 [cited by applicant]
US 20190102965A1 · Greyshock · 2019 [cited by applicant]
US 20190362178A1 · Huang · 2019 [cited by applicant]
US 20200117212A1 · Tian · 2020 [cited by applicant]
US 20200130961A1 · Diankov · 2020 [cited by applicant]
CA 2357271 · 2002 [cited by applicant]
CN 107530881 · 2008 [cited by applicant]
CN 107000208 · 2017 [cited by applicant]
CN 107088877 · 2017 [cited by applicant]
CN 108972564 · 2018 [cited by applicant]
CN 109255813 · 2021 [cited by applicant]
EP 1489025 · 2004 [cited by applicant]
EP 3349182 · 2018 [cited by applicant]
JP 2001184500 · 2001 [cited by applicant]
JP 2006302195 · 2006 [cited by applicant]
JP 2015135331 · 2015 [cited by applicant]
JP 5905549 · 2016 [cited by applicant]
WO 9823511 · 1998 [cited by applicant]
WO 2017195801 · 2017 [cited by applicant]
WO 20180130491 · 2018 [cited by applicant]
Chen et al. “Random Bin Picking with Multi-view Image Acquisition and CAD-Based Pose Estimation,” 2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC), USA, IEEE, Oct. 7, 2018, pp. 2218-2223 (Docume… [cited by applicant]
Gratal et al: “Scene Representation and Object Grasping Using Active Vision”, https://batavia.internal.epo.org/digital-file-repository/digital-file-repository-frontend-prod/dossier/ep1991600i?toc=L0gggv0b bhd69lv, Jan. … [cited by applicant]
Ji et al: “Autonomous 3D scene understanding and exploration in cluttered workspaces using point cloud data”, 2018 IEEE 15TH International Conference on Networking, Sensing and Control (ICNSC), IEEE, Mar. 27, 2018 (Mar.… [cited by applicant]
Kakiuchi et al: “Creating household environment map for environment manipulation using color range sensors on environment and robot”, Robotics and Automation (ICRA), 2011 IEEE International Conference on, IEEE, May 9, 2… [cited by applicant]
Kato et al. “Extraction of Reference Point Candidates from Three-Dimensional Point Group and Learning of Relative Position Concepts,” The 23rd Symposium on Sensing via Image Information SSII2017, Japan, SSII, Jun. 7, 20… [cited by applicant]
Nishida et al., “Object Classification Considering Movability Based on Shape Features with Parts Decomposition,” Transactions of the Society of Instrument and Control Engineers, Japan, The Society of Instrument and Cont… [cited by applicant]
Takeguchi et al., “Robust Object Recognition through Depth Aspect Image by Regular Voxels, ” The IEICE Transactions on Information and Systems, Japan, The Institute of Electronics, Information and Communication Engineer… [cited by applicant]
Takaaki Nishida et. al., “Object Classification Considering Movability Based on Shape Features with Parts Decomposition,” Proceedings of the Society of Instrument and Control Engineers, vol. 51, No. 5, Society of Instru… [cited by applicant]