IP Library › Granted Patent US 12,602,806
Granted Patent B2
US 12,602,806 · App. 17/963,156 · Granted Apr 14, 2026

Systems and methods for machine vision robotic processing

Inventors: Yang Tao (North Potomac, MD); Dongyi Wang (Fayetteville, AR); Mohamed Amr Ali (Lanham, MD)
Assignee: University of Maryland, College Park
G06T7/593B25J9/1697G06V10/141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,806
App. No.
17/963,156
Filed
Oct 10, 2022
Granted
Apr 14, 2026
Kind
B2
Art Unit
2663
USPC
382/154
Abstract

Systems, methods, and media for machine vision-guided robotic loading are provided. Such systems and methods can allow for more accurate object detection so as to allow for robotic picking or loading of an object from a pile or group of objects. In some embodiments, a camera is controlled to capture a background image of the object or group of objects. One or more illumination sources (such as, e.g., laser illuminators) are controlled to project immunization toward the object or group of objects within the field-of-view of the camera. The camera may be controlled to capture a scan image comprising the projected illumination. Based on the scan image, a depth map can be determined, which is provided to a machine learning model along with the background image. The output of the machine learning model can then provide information associated with one or more objects, such as location information, orientation information, and the like.

Claims (78)

1 . A system for detecting an object on a surface, comprising:

a camera oriented to include the surface in a field-of-view (FOV) of the camera;

a plurality of aimable illumination sources configured to project corresponding illuminations toward the surface within the FOV of the camera; and

a controller comprising a processor that is configured:

to control the camera to capture, with the camera, a first background image of the object on the surface;

to control the plurality of aimable illumination sources to project the corresponding illuminations simultaneously toward the surface to form a projected illumination on the surface;

to control the camera to capture a scan image comprising the projected illumination formed by simultaneously projected corresponding illuminations that intersect the surface at respective positions;

to determine a first depth map based on separating the scan image into a plurality of object scan mask sets, each object scan mask set representing a corresponding image formed only by an illumination of an aimable illumination source chosen from the plurality of aimable illumination sources;

to provide the first depth map and the first background image to a trained machine learning model executed by the processor; and

to receive an output from the machine learning model, wherein the output comprises information associated with the object;

wherein the plurality of aimable illumination sources comprises a first aimable illumination source and a second aimable illumination source, and wherein the first aimable illumination source and the second aimable illumination source have different spectra, and

wherein the controller is further configured to separate the scan image into the plurality of scan image sets each of which is formed only by light having a spectrum of an illumination of said aimable illumination source chosen from the plurality of aimable illumination sources.

2 . The system of claim 1 , wherein each of the aimable illumination sources comprises a line laser.

3 . The system of claim 2 , wherein each of the aimable illumination sources further comprises:

a motor; and

a mirror coupled to the motor,

wherein the mirror is positioned to reflect illumination from the line laser;

wherein the motor is controllable to rotate the mirror based on a received control input.

4 . The system of claim 1 , wherein the controller is further configured:

to controllably scan the projected illumination across the surface within the FOV in a plurality of scan steps by changing a position of a corresponding illumination of each of the plurality of aimable illumination sources on the surface;

to determine an object scan set, comprising a plurality of scan image sets, at each scan step of the plurality of scan steps; and

to determine the first depth map based on the object scan set.

5 . The system of claim 4 , wherein the controller is further configured:

to determine a baseline scan set for each of the plurality of aimable illumination sources; and

to determine the first depth map based on the object scan set by subtracting, pixel by pixel, the baseline scan set from the object scan set to identify a pixel shift in the projected illumination captured in a given object scan mask set as compared to a corresponding baseline scan mask set, for each scan step, wherein the shift is due to the projected illumination intersecting with the object.

6 . The system of claim 1 , wherein the controller is further configured to control a robotic picking device to acquire the object according to the output of the machine learning model.

7 . The system of claim 1 , wherein the controller is further configured:

to determine, based on the output of the machine learning model, a first object among a plurality of objects to prioritize; and

to provide a first signal to a robotic picking device to aid in acquiring the first object.

8 . The system of claim 7 , wherein the controller is further configured, after the first object has been removed from the surface:

to provide a second depth map and a second background image to the machine learning model;

to receive a second output from the machine learning model and, based on the second output, determine a second object among the plurality of objects to prioritize; and

to provide a second signal to the robotic picking device to aid in acquiring the second object.

9 . The system of claim 1 , wherein the machine learning model is configured:

to identify the object; and

wherein the information associated with the object comprises at least one of a mask, a key a point, a class, and a bounding box for the object identified by the machine learning model.

10 . A method for detecting an object, the method comprising:

capturing, with a camera, a first background image of the object on a surface that is positioned in a field-of-view (FOV) of the camera;

controlling a plurality of aimable illumination sources to simultaneously project corresponding illuminations toward the surface within the FOV of the camera to form a projected illumination on the surface;

capturing, with the camera, a scan image comprising the projected illumination formed by simultaneously projected corresponding illuminations that intersect the surface at respective positions;

determining a first depth map based on separating the scan image into plurality of object scan mask sets, each object scan mask set representing a corresponding image formed only by an illumination of an aimable illumination source chosen from the plurality of aimable illumination sources; providing the first depth map and the first background image to a machine learning model; and receiving an output from the machine learning model, wherein the output comprises information associated with the object,

wherein the plurality of aimable illumination sources comprises a first aimable illumination source and a second aimable illumination source, and wherein the first aimable illumination source and the second aimable illumination source have different spectra, and

wherein the controller is configured to separate the scan image into the plurality of scan image sets each of which is formed only by light having a spectrum of an illumination of said aimable illumination source chosen from the plurality of aimable illumination sources.

11 . The method of claim 10 , wherein each of the aimable illumination sources comprises a line laser.

12 . The method of claim 11 , wherein each of the aimable illumination sources further comprises:

a motor; and

a mirror coupled to the motor and positioned to reflect illumination from the line laser; and

wherein the controlling the plurality of aimable illumination sources comprises providing a control input to the motor of each aimable illumination source.

13 . The method of claim 10 , wherein the method further comprises:

controllably scanning the projected illumination across the surface within the FOV in a plurality of scan steps by changing a position, on the surface, of the corresponding illumination of an aimable illumination source of the plurality of aimable illumination sources,

determining an object scan set, comprising a plurality of scan image sets, at each scan step of the plurality of scan steps, and

determining the first depth map based on the object scan set.

14 . The method of claim 13 , wherein the method further comprises:

determining a baseline scan set for each of the plurality of aimable illumination sources; and

wherein the determining the first depth map based on the object scan set comprises subtracting, pixel by pixel, the baseline scan set from the object scan set to identify a pixel shift in the projected illumination captured in a given object scan mask set as compared to a corresponding baseline scan mask set, for each scan step, wherein the shift is due to the projected illumination intersecting with the object.

15 . The method of claim 10 , wherein the method further comprises generating a control signal for a robotic picking device to acquire the object according to the output of the machine learning model.

16 . The method of claim 10 , wherein the method further comprises:

determining, based on the output of the machine learning model, a first object among a plurality of objects on the surface to prioritize; and

controlling a robotic picking device to acquire the first object.

17 . The method of claim 16 , wherein the method further comprises determining a second object among the plurality of objects to prioritize, after the first object has been removed from the surface.

18 . The method of claim 10 , wherein the machine learning model is configured to identify the object; and wherein the information associated with the object comprises at least one of a mask, a key point, a class, and a bounding box for the object identified by the machine learning model.

19 . A system for detecting an object on a surface, comprising:

a camera oriented to include the surface in a field-of-view (FOV) of the camera;

a plurality of aimable illumination sources configured to project corresponding illuminations toward the surface within the FOV of the camera; and

a controller comprising a processor that is configured:

to control the camera to capture, with the camera, a first background image of the object on the surface;

to control the plurality of aimable illumination sources to project the corresponding illuminations simultaneously toward the surface to form a projected illumination on the surface;

to control the camera to capture a scan image comprising the projected illumination formed by simultaneously projected corresponding illuminations that intersect the surface at respective positions;

to determine a first depth map based on separating the scan image into a plurality of object scan mask sets, each object scan mask set representing an image formed only by an illumination of an aimable illumination source chosen from the plurality of aimable illumination sources;

to provide the first depth map and the first background image to a trained machine learning model executed by the processor; and

to receive an output from the machine learning model, wherein the output comprises information associated with the object;

wherein the controller is further configured:

to controllably scan the projected illumination across the surface within the FOV in a plurality of scan steps by changing a position of the corresponding illumination of each of the plurality aimable illumination source on the surface;

to determine an object scan set, comprising a plurality of scan image sets, at each scan step of the plurality of scan steps; and

to determine the first depth map based on the object scan set.

20 . The system of claim 19 , wherein the machine learning model is configured:

to identify the object; and

wherein the information associated with the object comprises at least one of a mask, a key point, a class, and a bounding box for the object identified by the machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2022
From: TAO, YANG; WANG, DONGYI; ALI, MOHAMED AMR
To: UNIVERSITY OF MARYLAND, COLLEGE PARK
Reel/Frame 062100/0558 →
Continuity (3)
Provisional Application 63414881 · Oct 10, 2022
Provisional Application 63262327 · Oct 8, 2021
Related Publication 20230196599A1 · Jun 22, 2023
References Cited (18)
US 6700669B1 · Geng · 2004 [cited by examiner]
US 12046010B2 · Balthasar · 2024 [cited by examiner]
US 20150224650A1 · Xu · 2015 [cited by examiner]
US 20200070370A1 · Wakabayashi · 2020 [cited by examiner]
US 20200166333A1 · Zhou · 2020 [cited by examiner]
US 20210081791A1 · Goodrich · 2021 [cited by examiner]
Yu, C., Chen, X., & Xi, J. (2017). Modeling and calibration of a novel one-mirror galvanometric laser scanner. Sensors, 17(1), 164. (Year: 2017). [cited by examiner]
Je, C., Lee, K. H., & Lee, S. W. (2013). Multi-projector color structured-light vision. Signal Processing: Image Communication, 28(9), 1046-1058. (Year: 2013). [cited by examiner]
Li, Y., Qu, X., Zhang, F., & Zhang, Y. (2020). Separation method of superimposed gratings in double-projector structured-light vision 3D measurement system. Optics Communications, 456, 124676. (Year: 2020). [cited by examiner]
Biskup et al., A Stereo Imaging System for Measuring Structural Parameters of Plant Canopies, Plant, Cell and Environment, 2007, 30:1299-1308. [cited by applicant]
Cai et al., Measurement of Potato Volume with Laser Triangulation and Three-Dimensional Reconstruction, IEEE Access, 2020, 8:176565-176574. [cited by applicant]
Geng, Structured-Light 3D Surface Imaging: A Tutorial, Advances in Optics and Photonics, 2011, 3:128-160. [cited by applicant]
Horn, Shape From Shading: A Method for Obtaining the Shape of a Smooth Opaque Object From One View, Technical Report 232, MIT Artificial Intelligence Laboratory, 1970, 198 pages. [cited by applicant]
Lenz et al., Techniques for Calibration of the Scale Factor and Image Center for High Accuracy 3-D Machine Vision Metrology, IEEE Transactions on Pattern Analysis and Machine Intelligence, 1988, 10(5):713-720. [cited by applicant]
Mertz et al., A Low-Power Structured Light Sensor for Outdoor Scene Reconstruction and Dominant Material Identification, In 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, pp.… [cited by applicant]
Misimi et al., GRIBBOT—Robotic 3D Vision-Guided Harvesting of Chicken Fillets, Computers and Electronics in Agriculture, 2016, 121:84-100. [cited by applicant]
Vazquez-Arellano et al., 3-D Imaging Systems for Agricultural Applications—A Review, Sensors, 2016, 16:618, pp. 1-24. [cited by applicant]
Westoby et al., ‘Structure-from-Motion’ Photogrammetry: A Low-Cost Effective Tool for Geoscience Applications, Geomorphology, 2012, 179:300-314. [cited by applicant]