IP Library › Granted Patent US 12,415,281
Granted Patent B2
US 12,415,281 · App. 18/557,125 · Granted Sep 16, 2025

Robotic meal-assembly systems and robotic methods for real-time object pose estimation of high-resemblance random food items

Inventors: Yee Seng Teoh (Singapore, SG); Yadan Zeng (Singapore, SG); Boon Heng Elvin Toh (Singapore, SG); Choon Yue Wong (Singapore, SG); I-Ming Chen (Singapore, SG); Guoniu Zhu (Singapore, SG)
Assignee: NANYANG TECHNOLOGICAL UNIVERSITY
B25J9/1697B25J11/0045B65B5/12B65B57/14B65G47/90G06T7/50G06T7/70G06V10/764G06V10/82G06V20/68G06T2207/30128G06V2201/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,415,281
App. No.
18/557,125
Granted
Sep 16, 2025
Kind
B2
Abstract

Methods, systems and computer readable media are provided for automatic kitting of items. The system for automatic kitting of items includes a robotic device, a first imaging device, a computing device and a controller. The robotic device includes an arm with a robotic gripper at one end. The first imaging device is focused on a device conveying kitted items. The computing device is coupled to the first imaging device and is configured to process image data from the first imaging device. The computing device includes item arrangement verification software configured to determine whether each item desired to be in the kitted items is present or absent in the kitted items in response to the processed image data from the first imaging device and generates data based on whether an item desired to be in the kitted items is absent from the kitted items. The controller is coupled to the computing device to receive the data from the computer representing whether an item desired to be in the kitted items is absent from the kitted items. The controller is also coupled to the robotic device for providing instructions to the robotic device to control movement of the arm and the robotic gripper, wherein at least some of the instructions provided to the robotic device are generated in response to the item desired to be in the kitted items being absent from the kitted items.

Claims (34)

1. A system for automatic kitting of items comprising:

a robotic device comprising an arm with a robotic gripper at one end;

a first imaging device focused on a device conveying kitted items;

a computing device coupled to the first imaging device and configured to process image data from the first imaging device, wherein the computing device includes item arrangement verification software configured to determine whether each item desired to be in the kitted items is present or absent in the kitted items in response to the processed image data from the first imaging device and generates data based on whether an item desired to be in the kitted items is absent from the kitted items; and

a controller coupled to the computing device to receive the data from the computing device representing whether the item desired to be in the kitted items is absent from the kitted items, the controller further coupled to the robotic device for providing instructions to the robotic device to control movement of the arm and the robotic gripper, wherein at least some of the instructions provided to the robotic device are generated in response to the item desired to be in the kitted items being absent from the kitted items,

wherein the controller controls the arm and the robotic gripper to pick an item corresponding to the item desired to be in the kitted items which is absent from the kitted items and place the picked item into the kitted items, and

wherein the computing device further includes object detection software configured to determine a type of each item in the kitted items and a position of that item in the kitted items in response to the processed image data from the first imaging device, and wherein the object detection software provides the type and position of each item present in the kitted items to the item arrangement verification software to determine whether each item desired to be in the kitted items is absent from the kitted items.

2. The system in accordance with claim 1 wherein the at least some of the instructions provided to the robotic device by the controller comprise one or more instructions to pick an item corresponding to the item desired to be in the kitted items which is absent from the kitted items and place the picked item into the kitted items.

3. The system in accordance with claim 1 wherein the item arrangement verification software is further configured to flag the kitted items for manual postprocessing in response to determining that an item in the kitted items is not positioned correctly.

4. The system in accordance with claim 1 wherein the system is a meal assembly system for automatic kitting of food items.

5. The system in accordance with claim 1 further comprising a second imaging device coupled to the computing device and focused on a device conveying items to be kitted.

6. The system in accordance with claim 5 wherein the items on the device conveying items to be kitted are conveyed in trays, and wherein the second imaging device comprises depth sensing capability to enable the robotic gripper and the arm to pick an item from the trays.

7. The system in accordance with claim 5 wherein the computing device is further configured to generate image data corresponding to predefined poses related to a type of item in response to a plurality of backgrounds and a plurality of poses of the item.

8. The system in accordance with claim 7 wherein the computing device is further configured to determine whether each item to be kitted is posed in accordance with one of the predefined poses corresponding to a location and a type of the item based on the type and the location received from the second imaging device.

9. The system in accordance with claim 8 wherein the computing device is configured to determine whether each item in the kitted items is posed in accordance with one of the predefined poses based on a small bounding box related to the location and the type of the item.

10. The system in accordance with claim 9 wherein the system is further configured to determine the small bounding box related to the location and the type of the item based on an original bounding box defined by a correct pose of the type of the item at the location which has been shrunk based on an adjustable tolerance factor.

11. The system in accordance with claim 10 wherein the generated image data further comprises labelled data, the labelled data comprising a range of orientation (ROO) for providing a rough direction to enable fast determination whether each item in the items to be kitted is posed in accordance with one of the predefined poses corresponding to the item to be kitted.

12. The system in accordance with claim 11 wherein the labelled data further comprises center point information and width and length information.

13. The system in accordance with claim 11 wherein the object detection software is configured to determine the type of each item in the kitted items and the position of that item in the kitted items by building a convolutional neural network (CNN) object detector model to classify the items in response to the processed image data from the second imaging device.

14. The system in accordance with claim 13 wherein the object detection software is further configured to determine the pose of the item in the device conveying items to be kitted in response to category and bounding box information and range of orientation (ROO) information received from the classification by the CNN object detector model.

15. The system in accordance with claim 14 wherein the pose is a six-dimensional (6D) pose, and wherein the object detection software is configured to determine the 6D pose of the item in the kitted items further in response to depth map information.

16. A robotic method for automatic kitting of items comprising:

imaging kitted items to generate first image data;

determining whether each item desired to be in the kitted items is present or absent in the kitted items in response to the first image data, comprising determining a type of each item in the kitted items and a position of that item in the kitted items in response to the first image data;

generating data based on whether an item desired to be in the kitted items is absent from the kitted items in response to the type and the position of each item present in the first image data;

generating robotic control instructions in response to the item desired to be in the kitted items being absent from the kitted items; and

providing the robotic control instructions to a robotic device, such that the robotic device controls movement of an arm and a robotic gripper of the robotic device to pick an item corresponding to the item desired to be in the kitted items which is absent from the kitted items and place the picked item into the kitted items.

17. The method in accordance with claim 16 , further comprising flagging the kitted items for manual postprocessing in response to determining that an item in the kitted items is not positioned correctly.

18. A non-transitory computer readable medium comprising instructions for automatic kitting of items in a robotic system, the instructions causing a controller in the robotic system to:

image kitted items to generate first image data;

determine whether each item desired to be in the kitted items is present or absent in the kitted items in response to the first image data, comprising determining a type of each item in the kitted items and a position of that item in the kitted items in response to the first image data;

generate data based on whether an item desired to be in the kitted items is absent from the kitted items in response to the type and the position of each item present in the first image data;

generate robotic control instructions in response to the item desired to be in the kitted items being absent from the kitted items; and

provide the robotic control instructions to a robotic device, such that the robot device controls movement of an arm and a robotic gripper of the robotic device, wherein the robotic control instructions comprise one or more instructions to pick an item corresponding to the item desired to be in the kitted items which is absent from the kitted items and place the picked item into the kitted items.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2024
From: TEOH, YEE SENG; ZENG, YADAN; TOH, BOON HENG ELVIN; WONG, CHOON YUE; CHEN, I-MING; ZHU, GUONIU
To: NANYANG TECHNOLOGICAL UNIVERSITY
Reel/Frame 066135/0508 →
Priority Claims (2)
SG 10202105019T · May 12, 2021 · national
SG 10202105020X · May 12, 2021 · national
Continuity (1)
Related Publication 20240246240A1 · Jul 25, 2024
References Cited (32)
US 11731792B2 · Menon · 2023 [cited by examiner]
US 20200094997A1 · Menon et al. · 2020 [cited by applicant]
US 20200095001A1 · Menon · 2020 [cited by examiner]
US 20200096599A1 · Hewett · 2020 [cited by examiner]
CN 205837254U · 2016 [cited by applicant]
CN 109919008A · 2019 [cited by applicant]
CN 110641787A · 2020 [cited by applicant]
CN 112061748A · 2020 [cited by applicant]
Y. Makiyama, Z. Wang and S. Hirai, “A Pneumatic Needle Gripper for Handling Shredded Food Products,” 2020 IEEE International Conference on Real-time Computing and Robotics (RCAR), Asahikawa, Japan, 2020, pp. 183-187. [cited by applicant]
R. Sam and N. Buniyamin, “A Bernoulli principle based flexible handling device for automation of food manufacturing processes,” 2012 International Conference on Control, Automation and Information Sciences (ICCAIS), Sai… [cited by applicant]
H. Wang, D. Sahoo, C. Liu, E. Lim and S. C. H. Hoi, “Learning Cross-Modal Embeddings With Adversarial Networks for Cooking Recipes and Food Images,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (C… [cited by applicant]
A. Collet, M. Martinez, and S. S. Srinivasa, “The moped framework: Object recognition and pose estimation for manipulation”, International Journal of Robotics Research (IJRR), 30(10):1284-1306, 2011. [cited by applicant]
A. Zeng et al., “Multi-view self-supervised deep learning for 6D pose estimation in the Amazon Picking Challenge,” in IEEE International Conference on Robotics and Automation (ICRA), Singapore, 2017, pp. 1386-1383. [cited by applicant]
W. Wu and J. Yang, “Fast food recognition from videos of eating for calorie estimation,” in 2009 IEEE International Conference on Multimedia and Expo, New York, NY, USA, 2009, pp. 1210-1213. [cited by applicant]
S. Rusinkiewicz and M. Levoy, “Efficient variants of the ICP algorithm,” Proceedings Third International Conference on 3-D Digital Imaging and Modeling, Quebec City, QC, Canada, 2001, pp. 145-152. [cited by applicant]
A. S. Periyasamy, M. Schwam and S. Behnke, “Robust 6D Object Pose Estimation in Cluttered Scenes Using Semantic Segmentation and Pose Regression Networks,” 2018 IEEE/RSJ International Conference on Intelligent Robots an… [cited by applicant]
K. Lai, L. Bo, X. Ren, and D. Fox, “A large-scale hierarchical multiview rgb-d object dataset,” in IEEE International Conference on Robotics and Automation (ICRA) , Shanghai, China, 2011, pp. 1817-1824. [cited by applicant]
P. Marion, P. R. Florence, L. Manuelli, and R. Tedrake, “LabelFusion: A pipeline for generating ground truth labels for real RGBD data of cluttered scenes,” arXivpreprint arXiv:1707.04796, 2017. [cited by applicant]
S. Sengupta, V. Jayaram, B. Curless, S. Seitz, and I. K. Shlizerman, “Background Matting: The World is Your Green Screen”, arXiv preprint arXiv: 2004.00626, 2020. [cited by applicant]
B. Bovcon, J. Muhovic, J. Pers and M. Kristan, “The MaSTrl325 dataset for training deep USV obstacle detection models,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 2019… [cited by applicant]
A. Radford, L. Metz, and S. Chintala, “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” arXiv preprint arXiv.T511.06434, Jan. 2016, 16 pages. [cited by applicant]
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein GAN,” arXiv preprint ar.Y / v: 1701.07875, Dec. 2017, 32 pages. [cited by applicant]
J. J. Lim, A. Khosla, and A. Torralba, “Fpm: Fine pose parts-based model with 3d cad models,” in European Conference on Computer Vision (ECCV), Springer, 2014, pp. 478-493. [cited by applicant]
G. Pavlakos, X. Zhou, A. Chan, K. G. Derpanis, and K. Daniilidis, “6-dof object pose from semantic keypoints,” in IEEE International Conference on Robotics and Automation (ICRA), Singapore, 2017, pp. 2011-2018. [cited by applicant]
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox, “Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes,” arXivprepFrint arXivAl11.00199, 2017. [cited by applicant]
J. Wu et al., “Real-Time Object Pose Estimation with Pose Interpreter Networks,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 2018, pp. 6798-6805. [cited by applicant]
H. Wang, S. Sridhar, J. Huang, J. Valentin, S. Song and L. J. Guibas, “Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,” 2019 IEEE/CVF Conference on Computer Vision and Pattern R… [cited by applicant]
S. Gu, A. Lugmayr, M. Danelljan, M. Fritsche, J. Lamour and R. Timofte, “DIV8K: Diverse 8K Resolution Image Dataset,” 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea (South), 201… [cited by applicant]
M. Shugrina et al., “Creative Flow+ Dataset,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 5379-5388. [cited by applicant]
M. Cordts et al., “The Cityscapes Dataset for Semantic Urban Scene Understanding,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 3213-3223. [cited by applicant]
S. R. Richter, V. Vineet, S. Roth, V. Koltun. “Playing for data: ground truth from computer games,” In European conference on computer vision (ECCV ), 2016. [cited by applicant]
A. Bochkovskiy, C. Wang, H. M. Liao, “YOLOv4: Optimal Speed and Accuracy of Object Detection,” arXiv preprint arXiv:2004.10934, 2020. [cited by applicant]