IP Library › Granted Patent US 12,725,442
Granted Patent B2
US 12,725,442 · App. 18/272,100 · Granted Sep 1, 2026

Object recognition method and time-of-flight object recognition circuitry

Inventors: Malte Ahl (Stuttgart, DE); David Dal Zot (Stuttgart, DE); Varun Arora (Stuttgart, DE)
Assignee: Sony Semiconductor Solutions Corporation
G06V40/10G01S17/894G06T11/00G06V10/14G06V10/774G06V10/82G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,442
App. No.
18/272,100
Granted
Sep 1, 2026
Kind
B2
Abstract

The present disclosure generally pertains to an object recognition method for time-of-flight camera data, including: recognizing a real object based on a pretrained algorithm, wherein the pretrained algorithm is trained based on time-of-flight training data, wherein the time-of-flight training data are generated based on a combination of real time-of-flight data being indicative of a background, and simulated time-of-flight data generated by applying a mask on synthetic overlay image data representing a simulated object, thereby generating a masked simulated object, the mask being generated based on the synthetic overlay image data.

Claims (25)

1 . An object recognition method, comprising:

recognizing a real object captured by an imaging processor of a first camera based on a pretrained algorithm, wherein the pretrained algorithm is trained based on training data having depth components, wherein the training data are generated based on a combination of real data having depth components derived from interpretation of data from a second camera and being indicative of a background, and simulated second data having depth components generated by applying a mask on synthetic overlay image data representing a simulated object, thereby generating a masked simulated object, the mask being generated based on the synthetic overlay image data,

wherein the training data have depth components representing both depth image data and confidence data of an image.

2 . The object recognition method of claim 1 , wherein the mask is based on at least one of a binarization of the simulated object, an erosion of the simulated object and a blurring of the simulated object.

3 . The object recognition method of claim 1 , wherein the mask is based on an application of at least one of the following to the simulated object: a random brightness change, a uniform brightness noise, and balancing the synthetic overlay image data based on the background.

4 . The object recognition method of claim 1 , wherein the pretrained algorithm is based on at least one of a generative adversarial network, a convolutional neural network, a recurrent neural network, and a convolutional neural network in combination with a neural network with a long short-term memory.

5 . The object recognition method of claim 1 , wherein the training data having depth components further include at least one of bounding box information and pixel precise masking information.

6 . The object recognition method of claim 1 , wherein the training data having depth components represent at least one of depth image data and confidence data of an image.

7 . The object recognition method of claim 1 , wherein the training data having depth components are further based on at least one of random data augmentation and hyperparameter tuning.

8 . The object recognition method of claim 1 , wherein the real data having depth components is time-of-flight data and the simulated second data is simulated time-of-flight data.

9 . The object recognition method of claim 1 , wherein the real object includes a hand.

10 . The object recognition method of claim 9 , the method further comprising: recognizing a gesture of the hand.

11 . Object recognition circuitry for recognizing an object in camera data, configured to:

recognize a real object based on a pretrained algorithm, wherein the pretrained algorithm is trained based on training data having depth components, wherein the training data are generated based on a combination of real data having depth components being indicative of a background, and simulated second data generated by applying a mask on synthetic overlay image data representing a simulated object, thereby generating a masked simulated object, the mask being generated based on the synthetic overlay image data,

wherein the training data have depth components representing both depth image data and confidence data of an image.

12 . The object recognition circuitry of claim 11 , wherein the mask is based on at least one of a binarization of the simulated object, an erosion of the simulated object and a blurring of the simulated object.

13 . The object recognition circuitry of claim 11 , wherein the mask is based on an application of at least one of the following to the simulated object: a random brightness change, a uniform brightness noise, and balancing the synthetic overlay image data based on the background.

14 . The object recognition circuitry of claim 11 , wherein the pretrained algorithm is based on at least one of a generative adversarial network, a convolutional neural network, a recurrent neural network, and a convolutional neural network in combination with a neural network with a long short-term memory.

15 . The object recognition circuitry of claim 11 , wherein the training data having depth components further include at least one of bounding box information and pixel precise masking information.

16 . The object recognition circuitry of claim 11 , wherein the training data having depth components represent at least one of depth image data and confidence data of an image.

17 . The object recognition circuitry of claim 11 , wherein the training data having depth components are further based on at least one of random data augmentation and hyperparameter tuning.

18 . The object recognition circuitry of claim 11 , wherein the pretrained algorithm is further trained based on early stopping.

19 . The object recognition circuitry of claim 11 , wherein the real object includes a hand.

20 . The object recognition circuitry of claim 19 , further configured to:

recognize a gesture of the hand.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2023
From: AHL, MALTE; DAL ZOT, DAVID; ARORA, VARUN
To: SONY SEMICONDUCTOR SOLUTIONS CORPORATION
Reel/Frame 065914/0581 →
Priority Claims (1)
EP 21151753 · Jan 15, 2021 · regional
Continuity (1)
Related Publication 20240071122A1 · Feb 29, 2024
References Cited (37)
US 10186038B1 · Kluckner · 2019 [cited by examiner]
US 10803310B2 · Estrada et al. · 2020 [cited by applicant]
US 11094134B1 · Fallin · 2021 [cited by examiner]
US 20140204013A1 · O'Prey · 2014 [cited by applicant]
US 20170060254A1 · Molchanov · 2017 [cited by examiner]
US 20170364733A1 · Estrada · 2017 [cited by examiner]
US 20180189951A1 · Liston · 2018 [cited by examiner]
US 20200057831A1 · Wu et al. · 2020 [cited by applicant]
US 20200117953A1 · Cooper · 2020 [cited by applicant]
US 20200167161A1 · Planche · 2020 [cited by applicant]
US 20200175759A1 · Russell et al. · 2020 [cited by applicant]
US 20200294201A1 · Planche · 2020 [cited by examiner]
US 20210383096A1 · White · 2021 [cited by examiner]
US 20220101047A1 · Puri · 2022 [cited by examiner]
JP 2016503220A · 2016 [cited by applicant]
JP 2018169690A · 2018 [cited by applicant]
JP 6719168B1 · 2020 [cited by applicant]
WO WO2019113510A1 · 2019 [cited by examiner]
WO 2020056532A1 · 2020 [cited by applicant]
WO WO2020117657A1 · 2020 [cited by applicant]
Song, Xibin, Fan Zhong, Yanke Wang, and Xueying Qin. “Estimation of kinect depth confidence through self-training.” The Visual Computer 30, No. 6 (2014): 855-865. (Year: 2014). [cited by examiner]
Aggarwal, Charu Neural Networks and Deep Learning 2018 Springer, ISBN: 3319944622; Abstract attached. [cited by applicant]
Goodfellow, Ian et al Deep Learning the MIT Press, 2016 ISBN: 0262035618; Abstract attached. [cited by applicant]
Haykin, Simon Neural Networks and Learning Machines Pearson, Nov. 18, 2008, ISBN: 0131471392; Abstract attached. [cited by applicant]
Kahn, Salman Guide to Convolutional Neural Networks for Computer Vision Morgan & Claypool, 2018, ISBN: 1681732785; Abstract attached. [cited by applicant]
Remondino, Fabio, et al TOF Range-Imaging Cameras 2013 Springer, ISBN 9783642275227; Abstract attached. [cited by applicant]
Yuki Hiramatsu et al., “Semantic segmentation using attention mechanism from class perspective”, SSII 2019 [USB] Picture sensing technology, Dec. 31, 2019. [cited by applicant]
Ishida Yutaro et al., “Deep Neural Networks for Object Detection and Classification on Domestic Service Robots”, IEICE Technical Report, The Institute of Electronics, Information and Communication Engineers, pp. 45-50. [cited by applicant]
Kim Wonjik et al., “ Automatic Labeled LiDAR Data Generation based on Precise Human Model”, 2019 International Conference on Robotics and Automation (ICRA) [online], IEEE, 2019 , pp. 43 49. [cited by applicant]
Ryosuke Kasahara, “The development method of a visual examination using image recognition”, Part1 Visual examination technology of Chapter 2 by the foundation and the machine learning of image recognition technology, a … [cited by applicant]
International Search Report and Written Opinion mailed on Apr. 28, 2022, received for PCT Application PCT/EP2022/050645, filed on Jan. 13, 2022, 13 pages. [cited by applicant]
Maximov et al., “Focus on defocus: bridging the synthetic to real domain gap for depth estimation”, arXiv:2005.09623v1 [cs.CV], May 19, 2020, 10 pages. [cited by applicant]
Rozantsev et al., “On Rendering Synthetic Images for Training an Object Detector”, Computer Vision and Image Understanding, Jan. 12, 2015, pp. 1-30. [cited by applicant]
Planche et al., “DepthSynth: Real-Time Realistic Synthetic Data Generation from CAD Models for 2.5D Recognition”, 2017 International Conference on 3D Vision (3DV), IEEE, Oct. 10, 2017, pp. 1-10. [cited by applicant]
Zanuttigh Pietro, “Time-of-Flight and Structured Light Depth Cameras: Technology and Applications”, Springer, XP055912778, Jan. 1, 2016, pp. 99-107. [cited by applicant]
Benjamin Planche et al, “DepthSynth:Real-Time Realistic Synthetic Data Generation from CAD Models for 2.5D Recognition”, Feb. 27, 2017. [cited by applicant]
Wonjik Kim et al, “Automatic Labeled LiDAR Data Generation based on Precise Human Model”, Feb. 14, 2019. [cited by applicant]