IP Library › Granted Patent US 12,725,306
Granted Patent B2
US 12,725,306 · App. 18/081,669 · Granted Sep 1, 2026

3D geometric and semantic awareness with deep learning for autonomous guidance

Inventors: Zhe Zhang (Sunnyvale, CA); Zhongwei Li (Beijing, CN); Peizhang Chen (Guangzhou, CN); Rui Xiang (Beijing, CN); Xu Han (Shanghai, CN)
Assignee: Trifo, Inc.
G06T7/80G05D1/0214G05D1/0246G05D1/0274G06T7/50A47L11/4011A47L2201/04G06T2207/10028G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,306
App. No.
18/081,669
Granted
Sep 1, 2026
Kind
B2
Abstract

The technology disclosed includes systems and methods for using a deep learning trained classifier to detect obstacles and pathways in an environment in which a robot moves. The system includes logic to receive image information captured by at least one visual spectrum-capable camera and location information captured by at least one depth measuring camera located on a mobile platform. The system includes logic to determine a three-dimensional 3D point cloud of points having 3D information. The system can determine, using an ensemble of trained neural network classifiers, an identity for objects. The system includes logic to determine an occupancy map of the environment. The system includes logic to provide the occupancy map to a process for initiating robot movement to avoid objects in the occupancy map of the environment.

Claims (63)

1 . A method for using a deep learning trained classifier to detect obstacles and pathways in an environment in which a robot moves, based upon image information as captured by at least one visual spectrum-capable camera that captures images in a visual spectrum (RGB) range and at least one depth measuring camera, the method comprising:

receiving image information captured by at least one visual spectrum-capable camera and location information captured by at least one depth measuring camera located on a mobile platform;

extracting, by a processor, from the image information, features in the environment;

determining, by a processor, a three-dimensional 3D point cloud of points having 3D information including location information from the depth camera and the at least one visual spectrum-capable camera, the points corresponding to the features in the environment as extracted;

determining, by a processor, using an ensemble of trained neural network classifiers, including first trained neural network classifiers, an identity for objects corresponding to the features as extracted from the images;

determining, by a processor, from the 3D point cloud and the identity for objects as determined using the ensemble of trained neural network classifiers, an occupancy map of the environment; and

providing the occupancy map to a process for initiating robot movement to avoid objects in the occupancy map of the environment.

2 . The method of claim 1 ,

wherein the depth camera is tightly coupled with the at least one visual spectrum-capable camera by (i) an overlapping of fields of view; (ii) a calibration of pixels per unit area of field of view; and (iii) a synchronous capture of images; thereby enabling locations and features of objects to correspond to one another in sets of images captured by the cameras.

3 . The method of claim 2 ,

wherein the calibration of pixels per unit area of field of view of the at least one visual spectrum-capable camera and the depth camera is 1:1.

4 . The method of claim 2 ,

wherein the calibration of pixels per unit area of field of view of the at least one visual spectrum-capable camera and the depth camera is 16:1.

5 . The method of claim 2 ,

wherein the calibration of pixels per unit area of field of view of the at least one visual spectrum-capable camera and the depth camera is 24:1.

6 . The method of claim 1 , further including:

annotating, by a processor, the occupancy map with annotations of object identities at locations corresponding to at least some of the points in the 3D point cloud; and

using the occupancy map as annotated to plan paths to avoid certain ones of objects based upon identity and location.

7 . The method of claim 6 , wherein the occupancy map is one of a 3D map and a 2D grid representation of a 3D map.

8 . The method of claim 1 , further including:

determining, by a processor, using second trained neural network classifiers, an identity for a room based upon objects identified that correspond to the features as extracted from the images;

annotating, by a processor, the occupancy map with annotations of room identities at locations corresponding to at least some of the points in the 3D point cloud; and

using the occupancy map as annotated to plan paths to remain within or to avoid certain ones of rooms based upon identity and location.

9 . The method of claim 1 ,

wherein the field of view of the at least one visual spectrum-capable camera is 1920×1080 and the depth camera is 224×172.

10 . The method of claim 1 ,

wherein the field of view of the depth camera is within one of: (i) a range including −20 and +20 about a principal axis of the depth camera; and (ii) a range of −30 and +30 about a principal axis of the depth camera.

11 . The method of claim 1 ,

wherein alignment for an angle between a principal axis of the depth camera and a principal axis of the robot along a horizontal plane includes values between 0 degrees and +/−30 degrees.

12 . The method of claim 1 ,

wherein alignment for an angle between a principal axis of the depth camera and a principal axis of the robot along a horizontal plane includes values between 0 degrees and +/−45 degrees.

13 . The method of claim 1 ,

wherein alignment for an angle between a principal axis of the depth camera and a principal axis of the robot along a horizontal plane includes values between 0 degrees and +/−90 degrees.

14 . The method of claim 1 ,

wherein alignment for an angle between a principal axis of the depth camera and a principal axis of the robot along a horizontal plane includes values between 0 degrees and +/−120 degrees.

15 . The method of claim 1 , wherein trained neural network classifiers implement convolutional neural networks (CNN).

16 . The method of claim 1 , further including employing trained neural network classifiers implementing recursive neural networks (RNN) for time-based information.

17 . The method of claim 1 , further including employing trained neural network classifiers implementing long short-term memory networks (LSTM) for time-based information.

18 . The method of claim 1 , wherein the ensemble of neural network classifiers includes:

80 levels in total, from an input to an output.

19 . The method of claim 1 , wherein the ensemble of neural network classifiers implements a multi-layer convolutional network.

20 . The method of claim 19 , wherein the multi-layer convolutional network includes:

60 convolutional levels.

21 . The method of claim 1 , wherein the ensemble of neural network classifiers includes:

normal convolutional levels and depth-wise convolutional levels.

22 . A robot system comprising:

a mobile platform having disposed thereon:

at least one visual spectrum-capable camera to capture images in a visual spectrum (RGB) range;

at least one depth measuring camera; and

an interface to a host including one or more processors coupled to a memory storing instructions to implement a method comprising:

receiving image information captured by at least one visual spectrum-capable camera and location information captured by at least one depth measuring camera located on a mobile platform;

extracting, by a processor, from the image information, features in an environment;

determining, by a processor, a three-dimensional 3D point cloud of points having 3D information including location information from the depth camera and the at least one visual spectrum-capable camera, the points corresponding to the features in the environment as extracted;

determining, by a processor, using an ensemble of trained neural network classifiers, including first trained neural network classifiers, an identity for objects corresponding to the features as extracted from the images;

determining, by a processor, from the 3D point cloud and the identity for objects as determined using the ensemble of trained neural network classifiers, an occupancy map of the environment; and

providing the occupancy map to a process for initiating robot movement to avoid objects in the occupancy map of the environment.

23 . A non-transitory computer readable medium comprising stored instructions, which when executed by a processor, cause the processor to implement a method comprising:

receiving image information captured by at least one visual spectrum-capable camera and location information captured by at least one depth measuring camera located on a mobile platform;

extracting, by a processor, from the image information, features in the environment;

determining, by a processor, a three-dimensional 3D point cloud of points having 3D information including location information from the depth camera and the at least one visual spectrum-capable camera, the points corresponding to the features in the environment as extracted;

determining, by a processor, using an ensemble of trained neural network classifiers, including first trained neural network classifiers, an identity for objects corresponding to the features as extracted from the images;

determining, by a processor, from the 3D point cloud and the identity for objects as determined using the ensemble of trained neural network classifiers, an occupancy map of the environment; and

providing the occupancy map to a process for initiating robot movement to avoid objects in the occupancy map of the environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2025
From: ZHANG, ZHE; LI, ZHONGWEI; CHEN, PEIZHANG; XIANG, RUI; HAN, XU
To: TRIFO, INC.
Reel/Frame 070225/0492 →
Priority Claims (1)
CN 202111613025.5 · Dec 27, 2021 · national
Continuity (2)
Provisional Application 63294901 · Dec 30, 2021
Related Publication 20230363609A1 · Nov 16, 2023
References Cited (150)
US 8475275B2 · Weston et al. · 2013 [cited by applicant]
US 8639644B1 · Hickman et al. · 2014 [cited by applicant]
US 8790180B2 · Barney et al. · 2014 [cited by applicant]
US 9983592B2 · Hong et al. · 2018 [cited by applicant]
US 10032276B1 · Liu et al. · 2018 [cited by applicant]
US 10366508B1 · Liu et al. · 2019 [cited by applicant]
US 10390003B1 · Liu et al. · 2019 [cited by applicant]
US 10410328B1 · Liu et al. · 2019 [cited by applicant]
US 10423861B2 · Gao et al. · 2019 [cited by applicant]
US 10571925B1 · Zhang et al. · 2020 [cited by applicant]
US 10571926B1 · Zhang et al. · 2020 [cited by applicant]
US D908992S · Jang · 2021 [cited by applicant]
US D908993S · Li et al. · 2021 [cited by applicant]
US D924522S · Jang · 2021 [cited by applicant]
US 11069082B1 · Ebrahimi Afrouzi et al. · 2021 [cited by applicant]
US 11223497B2 · Hong et al. · 2022 [cited by applicant]
US D961177S · Jang · 2022 [cited by applicant]
US 11592825B2 · Jung et al. · 2023 [cited by applicant]
US 11687092B2 · Thorne et al. · 2023 [cited by applicant]
US 11774983B1 · Zhang et al. · 2023 [cited by applicant]
US 12140954B2 · Hong et al. · 2024 [cited by applicant]
US 12175759B2 · Furuhata · 2024 [cited by applicant]
US 12197204B2 · Kim et al. · 2025 [cited by applicant]
US 20070156286A1 · Yamauchi · 2007 [cited by applicant]
US 20070192910A1 · Vu et al. · 2007 [cited by applicant]
US 20080086236A1 · Saito et al. · 2008 [cited by applicant]
US 20080154429A1 · Lee et al. · 2008 [cited by applicant]
US 20100161225A1 · Hyung et al. · 2010 [cited by applicant]
US 20100332128A1 · Ikeuchi et al. · 2010 [cited by applicant]
US 20110077802A1 · Halloran et al. · 2011 [cited by applicant]
US 20110205338A1 · Choi et al. · 2011 [cited by applicant]
US 20120095619A1 · Pack et al. · 2012 [cited by applicant]
US 20120182392A1 · Kearns et al. · 2012 [cited by applicant]
US 20120185091A1 · Field et al. · 2012 [cited by applicant]
US 20130056032A1 · Choe et al. · 2013 [cited by applicant]
US 20130245937A1 · DiBernardo et al. · 2013 [cited by applicant]
US 20130338525A1 · Allen · 2013 [cited by applicant]
US 20140316636A1 · Hong · 2014 [cited by examiner]
US 20140350839A1 · Pack et al. · 2014 [cited by applicant]
US 20150234398A1 · Harris et al. · 2015 [cited by applicant]
US 20160166126A1 · Morin et al. · 2016 [cited by applicant]
US 20160183752A1 · Morin et al. · 2016 [cited by applicant]
US 20170196196A1 · Trottier et al. · 2017 [cited by applicant]
US 20180012411A1 · Richey et al. · 2018 [cited by applicant]
US 20180075403A1 · Mascorro Medina et al. · 2018 [cited by applicant]
US 20180284802A1 · Tsai et al. · 2018 [cited by applicant]
US 20180286072A1 · Tsai et al. · 2018 [cited by applicant]
US 20180289579A1 · Agrawal · 2018 [cited by applicant]
US 20180317725A1 · Lee et al. · 2018 [cited by applicant]
US 20180364731A1 · Liu et al. · 2018 [cited by applicant]
US 20190026943A1 · Yan et al. · 2019 [cited by applicant]
US 20190102667A1 · Bashkirov · 2019 [cited by examiner]
US 20190156944A1 · Eriksson et al. · 2019 [cited by applicant]
US 20190235083A1 · Zhang et al. · 2019 [cited by applicant]
US 20190291277A1 · Oleynik · 2019 [cited by applicant]
US 20190324473A1 · Hillen · 2019 [cited by applicant]
US 20190346271A1 · Zhang et al. · 2019 [cited by applicant]
US 20190355173A1 · Gao · 2019 [cited by applicant]
US 20190370691A1 · Chae et al. · 2019 [cited by applicant]
US 20190377349A1 · van der Merwe et al. · 2019 [cited by applicant]
US 20190392240A1 · Araújo et al. · 2019 [cited by applicant]
US 20200004260A1 · Kim et al. · 2020 [cited by applicant]
US 20200008639A1 · Lee et al. · 2020 [cited by applicant]
US 20200019181A1 · Kim et al. · 2020 [cited by applicant]
US 20200029490A1 · Bertucci et al. · 2020 [cited by applicant]
US 20200039068A1 · Kim · 2020 [cited by applicant]
US 20200107008A1 · Hur et al. · 2020 [cited by applicant]
US 20200117212A1 · Tian · 2020 [cited by applicant]
US 20200117213A1 · Tian et al. · 2020 [cited by applicant]
US 20200117898A1 · Tian et al. · 2020 [cited by applicant]
US 20200156255A1 · Soltani Bozchalooi et al. · 2020 [cited by applicant]
US 20200192388A1 · Zhang et al. · 2020 [cited by applicant]
US 20200225673A1 · Ebrahimi Afrouzi et al. · 2020 [cited by applicant]
US 20200311971A1 · Corcodel et al. · 2020 [cited by applicant]
US 20200334843A1 · Kasuya et al. · 2020 [cited by applicant]
US 20200334855A1 · Higo et al. · 2020 [cited by applicant]
US 20200394410A1 · Zhang et al. · 2020 [cited by applicant]
US 20210000006A1 · Ellaboudy et al. · 2021 [cited by applicant]
US 20210019527A1 · Zhang et al. · 2021 [cited by applicant]
US 20210089040A1 · Ebrahimi Afrouzi et al. · 2021 [cited by applicant]
US 20210114213A1 · Lee et al. · 2021 [cited by applicant]
US 20210121035A1 · Kim et al. · 2021 [cited by applicant]
US 20210142788A1 · Choi et al. · 2021 [cited by applicant]
US 20210151043A1 · Lee et al. · 2021 [cited by applicant]
US 20210174097A1 · Tsai et al. · 2021 [cited by applicant]
US 20210232144A1 · Lee · 2021 [cited by applicant]
US 20210294328A1 · Dhayalkar · 2021 [cited by applicant]
US 20220066456A1 · Ebrahimi Afrouzi et al. · 2022 [cited by applicant]
US 20220095872A1 · Bassa · 2022 [cited by examiner]
US 20220156554A1 · Fu · 2022 [cited by examiner]
US 20220193888A1 · Rephaeli · 2022 [cited by examiner]
US 20220198248A1 · Wang · 2022 [cited by examiner]
US 20220287527A1 · Hong et al. · 2022 [cited by applicant]
US 20230363609A1 · Zhang et al. · 2023 [cited by applicant]
US 20230363610A1 · Zhang et al. · 2023 [cited by applicant]
US 20240310851A1 · Ebrahimi Afrouzi et al. · 2024 [cited by applicant]
US 20250086031A1 · Gao · 2025 [cited by examiner]
CN 106725135A · 2017 [cited by applicant]
CN 111714042A · 2020 [cited by applicant]
CN 111839375A · 2020 [cited by applicant]
CN 211749328U · 2020 [cited by applicant]
CN 212382573U · 2021 [cited by applicant]
CN 112869673A · 2021 [cited by applicant]
CN 112998605A · 2021 [cited by applicant]
CN 213309501U · 2021 [cited by applicant]
CN 214073161U · 2021 [cited by applicant]
CN 214073183U · 2021 [cited by applicant]
CN 214414759U · 2021 [cited by applicant]
CN 214804493U · 2021 [cited by applicant]
CN 215191283U · 2021 [cited by applicant]
CN 215457682U · 2022 [cited by applicant]
CN 215650867U · 2022 [cited by applicant]
CN 215838790U · 2022 [cited by applicant]
CN 216020830U · 2022 [cited by applicant]
CN 216569783U · 2022 [cited by applicant]
WO 2015063119A1 · 2015 [cited by applicant]
U.S. Appl. No. 18/081,672, filed Dec. 14, 2022, 20230363610, Nov. 16, 2023, Pending. [cited by applicant]
Zhang et al., ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices, dated Dec. 7, 2017, 9 pages. [cited by applicant]
Lin et. al., Network in Network, in Proc. of ICLR, 2014. [cited by applicant]
Sifre, Rigid-motion Scattering for Image Classification, Ph.D. thesis, 2014. [cited by applicant]
Sifre et. al., Rotation, Scaling and Deformation Invariant Scattering for Texture Discrimination, in Proc. of CVPR, 2013. [cited by applicant]
Chollet, Xception: Deep Learning with Depthwise Separable Convolutions, in Proc. of CVPR, 2017. 8 pages. [cited by applicant]
He et. al., Deep Residual Learning for Image Recognition, in Proc. of CVPR, 2016. [cited by applicant]
Xie et. al., Aggregated Residual Transformations for Deep Neural Networks, in Proc. of CVPR, 2017. [cited by applicant]
Howard et. al., Mobilenets: Efficient Convolutional Neural Networks for Mobile Vision Applications, 2017. [cited by applicant]
Sandler et. al., MobileNetV2: Inverted Residuals and Linear Bottlenecks, 2018. [cited by applicant]
Qin et. al., FD-MobileNet: Improved MobileNet with a Fast Downsampling Strategy, 2018. [cited by applicant]
Oord et al., Wavenet: A Generative Model for Raw Audio, dated Sep. 19, 2016, 15 pages. [cited by applicant]
Arik et al., Deep Voice: Real-time Neural Text-to-Speech, dated 2017, 17 pages. [cited by applicant]
Yu et al., Multi-Scale Context Aggregation by Dilated Convolutions, ICLR 2016, dated Apr. 30, 2016, 13 pages. [cited by applicant]
He et. al., Deep Residual Learning for Image Recognition, 2015. [cited by applicant]
Srivastava et al., Highway Networks, dated 2015, 6 pages. [cited by applicant]
Huang et al., Densely Connected Convolutional Networks, dated Aug. 17, 2017, 9 pages. [cited by applicant]
Szegedy et. al., Going Deeper with Convolutions, dated 2014, 12 pages. [cited by applicant]
Ioffe et al., Batch Normalization Accelerating Deep Network Training by Reducing Internal Covariate Shift, dated 2015, 11 pages. [cited by applicant]
Piqueras, Autoregressive Model Based on a Deep Convolutional Neural Network for Audio Generation, Tampere University of Technology, dated 2016, 58 pages. [cited by applicant]
Wu, Introduction to Convolutional Neural Networks, Nanjing University, dated 2017, 31 pages. [cited by applicant]
Goodfellow et al., Chapter 9—Convolutional Networks, Deep Learning, MIT Press, dated 2016, 41 pages. [cited by applicant]
Gu et. al., Recent Advances in Convolutional Neural Networks, dated Jan. 5, 2017, 37 pages. [cited by applicant]
Srivastava, Dropout A Simple Way to Prevent Neural Networks from Overfitting, 2014, 30 pages. [cited by applicant]
Chaubard et al., CS224D: Deep Learning for NLP, Lecture Notes Part 1, Spring 2015, Stanford University, 11 pages. [cited by applicant]
Chaubard et al., CS224D: Deep Learning for NLP, Lecture Notes Part 2, Spring 2015, Stanford University, 11 pages. [cited by applicant]
Chaubard et al., CS224D: Deep Learning for NLP, Lecture Notes Part 3, Spring 2015, Stanford University, 14 pages. [cited by applicant]
Chaubard et al., CS224D: Deep Learning for NLP, Lecture Notes Part 4, Spring 2015, Stanford University, 12 pages. [cited by applicant]
Chaubard et al., CS224D: Deep Learning for NLP, Lecture Notes Part 5, Stanford University, Spring 2015, 6 pages. [cited by applicant]
EP 22216758.7 Extended European Search Report dated May 6, 2023, 13 pages. [cited by applicant]
Sun Hao et al.: “Semantic mapping and semantics-boosted navigation with path creation on a mobile robot,” 2017 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE Conference on Robotics, … [cited by applicant]
Sun Hao et al.: “Scene Recognition and Object Detection in a Unified Convolutional Neural Network on a Mobile Manipulator,” 2018 IEEE International Conference on Robotics and Automation (ICRA), IEEE, May 21, 2018, pp. 1… [cited by applicant]
Song Shuran et al.: “Sun RGB-D: A RBG-D scene understanding benchmark suite,” 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 7, 2015, pp. 567-576. [cited by applicant]
Jiao Jichao et al.: “A Post-Recitifcation Approach of Depth Images of Kinect v2 for 3D Reconstruction of Indoor Scenes,” ISPRS International Journal of Geo-Information, vol. 6, No. 11, Nov. 13, 2017, p. 349. [cited by applicant]