IP Library › Granted Patent US 12,307,695
Granted Patent B2
US 12,307,695 · App. 18/335,912 · Granted May 20, 2025

System and method for self-supervised monocular ground-plane extraction

Inventors: Vitor Guizilini (Santa Clara, CA); Rares A. Ambrus (Santa Clara, CA); Adrien David Gaidon (San Francisco, CA)
Assignee: TOYOTA RESEARCH INSTITUTE, INC.
G06T7/521B60W60/001G06N3/08G06T17/05G06T17/10G06V20/588B60W2420/408
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,695
App. No.
18/335,912
Granted
May 20, 2025
Kind
B2
Abstract

A method for controlling an agent to navigate through an environment includes generating a depth map associated with a monocular image of the environment. The method also includes generating a group of surface normal. Each surface normal of the group of surface normals is associated with a respective polygon of a group of polygons associated with the depth map. The method further includes identifying one or more ground planes in the depth map based on the group of surface normal. The method further includes controlling the agent to navigate through the environment based on identifying the one or more ground planes.

Claims (44)

1. An apparatus for controlling an agent to navigate through an environment, comprising:

one or more processors; and

one or more memories coupled with the one or more processors and storing processor-executable code that, when executed by the one or more processors, is configured to cause the apparatus to:

generate a first depth map and a second depth map associated with a monocular image of the environment, the first depth map having a first resolution and the second depth map having a second resolution that is different than the first resolution;

select one of the first depth map or the second depth map for a ground plane extraction task;

generate a plurality of surface normals, each surface normal of the plurality of surface normals being a vector cross product of two edges from a respective polygon of a plurality of polygons associated with the selected depth map;

identify a group of planes in the selected depth map based on the plurality of surface normals, each one of the group of planes being associated with a respective set of surface normals, from the plurality of surface normals, facing a same direction, at least one plane of the group of planes indicating a ground plane, at least one plane of the group of planes being a second plane that is different from the ground plane;

generate a height map of the environment based on the identified group of planes; and

control the agent to navigate through the environment based on generating the height map.

2. The apparatus of claim 1 , wherein each of the plurality of polygons is a triangle.

3. The apparatus of claim 1 , wherein the set of surface normal associated with the ground plane are perpendicular with respect to a direction of the agent and/or a sensor associated with the agent.

4. The apparatus of claim 1 , wherein the ground plane corresponds to a road in the environment.

5. The apparatus of claim 1 , wherein:

the agent is an autonomous or semi-autonomous vehicle; and

the monocular image is captured via a sensor integrated with the agent.

6. The apparatus of claim 1 , wherein execution of the processor-executable code further causes the apparatus to train a depth model, in a self-supervising manner, on the first depth map, the second depth map, and the one or more ground planes.

7. A method for controlling an agent to navigate through an environment, comprising:

generating a first depth map and a second depth map associated with a monocular image of the environment, the first depth map having a first resolution and the second depth map having a second resolution that is different than the first resolution;

selecting one of the first depth map or the second depth map for a ground plane extraction task;

generating a plurality of surface normals, each surface normal of the plurality of surface normals being a vector cross product of two edges from a respective polygon of a plurality of polygons associated with the selected depth map;

identifying a group of planes in the selected depth map based on the plurality of surface normals, each one of the group of planes being associated with a respective set of surface normals, from the plurality of surface normals, facing a same direction, at least one plane of the group of planes indicating a ground plane, at least one plane of the group of planes being a second plane that is different from the ground plane;

generating a height map of the environment based on the identified group of planes; and

controlling the agent to navigate through the environment based on generating the height map.

8. The method of claim 7 , wherein each of the plurality of polygons is a triangle.

9. The method of claim 7 , wherein the set of surface normal associated with the ground plane are perpendicular with respect to a direction of the agent and/or a sensor associated with the agent.

10. The method of claim 7 , wherein the ground plane corresponds to a road in the environment.

11. The method of claim 7 , wherein:

the agent is an autonomous or semi-autonomous vehicle; and

the monocular image is captured via a sensor integrated with the agent.

12. The method of claim 7 , further comprising training a depth model, in a self-supervising manner, on the first depth map, the second depth map, and the one or more ground planes.

13. A non-transitory computer-readable medium having program code recorded thereon for controlling an agent to navigate through an environment, the program code executed by a processor and comprising:

program code to generate a first depth map and a second depth map associated with a monocular image of the environment, the first depth map having a first resolution and the second depth map having a second resolution that is different than the first resolution;

program code to select one of the first depth map or the second depth map for a ground plane extraction task;

program code to generate a plurality of surface normals, each surface normal of the plurality of surface normals being a vector cross product of two edges from a respective polygon of a plurality of polygons associated with the selected depth map;

program code to identify a group of planes in the selected depth map based on the plurality of surface normals, each one of the group of planes being associated with a respective set of surface normals, from the plurality of surface normals, facing a same direction, at least one plane of the group of planes indicating a ground plane, at least one plane of the group of planes being a second plane that is different from the ground plane;

program code to generate a height map of the environment based on the identified group of planes; and

program code to control the agent to navigate through the environment based on generating the height map.

14. The non-transitory computer-readable medium of claim 13 , wherein each of the plurality of polygons is a triangle.

15. The non-transitory computer-readable medium of claim 13 , wherein the ground plane corresponds to a road in the environment.

16. The non-transitory computer-readable medium of claim 13 , wherein:

the agent is an autonomous or semi-autonomous vehicle; and

the monocular image is captured via a sensor integrated with the agent.

17. The non-transitory computer-readable medium of claim 13 , wherein the set of surface normal associated with the ground plane are perpendicular with respect to a direction of the agent and/or a sensor associated with the agent.

18. The non-transitory computer-readable medium of claim 13 , wherein the program code further comprises program code to train a depth model, in a self-supervising manner, on the first depth map, the second depth map, and the one or more ground planes.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 071798/0114 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2023
From: GUIZILINI, VITOR; AMBRUS, RARES A.; GAIDON, ADRIEN DAVID
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 063967/0930 →
Continuity (2)
Continuation 16913238 · Jun 26, 2020
Related Publication 20230326055A1 · Oct 12, 2023
References Cited (7)
US 11734845B2 · Guizilini et al. · 2023 [cited by applicant]
US 20110216213A1 · Kawahata · 2011 [cited by examiner]
US 20180364717A1 · Douillard · 2018 [cited by examiner]
US 20200225673A1 · Ebrahimi Afrouzi · 2020 [cited by examiner]
US 20210144357A1 · Kim · 2021 [cited by examiner]
US 20210352262A1 · Buslaev · 2021 [cited by examiner]
Man, Yunze, et al. “GroundNet: Monocular ground plane normal estimation with geometric consistency.” Proceedings of the 27th ACM International Conference on Multimedia. 2019. (Year: 2019). [cited by examiner]