IP Library Granted Patent US 12694601
Granted Patent B1
US 12694601 · App. 19/352,247 · Granted Jul 28, 2026

Camera-invariant 3D property formulation

Inventors: Tudor Stefan Achim (Menlo Park, CA); Onur Ozyesil (Princeton, NJ); Vladislav Voroninski (Redwood City, CA); Christopher Blake Melgaard (Seattle, WA); Ryan Halabi (Colorado Springs, CO)
Assignee: Helm.ai, Inc.
G06T15/00G06T7/70G06V20/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694601
App. No.
19/352,247
Granted
Jul 28, 2026
Kind
B1
Abstract

Determining a 3D property of an object in a 2D image captured by a camera includes receiving the 2D image captured by the camera. It further includes providing the 2D image to a prediction model trained to predict a quantity that is invariant to parameters of the camera. It further includes determining the 3D property of the object in the 2D image based on a camera parameter and the predicted quantity that is invariant to parameters of the camera.

Claims (32)

1 . A system, comprising:

one or more processors configured to:

receive a 2D (two-dimensional) image captured by a camera;

use a prediction model to predict, for a vehicle object in the 2D image, a camera invariant distance value, wherein the prediction model was trained to produce predictions of camera invariant distance values for input 2D images at least in part by:

assigning, as target labels for pixels that a unique vehicle occupies in a training image, a same target camera invariant distance value comprising a smallest depth value of the unique vehicle normalized by pixel focal length; and

updating the prediction model at least in part by comparing predictions made for the pixels that the unique vehicle occupies in the training image against the target labels comprising the same target camera invariant distance value comprising the smallest depth value of the unique vehicle normalized by pixel focal length; and

determine an absolute depth estimate of the vehicle object in the 2D image based at least in part on the camera invariant distance value predicted for the vehicle object and a camera parameter; and

a memory coupled to the one or more processors and configured to provide the one or more processors with instructions.

2 . The system of claim 1 , wherein the camera parameter comprises a pixel focal length.

3 . The system of claim 1 , wherein the training image was captured using a fisheye lens.

4 . The system of claim 3 , wherein the prediction model was trained at least in part by performing distortion correction to generate an undistorted version of the training image, and wherein the undistorted version of the training image was provided as input to the prediction model.

5 . The system of claim 1 , wherein the prediction model was trained at least in part by generating augmented training images at least in part by applying one or more augmentations to the training image.

6 . The system of claim 5 , wherein the prediction model was trained on the generated augmented training images.

7 . The system of claim 5 , wherein an applied augmentation simulates increasing or decreasing camera pitch.

8 . The system of claim 5 , wherein an applied augmentation simulates vertical movement in an image domain.

9 . The system of claim 5 , wherein an applied augmentation comprises virtual camera rotation.

10 . The system of claim 5 , wherein an applied augmentation comprises zooming in or out.

11 . A method, comprising:

receiving a 2D (two-dimensional) image captured by a camera;

using a prediction model to predict, for a vehicle object in the 2D image, a camera invariant distance value, wherein the prediction model was trained to produce predictions of camera invariant distance values for input 2D images at least in part by:

assigning, as target labels for pixels that a unique vehicle occupies in a training image, a same target camera invariant distance value comprising a smallest depth value of the unique vehicle normalized by pixel focal length; and

updating the prediction model at least in part by comparing predictions made for the pixels that the unique vehicle occupies in the training image against the target labels comprising the same target camera invariant distance value comprising the smallest depth value of the unique vehicle normalized by pixel focal length; and

determining an absolute depth estimate of the vehicle object in the 2D image based at least in part on the camera invariant distance value predicted for the vehicle object and a camera parameter.

12 . The method of claim 11 , wherein the camera parameter comprises a pixel focal length.

13 . The method of claim 11 , wherein the training image was captured using a fisheye lens.

14 . The method of claim 13 , wherein the prediction model was trained at least in part by performing distortion correction to generate an undistorted version of the training image, and wherein the undistorted version of the training image was provided as input to the prediction model.

15 . The method of claim 11 , wherein the prediction model was trained at least in part by generating augmented training images at least in part by applying one or more augmentations to the training image.

16 . The method of claim 15 , wherein the prediction model was trained on the generated augmented training images.

17 . The method of claim 15 , wherein an applied augmentation simulates increasing or decreasing camera pitch.

18 . The method of claim 15 , wherein an applied augmentation simulates vertical movement in an image domain.

19 . The method of claim 15 , wherein an applied augmentation comprises virtual camera rotation.

20 . The method of claim 15 , wherein an applied augmentation comprises zooming in or out.