IP Library Granted Patent US 12,608,955
Granted Patent B2
US 12,608,955 · App. 18/340,819 · Granted Apr 21, 2026

Inferring intent using computer vision

Inventors: Kevin Sheu (Fremont, CA); Jie Mao (Santa Clara, CA)
Assignee: Pony AI Inc.
G06V20/584G06F18/24G06N5/04G06N20/00G06T7/11G06T11/20G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,955
App. No.
18/340,819
Granted
Apr 21, 2026
Kind
B2
Abstract

A system trains a model to infer an intent of an entity. The model includes one or more sensors to obtain frames of data, one or more processors, and a memory storing instructions that, when executed by the one or more processors, cause the system to perform steps. A first step includes determining, in each frame of the frames, one or more bounding regions, each of the bounding regions enclosing an entity. A second step includes identifying a common entity, the common entity being present in bounding regions corresponding to a plurality of the frames. A third step includes associating the common entity across the frames. A fourth step includes training a model to infer an intent of the common entity based on data outside of the bounding regions.

Claims (29)

1 . A system configured to train a model to infer an intent of an entity, comprising:

one or more sensors configured to obtain frames of data;

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, cause the system to perform:

determining, in each frame of the frames, one or more bounding regions, each of the bounding regions enclosing an entity; and

inferring, based on a trained model, an intent associated with the entity based on data outside of the bounding regions, wherein inferring an intent is based on semantic segmentation to predict a category or a classification associated with one or more pixels of the frames and based on instance segmentation to predict whether two pixels associated with a common category or a common classification belong to same or different instances.

2 . The system of claim 1 , wherein the trained model is trained based on Lidar data.

3 . The system of claim 1 , wherein the instructions further cause the system to perform:

rescaling a segmentation output to fit dimensions of the bounding regions.

4 . The system of claim 1 , wherein:

the one or more sensors comprise a camera;

the entity comprises a vehicle; and

the intent is associated with a turning or braking maneuver of the vehicle.

5 . The system of claim 4 , wherein the intent is associated with a left or right turn signal.

6 . The system of claim 1 , wherein the trained model infers an intent based on a probability of a left turn signal of a vehicle being on, a probability of a right turn signal of the vehicle being on, and a probability of a brake light of the vehicle being on.

7 . The system of claim 4 , wherein the trained model is trained based on cross entropy losses over the inferred intent, over left or right turn signals of the vehicle, and over the vehicle.

8 . The system of claim 1 , wherein the inferring of the intent comprises inferring the intent under different weather and lighting conditions.

9 . The system of claim 1 , wherein the trained model is trained based on a classification loss, a bounding box loss, and a mask prediction loss.

10 . The system of claim 1 , wherein the model is associated with a softmax layer that determines probabilities that each pixel of the frames belongs to a particular classification or category.

11 . A method comprising:

obtaining, using one or more sensors, frames of data;

determining, in each frame of the frames, one or more bounding regions, each of the bounding regions enclosing an entity; and

inferring, based on a trained model, an intent associated with the entity based on data outside of the bounding regions, wherein inferring an intent is based on semantic segmentation to predict a category or a classification associated with one or more pixels of the frames and based on instance segmentation to predict whether two pixels associated with a common category or a common classification belong to same or different instances.

12 . A system configured to train a model to infer an intent of an entity, comprising:

one or more sensors configured to obtain frames of data;

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, cause the system to perform:

determining, in each frame of the frames, one or more bounding regions, each of the bounding regions enclosing an entity; and

inferring, based on a trained model, an intent associated with the entity based on data outside of the bounding regions, wherein inferring an intent is based on a probability of a left turn signal of a vehicle being on, a probability of a right turn signal of the vehicle being on, and a probability of a brake light of the vehicle being on, wherein the trained model is trained based on cross entropy losses over the inferred intent, over left or right turn signals of the vehicle, and over the vehicle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SHEU, KEVIN; MAO, JIE
To: PONY AI INC.
Reel/Frame 064048/0956 →
Continuity (2)
Continuation 17011901 · Sep 3, 2020
Related Publication 20230351775A1 · Nov 2, 2023
References Cited (24)
US 7239958B2 · Grougan et al. · 2007 [cited by applicant]
US 8502860B2 · Dernirdjian · 2013 [cited by applicant]
US 10486707B2 · Zelman et al. · 2019 [cited by applicant]
US 10824155B2 · Hong · 2020 [cited by examiner]
US 11003928B2 · Littman et al. · 2021 [cited by applicant]
US 11048982B2 · Steelberg et al. · 2021 [cited by applicant]
US 11126873B2 · Lee · 2021 [cited by examiner]
US 11370423B2 · Casas et al. · 2022 [cited by applicant]
US 11458987B2 · Li · 2022 [cited by examiner]
US 11518413B2 · Anthony · 2022 [cited by examiner]
US 11527078B2 · Littman et al. · 2022 [cited by applicant]
US 11587329B2 · Ranga · 2023 [cited by examiner]
US 11688179B2 · Sheu · 2023 [cited by examiner]
US 11851055B2 · Gutmann · 2023 [cited by examiner]
US 20170154529A1 · Zhao et al. · 2017 [cited by applicant]
US 20190180624A1 · Hassan-Shafique et al. · 2019 [cited by applicant]
US 20200073399A1 · Tateno et al. · 2020 [cited by applicant]
US 20200410853A1 · Akella et al. · 2020 [cited by applicant]
US 20210182605A1 · Anthony · 2021 [cited by examiner]
US 20210192239A1 · Ma · 2021 [cited by examiner]
US 20210197720A1 · Houston et al. · 2021 [cited by applicant]
US 20210319252A1 · Ha et al. · 2021 [cited by applicant]
US 20220383620A1 · Yoo et al. · 2022 [cited by applicant]
US 20250028322A1 · Armstrong-Crews · 2025 [cited by examiner]