IP Library Granted Patent US 11,551,030
Granted Patent B2
US 11,551,030 · App. 17/067,470 · Granted Jan 10, 2023

Visualizing machine learning predictions of human interaction with vehicles

Inventor: Stephen Cope (Somerville, MA)
Assignee: Perceptive Automata, Inc.
G06K9/6253G06F3/04842G06K9/6261G06N20/00G06T11/206G06V20/41G06V20/56G06V40/20G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,030
App. No.
17/067,470
Granted
Jan 10, 2023
Kind
B2
Abstract

A computing device accesses video data displaying one or more traffic entities and generates a plurality of sequences from the video data. For each sequence, the computing device identifies a plurality of stimuli in the sequence and applies a machine learning model to generate an output describing the traffic entity. The computing device generates a data structure for storing, for each sequence, information describing the sequence and linking frame indexes of stimuli from the sequence to outputs of the machine learning model. The computing device stores the data structure in association with the video data. Responsive to receiving a selection of a sequence, the computing device loads video data for the sequence. Responsive to receiving a selection of a traffic entity within the video data, the computing device generates a graphical display element including the machine learning model output for the selected traffic entity.

Claims (49)

1. A method comprising:

accessing video data, the video data displaying one or more traffic entities;

generating, from the video data, a plurality of sequences of video frames, each sequence of video frames corresponding to a portion of the video data displaying a traffic entity of the one or more traffic entities;

for a sequence of video frames of the plurality of sequences of video frames:

identifying a plurality of stimuli in the sequence of video frames, each stimulus including a subset of video frames of the sequence of video frames; and

applying a machine learning model to each stimulus of the plurality of stimuli, the machine learning model generating an output describing a state of mind of a traffic entity displayed in video frames of the subset of video frames;

generating a data structure storing:

information describing the sequence of video frames, and information linking frame indexes of the plurality of stimuli in the sequence of video frames to corresponding outputs of the machine learning model; and

storing the data structure in association with the video data; and

presenting, via a user interface, an output of the machine learning model for a traffic entity of interest displayed in the video data based on data structures associated with the traffic entity of interest.

2. The method of claim 1 , wherein the output of the machine learning model represents a likelihood that the traffic entity of interest will cross a street intersecting with a path of a vehicle having a position and motion corresponding to the video data.

3. The method of claim 1 , wherein the output of the machine learning model includes a likelihood that the traffic entity of interest is aware of a vehicle having a position and motion corresponding to the video data.

4. The method of claim 1 , wherein the output of the machine learning model includes a likelihood that the traffic entity of interest intends to cross a street.

5. The method of claim 1 , wherein the output of the machine learning model is a probabilistic distribution.

6. The method of claim 5 , wherein the probabilistic distribution is represented in one or more graphs, charts, and tables summarizing the outputs of the machine learning model for video frames in the sequence of video frames.

7. A non-transitory computer readable medium storing instructions that when executed by one or more computer processors cause the one or more computer processors to perform steps comprising:

accessing video data, the video data displaying one or more traffic entities;

generating, from the video data, a plurality of sequences of video frames, each sequence of video frames corresponding to a portion of the video data displaying a traffic entity of the one or more traffic entities;

for a sequence of video frames of the plurality of sequences of video frames:

identifying a plurality of stimuli in the sequence of video frames, each stimulus including a subset of video frames of the sequence of video frames; and

applying a machine learning model to each stimulus of the plurality of stimuli, the machine learning model generating an output describing a state of mind of a traffic entity displayed in video frames of the subset of video frames;

generating a data structure storing:

information describing the sequence of video frames, and

information linking frame indexes of the plurality of stimuli in the sequence of video frames to corresponding outputs of the machine learning model; and

storing the data structure in association with the video data; and

presenting, via a user interface, an output of the machine learning model for a traffic entity of interest displayed in the video data based on data structures associated with the traffic entity of interest.

8. The non-transitory computer readable medium of claim 7 , wherein the output of the machine learning model represents a likelihood that the traffic entity of interest will cross a street intersecting with a path of a vehicle having a position and motion corresponding to the video data.

9. The non-transitory computer readable medium of claim 7 , wherein the output of the machine learning model includes a likelihood that the traffic entity of interest is aware of a vehicle having a position and motion corresponding to the video data.

10. The non-transitory computer readable medium of claim 7 , wherein the output of the machine learning model includes a likelihood that the traffic entity of interest intends to cross a street.

11. The non-transitory computer readable medium of claim 7 , wherein the output of the machine learning model is a probabilistic distribution.

12. The non-transitory computer readable medium of claim 11 , wherein the probabilistic distribution is represented in one or more graphs, charts, and tables summarizing the outputs of the machine learning model for video frames in the sequence of video frames.

13. A computer system comprising:

one or more computer processors; and

a non-transitory computer readable medium storing instructions that when executed by the one or more computer processors cause the one or more computer processors to perform steps comprising:

accessing video data, the video data displaying one or more traffic entities;

generating, from the video data, a plurality of sequences of video frames, each sequence of video frames corresponding to a portion of the video data displaying a traffic entity of the one or more traffic entities;

for a sequence of video frames of the plurality of sequences of video frames:

identifying a plurality of stimuli in the sequence of video frames, each stimulus including a subset of video frames of the sequence of video frames; and

applying a machine learning model to each stimulus of the plurality of stimuli, the machine learning model generating an output describing a state of mind of a traffic entity displayed in video frames of the subset of video frames;

generating a data structure storing:

information describing the sequence of video frames, and

information linking frame indexes of the plurality of stimuli in the sequence of video frames to corresponding outputs of the machine learning model; and

storing the data structure in association with the video data; and

presenting, via a user interface, an output of the machine learning model for a traffic entity of interest displayed in the video data based on data structures associated with the traffic entity of interest.

14. The computer system of claim 13 , wherein the output of the machine learning model represents a likelihood that the traffic entity of interest will cross a street intersecting with a path of a vehicle having a position and motion corresponding to the video data.

15. The computer system of claim 13 , wherein the output of the machine learning model includes a likelihood that the traffic entity of interest is aware of a vehicle having a position and motion corresponding to the video data.

16. The computer system of claim 13 , wherein the output of the machine learning model includes a likelihood that the traffic entity of interest intends to cross a street.

17. The computer system of claim 13 , wherein the output of the machine learning model is a probabilistic distribution.

18. The computer system of claim 17 , wherein the probabilistic distribution is represented in one or more graphs, charts, and tables summarizing the outputs of the machine learning model for video frames in the sequence of video frames.

Assignments (3)
PATENT SECURITY AGREEMENT Recorded Mar 25, 2025
From: PERCEPTIVE AUTOMATA LLC
To: PICCADILLY PATENT FUNDING LLC, AS SECURITY HOLDER
Reel/Frame 070614/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2025
From: PERCEPTIVE AUTOMATA, INC.
To: PERCEPTIVE AUTOMATA LLC
Reel/Frame 070267/0727 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2020
From: COPE, STEPHEN
To: PERCEPTIVE AUTOMATA, INC.
Reel/Frame 054277/0157 →
Continuity (2)
Provisional Application 62914393 · Oct 11, 2019
Related Publication 20210110203A1 · Apr 15, 2021