IP Library Granted Patent US 12675733
Granted Patent B2
US 12675733 · App. 17/733,476 · Granted Jul 7, 2026

Systems and methods for generating uniform frames having sensor and agent data

Inventors: Chao Fang (Sunnyvale, CA); Charles Christopher Ochoa (San Francisco, CA); Kuan-Hui Lee (San Jose, CA); Kun-Hsin Chen (San Francisco, CA); Visak Kumar (San Francisco, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
G06N20/00G07C5/0841
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675733
App. No.
17/733,476
Granted
Jul 7, 2026
Kind
B2
Abstract

System, methods, and other embodiments described herein relate to a manner of generating and relating frames that improves the retrieval of sensor and agent data for processing by different vehicle tasks. In one embodiment, a method includes acquiring sensor data by a vehicle. The method also includes generating a frame including the sensor data and agent perceptions determined from the sensor data at a timestamp, the agent perceptions including multi-dimensional data that describes features for surrounding vehicles of the vehicle. The method also includes relating the frame to other frames of the vehicle by track, the other frames having processed data from various times and the track having a predetermined window of scene information associated with an agent. The method also includes training a learning model using the agent perceptions accessed from the track.

Claims (37)

1 . A control system for improving information retrieval and access, comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the processor to:

acquire sensor data by a vehicle;

generate a data frame including the sensor data and agent perceptions determined from the sensor data at a timestamp, the agent perceptions including multi-dimensional data that associates features of a scene retrieved from surrounding vehicles of the vehicle in a data slice that is associated with different agents that have a same dimensional type and context, and the different agents are vehicles;

relate the data frame to other data frames of the vehicle by a track, the other data frames include having processed data from various times and the track having a predetermined window of the scene associated with an agent and the data frame, and the data slice spans the track and another track that include both two-dimensional (2D) agent data and three-dimensional (3D) agent data associated with a same agent that is moving; and

train a learning model using the agent perceptions accessed from the track.

2 . The control system of claim 1 , wherein the instructions to train the learning model further include instructions to access, by an agent encoder, the agent perceptions from the data slice, and wherein the data slice includes the data frame and another data frame at the timestamp and the agent encoder extracts features for forming the multi-dimensional data.

3 . The control system of claim 2 , wherein the data slice includes the 3D agent data and the 2D agent data at the timestamp.

4 . The control system of claim 1 , wherein the instructions to generate the data frame further include instructions to form a map datum having traffic control data for the data frame at the timestamp for forming an agent schema where data access is system agnostic.

5 . The control system of claim 4 , wherein the instructions to train the learning model further include instructions to input the map datum to a map encoder, and wherein the learning model includes an agent encoder that processes the agent perceptions, and the map datum is time-varying and describes states of bulb groups associated with a traffic light in the scene.

6 . The control system of claim 4 , wherein the sensor data forms a depth datum with 3D data and an image datum with 2D data associated with characteristics of objects in the scene surrounding the vehicle.

7 . The control system of claim 6 , further including instructions to access the depth datum and the image datum non-sequentially and randomly associated with the data slice according to a request from a processing task for the vehicle.

8 . The control system of claim 1 , further including instructions to:

synchronize agent information of the data frame and the other data frames in the track according to a predetermined time span; and

retrieve, by the learning model, agent features from the data frame and the other data frames randomly across multiple tracks, wherein the agent features are associated with a 3D type and a 2D type.

9 . A non-transitory computer-readable medium comprising:

instructions that when executed by a processor cause the processor to:

acquire sensor data by a vehicle;

generate a data frame including the sensor data and agent perceptions determined from the sensor data at a timestamp, the agent perceptions including multi-dimensional data that associates features of a scene retrieved from surrounding vehicles of the vehicle in a data slice that is associated with different agents that have a same dimensional type and context, and the different agents are vehicles;

relate the data frame to other data frames of the vehicle by a track, the other data frames include having processed data from various times and the track having a predetermined window of the scene associated with an agent and the data frame, and the data slice spans the track and another track that include both two-dimensional (2D) agent data and three-dimensional (3D) agent data associated with a same agent that is moving; and

train a learning model using the agent perceptions accessed from the track.

10 . The non-transitory computer-readable medium of claim 9 , wherein the instructions to train the learning model further include instructions to access, by an agent encoder, the agent perceptions from the data slice, and wherein the data slice includes the data frame and another data frame at the timestamp and the agent encoder extracts features for forming the multi-dimensional data.

11 . A method comprising:

acquiring sensor data by a vehicle;

generating a data frame including the sensor data and agent perceptions determined from the sensor data at a timestamp, the agent perceptions including multi-dimensional data that associates features of a scene retrieved from surrounding vehicles of the vehicle in a data slice that is associated with different agents that have a same dimensional type and context, and the different agents are vehicles;

relating the data frame to other data frames of the vehicle by a track, the other data frames include having processed data from various times and the track having a predetermined window of the scene associated with an agent and the data frame, and the data slice spans the track and another track that include both two-dimensional (2D) agent data and three-dimensional (3D) agent data associated with a same agent that is moving; and

training a learning model using the agent perceptions accessed from the track.

12 . The method of claim 11 , wherein training the learning model further comprises accessing, by an agent encoder, the agent perceptions from the data slice, and wherein the data slice includes the data frame and another data frame at the timestamp and the agent encoder extracts features for forming the multi-dimensional data.

13 . The method of claim 12 , wherein the data slice includes the 3D agent data and the 2D agent data at the timestamp.

14 . The method of claim 11 , wherein generating the data frame further comprises forming a map datum having traffic control data for the data frame at the timestamp for forming an agent schema where data access is system agnostic.

15 . The method of claim 14 , wherein training the learning model further comprises inputting the map datum to a map encoder, and wherein the learning model includes an agent encoder that processes the agent perceptions, and the map datum is time-varying and describes states of bulb groups associated with a traffic light in the scene.

16 . The method of claim 14 , wherein the sensor data forms a depth datum with 3D data and an image datum with 2D data associated with characteristics of objects in the scene surrounding the vehicle.

17 . The method of claim 16 , further comprising accessing the depth datum and the image datum non-sequentially and randomly associated with the data slice according to a request from a processing task for the vehicle.

18 . The method of claim 11 , further comprising:

synchronizing agent information of the data frame and the other data frames in the track according to a predetermined time span; and

retrieving, by the learning model, agent features from the data frame and the other data frames randomly across multiple tracks, wherein the agent features are associated with a 3D type and a 2D type.