IP Library Granted Patent US 12,731,407
Granted Patent B2
US 12,731,407 · App. 18/348,590 · Granted Sep 8, 2026

Method and system for in-vehicle self-supervised training of perception functions for an automated driving system

Inventors: Magnus Gyllenhammar (Pixbo, SE); Adam Tonderski (Västra Frölunda, SE)
Assignee: ZENSEACT AB
G06V20/56G06V10/7792
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,407
App. No.
18/348,590
Granted
Sep 8, 2026
Kind
B2
Abstract

A computer-implemented method for updating a perception function of a vehicle having an Automated Driving System (ADS) is disclosed. The ADS has a machine-learning algorithm for: generating an attention map or a feature map based on one or more ingested images and for providing one or more in-vehicle perception functions based on one or more ingested images. The method comprises obtaining one or more images of a scene in a surrounding environment of the vehicle, and updating one or more model parameters of the self-supervised machine-learning algorithm in accordance with a self-supervised machine learning process based on the obtained one or more images. The method further comprises generating a first output comprising an attention map or a feature map by processing the obtained one or more images by using the self-supervised machine-learning algorithm, and generating a supervisory signal for a supervised learning process based on the first output.

Claims (38)

1 . A computer-implemented method for updating a perception function of a vehicle having an Automated Driving System (ADS) having a self-supervised machine-learning algorithm configured to generate an output based on one or more ingested images and a machine-learning algorithm for an in-vehicle perception module trained to provide one or more in-vehicle perception functions based on one or more ingested images, the method comprising:

obtaining one or more images of a scene in a surrounding environment of the vehicle;

updating one or more model parameters of the self-supervised machine-learning algorithm in accordance with a self-supervised machine learning process based on the obtained one or more images;

generating a first output by processing the obtained one or more images by using the self-supervised machine-learning algorithm;

generating a supervisory signal for a supervised learning process based on the first output; and

updating one or more model parameters of the machine-learning algorithm for the in-vehicle perception module based on the obtained one or more images and the generated supervisory signal in accordance with the supervised learning process.

2 . The method according to claim 1 , wherein the generated supervisory signal comprises the generated first output, and wherein the obtained one or more images and the generated supervisory signal forms training data for the machine-learning algorithm for the in-vehicle perception module.

3 . The method according to claim 1 , wherein the generating of the supervisory signal comprises processing the first output by using a secondary machine learning algorithm trained to generate a second output based on the generated first output;

wherein the supervisory signal comprises the second output;

wherein the obtained one or more images and the supervisory signal forms training data for the machine-learning algorithm for the in-vehicle perception module.

4 . The method according to claim 3 , wherein the second output comprises at least one of object classification, depth estimation, bounding box, segmentation mask, and object trajectory.

5 . The method according to claim 1 , further comprising:

detecting anomalous image data by using a machine-learning classification system trained to distinguish new experiences from experiences known to the self-supervised machine-learning algorithm in the obtained one or more images and to output an anomaly value;

adding a weight to the supervisory signal based on the anomaly value.

6 . The method according to claim 5 , wherein the machine-learning classification system comprises an autoencoder trained on the same dataset as the self-supervised machine-learning algorithm, and wherein the anomaly value is a reconstruction error.

7 . The method according to claim 1 , further comprising:

transmitting the updated one or more model parameters of the self-supervised machine-learning algorithm and the updated one or more model parameters of the machine-learning algorithm for the in-vehicle perception module to a remote entity;

receiving a set of globally updated one or more model parameters of the self-supervised machine-learning algorithm from the remote entity, wherein the set of globally updated one or more model parameters of the self-supervised machine-learning algorithm are based on information obtained from a plurality of vehicles comprising a corresponding self-supervised machine-learning algorithm;

receiving a set of globally updated one or more model parameters of the machine-learning algorithm for the in-vehicle perception module from the remote entity, wherein the set of globally updated one or more model parameters of the machine-learning algorithm for the in-vehicle perception module are based on information obtained from a plurality of vehicles comprising a corresponding machine-learning algorithm for the in-vehicle perception module;

updating the self-supervised machine-learning algorithm based on the received set of globally updated one or more model parameters of the self-supervised machine-learning algorithm; and

updating the machine-learning algorithm for the in-vehicle perception module based on the received set of globally updated one or more model parameters of the machine-learning algorithm for the in-vehicle perception module.

8 . The method according to claim 1 , wherein the self-supervised machine-learning algorithm is a Masked Autoencoder (MAE).

9 . The method according to claim 1 , wherein the one or more in-vehicle perception functions comprises at least one of:

a semantic segmentation function, an instance segmentation function, an object classification function, an object detection function, a free-space estimation function, and a tracking function, an object trajectory prediction function.

10 . A non-transitory computer-readable storage medium comprising instructions which, when executed by a computing device, causes the computer to carry out the method according to claim 1 .

11 . A system for updating a perception function of a vehicle having an Automated Driving System (ADS) having a self-supervised machine-learning algorithm configured to generate an output based on one or more ingested images and a machine-learning algorithm for an in-vehicle perception module trained to provide one or more perception functions based on one or more ingested images, the system comprising control circuitry configured to:

obtain one or more images of a scene in a surrounding environment of the vehicle;

update one or more model parameters of the self-supervised machine-learning algorithm in accordance with a self-supervised machine learning process based on the obtained one or more images;

generate a first output by processing the obtained one or more images by using the self-supervised machine-learning algorithm;

generate a supervisory signal for a supervised learning process based on the first output; and

update one or more model parameters of the machine-learning algorithm for the perception module based on the obtained one or more images and the generated supervisory signal in accordance with the supervised learning process.

12 . The system according to claim 11 , wherein the control circuitry is further configured to:

detect anomalous image data by using a machine-learning classification system trained to distinguish new experiences from experiences known to the self-supervised machine-learning algorithm in the obtained one or more images and to output an anomaly value; and

adding a weight to the supervisory signal based on the anomaly value.

13 . The system according to claim 12 , wherein the machine-learning classification system comprises an autoencoder trained on the same dataset as the self-supervised machine-learning algorithm, and wherein the anomaly value is a reconstruction error.

14 . A vehicle comprising:

one or more sensors configured to capture images of a scene in a surrounding environment of the vehicle; and

a system according to claim 11 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2023
From: GYLLENHAMMAR, MAGNUS; TONDERSKI, ADAM
To: ZENSEACT AB
Reel/Frame 064609/0255 →
Priority Claims (1)
EP 22184043 · Jul 11, 2022 · regional
Continuity (1)
Related Publication 20240010227A1 · Jan 11, 2024
References Cited (32)
US 11656927B1 · Rosenkranz · 2023 [cited by examiner]
US 11684241B2 · Crosby · 2023 [cited by examiner]
US 12236628B1 · Karuppusamy · 2025 [cited by examiner]
US 20190042937A1 · Sheller et al. · 2019 [cited by applicant]
US 20190197236A1 · Niculescu-Mizil · 2019 [cited by examiner]
US 20200065711A1 · Clément · 2020 [cited by examiner]
US 20200210696A1 · Hou · 2020 [cited by examiner]
US 20200249685A1 · Elluswamy · 2020 [cited by examiner]
US 20210112441A1 · Sabella et al. · 2021 [cited by applicant]
US 20210304002A1 · George · 2021 [cited by examiner]
US 20210325901A1 · Gyllenhammar et al. · 2021 [cited by applicant]
US 20210350150A1 · An et al. · 2021 [cited by applicant]
US 20220180193A1 · Caine · 2022 [cited by examiner]
US 20230206067A1 · Wang · 2023 [cited by examiner]
US 20230342598A1 · Tatsubori · 2023 [cited by examiner]
US 20240010227A1 · Gyllenhammar · 2024 [cited by examiner]
US 20240134085A1 · Lathrop · 2024 [cited by examiner]
CN 112700639A · 2021 [cited by applicant]
WO 2021055088A1 · 2021 [cited by applicant]
Tan et al, MGAE: Masked Autoencoders for Self-Supervised Learning on Graphs, Jan. 7, 2022 (Year: 2022). [cited by examiner]
Hess, G., et al.; “Masked Autoencoders for Self-Supervised Learning on Automotive Point Clouds”; Cornell University Library; Jul. 1, 2022; 13 pages. [cited by applicant]
Shu, L. et al.; “Kernel-Based Transductive Learning with Nearest Neighbors”, Apr. 2, 2009; SAT 18th International Conference; Sep. 24-27, 2015; Austin, Texas; 12 pages. [cited by applicant]
Diao, E., et al.; “SemiFL: Communication Efficient Semi-Supervised Federated Learning with Unlabeled Clients”; Cornell University Library; Jan. 29, 202; 24 pages. [cited by applicant]
Saeed, A., et al.; “Federated Self-Supervised Learning of Multi-Sensor Representations for Embedded Intelligence”; IEEE Internet of Things Journal; Jul. 25, 2020; 11 pages. [cited by applicant]
Extended European Search Report mailed Jan. 4, 2023 for European Application No. 22184043.2, 10 pages. [cited by applicant]
Bao, H., et al.; “Beit: Bert Pre-Training of Image Transformers”; Cornell University Library; Sep. 3, 2022; 18 pages. [cited by applicant]
He, K., et al.; “Masked Autoencoders Are Scalable Vision Learners”; Cornell University Library; Dec. 19, 2021; 14 pages. [cited by applicant]
Xie, Z., et al.; “SimMIM: A Simple Framework for Masked Image Modeling”; Cornell University Library; Apr. 17, 2022; 13 pages. [cited by applicant]
Zhao, J. et al.; “iBOT: Image BERT Pre-Training with Online Tokenizer”; Cornell University Library; Jan. 27, 2022; ICLR 2022; Cornell University Library; 29 pages. [cited by applicant]
Office Action, European Patent Application No. 22184043.2, mailed May 26, 2026, 6 pages. [cited by applicant]
Akiyoshi Kurobe et al., “Audio-Visual Self-Supervised Terrain Type Recognition for Ground Mobile Platforms”, Department of Science and Technology, Keio University, Yokohama, Japan, XP11839710A, Feb. 16, 2021, 10 pages. [cited by applicant]
Pedro Savarese et al., “Information-Theoretic Segmentation by Inpainting Error Maximization”, XP81837388A, Dec 14, 2020, 10 pages. [cited by applicant]