IP Library › Granted Patent US 12,597,233
Granted Patent B2
US 12,597,233 · App. 18/183,030 · Granted Apr 7, 2026

System and method for training a machine learning model

Inventors: Amir Afrasiabi (Fircrest, WA); Sina Rafati (Cedar Park, TX); Matthew David Johnson (Bryn Mawr, PA)
Assignee: The Boeing Company
G06V10/7625G06V10/26G06V10/761G06V10/7715G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,233
App. No.
18/183,030
Granted
Apr 7, 2026
Kind
B2
Abstract

A computing system is configured to collect input data related to at least a portion of an object or an environment from the input sensor, execute a feature extractor to extract features for data elements of the input data, execute a clustering model configured to cluster the data elements of the input data into feature clusters based on similarities of the extracted features to each other, label a target clusters of the feature clusters and data elements of the target clusters with respective predetermined labels, generate a training dataset including the data elements of the target clusters, and train a machine learning model using the training dataset to predict a label for an inference time input data element at inference time. The respective predetermined labels of the target clusters correspond to prediction labels of the machine learning model.

Claims (71)

1 . A computing system for training a machine learning model configured to recognize elements of an object or an environment, the system comprising:

an input sensor;

a processor; and

a memory storing executable instructions that, in response to execution by the processor cause the processor to:

collect input data related to at least a portion of the object or the environment from the input sensor;

execute a feature extractor to extract features for a plurality of data elements of the input data;

execute a clustering model configured to cluster the plurality of data elements of the input data into a plurality of feature clusters based on similarities of the extracted features to each other;

apply a clustering filter on the plurality of feature clusters to output a plurality of target clusters from among the plurality of feature clusters;

label a plurality of target data elements of the plurality of target clusters with respective ground truth predetermined labels for each target cluster;

generate a training dataset by:

matching clusters of the machine learning model with the plurality of target clusters using an inference algorithm, running each target data element of the plurality of target clusters through the machine learning model in inference mode to generate prediction labels; and

replacing the ground truth predetermined labels of the target clusters with the prediction labels generated by the machine learning model, so as to generate the training dataset comprising the plurality of target data elements of the plurality of target clusters labeled with the prediction labels generated by the machine learning model; and

update a training of the machine learning model in a training mode using the training dataset by performing hyper-parameter fine-tuning on the machine learning model for the plurality of target clusters in the training dataset.

2 . The computing system of claim 1 , wherein the input sensor is a camera; and

the plurality of data elements are images.

3 . The computing system of claim 2 , wherein

an object detector comprising multiple convolutional neural networks is used to crop the images to generate cropped images capturing detected objects; and

the feature extractor is executed to extract features from the cropped images.

4 . The computing system of claim 3 , wherein a classification layer is omitted from the object detector.

5 . The computing system of claim 3 , wherein

the cropped images are filtered before the feature extractor extracts the features from the cropped images; and

the cropped images are filtered based on aspect ratio and/or pixel size.

6 . The computing system of claim 1 , wherein the data elements are clustered into the plurality of feature clusters by calculating pairwise similarity values between the extracted features of the data elements.

7 . The computing system of claim 6 , wherein the similarity values are calculated using a cosine similarity metrics function.

8 . The computing system of claim 1 , wherein the data elements are clustered into the plurality of feature clusters using a hierarchical clustering method.

9 . The computing system of claim 8 , wherein

the data elements are linked to each other via linkages; and

when a given set of linkages is within a predetermined height ratio range, data elements associated with the given set of linkages are assigned into a target cluster of the plurality of target clusters.

10 . A method for training a machine learning model configured to recognize elements of an object or an environment, the method comprising steps to:

collect input data related to at least a portion of the object or the environment from an input sensor;

execute a feature extractor to extract features for a plurality of data elements of the input data;

execute a clustering model configured to cluster the plurality of data elements of the input data into a plurality of feature clusters based on similarities of the extracted features to each other;

apply a clustering filter on the plurality of feature clusters to output a plurality of target clusters from among the plurality of feature clusters;

label a plurality of target data elements of the plurality of target clusters with respective ground truth predetermined labels for each target cluster;

generate a training dataset by:

matching clusters of the machine learning model with the plurality of target clusters using an inference algorithm, running each target data element of the plurality of target clusters through the machine learning model in inference mode to generate prediction labels; and

replacing the ground truth predetermined labels of the target clusters with the prediction labels generated by the machine learning model, so as to generate the training dataset comprising the plurality of target data elements of the plurality of target clusters labeled with the prediction labels generated by the machine learning model; and

update a training of the machine learning model in a training mode using the training dataset by performing hyper-parameter fine-tuning on the machine learning model for the plurality of target clusters in the training dataset.

11 . The method of claim 10 , wherein

the input sensor is a camera; and

the plurality of data elements are images.

12 . The method of claim 11 , wherein

object detection is performed to crop the images to generate cropped images capturing detected objects; and

features are extracted from the cropped images.

13 . The method of claim 12 , wherein

the input data includes 3-D objects; and

a view augmentation algorithm is used to process the 3-D objects and generate stacked view images as the plurality of data elements.

14 . The method of claim 12 , wherein

the cropped images are filtered before the feature extractor extracts the features from the cropped images; and

the cropped images are filtered based on aspect ratio and/or pixel size.

15 . The method of claim 10 , wherein the data elements are clustered into the plurality of feature clusters by calculating pairwise similarity values between the extracted features of the data elements.

16 . The method of claim 15 , wherein the similarity values are calculated using a cosine similarity metrics function.

17 . The method of claim 10 , wherein the data elements are clustered into the plurality of feature clusters using a hierarchical clustering method.

18 . A computing system for training an object detection machine learning model configured to recognize objects in images of an object or an environment, the system comprising:

a camera;

a processor; and

a memory storing executable instructions that, in response to execution by the processor cause the processor to:

collect image data related to at least a portion of the object or the environment from the camera;

perform object detection to crop images of the image data to generate a plurality of cropped images capturing detected objects;

execute a feature extractor to extract features for the plurality of cropped images;

execute a clustering model configured to cluster the plurality of cropped images of the image data into a plurality of feature clusters based on similarities of the extracted features to each other;

apply a clustering filter on the plurality of feature clusters to output a plurality of target clusters from among the plurality of feature clusters;

label a plurality of cropped target images of the plurality of target clusters with respective ground truth predetermined object labels for each target cluster;

generate a training dataset by:

matching clusters of the object detection machine learning model with the plurality of target clusters using an inference algorithm, running each cropped target image of the plurality of target clusters through the object detection machine learning model in inference mode to generate prediction labels; and

replacing the ground truth predetermined object labels of the target clusters with the prediction labels generated by the object detection machine learning model, so as to generate the training dataset comprising the plurality of cropped target images of the plurality of target clusters labeled with the prediction labels generated by the object detection machine learning model; and

update a training of the object detection machine learning model in a training mode using the training dataset by performing hyper-parameter fine-tuning on the object detection machine learning model for the plurality of target clusters in the training dataset.

19 . The computing system of claim 18 , wherein

an object detector comprising multiple convolutional neural networks is used to crop the images to generate cropped images capturing detected objects; and

the feature extractor is executed to extract features from the cropped images.

20 . The computing system of claim 19 , wherein a classification layer is omitted from the object detector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2023
From: AFRASIABI, AMIR; RAFATI, SINA; JOHNSON, MATTHEW DAVID
To: THE BOEING COMPANY
Reel/Frame 062966/0419 →
Continuity (2)
Provisional Application 63481151 · Jan 23, 2023
Related Publication 20240249500A1 · Jul 25, 2024
References Cited (13)
US 20200134510A1 · Basel · 2020 [cited by examiner]
US 20210089841A1 · Mithun · 2021 [cited by examiner]
US 20210097345A1 · Goldstein · 2021 [cited by examiner]
US 20210174196A1 · Desmond · 2021 [cited by examiner]
US 20210182618A1 · Hoffmann · 2021 [cited by examiner]
US 20210264300A1 · Staudinger · 2021 [cited by examiner]
US 20220076164A1 · Conort · 2022 [cited by examiner]
US 20220301173A1 · Cheng · 2022 [cited by examiner]
US 20230095533A1 · Wong · 2023 [cited by examiner]
US 20240232699A1 · Shanker · 2024 [cited by examiner]
Hossain et al., Machine Learning Model Optimization with HyperParameter Tuning Approach, 2021, Global J. of Comp. Scie. and Tech. D. Neural & Artificial Intelligence , 21(2): 1-8. (Year: 2021). [cited by examiner]
Hsu, K. et al., “Unsupervised Learning via Meta-Learning,” Proceedings of “Workshop on Meta-Learning at NeurIPS”, Dec. 8, 2018, Montréal, Canada, 5 pages. [cited by applicant]
European Patent Office, Extended European Search Report Issued in Application No. 23217684.2, May 24, 2024, Germany, 6 pages. [cited by applicant]