IP Library Granted Patent US 12,670,700
Granted Patent B2
US 12,670,700 · App. 17/660,973 · Granted Jun 30, 2026

Active data collection, sampling, and generation for use in training machine learning models for automotive or other applications

Inventors: Jongmoo Choi (Gardena, CA); Kilsoo Kim (Hermosa Beach, CA); Qi Peng (Torrance, CA); Lei Cao (Torrance, CA); Mayukh Sattiraju (Redondo Beach, CA); Shanmukha M. Bhumireddy (Torrance, CA); Emilio Aron Moyers Barrera (Torrance, CA); Phillip Vu (Calabasas, CA)
Assignee: WHS ENERGY SOLUTIONS, LLC
G06V10/7753G06V20/56G06V20/59
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,700
App. No.
17/660,973
Granted
Jun 30, 2026
Kind
B2
Abstract

A method includes identifying one or more edge cases associated with at least one trained machine learning model, where the at least one trained machine learning model is configured to perform at least one function related to one or more vehicles. The method also includes obtaining raw data associated with the one or more edge cases from at least one of the one or more vehicles and selecting a subset of the raw data. The method further includes generating synthetic data associated with the one or more edge cases. In addition, the method includes at least one of: retraining the at least one trained machine learning model and training at least one new machine learning model using the selected subset of raw data and the synthetic data.

Claims (68)

1 . A method comprising:

identifying one or more edge cases associated with at least one trained machine learning model, the at least one trained machine learning model configured to perform at least one function related to one or more vehicles;

obtaining non-edge case data having features that are visually similar to the identified one or more edge cases;

obtaining raw data associated with the one or more edge cases and the non-edge case data from at least one of the one or more vehicles;

selecting a subset of the raw data, including identifying a target distribution of data related to both the one or more edge cases and the non-edge case data;

generating synthetic data associated with both the one or more edge cases and the non-edge case data; and

performing at least one of:

retraining the at least one trained machine learning model; or

training at least one new machine learning model using the selected subset of the raw data and the synthetic data.

2 . The method of claim 1 , wherein:

the at least one trained machine learning model represents at least one model previously trained using a training dataset that comprises multiple first images;

the raw data comprises multiple second images captured using one or more cameras of the at least one of the one or more vehicles; and

the synthetic data comprises multiple third images representing at least one of artificial images and modified versions of captured images associated with the one or more edge cases.

3 . The method of claim 2 , wherein:

the at least one trained machine learning model is configured to perform at least one camera perception function in the one or more vehicles; and

the second images comprise images of one or more drivers.

4 . The method of claim 2 , further comprising:

generating one or more datasets containing at least some of the first, second, and third images,

wherein the retraining of the at least one trained machine learning model or the training of the at least one new machine learning model is based on the one or more datasets.

5 . The method of claim 1 , wherein the one or more edge cases represent one or more instances in which the at least one trained machine learning model provides incorrect results or provides correct results having a confidence below a threshold.

6 . The method of claim 1 , wherein obtaining the raw data associated with the one or more edge cases comprises obtaining the raw data from multiple vehicles, the raw data satisfying at least one target distribution of vehicle-related or driver-related data.

7 . The method of claim 1 , wherein the at least one trained machine learning model is trained to perform at least one of object detection, object classification, semantic segmentation, interesting point detection, pose estimation, and driver monitoring for the one or more vehicles.

8 . An apparatus comprising:

at least one processing device configured to:

identify one or more edge cases associated with at least one trained machine learning model, the at least one trained machine learning model configured to perform at least one function related to one or more vehicles;

obtain non-edge case data having features that are visually similar to the identified one or more edge cases;

obtain raw data associated with the one or more edge cases and the non-edge case data from at least one of the one or more vehicles;

select a subset of the raw data, wherein the at least one processing device is configured to identify a target distribution of data related to both the one or more edge cases and the non-edge case data;

generate synthetic data associated with both the one or more edge cases and the non-edge case data; and

perform at least one of:

retrain the at least one trained machine learning model; or

train at least one new machine learning model using the selected subset of the raw data and the synthetic data.

9 . The apparatus of claim 8 , wherein:

the at least one trained machine learning model represents at least one model previously trained using a training dataset that comprises multiple first images;

the raw data comprises multiple second images captured using one or more cameras of the at least one of the one or more vehicles; and

the synthetic data comprises multiple third images representing at least one of artificial images and modified versions of captured images associated with the one or more edge cases.

10 . The apparatus of claim 9 , wherein:

the at least one trained machine learning model is configured to perform at least one camera perception function in the one or more vehicles; and

the second images comprise images of one or more drivers.

11 . The apparatus of claim 9 , wherein:

the at least one processing device is further configured to generate one or more datasets containing at least some of the first, second, and third images; and

the at least one processing device is configured to retrain the at least one trained machine learning model or train the at least one new machine learning model based on the one or more datasets.

12 . The apparatus of claim 8 , wherein the one or more edge cases represent one or more instances in which the at least one trained machine learning model provides incorrect results or provides correct results having a confidence below a threshold.

13 . The apparatus of claim 8 , wherein the at least one processing device is configured to obtain the raw data associated with the one or more edge cases from multiple vehicles, the raw data satisfying at least one target distribution of vehicle-related or driver-related data.

14 . The apparatus of claim 8 , wherein the at least one trained machine learning model is trained to perform at least one of object detection, object classification, semantic segmentation, interesting point detection, pose estimation, and driver monitoring for the one or more vehicles.

15 . A non-transitory machine-readable medium containing instructions that when executed cause at least one processor to:

identify one or more edge cases associated with at least one trained machine learning model, the at least one trained machine learning model configured to perform at least one function related to one or more vehicles;

obtain non-edge case data having features that are visually similar to the identified one or more edge cases;

obtain raw data associated with the one or more edge cases and the non-edge case data from at least one of the one or more vehicles;

select a subset of the raw data, including performing an identification of a target distribution of data related to both the one or more edge cases and the non-edge case data;

generate synthetic data associated with both the one or more edge cases and the non-edge case data; and

perform at least one of:

retrain the at least one trained machine learning model; or

train at least one new machine learning model using the selected subset of the raw data and the synthetic data.

16 . The non-transitory machine-readable medium of claim 15 , wherein:

the at least one trained machine learning model represents at least one model previously trained using a training dataset that comprises multiple first images;

the raw data comprises multiple second images captured using one or more cameras of the at least one of the one or more vehicles; and

the synthetic data comprises multiple third images representing at least one of artificial images and modified versions of captured images associated with the one or more edge cases.

17 . The non-transitory machine-readable medium of claim 16 , wherein:

the at least one trained machine learning model is configured to perform at least one camera perception function in the one or more vehicles; and

the second images comprise images of one or more drivers.

18 . The non-transitory machine-readable medium of claim 15 , wherein the one or more edge cases represent one or more instances in which the at least one trained machine learning model provides incorrect results or provides correct results having a confidence below a threshold.

19 . The non-transitory machine-readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to obtain the raw data associated with the one or more edge cases comprises:

instructions that when executed cause the at least one processor to obtain the raw data associated with the one or more edge cases from multiple vehicles, the raw data satisfying at least one target distribution of vehicle-related or driver-related data.

20 . The non-transitory machine-readable medium of claim 15 , wherein the at least one trained machine learning model is trained to perform at least one of object detection, object classification, semantic segmentation, interesting point detection, pose estimation, and driver monitoring for the one or more vehicles.

21 . The non-transitory machine-readable medium of claim 15 , further containing instructions that when executed cause the at least one processor to continue an active learning process that includes identifying the one or more edge cases, obtaining the raw data, selecting the subset of the raw data, generating the synthetic data, and at least one of retraining the at least one trained machine learning model and training the at least one new machine learning model until a decision is made to stop the active learning process.

22 . The non-transitory machine-readable medium of claim 21 , wherein the decision to stop the active learning process is based on at least one of:

a determination that an uncertainty value associated with a prediction by the at least one trained machine learning model based on unlabeled training data is below a threshold value; and a determination that labels generated by the at least one trained machine learning model based on unlabeled training data during multiple consecutive training iterations match one another.