IP Library Granted Patent US 12,277,752
Granted Patent B2
US 12,277,752 · App. 18/364,569 · Granted Apr 15, 2025

Systems and methods for intelligent selection of data for building a machine learning model

Inventors: Jelena Frtunikj (Bavaria, DE); Daniel Alfonsetti (Auburn, ME)
Assignee: Volkswagen Group of America Investments, LLC
G06V10/774G06F18/2115G06F18/2148G06N20/00G06V10/255G06V10/772G06V10/82G06V20/56G06V20/58G06V20/588G06V20/64
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,752
App. No.
18/364,569
Granted
Apr 15, 2025
Kind
B2
Abstract

Systems and methods for selecting data for training a machine learning model using active learning are disclosed. The methods include receiving a plurality of unlabeled sensor data logs corresponding to surroundings of an autonomous vehicle and identifying one or more trends associated with a training dataset comprising a plurality of labeled data logs. The methods also include selecting a subset of the plurality of unlabeled sensor data logs that have an importance score greater than a threshold, the importance score being determined based on the one or more trends. The subset of the plurality of unlabeled sensor data logs is used for updating the machine learning model to generate an updated model.

Claims (48)

1. A method for selecting data for building a machine learning model, the method comprising:

receiving, from a plurality of sensors, a plurality of unlabeled sensor data logs corresponding to surroundings of an autonomous vehicle;

identifying one or more trends associated with a training dataset comprising a plurality of labeled data logs, the training dataset being used for training a machine learning model;

selecting a subset of the plurality of unlabeled sensor data logs that have an importance score greater than a threshold, the importance score being assigned to each of the plurality of unlabeled sensor data logs in the subset using a function determined using the one or more trends; and

using the subset of the plurality of unlabeled sensor data logs for updating the machine learning model to generate an updated model.

2. The method of claim 1 , further comprising:

identifying, using the machine learning model, within each of the plurality of unlabeled sensor data logs, a classification label;

determining a confidence score of the machine learning model trained for identifying the classification label; and

using the confidence score and the one or more trends for identifying the function, the function being configured for generating the importance score for each of the plurality of unlabeled sensor data logs.

3. The method of claim 2 , wherein the machine learning model is an object detection model.

4. The method of claim 2 , wherein the confidence score is determined based on at least one of the following associated with the machine learning model: an actual accuracy; a desired accuracy; a false positive rate; a false negative rate; a convergence; an output; a statistical fit; identification of a problem being solved; or a training status.

5. The method of claim 2 , further comprising:

receiving metadata associated with each of the plurality of labeled data logs in the training data sets; and

using the metadata for identifying the one or more trends.

6. The method of claim 5 , wherein the metadata comprises at least one of the following: a time of capture of a data log; a location of capture of a data log; information relating to environmental conditions at a location or time of capture of a data log; or information relating to one or more events a location or time of capture of a data log.

7. The method of claim 2 , wherein the one or more trends are used to identify one or more characteristics in the plurality labeled data logs of the training dataset that are underrepresented.

8. The method of claim 7 , wherein the one or more characteristics are selected from at least one of the following: a label class count; a label count by time; a label count by location; an event class count; an event class count by time; an event class count by location; an instance count by environmental conditions; or a machine leaning model uncertainty.

9. The method of claim 2 , further comprising:

receiving metadata associated with the unlabeled sensor data logs, the metadata comprising at least one of the following: a time of capture of a data log, a location of capture of a data log, information relating to environmental conditions at a location or time of capture of a data log, or information relating to one or more events a location or time of capture of a data log;

wherein the importance score is assigned to each of the plurality of unlabeled sensor data logs by assigning a score to one or more of a plurality of characteristics of each of the plurality of unlabeled sensor data logs, the plurality of characteristics comprising at least one of the following: a label class count; a label count by time; a label count by location; an event class count; an event class count by time; an event class count by location; an instance count by environmental conditions; or a machine leaning model uncertainty.

10. The method of claim 1 , further comprising discarding one or more of the plurality of unlabeled sensor data logs that have an importance score less than the threshold.

11. A system for selecting data for training a machine learning model using active learning, the system comprising:

a processor; and

a non-transitory computer-readable medium comprising one or more programming instructions that when executed by the processor, cause the processor to:

receive, from a plurality of sensors, a plurality of unlabeled sensor data logs corresponding to surroundings of an autonomous vehicle,

identify one or more trends associated with a training dataset comprising a plurality of labeled data logs, the training dataset being used for training a machine learning model,

select a subset of the plurality of unlabeled sensor data logs that have an importance score greater than a threshold, the importance score being assigned to each of the plurality of unlabeled sensor data logs in the subset using a function determined using the one or more trends, and

use the subset of the plurality of unlabeled sensor data logs for updating the machine learning model to generate an updated model.

12. The system of claim 11 , further comprising programming instructions that when executed by the processor, cause the processor to:

identify, using the machine learning model, within each of the plurality of unlabeled sensor data logs, a classification label;

determine a confidence score of the machine learning model trained for identifying the classification label; and

use the confidence score and the one or more trends for identifying the function, the function being configured for generating the importance score for each of the plurality of unlabeled sensor data logs.

13. The system of claim 12 , wherein the machine learning model is an object detection model.

14. The system of claim 12 , wherein the confidence score is determined based on at least one of the following associated with the machine learning model trained using the training dataset: an actual accuracy; a desired accuracy; a false positive rate; a false negative rate; a convergence; an output; a statistical fit; identification of a problem being solved; or a training status.

15. The system of claim 12 , further comprising programming instructions that when executed by the processor, cause the processor to:

receive metadata associated with each of the plurality of labeled data logs in the training data sets; and

use the metadata for identifying the one or more trends.

16. The system of claim 15 , wherein the metadata comprises at least one of the following: a time of capture of a data log; a location of capture of a data log; information relating to environmental conditions at a location or time of capture of a data log; or information relating to one or more events a location or time of capture of a data log.

17. The system of claim 12 , wherein the one or more trends are used to identify one or more characteristics in the plurality labeled data logs of the training dataset that are underrepresented.

18. The system of claim 17 , wherein the one or more characteristics are selected from at least one of the following: a label class count; a label count by time; a label count by location; an event class count; an event class count by time; an event class count by location; an instance count by environmental conditions; or a machine leaning model uncertainty.

19. The system of claim 12 , further comprising programming instructions that when executed by the processor, cause the processor to:

receive metadata associated with the unlabeled sensor data logs, the metadata comprising at least one of the following: a time of capture of a data log, a location of capture of a data log, information relating to environmental conditions at a location or time of capture of a data log, or information relating to one or more events a location or time of capture of a data log;

wherein the importance score is assigned to each of the plurality of unlabeled sensor data logs by assigning a score to one or more of a plurality of characteristics of each of the plurality of unlabeled sensor data logs, the plurality of characteristics comprising at least one of the following: a label class count; a label count by time; a label count by location; an event class count; an event class count by time; an event class count by location; an instance count by environmental conditions; or a machine leaning model uncertainty.

20. A computer program product comprising a non-transitory computer-readable medium that stores instructions that, when executed by a computing device, will cause the computing device to perform operations comprising:

receiving, from a plurality of sensors, a plurality of unlabeled sensor data logs corresponding to surroundings of an autonomous vehicle;

identifying one or more trends associated with a training dataset comprising a plurality of labeled data logs, the training dataset being used for training a machine learning model;

selecting a subset of the plurality of unlabeled sensor data logs that have an importance score greater than a threshold, the importance score being assigned to each of the plurality of unlabeled sensor data logs in the subset using a function determined using the one or more trends; and

using the subset of the plurality of unlabeled sensor data logs for updating the machine learning model to generate an updated model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2024
From: ARGO AI, LLC
To: VOLKSWAGEN GROUP OF AMERICA INVESTMENTS, LLC
Reel/Frame 069177/0099 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2023
From: FRTUNIKJ, JELENA; ALFONSETTI, DANIEL
To: ARGO AI, LLC
Reel/Frame 064480/0116 →
Continuity (2)
Continuation 17101633 · Nov 23, 2020
Related Publication 20230377317A1 · Nov 23, 2023
References Cited (45)
US 7793188B2 · Mukhopadhyay et al. · 2010 [cited by applicant]
US 9031317B2 · Yakubovich et al. · 2015 [cited by applicant]
US 10379842B2 · Malladi et al. · 2019 [cited by applicant]
US 10489222B2 · Sathyanarayana et al. · 2019 [cited by applicant]
US 10499283B2 · Elias et al. · 2019 [cited by applicant]
US 11657591B2 · Muehlenstaedt et al. · 2023 [cited by applicant]
US 20060046711A1 · Jung et al. · 2006 [cited by applicant]
US 20160321561A1 · Röder et al. · 2016 [cited by applicant]
US 20170011281A1 · Dijkman et al. · 2017 [cited by applicant]
US 20170083792A1 · Rodríguez-Serrano et al. · 2017 [cited by applicant]
US 20190258904A1 · Ma et al. · 2019 [cited by applicant]
US 20200005083A1 · Collins · 2020 [cited by applicant]
US 20200081426A1 · Kane et al. · 2020 [cited by applicant]
US 20200111005A1 · Ghosh et al. · 2020 [cited by applicant]
US 20200202171A1 · Hughes · 2020 [cited by examiner]
US 20200272854A1 · Caesar · 2020 [cited by applicant]
US 20200286615A1 · Hartung · 2020 [cited by examiner]
US 20200380285A1 · Price · 2020 [cited by examiner]
US 20200382361A1 · Chandrasekhar · 2020 [cited by examiner]
US 20210125423A1 · Isaac · 2021 [cited by applicant]
US 20210129845A1 · Bonk · 2021 [cited by applicant]
US 20210134379A1 · Zhao · 2021 [cited by applicant]
US 20210156963A1 · Popov et al. · 2021 [cited by applicant]
US 20210295201A1 · Kim et al. · 2021 [cited by applicant]
US 20220164602A1 · Frtunikj et al. · 2022 [cited by applicant]
US 20230040068A1 · Gandhi · 2023 [cited by examiner]
CN 111914944A · 2020 [cited by applicant]
WO 2020159568A1 · 2020 [cited by applicant]
D. Eldowa, K. Elgazzar, H. S. Hassanein, T. Sharaf and S. Shah, “Assessing the Integrity of Traffic Data through Short Term State Prediction,” 2019 IEEE Global Communications Conference (GLOBECOM), Waikoloa, HI, USA, 20… [cited by applicant]
A A Sodemann, M. P. Ross and B. J. Borghetti, “A Review of Anomaly Detection in Automated Surveillance,” in IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), vol. 42, No. 6, pp. 1257… [cited by applicant]
Castano, F.; Beruvides, G.; Villalonga, A; Haber, R.E. Self-Tuning Method for Increased Obstacle Detection Reliability Based on Internet of Things LiDAR Sensor Models. Sensors 2018, 18, 1508. https://doi.org/10.3390/s18… [cited by applicant]
S. Thornton and S. Dey, “Machine Learning Techniques for Vehicle Matching with Non-Overlapping Visual Features,” 2020 IEEE 3rd Connected and Automated Vehicles Symposium (CAVS), Victoria, BC, Canada, 2020, pp. 1-6, doi:… [cited by applicant]
Dave, H., “Active Learning Sampling Strategies”, available at https://towardsdatascience.com/active-learning-sampling-strategies-f8d8ac7037c8 accessed on Nov. 17, 2020. [cited by applicant]
Nvidia AI, “Scalable Active Learning for Autonomous Driving”, Dec. 18, 2019, available at https://medium.com/nvidia-ai/scalable-active-learning-for-autonomous-driving-a-practical-implementation-and-a-b-test-4d315ed04b5f… [cited by applicant]
Wei, X., “Deep Active Learing for 3D Object Detection for Autonomous Driving”, Degree Project in Computer Science and Engineering, Second Cycle, Stockholm, Sweden, Nov. 22, 2019, available at http://www.diva-portal.org/… [cited by applicant]
Liu, S. et al., “Edge Computing for Autonomous Driving: Opportunities and Challenges”, Proceedings of the IEEE, vol. 107, No. 8, Aug. 2019, available at http://weisong.eng.wayne.edu/_resources/pdfs/liu19-EdgeAV.pdf. [cited by applicant]
Cai, S. et al., “Data Collection in Underwater Sensor Networks Based on Mobile Edge Computing”, IEEE Access, (2019) vol. 7, pp. 65357-65367, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=8719985. [cited by applicant]
Information about Related Patents and Patent Applications, see section 6 of the accompanying Information Disclosure Statement Letter, which concerns Related Patents and Patent Applications. [cited by applicant]
Huang, H et al., “On the improvement of reinforcement active learning with the involvement of cross entropy to address one-shot learning problem” PLoS One (2019) 14(6): e0217408. available at https://journals.plos.org/p… [cited by applicant]
Yoo, D. et al., “Learning Loss for Active Learning”, arXiv:1905.03677 [cs.CV], May 9, 2019. available at: https://openaccess.thecvf.com/content_CVPR_2019/papers/Yoo_Learning_Loss_for_Active_Learning_CVPR_2019_paper.pdf. [cited by applicant]
Aussel, N. et al., “Combining federated and active learning for communication-efficient distributed failure prediction in aeronautics” 2019, hal-02446200f, available at https://hal.archives-ouvertes.fr/hal-02446200/docu… [cited by applicant]
International Search Report and Written Opinion mailed Apr. 7, 2022, issued in International Application No. PCT/US2021/072897 (10 pages). [cited by applicant]
Zhang et al., Edge-to-edge cooperative artificial intelligence in smart cities with on-demand learning offloading, 2019. [cited by applicant]
Claviere et al., Trajectory tracking control for robotic vehicles using counterexample guided training of neural networks, 2019. [cited by applicant]
Bobbili, N.P. et al., Jun. 2018, Adaptive weighting with smote for learning from imbalanced datasets: a case study for traffic offence prediction, 2018 IEEE International Conference on CIVEMSA, p. 1-6. [cited by applicant]