IP Library › Granted Patent US 12,265,891
Granted Patent B2
US 12,265,891 · App. 17/116,698 · Granted Apr 1, 2025

Methods and apparatus for automatic attribute extraction for training machine learning models

Inventors: Sakib Abdul Mondal (Bangalore, IN); Tuhin Bhattacharya (Kolkata, IN); Abhijit Mondal (Bangalore, IN)
Assignee: Walmart Apollo, LLC
G06N20/00G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,891
App. No.
17/116,698
Granted
Apr 1, 2025
Kind
B2
Abstract

This application relates to apparatus and methods for training machine learning models using supervised, or semi-supervised, learning. In some examples, a computing device obtains training data that includes labelled, and unlabeled, data for training a machine learning model. The computing device applies the machine learning model to the training data to generate output data. The machine learning model executes with a plurality of coefficients applied to a plurality of hyperparameters. The computing device further applies a loss model to the training data and the output data to generate a loss value. Based on the loss values, the computing device determines updated values for the plurality of coefficients of the machine learning model. The computing device may continue to determine updated values for the plurality of coefficients until one or more conditions are satisfied. The computing device may then store the final coefficient values in a data repository.

Claims (67)

1. A system, comprising:

a computing device configured to:

obtain training data for training a machine learning model comprising a plurality of hyperparameters and a plurality of coefficients;

generate a first training dataset based on the obtained training data;

train the machine learning model by applying the machine learning model to the first training dataset to generate first output data;

apply a loss algorithm to the first training dataset and the first output data to generate a first loss value;

determine a first updated value for each of the plurality of coefficients based on the first loss value and a corresponding previous value of each of the plurality of coefficients;

in accordance with a determination that the first updated value for each of the plurality of coefficients is outside of a threshold of a last value:

generate a second training dataset based on the obtained training data, wherein the second training dataset is different from the first training dataset;

retrain the machine learning model by applying the machine learning model with the first updated values of the plurality of coefficients to the second training dataset to generate second output data;

apply the loss algorithm to the second training dataset and the second output data to generate a second loss value;

determine a second updated value for each of the plurality of coefficients based on the second loss value and the first updated value of each of the plurality of coefficients;

generate a comparison value based on the first updated value for each of the plurality of coefficients of the machine learning model and the second updated value for each of the plurality of coefficients for each of the plurality of coefficients of the machine learning model;

determine that the comparison value is less than or equal a predetermined threshold; and

store, in a database associated with the computing device, the first updated value and the second updated value for each of the plurality of coefficients in a data repository; and

apply the first updated value and the second update value for each of the plurality of coefficient in a data repository from a source external to the computing device to generate predicted outputs.

2. The system of claim 1 , wherein the computing device is further configured to determine that updating of the plurality of coefficients is not complete based on the first updated value and the corresponding previous value for at least one of the plurality of coefficients.

3. The system of claim 2 , wherein determining that updating of the plurality of coefficients is not complete comprises comparing the first updated value and the corresponding previous value for the at least one of the plurality of coefficients, and determining that a threshold is not met.

4. The system of claim 1 , wherein the loss algorithm is based on logistics regression.

5. The system of claim 1 , wherein the computing device is configured to:

receive the machine learning model from an item recommendation system; and

transmit the first updated value for each of the plurality of coefficients to the item recommendation system.

6. The system of claim 1 , wherein the training data comprises labelled and unlabeled training data.

7. The system of claim 1 , wherein the computing device is further configured to validate the machine learning model by:

applying the machine learning model to validation data to generate second output data; and

determining that a metric is satisfied based on the second output data.

8. The system of claim 1 , wherein the computing device is configured to randomly determine an initial value for each of the plurality of coefficients.

9. The system of claim 1 , wherein the computing device is configured to train the machine learning model by applying the machine learning model to the first training dataset to generate the first output data by using manifold learning-based techniques that leverage a framework for data-dependent regularization, and a geometry of a probability distribution.

10. The system of claim 1 , wherein the predetermined threshold is preconfigurable by a user via a user interface.

11. The system of claim 10 , wherein the user interface enables the user to select labelled and/or unlabeled data to be used for the training the machine learning model.

12. The system of claim 10 , wherein the user interface enables the user to select whether to train the machine learning model using supervised or semi-supervised learning.

13. The system of claim 12 , wherein the computing device comprises a processor and a non-transitory memory storing instructions, that when executed, cause the processor to apply a manifold learning-based algorithm that uses logistics regression as a classifier for labelled training data and/or unlabeled training data.

14. A method comprising:

obtaining training data for training a machine learning model comprising a plurality of hyperparameters and a plurality of coefficients;

generating a first training dataset based on the obtained training data;

training the machine learning model by applying the machine learning model to the first training dataset to generate first output data;

applying a loss algorithm to the first training dataset and the first output data to generate a first loss value;

determining a first updated value for each of the plurality of coefficients based on the first loss value and a corresponding previous value of each of the plurality of coefficients;

in accordance with a determination that the first updated value for each of the plurality of coefficients is outside of a threshold of a last value:

generating a second training dataset based on the obtained training data, wherein the second training dataset is different from the first training dataset;

retraining the machine learning model by applying the machine learning model with the first updated values of the plurality of coefficients to the second training dataset to generate second output data;

applying the loss algorithm to the second training dataset and the second output data to generate a second loss value;

determining a second updated value for each of the plurality of coefficients based on the second loss value and the first updated value of each of the plurality of coefficients;

generating a comparison value based on the first updated value for each of the plurality of coefficients of the machine learning model and the second updated value for each of the plurality of coefficients of the machine learning model;

determining that the comparison value is less than or equal a predetermined threshold; and

storing, in a database associated with a computing device, the first updated value and the second updated value for each of the plurality of coefficients in a data repository; and

applying the first updated value and the second update value for each of the plurality of coefficient in a data repository on data received from a source external to the computing device to generate predicted outputs.

15. The method of claim 14 wherein the training data comprises labelled and unlabeled training data.

16. The method of claim 14 , further comprising validating the machine learning model by:

applying the machine learning model to validation data to generate second output data; and

determining that a metric is satisfied based on the second output data.

17. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by at least one processor, cause a device to perform operations comprising:

obtaining training data for training a machine learning model comprising a plurality of hyperparameters and a plurality of coefficients;

generating a first training dataset based on the obtained training data;

training the machine learning model by applying the machine learning model to the first training dataset to generate first output data;

applying a loss algorithm to the first training dataset and the first output data to generate a first loss value;

determining a first updated value for each of the plurality of coefficients based on the first loss value and a corresponding previous value of each of the plurality of coefficients;

in accordance with a determination that the first updated value for each of the plurality of coefficients is outside of a threshold of a last value:

generating a second training dataset based on the obtained training data, wherein the second training dataset is different from the first training dataset;

retraining the machine learning model by applying the machine learning model with the first updated values of the plurality of coefficients to the second training dataset to generate second output data;

applying the loss algorithm to the second training dataset and the second output data to generate a second loss value;

determining a second updated value for each of the plurality of coefficients based on the second loss value and the first updated value of each of the plurality of coefficients;

generating a comparison value based on the first updated value for each of the plurality of coefficients of the machine learning model and the second updated value for each of the plurality of coefficients of the machine learning model;

determining that the comparison value is less than or equal a predetermined threshold; and

storing, in a database associated with the device, the first updated value and the second updated value for each of the plurality of coefficients in a data repository; and

applying the first updated value and the second update value for each of the plurality of coefficient in a data repository on data received from a source external to the device to generate predicted outputs.

18. The non-transitory computer readable medium of claim 17 , wherein the training data comprises labelled and unlabeled training data.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE POSTAL CODE PREVIOUSLY RECORDED AT REEL: 54597 FRAME: 191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 3, 2025
From: MONDAL, SAKIB ABDUL; BHATTACHARYA, TUHIN; MONDAL, ABHIJIT
To: WALMART APOLLO, LLC
Reel/Frame 070380/0739 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2020
From: MONDAL, SAKIB ABDUL; BHATTACHARYA, TUHIN; MONDAL, ABHIJIT
To: WALMART APOLLO, LLC
Reel/Frame 054597/0191 →
Continuity (1)
Related Publication 20220180246A1 · Jun 9, 2022
References Cited (15)
US 10235449B1 · Viswanathan et al. · 2019 [cited by applicant]
US 20190066185A1 · More et al. · 2019 [cited by applicant]
US 20190385052A1 · Bauer, Jr. · 2019 [cited by examiner]
US 20200057958A1 · Moore · 2020 [cited by examiner]
US 20200372342A1 · Nair · 2020 [cited by examiner]
US 20210232980A1 · Velagapudi · 2021 [cited by examiner]
Biadsy, Naseem, Lior Rokach, and Armin Shmilovici. “Transfer learning for content-based recommender systems using tree matching.” Availability, Reliability, and Security in Information Systems and HCI: IFIP WG 8.4, 8.9,… [cited by examiner]
Hsu, Chih-Chung, and Chia-Wen Lin. “Unsupervised convolutional neural networks for large-scale image clustering.” 2017 IEEE International Conference on Image Processing (ICIP). IEEE, 2017. (Year: 2017). [cited by examiner]
Guo, Haonan, et al. “Learning automata based incremental learning method for deep neural networks.” IEEE Access 7 (2019): 41164-41171. (Year: 2019). [cited by examiner]
Srinivasan, Aishwarya “Stochastic Gradient Descent Clearly Explained” (2019), Medium, Towards Data Science. (Year: 2019). [cited by examiner]
Li, Mu, et al. “Scaling distributed machine learning with the parameter server.” 11th USENIX Symposium on operating systems design and implementation (OSDI 14). 2014. (Year: 2014). [cited by examiner]
Wang et al., “Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning”, (2019) p. 1-34. [cited by applicant]
Lin et al., “Shoestring: Graph-Based Semi-Supervised Classification with Severely Limited Labeled Data”, (2020) p. 4321-4329. [cited by applicant]
Kickview Tech Blog, “No GPU Required: Real-Time Inference with Optimized Networks in kvSonata”, (2019) p. 1-5. [cited by applicant]
Belkin et al., “Manifold Regularization: A Geometric Framework for Learning from Labeled and Unlabeled Examples”, (2006) p. 1-36. [cited by applicant]