IP Library › Granted Patent US 11,551,036
Granted Patent B2
US 11,551,036 · App. 16/112,637 · Granted Jan 10, 2023

Methods and apparatuses for building data identification models

Inventors: Xiaoyan Jiang (Hangzhou, CN); Xu Yang (Hangzhou, CN); Bin Dai (Hangzhou, CN); Wei Chu (Hangzhou, CN)
Assignee: Alibaba Group Holding Limited
G06K9/6257G06K9/6262G06K9/6277G06N3/08G06N20/00G06Q30/018G06Q30/0201G06Q30/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,036
App. No.
16/112,637
Granted
Jan 10, 2023
Kind
B2
Abstract

The present disclosure provides methods and an apparatuses for building a data identification model. One exemplary method for building a data identification model includes: performing logistic regression training using training samples to obtain a first model, the training samples comprising positive and negative samples; sampling the training samples proportionally to obtain a first training sample set; identifying the positive samples using the first model, and selecting a second training sample set from positive samples that have identification results after being identified using the first model; and performing Deep Neural Networks (DNN) training using the first training sample set and the second training sample set to obtain a final data identification model. The methods and the apparatuses of the present disclosure improve the stability of data identification models.

Claims (49)

1. A method for building a data identification model, comprising:

performing logistic regression training using training samples to obtain a first model, the training samples comprising positive samples and negative samples;

sampling the training samples proportionally to obtain a first training sample set;

identifying the positive samples using the first model, and selecting a second training sample set from positive samples that have identification results after being identified using the first model; and

performing Deep Neural Networks (DNN) training using the first training sample set and the second training sample set to obtain a final data identification model.

2. The method for building a data identification model of claim 1 , wherein prior to sampling the training samples proportionally or performing logistic regression training, the method further comprises:

performing feature engineering preprocessing on the training samples.

3. The method for building a data identification model of claim 2 , wherein prior to performing logistic regression training using the training samples to obtain a first model, the method further comprises:

performing feature screening on the training samples to remove features having information values that satisfy a predetermined threshold condition.

4. The method for building a data identification model of claim 1 , wherein prior to selecting the second training sample set from positive samples that have identification results after being identified using the first model, the method further comprises:

performing DNN training using the first training sample set to obtain a second model.

5. The method for building a data identification model of claim 4 , wherein selecting the second training sample set from positive samples that have identification results after being identified using the first model further comprises:

evaluating the first model and obtaining an ROC curve corresponding to the first model;

evaluating the second model and obtaining an ROC curve corresponding to the second model; and

selecting, based on a threshold probability corresponding to an intersection point of the ROC curves of the first model and the second model, samples having probabilities satisfying a threshold probability condition from positive samples that have identification results after being identified using the first model to serve as the second training sample set.

6. An apparatus for building a data identification model, comprising:

a memory storing a set of instructions; and

a processor configured to execute the set of instructions to cause the apparatus to perform:

logistic regression training using training samples to obtain a first model, the training samples comprising positive samples and negative samples;

sampling the training samples proportionally to obtain a first training sample set;

identifying the positive samples using the first model, and selecting a second training sample set from positive samples that have identification results after being identified using the first model; and

Deep Neural Networks (DNN) training using the first training sample set and the second training sample set to obtain a final data identification model.

7. The apparatus for building a data identification model of claim 6 , wherein the processor is configured to execute the set of instructions to cause the apparatus to further perform:

prior to sampling the training samples proportionally or performing logistic regression training, feature engineering preprocessing on the training samples.

8. The apparatus for building a data identification model of claim 7 , wherein the processor is configured to execute the set of instructions to cause the apparatus to further perform:

prior to performing logistic regression training using the training samples to obtain a first model, feature screening on the training samples to remove features having information values that satisfy a predetermined threshold condition.

9. The apparatus for building a data identification model of claim 6 , wherein the processor is configured to execute the set of instructions to cause the apparatus to further perform:

DNN training using the first training sample set to obtain a second model.

10. The apparatus for building a data identification model of claim 9 , wherein selecting the second training sample set from positive samples that have identification results after being identified using the first model further comprises:

evaluating the first model and obtaining an ROC curve corresponding to the first model;

evaluating the second model and obtaining an ROC curve corresponding to the second model; and

selecting, based on a threshold probability corresponding to an intersection point of the ROC curves of the first model and the second model, samples having probabilities satisfying a threshold probability condition from positive samples that have identification results after being identified using the first model to serve as the second training sample set.

11. A non-transitory computer readable medium that stores a set of instructions that is executable by at least one processor of a computer to cause the computer to perform a method for building a data identification model, the method comprising:

performing logistic regression training using training samples to obtain a first model, the training samples comprising positive samples and negative samples;

sampling the training samples proportionally to obtain a first training sample set;

identifying the positive samples using the first model, and selecting a second training sample set from positive samples that have identification results after being identified using the first model; and

performing Deep Neural Networks (DNN) training using the first training sample set and the second training sample set to obtain a final data identification model.

12. The non-transitory computer readable medium of claim 11 , wherein the processor is configured to execute the set of instructions to cause the computer to further perform:

prior to sampling the training samples proportionally or performing logistic regression training, feature engineering preprocessing on the training samples.

13. The non-transitory computer readable medium of claim 12 , wherein the processor is configured to execute the set of instructions to cause the computer to further perform:

prior to performing logistic regression training using the training samples to obtain a first model, feature screening on the training samples to remove features having information values that satisfy a predetermined threshold condition.

14. The non-transitory computer readable medium of claim 13 , wherein the predetermined threshold condition is satisfied when the features have information values less than a predetermined threshold.

15. The non-transitory computer readable medium of claim 11 , wherein the processor is configured to execute the set of instructions to cause the computer to further perform:

DNN training using the first training sample set to obtain a second model.

16. The non-transitory computer readable medium of claim 15 , selecting the second training sample set from positive samples that have identification results after being identified using the first model further comprises:

evaluating the first model and obtain an ROC curve corresponding to the first model;

evaluating the second model and obtain an ROC curve corresponding to the second model; and

selecting, based on a threshold probability corresponding to an intersection point of the ROC curves of the first model and the second model, samples having probabilities satisfying a threshold probability condition from positive samples that have identification results after being identified using the first model to serve as the second training sample set.

17. The non-transitory computer readable medium of claim 16 , wherein the threshold probability condition is satisfied when the samples have probabilities less than a threshold probability.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: JIANG, XIAOYAN; YANG, XU; DAI, BIN; CHU, WEI
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 054538/0811 →
Priority Claims (1)
CN 201610110817.3 · Feb 26, 2016 · national
Continuity (2)
Continuation PCTCN2017073444 · Feb 14, 2017
Related Publication 20180365522A1 · Dec 20, 2018