IP Library Granted Patent US 9,619,758
Granted Patent B2
US 9,619,758 · App. 14/586,902 · Granted Apr 11, 2017

Method and apparatus for labeling training samples

Inventors: Huige Cheng (Beijing, CN); Yaozong Mao (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD
G06N99/005G06N7/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,619,758
App. No.
14/586,902
Granted
Apr 11, 2017
Kind
B2
Abstract

Provided in the present invention are a method and apparatus for labeling training samples. In the embodiments of the present invention, two mutually independent classifiers, i.e. a first classifier and a second classifier, are used to perform collaborative forecasting on M unlabeled first training samples to obtain some of the labeled first training samples, without the need for the participation of operators; the operation is simple and the accuracy is high, thereby improving the efficiency and reliability of labeling training samples.

Claims (45)

1. A method for labeling training samples, comprising:

inputting M unlabeled first training samples into a first classifier to obtain a first forecasting result of each first training sample in the M first training samples, M being an integer greater than or equal to 1;

selecting N first training samples as second training samples from the M first training samples based upon the first forecasting result of each first training sample, N being an integer greater than or equal to 1 and less than or equal to M;

inputting the N second training samples into a second classifier to obtain a second forecasting result of each second training sample in the N second training samples, the first classifier and the second classifier being independent of each other;

selecting P second training samples from the N second training samples based upon the second forecasting result of each second training sample, P being an integer greater than or equal to 1 and less than or equal to N;

selecting Q first training samples from other first training samples based upon the first forecasting results of other first training samples in the M first training samples apart from the N second training samples and the value of P, Q being an integer greater than or equal to 1 and less than or equal to a difference between M and N; and

generating P labeled second training samples based upon second forecasting results of the P second training samples and each of the second training samples therein; and

generating Q labeled first training samples based upon first forecasting results of the Q first training samples and each of the first training samples therein.

2. The method of claim 1 , wherein, said selecting the N first training samples comprises:

obtaining a first probability that said first training samples indicated by the first forecasting result are of a preselected type; and

selecting, from the M first training samples, the N first training samples of which the first probability satisfies a pre-set first training condition as the second training samples.

3. The method of claim 2 , wherein the first training condition comprises a probability that the first training samples indicated by the first forecasting result are of a designated type is greater than or equal to a first threshold value and is less than or equal to a second threshold value.

4. The method of claim 2 , wherein said selecting the P second training samples comprises:

obtaining a second probability that the second training samples indicated by the second forecasting result are of a designated type; and

selecting, from the N second training samples, the P second training samples of which the second probability satisfies a pre-set second training condition.

5. The method of claim 4 , wherein the second training condition comprises a designated number with a minimum probability that the second training samples indicated by the second forecasting result are of a designated type.

6. The method of claim 4 , wherein said selecting the Q first training samples comprises:

selecting, from the other first training samples, P first training samples of which a third probability that the first training samples indicated by the first forecasting result are of a designated type satisfies a pre-set third training condition; and

selecting, from the other first training samples, Q−P first training samples of which the third probability satisfies a pre-set fourth training condition.

7. The method of claim 6 , wherein the third training condition comprises a designated number with a minimum probability that the first training samples indicated by the first forecasting result are of a designated type.

8. The method of claim 6 , wherein the fourth training condition comprises a designated number with the maximum probability that the first training samples indicated by the first forecasting result are of a designated type.

9. The method of claim 6 , wherein a ratio of Q−P to 2P is a golden ratio.

10. The method of claim 2 , wherein the designated type comprises at least one of a positive-example type and a counter-example type.

11. An apparatus for labeling training samples, comprising:

a classification system for inputting M unlabeled first training samples into a first classifier to obtain a first forecasting result of each first training sample in the M first training samples, M being an integer greater than or equal to 1;

a selection system for, according to the first forecasting result of each first training sample, selecting, from the M first training samples, N first training samples as second training samples, N being an integer greater than or equal to 1 and less than or equal to M;

said classification system further being for inputting the N second training samples into a second classifier to obtain a second forecasting result of each second training sample in the N second training samples, the first classifier and the second classifier being independent of each other;

said selection system further being for, according to the second forecasting result of each second training sample, selecting, from said N second training samples, P second training samples, P being an integer greater than or equal to 1 and less than or equal to N;

said selection system further being for, according to first forecasting results of other first training samples in the M first training samples, apart from the N second training samples, and the value of P, selecting, from the other first training samples, Q first training samples, Q being an integer greater than or equal to 1 and less than or equal to a difference of M-N; and

a processing system for generating, according to second forecasting results of the P second training samples and each of the second training samples, P labeled second training samples, and generating, according to first forecasting results of the Q first training samples and each of the first training samples therein, Q labeled first training samples.

12. The apparatus of claim 11 , wherein said selection system is configured for:

obtaining a first probability that said first training samples indicated by the first forecasting result are of a preselected type; and

selecting, from the M first training samples, the N first training samples of which the first probability satisfies a pre-set first training condition as the second training samples.

13. The apparatus of claim 12 , wherein the first training condition comprises a probability that the first training samples indicated by the first forecasting result are of a designated type is greater than or equal to a first threshold value and is less than or equal to a second threshold value.

14. The apparatus of claim 12 , wherein said selection system is configured for:

obtaining a second probability that the second training samples indicated by the second forecasting result are of a designated type; and

selecting, from the N second training samples, the P second training samples of which the second probability satisfies a pre-set second training condition.

15. The apparatus of claim 14 , wherein the second training condition comprises a designated number with a minimum probability that the second training samples indicated by the second forecasting result are of a designated type.

16. The apparatus of claim 14 , wherein said selection system is configured for:

selecting, from the other first training samples, P first training samples of which a third probability that the first training samples indicated by the first forecasting result are of a designated type satisfies a pre-set third training condition; and

selecting, from the other first training samples, Q−P first training samples of which the third probability satisfies a pre-set fourth training condition.

17. The apparatus of claim 16 , wherein the third training condition comprises a designated number with a minimum probability that the first training samples indicated by the first forecasting result are of a designated type.

18. The apparatus of claim 16 , wherein the fourth training condition comprises a designated number with a maximum probability that the first training samples indicated by the first forecasting result are of a designated type.

19. The apparatus of claim 16 , wherein the ratio of Q−P to 2P is a golden ratio.

20. The apparatus of claim 19 , wherein the designated type comprises at least one of a positive-example type and a counter-example type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2015
From: CHENG, HUIGE; MAO, YAOZONG
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD
Reel/Frame 035363/0910 →
Priority Claims (1)
CN 2014 1 0433020 · Aug 28, 2014 · national
Continuity (1)
Related Publication 20160063395A1 · Mar 3, 2016