Data annotation method and apparatus, and fine-grained recognition method and apparatus
View Patent ↗This application relates to the field of image annotation and recognition in the field of artificial intelligence technologies, and in particular, to a data annotation method. The method includes: using at least two different classification models; pretraining one of the classification models as an initial classification model, and annotating a label for data in a to-be-annotated source dataset as initial data by using the pretrained classification model; and controlling the classification models to perform alternating training and data annotation a quantity of times. Operations of current training and current data annotation include: obtaining data that is re-annotated with a label by a previously trained classification model, selecting a first part of the data to train a current classification model, and re-annotating, by the trained current classification model, a label for a second part of data that is not selected.
1 . A data annotation method, comprising:
using at least two classification models with different structures, the at least two classification models comprising a first classification model and a second classification model;
pretraining the first classification model using a target dataset with a target annotation type label, wherein the target annotation type label includes a label from the first classification model and a label from the second classification model;
annotating a label for data in a to-be-annotated source dataset using the first classification model to produce initial data annotated with the label from the first classification model;
performing alternating training and data annotation a quantity of times, comprising using the initial data annotated with the label from the first classification model to train the second classification model, and comprising using a data annotated with the label from the second classification model to train the first classification model; and
in the alternating training and data annotation, current training and current data annotation performed by a currently trained classification model, comprising at least one of the first classification model or the second classification model, comprise:
obtaining data that is re-annotated with a label by a previously trained classification model, comprising at least one of the first classification model or the second classification model, that is different than the currently trained classification model;
selecting a first part of the data to train the currently trained classification model, wherein the selecting the first part of the data is performed based on stability of an annotation of each piece of data, and wherein the stability is measured by using information entropy of soft label probabilities output by the previously trained classification model, and the selecting the first part of the data comprises:
calculating an information entropy value for each piece of the data based on the soft label probabilities for that label;
ordering the information entropy values in ascending order such that lower entropy indicates higher annotation stability; and
selecting the first part of the data from a front of the ascending order; and
re-annotating, by the currently trained classification model, respective labels for remaining data that is not included in the first part of the data.
2 . The method according to claim 1 , wherein the source dataset and target dataset have labels of a same basic classification; and
the target annotation type label is a label of a further fine-grained classification in the basic classification.
3 . A data annotation method, comprising:
using at least two classification models with different structures, the at least two classification models comprising a first classification model and a second classification model;
performing alternating training and data annotation a quantity of times, wherein, in the alternating training and data annotation, a part of data used for training the first classification model has a target annotation type label, wherein the target annotation type label includes a label from the first classification model and a label from the second classification model, and a data annotated with the label from the second classification model is used to train the first classification model; and
in the alternating training and data annotation, current training and current data annotation performed by a currently trained classification model, comprising at least one of the first classification model or the second classification model, comprise:
obtaining data that is re-annotated with a label by a previously trained classification model, comprising at least one of the first classification model or the second classification model, that is different than the currently trained classification model;
selecting a first part of the data to train the currently trained classification model, wherein the selecting the first part of the data is performed based on stability of an annotation of each piece of data, and wherein the stability is measured by using information entropy of soft label probabilities output by the previously trained classification model, and the selecting the first part of the data comprises:
calculating the information entropy value for each piece of the data based on the soft label probabilities for that label;
ordering the information entropy values in ascending order such that lower entropy indicates higher annotation stability; and
selecting the first part of the data from a front of the ascending order; and
re-annotating, by the currently trained classification model, respective labels for remaining data that is not included in the first part of the data.
4 . The method according to claim 3 , wherein before the alternating training and data annotation are performed, the method further comprises: pretraining the first classification model by using a target dataset with the target annotation type label.
5 . The method according to claim 3 , wherein the data used to train the currently trained classification model has labels of a same basic classification; and
the target annotation type label is a label of a further fine-grained classification in the basic classification.
6 . A computer device, comprising:
a bus;
a communications interface, wherein the communications interface is connected to the bus;
at least one processor, wherein the at least one processor is connected to the bus; and at least one memory, wherein the at least one memory is connected to the bus and stores program instructions, and the at least one processor executes the program instructions to:
use at least two classification models with different structures, the at least two classification models comprising a first classification model and a second classification model;
pretrain the first classification model using a target dataset with a target annotation type label, wherein the target annotation type label includes a label from the first classification model and a label from the second classification model;
annotate a label for data in a to-be-annotated source dataset using the first classification model to produce initial data annotated with the label from the first classification model;
perform alternating training and data annotation a quantity of times, comprising using the initial data annotated with the label from the first classification model to train the second classification model, and comprising using a data annotated with the label from the second classification model to train the first classification model; and
in the alternating training and data annotation process, current training and current data annotation performed by a currently trained classification model, comprising at least one of the first classification model or the second classification model, comprise:
obtain data that is re-annotated with a label by a previously trained classification model, comprising at least one of the first classification model or the second classification model, that is different than the currently trained classification model;
select a first part of the data to train the currently trained classification model, wherein the selecting the first part of the data is performed based on stability of an annotation of each piece of data, and wherein the stability is measured by using information entropy of soft label probabilities output by the previously trained classification model, and the selecting the first part of the data comprises:
calculate an information entropy value for each piece of the data based on the soft label probabilities for that label;
order the information entropy values in ascending order such that lower entropy indicates higher annotation stability; and
select the first part of the data from a front of the ascending order; and
re-annotate, by the currently trained classification model, respective labels for remaining data that is not included in the first part of the data.
7 . The computer device according to claim 6 , wherein the source dataset and target dataset have labels of a same basic classification; and
the target annotation type label is a label of a further fine-grained classification in the basic classification.
8 . A computer device, comprising:
a bus;
a communications interface, wherein the communications interface is connected to the bus;
at least one processor, wherein the at least one processor is connected to the bus; and
at least one memory, wherein the at least one memory is connected to the bus and stores program instructions, and the at least one processor executes the program instructions to:
use at least two classification models with different structures, the at least two classification models comprising a first classification model and a second classification model;
perform alternating training and data annotation a quantity of times, wherein in the alternating training and data annotation, a part of data used for training the first classification model has a target annotation type label, wherein the target annotation type label includes a label from the first classification model and a label from the second classification model, and a data annotated with the label from the second classification model is used to train the first classification model; and
in the alternating training and data annotation, current training and current data annotation performed by a currently trained classification model, comprising at least one of the first classification model or the second classification model, comprise:
obtain data that is re-annotated with a label by a previously trained classification model, comprising at least one of the first classification model or the second classification model, that is different than the currently trained classification model;
select a first part of the data to train the currently trained classification model, wherein the selecting the first part of the data is performed based on stability of an annotation of each piece of data, and wherein the stability is measured by using information entropy of soft label probabilities output by the previously trained classification model, and the selecting the first part of the data comprises:
calculate an information entropy value for each piece of the data based on the soft label probabilities for that label;
order the information entropy values in ascending order such that lower entropy indicates higher annotation stability; and
select the first part of the data from a front of the ascending order; and
re-annotate, by the currently trained classification model, respective labels for remaining data that is not included in the first part of the data.
9 . The computer device according to claim 8 , wherein before the alternating training and data annotation are performed, the at least one processor executes the program instructions to pretrain the first classification model by using a target dataset with the target annotation type label.
10 . The computer device according to claim 8 , wherein the data used to train the currently trained classification model has labels of a same basic classification; and
the target annotation type label is a label of a further fine-grained classification in the basic classification.