IP Library › Granted Patent US 11,741,168
Granted Patent B1
US 11,741,168 · App. 16/588,595 · Granted Aug 29, 2023

Multi-label document classification for documents from disjoint class sets

Inventors: Sravan Babu Bodapati (Bellevue, WA); Rishita Rajal Anubhai (Seattle, WA); Yahor Pushkin (Redmond, WA)
Assignee: Amazon Technologies, Inc.
G06F16/93G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,168
App. No.
16/588,595
Granted
Aug 29, 2023
Kind
B1
Abstract

Techniques for multi-label document classification are described. Clustering is used to cluster labels in a set. A machine learning model including a multi-label classifier for each cluster is created, the multi-label classifier for a given cluster to classify a document with one or more of the labels in the cluster.

Claims (67)

1 . A computer-implemented method comprising:

generating, by a machine learning service of a provider network, a plurality of label embeddings corresponding to a plurality of labels using a pre-trained model;

clustering the plurality of label embeddings into a plurality of clusters, wherein a number of clusters in the plurality of clusters is less than a number of labels in the plurality of labels, wherein a first cluster of the plurality of clusters corresponds to a first document class and a second cluster of the plurality of clusters corresponds to a second document class;

creating a multi-label classifier model that includes a neural network-based classifier for each cluster in the plurality of clusters, wherein the neural network-based classifier for each cluster includes an encoder to generate a representation of the unlabeled document and, for each label corresponding to a label embedding in the cluster, a decoder to classify the representation of the unlabeled document with the label;

training a set of parameters of each neural network-based classifier using a training document set, the training document set including a plurality of documents and, for each document, an associated one or more labels from the plurality of labels;

classifying, using the multi-label classifier model, an unlabeled document identified by a user with multiple labels of the plurality of labels; and

storing the one or more labels in a data store.

2 . The computer-implemented method of claim 1 , wherein the multi-label classifier model further includes a group classifier and further comprising:

providing, by the group classifier to the neural network-based classifier for a given cluster of the plurality of clusters, an indication of whether the unlabeled document belongs to the given cluster.

3 . The computer-implemented method of claim 1 , wherein the neural network-based classifier for each cluster includes an encoder to generate a fixed-length representation of the unlabeled document and, for each label corresponding to a label embedding in the cluster, a decoder to classify the fixed-length representation of the unlabeled document with the label.

4 . A computer-implemented method comprising:

generating a plurality of label embeddings corresponding to a plurality of labels;

clustering the plurality of label embeddings into a plurality of clusters, wherein a number of clusters in the plurality of clusters is less than a number of labels in the plurality of labels;

creating a machine learning model that includes a neural network-based classifier for each cluster in the plurality of clusters, wherein the neural network-based classifier for each cluster includes an encoder to generate a representation of the unlabeled document and, for each label corresponding to a label embedding in the cluster, a decoder to classify the representation of the unlabeled document with the label;

training a set of parameters of each neural network-based classifier in the machine learning model using a training document set, the training document set including a plurality of documents and, for each document, an associated one or more labels from the plurality of labels;

classifying, using the machine learning model, an unlabeled document identified by a user with one or more labels of the plurality of labels; and

storing the one or more labels in a data store.

5 . The computer-implemented method of claim 4 , wherein the machine learning model further includes a group classifier and further comprising:

providing, by the group classifier to the neural network-based classifier for a given cluster of the plurality of clusters, an indication of whether the unlabeled document belongs to the given cluster.

6 . The computer-implemented method of claim 4 , wherein the clustering is performed using k-means clustering.

7 . The computer-implemented method of claim 4 , further comprising selecting the plurality of clusters, wherein the selecting includes:

generating a first metric based at least in part on a distance between a first label embedding in a first cluster of the plurality of clusters to a cluster mean for the first cluster;

clustering the plurality of label embeddings into another plurality of clusters, wherein a number of clusters in the other plurality of clusters is different than the number of clusters in the plurality of clusters;

generating a second metric based at least in part on a distance between a second label embedding in a second cluster of the other plurality of clusters to a cluster mean for the second cluster; and

selecting the plurality of clusters based at least in part on a comparison of the first metric to the second metric.

8 . The computer-implemented method of claim 4 , further comprising:

clustering the plurality of label embeddings into another plurality of clusters, wherein a number of clusters in the other plurality of clusters is different than the number of clusters in the plurality of clusters;

creating another machine learning model that includes a neural network-based classifier for each cluster in the other plurality of clusters;

generating a first metric representing a classification performance of the machine learning model on a validation document set;

training a set of parameters of each neural network-based classifier in the other machine learning model using the training document set;

generating a second metric representing a classification performance of the other machine learning model on the validation document set; and

selecting the machine learning model to classify the unlabeled document based at least in part on a comparison of the first metric to the second metric.

9 . The computer-implemented method of claim 4 , wherein the plurality of label embeddings corresponding to the plurality of labels are generated using a pre-trained model.

10 . The computer-implemented method of claim 4 , further comprising:

training a set of parameters of each neural network-based classifier using a training document set, the training document set including a plurality of training documents and, for each document, an associated one or more labels from the plurality of labels.

11 . The computer-implemented method of claim 10 :

wherein a label embedding corresponding to a given label in the plurality of label embeddings includes a document embedding of at least one training document labeled with the given label, and

wherein the document embedding is generated using a pre-trained model.

12 . A system comprising:

a first one or more electronic devices of a provider network to implement a data store; and

a second one or more electronic devices of the provider network to implement a machine learning service, the machine learning service including instructions that upon execution cause the machine learning service to:

generate a plurality of label embeddings corresponding to a plurality of labels;

cluster the plurality of label embeddings into a plurality of clusters, wherein a number of clusters in the plurality of clusters is less than a number of labels in the plurality of labels;

create a machine learning model that includes a neural network-based classifier for each cluster in the plurality of clusters, wherein the neural network-based classifier for each cluster includes an encoder to generate a representation of the unlabeled document and, for each label corresponding to a label embedding in the cluster, a decoder to classify the representation of the unlabeled document with the label;

train a set of parameters of each neural network-based classifier in the machine learning model using a training document set, the training document set including a plurality of documents and, for each document, an associated one or more labels from the plurality of labels;

classify, using the machine learning model, an unlabeled document identified by a user with one or more labels of the plurality of labels; and

store the one or more labels in the data store.

13 . The system of claim 12 , wherein the machine learning model further includes a group classifier, and wherein the machine learning service includes further instructions that upon execution cause the machine learning service to:

provide, by the group classifier to the neural network-based classifier for a given cluster of the plurality of clusters, an indication of whether the unlabeled document belongs to the given cluster.

14 . The system of claim 12 , wherein the neural network-based classifier for each cluster includes an encoder to generate a representation of the unlabeled document and, for each label corresponding to a label embedding in the cluster, a decoder to classify the representation of the unlabeled document with the label.

15 . The system of claim 12 , wherein the machine learning service includes further instructions that upon execution cause the machine learning service to:

generate a first metric based at least in part on a distance between a first label embedding in a first cluster of the plurality of clusters to a cluster mean for the first cluster;

cluster the plurality of label embeddings into another plurality of clusters, wherein a number of clusters in the other plurality of clusters is different than the number of clusters in the plurality of clusters;

generate a second metric based at least in part on a distance between a second label embedding in a second cluster of the other plurality of clusters to a cluster mean for the second cluster; and

select the plurality of clusters based at least in part on a comparison of the first metric to the second metric.

16 . The system of claim 12 , wherein the machine learning service includes further instructions that upon execution cause the machine learning service to:

cluster the plurality of label embeddings into another plurality of clusters, wherein a number of clusters in the other plurality of clusters is different than the number of clusters in the plurality of clusters;

create another machine learning model that includes a neural network-based classifier for each cluster in the other plurality of clusters;

generate a first metric representing a classification performance of the machine learning model on a validation document set;

train a set of parameters of each neural network-based classifier in the other machine learning model using the training document set;

generate a second metric representing a classification performance of the other machine learning model on the validation document set; and

select the machine learning model to classify the unlabeled document based at least in part on a comparison of the first metric to the second metric.

17 . The system of claim 12 , wherein the plurality of label embeddings corresponding to the plurality of labels are generated using a pre-trained model.

18 . The system of claim 12 , wherein the machine learning service includes further instructions that upon execution cause the machine learning service to:

train a set of parameters of each neural network-based classifier using a training document set, the training document set including a plurality of training documents and, for each document, an associated one or more labels from the plurality of labels,

wherein a label embedding corresponding to a given label in the plurality of label embeddings includes a document embedding of at least one training document labeled with the given label, and

wherein the document embedding is generated using a pre-trained model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2019
From: BODAPATI, SRAVAN BABU; ANUBHAI, RISHITA RAJAL; PUSHKIN, YAHOR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050921/0422 →
Cited By (3)
US 12,197,855 US 12,211,303 US 12,417,222