IP Library › Granted Patent US 11,157,829
Granted Patent B2
US 11,157,829 · App. 15/653,007 · Granted Oct 26, 2021

Method to leverage similarity and hierarchy of documents in NN training

Inventor: Gakuto Kurata (Tokyo, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/00G06F16/3344G06F16/93G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,157,829
App. No.
15/653,007
Granted
Oct 26, 2021
Kind
B2
Abstract

A computer-implemented method for training a natural language-based classifier, includes obtaining a query and a first label which is a binary vector, each of a plurality of elements of the binary vector being associated with one of a plurality of instances, the first label indicating that the query is classified into a specific instance of the plurality of instances by a value set to a specific element associated with the specific instance, estimating relationships between the specific instance and instances other than the specific instance of the plurality of instances, generating a second label which is a continuous-valued vector from the first label by distributing the value set to the specific element to elements other than the specific element of the plurality of elements according to the relationships, and training the natural language-based classifier using the query and the second label.

Claims (50)

1. A computer-implemented method for training a natural language-based classifier, the method comprising:

obtaining a query and a document label represented by a binary vector, each of a plurality of elements of the binary vector being associated with at least one instance from a plurality of instances, the document label indicating that the query is classified into a specific instance from the plurality of instances by a value set to a specific element associated with the specific instance;

estimating relationships between the specific instance and instances other than the specific instance from the plurality of instances;

generating a relation label represented by a continuous-valued vector from the document label by distributing the value set to the specific element to elements other than the specific element from the plurality of elements according to the relationships;

training the natural language-based classifier using the query and the relation label; and

detecting probabilities each indicating that a corresponding instance is associated with a selected training query to output predicted labels.

2. The method of claim 1 , wherein the relationships are similarities.

3. The method of claim 2 , wherein the similarities include a cosine similarity between two documents among the plurality of documents.

4. The method of claim 3 , wherein the similarities are based on a number of words commonly appearing in the two documents.

5. The method of claim 1 , wherein training includes training the natural language-based classifier using the document label.

6. The method of claim 5 , wherein:

training includes training the natural language-based classifier using two loss functions; and

the two loss functions are a loss function which is cross-entropy based on the document label, and a loss function which is cross-entropy based on the relation label.

7. The method of claim 5 , wherein:

training includes training the natural language-based classifier using one loss function; and

the one loss function is cross-entropy based on the document label and the relation label.

8. An apparatus for training a natural language-based classifier, the apparatus comprising:

a processor; and

a memory coupled to the processor, wherein:

the memory comprises program instructions executable by the processor to cause the processor to perform a method comprising:

obtaining a query and a document label represented by a binary vector, each of a plurality of elements of the binary vector being associated with at least one instance from a plurality of instances, the document label indicating that the query is classified into a specific instance from the plurality of instances by a value set to a specific element associated with the specific instance;

estimating relationships between the specific instance and instances other than the specific instance from the plurality of instances;

generating a relation label represented by a continuous-valued vector from the document label by distributing the value set to the specific element to elements other than the specific element from the plurality of elements according to the relationships;

training the natural language-based classifier using the query and the relation label; and

detecting probabilities each indicating that a corresponding instance is associated with a selected training query to output predicted labels.

9. The apparatus of claim 8 , wherein the relationships are similarities.

10. The apparatus of claim 9 , wherein the similarities include a cosine similarity between two documents among the plurality of documents.

11. The apparatus of claim 10 , wherein the similarities are based on a number of words commonly appearing in the two documents.

12. The apparatus of claim 8 , wherein training includes training the natural language-based classifier using the document label.

13. The apparatus of claim 12 , wherein:

training includes training the natural language-based classifier using two loss functions; and

the two loss functions are a loss function which is cross-entropy based on the document label, and a loss function which is cross-entropy based on the relation label.

14. The apparatus of claim 12 , wherein:

training includes training the natural language-based classifier using one loss function; and

the one loss function is cross-entropy based on the document label and the relation label.

15. A computer program product for training a natural language-based classifier, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

obtaining a query and a document label represented by a binary vector, each of a plurality of elements of the binary vector being associated with at least one instance from a plurality of instances, the document label indicating that the query is classified into a specific instance from the plurality of instances by a value set to a specific element associated with the specific instance;

estimating relationships between the specific instance and instances other than the specific instance from the plurality of instances;

generating a relation label represented by a continuous-valued vector from the document label by distributing the value set to the specific element to elements other than the specific element from the plurality of elements according to the relationships;

training the natural language-based classifier using the query and the relation label; and

detecting probabilities each indicating that a corresponding instance is associated with a selected training query to output predicted labels.

16. The computer program product of claim 15 , wherein the relationships are similarities.

17. The computer program product of claim 16 , wherein the similarities include a cosine similarity between two documents among the plurality of documents.

18. The computer program product of claim 15 , wherein training includes training the natural language-based classifier using the document label.

19. The computer program product of claim 18 , wherein:

training includes training the natural language-based classifier using two loss functions; and

the two loss functions are a loss function which is cross-entropy based on the document label, and a loss function which is cross-entropy based on the relation label.

20. The computer program product of claim 18 , wherein:

training includes training the natural language-based classifier using one loss function; and

the one loss function is cross-entropy based on the document label and the relation label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2017
From: KURATA, GAKUTO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 043035/0290 →
Continuity (1)
Related Publication 20190026646A1 · Jan 24, 2019
Cited By (1)
US 12,572,746