IP Library › Granted Patent US 12,216,696
Granted Patent B2
US 12,216,696 · App. 18/276,378 · Granted Feb 4, 2025

Classification system, method, and program

Inventors: Taro Yano (Tokyo, JP); Kunihiro Takeoka (Tokyo, JP); Masafumi Oyamada (Tokyo, JP)
Assignee: NEC CORPORATION
G06F16/35G06F16/383
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,216,696
App. No.
18/276,378
Granted
Feb 4, 2025
Kind
B2
Abstract

The input means 181 accepts inputs of test data, a hierarchical structure in which a node of bottom layer represents a target class, and a classification score of a seen class as the classification score indicating a probability that the test data is classified into each class. The unseen class score calculation means 182 calculates the classification score of an unseen class based on uniformity of the classification score of each seen class. The matching score calculation means 183 calculates a matching score indicating similarity between the test data and each class label. The final classification score calculation means 184 calculates a final classification score indicating a probability that the test data is classified into the class so that the larger the classification score of each class, and the matching score, the larger the final classification score.

Claims (26)

1. A classification system comprising:

a memory storing instructions; and

one or more processors configured to execute the instructions to:

accept input of test data that is document data to be classified, a hierarchical structure in which a node of a bottom layer represents a target class, and a classification score of each of a plurality of seen classes for the document data, wherein the classification score of each seen class indicates a probability that the test data is correctly classified into the each seen class;

calculate a classification score of each of a plurality of unseen classes for the document data based on uniformity of the classification score of each seen class;

allocate the classification scores of the seen classes under a parent node of the unseen classes to the classification scores of the unseen class such that a sum of the classification scores of the seen classes and the unseen classes under the parent node are equal to the classification scores of the parent node;

for each class of the seen classes and the unseen classes, calculate a matching score indicating similarity between the test data and a class label of the each class, by applying the class label of each class and the test data to a matcher which inputs the class label indicating linguistic meaning of the each class and a document sample and outputs the matching score corresponding to a similarity between the class label and the document sample; and

calculate a final classification score indicating a probability that the test data is classified into a class selected from the seen classes and the unseen classes such that the larger the classification score of the selected class and the larger the matching score for the selected class are, the larger the final classification score is.

2. The classification system according to claim 1 , wherein the processor is configured to execute the instructions to calculate the classification score of each unseen class so that the more uniform the classification score of each seen class is, the higher the classification score of the each unseen class is.

3. The classification system according to claim 1 , wherein the processor is configured to execute the instructions to calculate the matching score by applying the class label of each class and the test data to a matcher which calculates the matching score using a sigmoid function that takes as an argument a weighted linear sum of at least one of similarities of either or both semantic similarities and similarities of included character strings.

4. The classification system according to claim 1 , wherein the processor is configured to execute the instructions to calculate the final classification score by multiplying a value calculated by applying the classification score for each class and the matching score for the each class to a function that takes the classification score as an input and outputs the value that is determined based on a magnitude of the classification score.

5. The classification system according to claim 1 , wherein the processor is configured to execute the instructions to allocate the classification scores of the seen classes equally to the classification scores of the unseen classes.

6. A classification method performed by a computer and comprising:

accepting input of test data that is document data to be classified, a hierarchical structure in which a node of a bottom layer represents a target class, and a classification score of each of a plurality of seen classes for the document data, wherein the classification score of each seen class indicates a probability that the test data is correctly classified into the each seen class;

calculating a classification score of each of a plurality of unseen classes for the document data based on uniformity of the classification score of each seen class;

allocating the classification scores of the seen classes under a parent node of the unseen classes to the classification scores of the unseen class such that a sum of the classification scores of the seen classes and the unseen classes under the parent node are equal to the classification scores of the parent node;

for each class of the seen classes and the unseen classes, calculating a matching score indicating similarity between the test data and a class label of the each class, by applying the class label of each class and the test data to a matcher which inputs the class label indicating linguistic meaning of the each class and a document sample and outputs the matching score corresponding to a similarity between the class label and the document sample; and

calculating a final classification score indicating a probability that the test data is classified into a class selected from the seen classes and the unseen classes such that the larger the classification score of the selected class and the larger the matching score for the selected class are, the larger the final classification score is.

7. The classification method according to claim 6 , wherein the classification score of each unseen class is calculated so that the more uniform the classification score of each seen class is, the higher the classification score of the each unseen class is.

8. A non-transitory computer readable information recording medium storing a classification program executable by a processor to perform a method comprising:

accepting input of test data that is document data to be classified, a hierarchical structure in which a node of a bottom layer represents a target class, and a classification score of each of a plurality of seen classes for the document data, wherein the classification score of each seen class indicates a probability that the test data is correctly classified into the each seen class;

calculating a classification score of each of a plurality of unseen classes for the document data based on uniformity of the classification score of each seen class;

allocating the classification scores of the seen classes under a parent node of the unseen classes to the classification scores of the unseen class such that a sum of the classification scores of the seen classes and the unseen classes under the parent node are equal to the classification scores of the parent node;

for each class of the seen classes and the unseen classes, calculating a matching score indicating similarity between the test data and a class label of the each class, by applying the class label of each class and the test data to a matcher which inputs the class label indicating linguistic meaning of the each class and a document sample and outputs the matching score corresponding to a similarity between the class label and the document sample; and

calculating a final classification score indicating a probability that the test data is classified into a class selected from the seen classes and the unseen classes such that the larger the classification score of the selected class and the larger the matching score for the selected class are, the larger the final classification score is.

9. The non-transitory computer readable information recording medium according to claim 8 , wherein the method further comprises calculating the classification score of each unseen class so that the more uniform the classification score of each seen class is, the higher the classification score of the each unseen class is.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2023
From: YANO, TARO; TAKEOKA, KUNIHIRO; OYAMADA, MASAFUMI
To: NEC CORPORATION
Reel/Frame 064526/0054 →
Continuity (1)
Related Publication 20240119079A1 · Apr 11, 2024
References Cited (12)
US 9928448B1 · Merler · 2018 [cited by examiner]
US 20120269436A1 · Mensink · 2012 [cited by examiner]
US 20130304743A1 · Kurokawa · 2013 [cited by applicant]
US 20210056364A1 · Toizumi · 2021 [cited by applicant]
US 20210357677A1 · Hirayama · 2021 [cited by examiner]
JP H09006799A · 1997 [cited by applicant]
JP 2012155524A · 2012 [cited by applicant]
JP 2020052644A · 2020 [cited by applicant]
JP 2021022343A · 2021 [cited by applicant]
WO 2019171416A1 · 2019 [cited by applicant]
International Search Report for PCT Application No. PCT/JP2021/007385, mailed on May 18, 2021. [cited by applicant]
English translation of Written opinion for PCT Application No. PCT/JP2021/007385, mailed on May 18, 2021. [cited by applicant]