IP Library › Granted Patent US 12,118,308
Granted Patent B2
US 12,118,308 · App. 17/041,299 · Granted Oct 15, 2024

Document classification device and trained model

Inventor: Taishi Ikeda (Chiyoda-ku, JP)
Assignee: NTT DOCOMO, INC.
G06F40/284G06F16/285G06F16/93G06F40/289G06N3/08G06V30/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,118,308
App. No.
17/041,299
Filed
Sep 24, 2020
Granted
Oct 15, 2024
Kind
B2
Art Unit
2671
USPC
382/161
Abstract

A document classification device is a device that generates a document classification model which outputs identification information for identifying a result of classification on the basis of an input document by machine learning, and includes: an acquisition unit configured to acquire learning data including a document and the identification information correlated with the document; a feature extracting unit configured to extract words included in the document and character information which is a character string including one character of characters constituting the words or a plurality of characters consecutive in the words and which is one or more pieces of information capable of being extracted from the words as features; and a model generating unit configured to perform machine learning on the basis of the feature extracted from the document and the identification information correlated with the document and to generate the document classification model.

Claims (19)

1. A document classification device that generates a document classification model which is a model for classifying a document and which outputs identification information for identifying a result of classification on the basis of an input document by machine learning, the document classification device comprising circuitry configured to:

acquire learning data including a document and the identification information correlated with the document;

extract words included in the document and character information which is a character string including one character of characters constituting the words or a plurality of characters consecutive in the words and which is one or more pieces of information capable of being extracted from the words as features; and

perform machine learning on the basis of the feature extracted from the document and the identification information correlated with the document and to generate the document classification model,

wherein the circuitry is configured to extract the character information from a word satisfying a predetermined condition out of the words included in the document.

2. The document classification device according to claim 1 , wherein the circuitry is configured to exclude a word corresponding to a second part of speech out of the words included in the document from the words from which the character information is extracted.

3. The document classification device according to claim 2 , wherein the circuitry is configured to extract the character information which is included in a predetermined frequency or less in the learning data acquired by the circuitry as features.

4. The document classification device according to claim 2 , wherein the circuitry is configured to remove repeated character information from the character information extracted from the words included in the document.

5. The document classification device according to claim 1 , wherein the circuitry is configured to extract the character information which is included in a predetermined frequency or less in the learning data acquired by the circuitry as features.

6. The document classification device according to claim 5 , wherein the circuitry is configured to remove repeated character information from the character information extracted from the words included in the document.

7. The document classification device according to claim 1 , wherein the circuitry is configured to remove repeated character information from the character information extracted from the words included in the document.

8. The document classification device according to claim 1 , wherein the circuitry is configured to extract the character information which is included in a predetermined frequency or less in the learning data acquired by the circuitry as features.

9. The document classification device according to claim 1 , wherein the circuitry is configured to remove repeated character information from the character information extracted from the words included in the document.

10. A method, implemented by circuitry of a document classification device that generates a document classification model which is a model for classifying a document and which outputs identification information for identifying a result of classification on the basis of an input document by machine learning, the method comprising:

acquiring learning data including a document and the identification information correlated with the document;

extracting words included in the document and character information which is a character string including one character of characters constituting the words or a plurality of characters consecutive in the words and which is one or more pieces of information capable of being extracted from the words as features; and

performing machine learning on the basis of the feature extracted from the document and the identification information correlated with the document and to generate the document classification model,

wherein the method includes extracting the character information from a word satisfying a predetermined condition out of the words included in the document,

wherein the circuitry extracts the character information from a word corresponding to a first part of speech out of the words included in the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2020
From: IKEDA, TAISHI
To: NTT DOCOMO, INC.
Reel/Frame 053875/0798 →
Priority Claims (1)
JP 2018-138453 · Jul 24, 2018 · national
Continuity (1)
Related Publication 20210026874A1 · Jan 28, 2021