IP Library Granted Patent US 12,468,804
Granted Patent B2
US 12,468,804 · App. 17/283,254 · Granted Nov 11, 2025

Data classification device, data classification method, and data classification program

Inventors: Toshiki Shibahara (Musashino, JP); Daiki Chiba (Musashino, JP); Mitsuaki Akiyama (Musashino, JP); Kunio Hato (Musashino, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F21/56G06F18/10G06F18/241G06F18/2431G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,804
App. No.
17/283,254
Granted
Nov 11, 2025
Kind
B2
Abstract

A data classification device includes: a known data input unit that receives an input of known data, the known data being data already classified into a class and a subclass subordinate to the class; a feature extraction unit that extracts, from features included in the known data, a feature that causes classification of the known data belonging to the same class into a subclass using the feature to fail; and a classification unit that classifies classification target data into a class using the feature extracted by the feature extraction unit.

Claims (24)

1 . A data classification device, comprising:

a memory; and

a processor coupled to the memory and programmed to execute a process comprising:

receiving an input of known data, the known data being data already classified into a class and a subclass subordinate to the class;

extracting features from the known data, wherein the features which are extracted are features shared between the subclasses in the class;

determining whether classification of the known data belonging to the class into a subclass existing within the subclasses using a feature of the features which are extracted is a success or failure;

outputting the feature if it is determined that the feature causes the classification to fail; and

classifying classification target data into a class indicating malicious or not using the output feature for detecting an attack and notifying a terminal which sent the classification target data of a result of the classifying.

2 . The data classification device according to claim 1 , wherein the extracting extracts, from features included in the known data, a feature that causes classification of the known data belonging to a same class into subclasses similar to one another using the feature to fail.

3 . The data classification device according to claim 1 , wherein the extracting extracts, from features included in the known data, a feature that causes classification of the known data of a same class into a subclass using the feature to fail, and that causes classification of the known data into a class using the feature to succeed.

4 . The data classification device according to claim 1 , wherein the extracting, when known data of a same class is classified into a subclass using a feature, calculates a predictive probability predicting into which subclass the known data is classified, and calculates a value by smoothing the calculated predictive probability between subclasses, and extracts, from features included in the known data, a feature that makes a result of classification of the known data into a subclass using the feature close to a value of the smoothed predictive probability.

5 . The data classification device according to claim 1 , wherein a data group belonging to the same subclass is a malicious data group belonging to a malicious class and created using a same malicious tool.

6 . A data classification method executed by a data classification device, the data classification method comprising:

receiving an input of known data, the known data being data already classified into a class and a subclass subordinate to the class;

extracting, features from the known data, wherein the features which are extracted are features shared between the subclasses in the class:

determining whether classification of the known data belonging to the class into a subclass existing within the subclasses using a feature of the features which are extracted is a success or failure;

outputting the feature if it is determined that the feature causes the classification to fail; and

classifying classification target data into a class indicating malicious or not using the output feature for detecting an attack and notifying a terminal which sent the classification target data of a result of the classifying.

7 . A non-transitory computer-readable recording medium having stored therein data classification program that causes a computer to execute a process comprising:

receiving an input of known data, the known data being data already classified into a class and a subclass subordinate to the class;

extracting features from the known data, wherein the features which are extracted are features shared between the subclasses in the class;

determining whether classification of the known data belonging to the class into a subclass existing within the subclasses using a feature of the features which are extracted is a success or failure;

outputting the feature if it is determined that the feature causes the classification to fail; and

classifying classification target data into a class indicating malicious or not using the output feature for detecting an attack and notifying a terminal which sent the classification target data of a result of the classifying.

Assignments (2)
CHANGE OF NAME Recorded Aug 20, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072556/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2021
From: SHIBAHARA, TOSHIKI; CHIBA, DAIKI; AKIYAMA, MITSUAKI; HATO, KUNIO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 055847/0882 →
Priority Claims (1)
JP 2018-191174 · Oct 9, 2018 · national
Continuity (1)
Related Publication 20210342651A1 · Nov 4, 2021
References Cited (35)
US 8078625B1 · Zhang et al. · 2011 [cited by applicant]
US 20120158626A1 · Zhu et al. · 2012 [cited by applicant]
US 20120254333A1 · Chandramouli · 2012 [cited by examiner]
US 20150156211A1 · Chi Tin · 2015 [cited by examiner]
US 20170244741A1 · Ferrer · 2017 [cited by examiner]
US 20180007003A1 · Hodgman et al. · 2018 [cited by applicant]
CN 107534675A · 2018 [cited by examiner]
CN 103927486B · 2018 [cited by examiner]
JP 2004537916A · 2004 [cited by examiner]
JP 2017142552A · 2017 [cited by examiner]
JP 2018517999A · 2018 [cited by examiner]
Shun Tobiyama, etc., “Malware Detection with Deep Neural Network Using Process Behavior”, published via 2016 IEEE 40th Annual Computer Software and Applications Conference, retrieved Aug. 9, 2024. (Year: 2016). [cited by examiner]
TaeGuen Kim, etc., “A Multimodal Deep Learning Method for Android Malware Detection using Various Features”, published in IEEE Transactions on Information Forensics and Security, 14(3), 773-788 (print publication Aug. 2… [cited by examiner]
Rajesh Kumar, etc., “Malware Detection Modeling Systems”, published via 2018 International Conference on Recent Trends in Advance Computing, pp. 187-192 (publication date Sep. 1, 2018), retrieved Aug. 9, 2024. (Year: 20… [cited by examiner]
Hye Min Kim, etc., “Andro-Simnet: Android Malware Family Classification using Social Network Analysis”, published via 2018 16th Annual Conference on Privacy, Security and Trust, pp. 1-8 (publication date Aug. 1, 2018), … [cited by examiner]
Zhang Fuyong, etc., “Malware Detection and Classification Based on n-grams Attribute Similarity”, published via 2017 IEEE International Conference on Computational Science and Engineering (CSE) and IEEE International Co… [cited by examiner]
Mohammad Imran, etc., “Similarity-based Malware Classification using Hidden Markov Model”, published via 2015 Fourth International Conference on Cyber Security, Cyber Warfare, and Digital Forensic, pp. 129-134, retrieve… [cited by examiner]
Zhihua Cui, etc., “Detection of Malicious Code Variants Based on Deep Learning”, published via IEEE Transactions On Industrial Informatics, vol. 14, No. 7, p. 3187, Jul. 2018, retrieved Aug. 9, 2024. (Year: 2018). [cited by examiner]
Edmar Rezende, etc., “Malicious Software Classification using Transfer Learning of ResNet-50 Deep Neural Network”, published via 2017 16th IEEE International Conference on Machine Learning and Applications, retrieved Au… [cited by examiner]
William Fleshman, etc., “Static Malware Detection & Subterfuge: Quantifying the Robustness of Machine Learning and Current Anti-Virus”, published via 2018 13th International Conference on Malicious and Unwanted Software… [cited by examiner]
Chang-Bin Zhang, etc., “Delving Deep into Label Smoothing”, published via Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, retrieved Aug. 9, 2024. (Year: 2015). [cited by examiner]
“1.16 Probability calibration”, published on Aug. 27, 2016 to https://scikit-learn.org/stable/modules/calibration.html, retrieved Aug. 9, 2024. ( Year: 2016). [cited by examiner]
Dario Garcia-Gasulla, etc., “On the Behavior of Convolutional Nets for Feature Extraction”, published on Jan. 29, 2018 to arXiv, retrieved Feb. 22, 2025. (Year: 2018). [cited by examiner]
CS231n Convolutional Neural Networks for Visual Recognition, published on Feb. 10, 2015 to https://cs231n.github.io/transfer-learning, retrieved Feb. 22, 2025. (Year: 2015). [cited by examiner]
Lars Hulstaert, “Transfer Learning: Leverage Insights from Big Data”, published on Jan. 19, 2018 to https:/ /www.datacamp.com/tutorial /transfer-learning, retrieved Feb. 22, 2025. (Year: 2018). [cited by examiner]
Kateryna Chumachenko, “Machine Learning Methods for Malware Detection and Classification”, published in 2017 to https://core.ac.uk/download/pdf/80994982.pdf, retrieved Feb. 22, 2025. (Year: 2017). [cited by examiner]
Yong Jin, etc., “A Client Based Anomaly Traffic Detection and Blocking Mechanism by Monitoring DNS Name Resolution with User Alerting Feature”, published via 2018 International Conference on Cyberworlds (CW) (2018, pp. … [cited by examiner]
Google Workspace Admin Help, “Advanced phishing and malware protection”, published Mar. 21, 18 to https://support.google.com/a/answer/9157861?hl=en, retrieved Jun. 16, 2025. (Year: 2018). [cited by examiner]
“Malware Incident Response Steps on Windows, and Determining if the Threat Is Truly Gone”, published Mar. 21, 2017 to https://www.rapid7.com/blog/post/2017/03/21/responding-to-malware-events-on-windows-determining-if-th… [cited by examiner]
Yahoo Help, “Suspicious activity alert received when an email is sent”, published on May 22, 2015 to http://help.yahoo.com/kb/SLN3406.html, retrieved Jun. 16, 25. (Year: 2015). [cited by examiner]
Arp et al., “DREBIN: Effective and Explainable Detection of Android Malware in Your Pocket”, Proceedings of the 2014 Network and Distributed System Security Symposium, Feb. 2014, pp. 1-15. [cited by applicant]
Xu et al., “Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity Detection”, Proceedings of the 24th ACM Conference on Computer and Communications Security, Oct. 2017, pp. 363-376. [cited by applicant]
Bartos et al., “Optimized Invariant Representation of Network Traffic for Detecting Unseen Malware Variants”, Proceedings of the 25th USENIX Security Symposium, Aug. 10-12, 2016, pp. 807-822. [cited by applicant]
Jordaney et al., “Transcend: Detecting Concept Drift in Malware Classification Models”, Proceedings of the 26th USENIX Security Symposium, Aug. 16-18, 2017, pp. 625-642. [cited by applicant]
Bousmalis et al., “Domain Separation Networks”, Proceedings of the 29th Advances in Neural Information Processing Systems, Dec. 2016, pp. 343-351. [cited by applicant]