IP Library Granted Patent US 12,626,187
Granted Patent B2
US 12,626,187 · App. 17/919,542 · Granted May 12, 2026

Learning device, learning method, and learning program

Inventors: Masanori Yamada (Musashino, JP); Sekitoshi Kanai (Musashino, JP); Tomokatsu Takahashi (Musashino, JP); Yuki Yamanaka (Musashino, JP)
Assignee: NTT, Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,187
App. No.
17/919,542
Granted
May 12, 2026
Kind
B2
Abstract

A learning device includes processing circuitry configured to acquire data of which a label is predicted, and reduce, in a model representing a probability distribution of the label of the acquired data, a rank of a Fisher information matrix for the data to a value less than a predetermined value and learn the model.

Claims (28)

1 . A learning device comprising:

processing circuitry configured to:

acquire training data of which a label is predicted;

generate a model, representing a probability distribution of the label of the training data, to resist an adversarial example by:

increasing a temperature in a Boltzmann distribution to a value greater than 1 in the probability distribution, and

learning the model;

receive input data of which a label is to be predicted; and

predict the label of the input data by using the learned model.

2 . The learning device of claim 1 , wherein the processing circuitry is further configured to perform the learning using the adversarial example that is generated by superimposing noise on the training data.

3 . The learning device of claim 1 , wherein the processing circuitry is further configured to perform the learning by generating the adversarial example using a first temperature value in the probability distribution and updating the model using a loss function generated with a second temperature value that is greater than the first temperature value.

4 . The learning device of claim 3 , wherein the first temperature value is 1.

5 . The learning device of claim 1 , wherein the processing circuitry is further configured to predict the label of the input data by setting a temperature value in the probability distribution to 1.

6 . The learning device of claim 1 , wherein the processing circuitry is further configured to perform the learning by performing an iterative process of generating the adversarial example and updating the model until a loss function converges.

7 . The learning device of claim 1 , wherein the processing circuitry is further configured to perform the learning by reducing a loss function that is based on a true probability of the label of the training data and a probability predicted by the model.

8 . A learning method which is executed in a learning device, the learning method comprising:

acquiring training data of which a label is predicted;

generating a model, representing a probability distribution of the label of the training data, to resist an adversarial example by:

increasing a temperature in a Boltzmann distribution to a value greater than 1 in the probability distribution, and

learning the model;

receiving input data of which a label is to be predicted; and

predicting the label of the input data by using the learned model.

9 . A non-transitory computer-readable recording medium storing therein a learning program that causes a computer to execute a process comprising:

acquiring training data of which a label is predicted;

generating a model, representing a probability distribution of the label of the training data, to resist an adversarial example by:

increasing a temperature in a Boltzmann distribution to a value greater than 1 in the probability distribution, and

learning the model;

receiving input data of which a label is to be predicted; and

predicting the label of the input data by using the learned model.

Assignments (2)
CHANGE OF NAME Recorded Aug 20, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072556/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: YAMADA, MASANORI; KANAI, SEKITOSHI; TAKAHASHI, TOMOKATSU; YAMANAKA, YUKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 061454/0705 →
Continuity (1)
Related Publication 20230162085A1 · May 25, 2023
References Cited (19)
US 20160328253A1 · Majumdar · 2016 [cited by examiner]
US 20190220733A1 · Fisher · 2019 [cited by examiner]
US 20190378017A1 · Kung · 2019 [cited by examiner]
US 20200065703A1 · Nag · 2020 [cited by examiner]
US 20210256351A1 · Cao · 2021 [cited by examiner]
JP 2008215881A · 2008 [cited by examiner]
Machta et al., “Parameter Space Compression Underlies Emergent Theories and Predictive Models”, ARXIV ID: 1303.6738, Mar. 27, 2013, pp. 1-14. (Year: 2013). [cited by examiner]
Katsoulakis, et al., “Scalable Information Inequalities for Uncertainty Quantification”, ARXIV ID: 1605.04184, May 13, 2016, pp. 1-35. ( Year: 2016). [cited by examiner]
Fisher et al., “Boltzmann Encoded Adversarial Machines”, ARXIV ID: 1804.08682, Apr. 23, 2018, pp. 1-17. (Year: 2018). [cited by examiner]
Zhao et al., “A Confident Information First Principle for Parameter Reduction and Model Selection of Boltzmann Machines”, IEEE Transactions on Neural Networks and Learning Systems, vol. 29, No. 5, May 2018, pp. 1608-162… [cited by examiner]
Faury et al., “Distributionally Robust Counterfactual Risk Minimization”, ARXIV ID: 1906.06211, Jun. 14, 2019, pp. 1-13. (Year: 2019). [cited by examiner]
Goibert et al., “Adversarial Robustness via Adversarial Label-Smoothing”, ARXIV ID: 1906.11567, Jun. 27, 2019, pp. 1-12. (Year: 2019). [cited by examiner]
Han, “Scalable Approximate Inference and Some Applications”, A Thesis Submitted to the Faculty in partial fulllment of the requirements for the degree of Doctor of Philosophy in Computer Science, Dartmouth College, Hano… [cited by examiner]
Dixit, “TMI: Thermodynamic inference of data manifolds”, ARXIV ID: 1911.09776, Nov. 21, 2019, pp. 1-11. (Year: 2019). [cited by examiner]
Kingma et al., “Auto-Encoding Variational Bayes”, arXiv:1312.6114v10, Available Online On: https://arxiv.org/pdf/1312.6114.pdf, May 1, 2014, pp. 1-14. [cited by applicant]
Zhang et al., “The Limitations Of Adversarial Training and The Blind-Spot Attack”, arXiv:1901.04684v1, Available Online On: https://arxiv.org/pdf/1901.04684.pdf, Jan. 15, 2019, pp. 1-16. [cited by applicant]
Tramèr et al., “Adversarial Training and Robustness for Multiple Perturbations”, arXiv:1904.13000v1, Available Online On: https://arxiv.org/pdf/1904.13000v1.pdf, Apr. 30, 2019, pp. 1-22. [cited by applicant]
Belghazi et al., “Mutual Information Neural Estimation”, arXiv:1801.04062v4, Available Online On: https://arxiv.org/pdf/1801.04062.pdf, Jun. 7, 2018, 18 pages. [cited by applicant]
International Search Report and Written Opinion mailed on Oct. 13, 2020, received for PCT Application PCT/JP2020/017115, filed on Apr. 20, 2020, 7 pages including English Translation. [cited by applicant]