IP Library Granted Patent US 11,335,337
Granted Patent B2
US 11,335,337 · App. 16/675,692 · Granted May 17, 2022

Information processing apparatus and learning method

Inventors: Shoji Hayakawa (Akashi, JP); Shouji Harada (Kawasaki, JP)
Assignee: FUJITSU LIMITED
G10L15/187G10L15/1807G10L15/22G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,335,337
App. No.
16/675,692
Granted
May 17, 2022
Kind
B2
Abstract

An information processing apparatus includes a memory; and a processor coupled to the memory and the processor configured to: generate phoneme string information in which a plurality of phonemes included in voice information is arranged in time series, based on a recognition result of the phonemes for the voice information; and learn parameters of a network such that when the phoneme string information is input to the network, output information that is output from the network approaches correct answer information that indicates whether a predetermined conversation situation is included in the voice information that corresponds to the phoneme string information.

Claims (69)

1. An information processing apparatus, comprising:

a memory; and

a processor coupled to the memory and the processor configured to:

generate phoneme string information in which a plurality of phonemes included in voice information is arranged in time series, based on a recognition result of the phonemes for the voice information; and

learn parameters of a network such that when the phoneme string information is input to the network, output information that is output from the network approaches correct answer information that indicates whether a predetermined conversation situation is included in the voice information, that corresponds to the phoneme string information, by calculating an internal vector of the phoneme string information.

2. The information processing apparatus according to claim 1 , wherein

the network includes a first network that has a recursive path and a second network that has no recursive path, and

the processor is further configured to:

calculate the internal vector by inputting the phoneme string information to the first network;

calculate the output information by inputting information related to the internal vector to the second network; and

learn a first set of parameters of the first network and a second set of parameters of the second network such that the output information approaches the correct answer information.

3. The information processing apparatus according to claim 2 , wherein the first network is a long short term memory (LSTM).

4. The information processing apparatus according to claim 2 , wherein the processor is further configured to:

calculate statistical information of a plurality of internal vectors output from the first network; and

calculate the output information by inputting the statistical information to the second network.

5. The information processing apparatus according to claim 4 , wherein the processor is further configured to:

calculate the statistical information based on weight parameters in a time direction; and

learn the first set of parameters, the second set of parameters, and the weight parameters such that the output information approaches the correct answer information.

6. The information processing apparatus according to claim 4 , wherein the processor is further configured to:

calculate an average vector of the plurality of internal vectors output from the first network; and

calculate the output information by inputting the average vector to the second network.

7. The information processing apparatus according to claim 2 , wherein the processor is further configured to:

extract a feature amount that includes at least one of a stress evaluation value or a conversation time based on the voice information;

generate a connected vector by connecting the internal vector and a vector of the feature amount; and

calculate the output information by inputting the connected vector to the second network.

8. The information processing apparatus according to claim 1 , wherein the processor is further configured to:

set the learned parameters in the network;

generate first phoneme string information in which a plurality of phonemes included in input voice information is arranged in time series, based on a result of recognition of phonemes for the input voice information; and

determine whether a predetermined conversation situation is included in the input voice information, by inputting the first phoneme string information to the network.

9. A learning method, comprising:

generating, by a computer, phoneme string information in which a plurality of phonemes included in voice information is arranged in time series, based on a recognition result of the phonemes for the voice information; and

learning parameters of a network such that when the phoneme string information is input to the network, output information that is output from the network approaches correct answer information that indicates whether a predetermined conversation situation is included in the voice information that corresponds to the phoneme string information, by calculating an internal vector of the phoneme string information.

10. The learning method according to claim 9 , wherein

the network includes a first network that has a recursive path and a second network that has no recursive path, and

the learning method further comprises:

calculating the internal vector by inputting the phoneme string information to the first network;

calculating the output information by inputting information related to the internal vector to the second network; and

learning a first set of parameters of the first network and a second set of parameters of the second network such that the output information approaches the correct answer information.

11. The learning method according to claim 10 , wherein the first network is a long short term memory (LSTM).

12. The learning method according to claim 10 , further comprising:

calculating statistical information of a plurality of internal vectors output from the first network; and

calculating the output information by inputting the statistical information to the second network.

13. A non-transitory computer-readable recording medium having stored therein a program that causes a computer to execute a process, the process comprising:

generating phoneme string information in which a plurality of phonemes included in voice information is arranged in time series, based on a recognition result of the phonemes for the voice information; and

learning parameters of a network such that when the phoneme string information is input to the network, output information that is output from the network approaches correct answer information that indicates whether a predetermined conversation situation is included in the voice information that corresponds to the phoneme string information, by calculating an internal vector of the phoneme string information.

14. The non-transitory computer-readable recording medium according to claim 13 , wherein

the network includes a first network that has a recursive path and a second network that has no recursive path, and

the process further comprises:

calculating the internal vector by inputting the phoneme string information to the first network;

calculating the output information by inputting information related to the internal vector to the second network; and

learning a first set of parameters of the first network and a second set of parameters of the second network such that the output information approaches the correct answer information.

15. The non-transitory computer-readable recording medium according to claim 14 , wherein the first network is a long short term memory (LSTM).

16. The non-transitory computer-readable recording medium according to claim 14 , the process further comprising:

calculating statistical information of a plurality of internal vectors output from the first network; and

calculating the output information by inputting the statistical information to the second network.

17. The non-transitory computer-readable recording medium according to claim 16 , the process further comprising:

calculating the statistical information based on weight parameters in a time direction; and

learning the first set of parameters, the second set of parameters, and the weight parameters such that the output information approaches the correct answer information.

18. The non-transitory computer-readable recording medium according to claim 16 , the process further comprising:

calculating an average vector of the plurality of internal vectors output from the first network; and

calculating the output information by inputting the average vector to the second network.

19. The non-transitory computer-readable recording medium according to claim 14 , the process further comprising:

extracting a feature amount that includes at least one of a stress evaluation value or a conversation time based on the voice information;

generating a connected vector by connecting the internal vector and a vector of the feature amount; and

calculating the output information by inputting the connected vector to the second network.

20. The non-transitory computer-readable recording medium according to claim 13 , the process further comprising:

setting the learned parameters in the network;

generating first phoneme string information in which a plurality of phonemes included in input voice information is arranged in time series, based on a result of recognition of phonemes for the input voice information; and

determining whether a predetermined conversation situation is included in the input voice information, by inputting the first phoneme string information to the network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2019
From: HAYAKAWA, SHOJI; HARADA, SHOUJI
To: FUJITSU LIMITED
Reel/Frame 050992/0296 →
Priority Claims (1)
JP JP2018-244932 · Dec 27, 2018 · national
Continuity (1)
Related Publication 20200211535A1 · Jul 2, 2020