IP Library Granted Patent US 11,640,819
Granted Patent B2
US 11,640,819 · App. 17/085,630 · Granted May 2, 2023

Information processing apparatus and update method

Inventor: Naoshi Matsuo (Yokohama, JP)
Assignee: FUJITSU LIMITED
G10L15/063G06F16/2282G06F16/2379G10L15/02G10L15/16G10L15/22G10L19/038H04M3/51G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,819
App. No.
17/085,630
Granted
May 2, 2023
Kind
B2
Abstract

A non-transitory computer-readable recording medium having stored therein an update program that causes a computer to execute a procedure, the procedure includes calculating a selection rate of each of a plurality of quantization points included in a quantization table, based on quantization data obtained by quantizing features of a plurality of utterance data, and updating the quantization table by updating the plurality of quantization points based on the selection rate.

Claims (50)

1. A non-transitory computer-readable recording medium having stored therein an update program that causes a computer to execute a procedure including voice processing after training of a neural network not identifying words in utterance data, the procedure comprising:

calculating a selection rate of each of a plurality of quantization points included in a quantization table according to white noise included in the quantization table as an initial value, based on quantization data obtained by quantizing features of a plurality of utterance data;

updating the quantization table by updating the plurality of quantization points based on the selection rate;

setting parameters of at least one neural network based on the quantization table;

receiving voice data; and

detecting a result of a specific conversation situation by the at least one neural network based on the voice data after said setting of the parameters.

2. The non-transitory computer-readable recording medium according to claim 1 , wherein the updating of the quantization table includes

excluding, from the quantization table, quantization points whose selection rate is equal to or less than a predetermined reference, among the plurality of quantization points,

adding, to the quantization table, quantization points different from the excluded quantization points, and

updating quantization points other than the quantization points whose selection rate is equal to or less than the predetermined reference, based on a result of a selection of the quantization points.

3. The non-transitory computer-readable recording medium according to claim 2 , wherein the calculating of the selection rate includes:

calculating a distance between each of the quantization data based on the feature of each of the plurality of utterance data and each of the plurality of quantization points included in the quantization table, and

selecting a quantization point of the plurality of quantization points that minimizes the distance, as the result of the selection of each of the quantization point.

4. The non-transitory computer-readable recording medium according to claim 3 , wherein the updating of the quantization table includes:

excluding, from the quantization table, the quantization points whose the selection rate is equal to or less than the predetermined reference,

adding, to the quantization table, quantization points before update, whose selection rate is equal to or more than the predetermined reference, and

updating each of the quantization points whose selection rate is equal to or less than the predetermined reference to an average value of selected quantization data.

5. The non-transitory computer-readable recording medium according to claim 4 ,

the procedure further comprising:

calculating the selection rate of each of the plurality of quantization points based on each of the quantization data generated from a first utterance data of the plurality of utterance data; and

specifying the quantization point that includes a highest selection rate among the plurality of quantization points, as a quantization point equivalent to silence,

wherein the updating of the quantization table includes:

updating the quantization point equivalent to silence to an average value of the selected quantization data, and

updating quantization points other than the quantization point equivalent to silence, based on the selection rate.

6. The non-transitory computer-readable recording medium according to claim 5 , wherein the updating of the quantization table includes:

calculating a quantization error based on quantization points after update excluding the quantization point equivalent to silence,

when the quantization error is equal to or larger than a threshold value,

updating the quantization point equivalent to silence by using a second utterance data of the plurality of utterance data, and

updating quantization points other than the quantization point equivalent to silence, based on the selection rate, and

when the quantization error is smaller than the threshold value, outputting the quantization table after update.

7. The non-transitory computer-readable recording medium according to claim 1 , the procedure further comprising:

generating a quantization result associated with a quantization point corresponding to a feature of the voice information, based on vector quantization on input voice information and the quantization table after update that includes the updated plurality of quantization points; and

performing learning of a model to which a neural network is applied, when the quantization result is input into the model so that output information output from the model approaches correct answer information for indicating whether the voice information corresponding to the quantization result includes a predetermined conversation situation.

8. The non-transitory computer-readable recording medium according to claim 7 , the procedure further comprising: determining whether the predetermined conversation situation is included in utterance data to be determined, based on the output information acquired by inputting quantization data obtained by quantizing the feature of the utterance data to be determined, into the model that has been subjected to the learning.

9. The non-transitory computer-readable recording medium according to claim 1 , wherein the selection rate is a ratio of a total of selected features to a number of selection of the quantization point.

10. The non-transitory computer-readable recording medium according to claim 1 , wherein the features are information extracted from utterance data and to be used for speech recognition.

11. A method comprising:

calculating a selection rate of each of a plurality of quantization points included in a quantization table according to white noise included in the quantization table as an initial value, based on quantization data obtained by quantizing features of a plurality of utterance data and without identifying target data among the utterance data;

updating the quantization table by updating the plurality of quantization points based on the selection rate, by a processor;

setting parameters of at least one neural network based on the quantization table;

receiving voice data; and

detecting a result of a specific conversation situation by the at least one neural network based on the voice data after said setting of the parameters.

12. An information processing apparatus comprising:

a memory; and

a processor coupled to the memory and configured to:

calculate a selection rate of each of a plurality of quantization points included in a quantization table according to white noise included in the quantization table as an initial value, based on quantization data obtained by quantizing features of a plurality of utterance data and independent of a target in the utterance data;

update the quantization table by updating the plurality of quantization points based on the selection rate;

setting parameters of at least one neural network based on the quantization table;

receiving voice data; and

detecting a result of a specific conversation situation by the at least one neural network based on the voice data after said setting of the parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2020
From: MATSUO, NAOSHI
To: FUJITSU LIMITED
Reel/Frame 054228/0032 →
Priority Claims (1)
JP JP2019-233503 · Dec 24, 2019 · national
Continuity (1)
Related Publication 20210193120A1 · Jun 24, 2021