IP Library › Granted Patent US 11,495,235
Granted Patent B2
US 11,495,235 · App. 16/296,410 · Granted Nov 8, 2022

System for creating speaker model based on vocal sounds for a speaker recognition system, computer program product, and controller, using two neural networks

Inventor: Hiroshi Fujimura (Kanagawa, JP)
Assignee: Kabushiki Kaisha Toshiba
G10L17/00G06N3/08G10L15/075G10L15/16G10L17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,495,235
App. No.
16/296,410
Granted
Nov 8, 2022
Kind
B2
Abstract

According to one embodiment, a system for creating a speaker model includes one or more processors. The processors change a part of network parameters from an input layer to a predetermined intermediate layer based on a plurality of patterns and inputs a piece of speech into each of neural networks so as to obtain a plurality of outputs from the intermediate layer. The part of network parameters of the each of the neural networks is changed based on one of the plurality of patterns. The processors create a speaker model with respect to one or more words detected from the speech based on the outputs.

Claims (42)

1. A system for creating a speaker model, the system comprising:

one or more processors configured to:

generate second neural networks, each of the second neural networks being generated by changing a part of network parameters from an input layer to a predetermined intermediate layer of a first neural network based on one of a plurality of patterns, the first neural network being a neural network for detecting one or more words without recognizing a speaker;

input a piece of speech into the each of the second neural networks so as to obtain a plurality of outputs from the intermediate layer; and

create a speaker model that receives a speaker feature comprising vocal sounds as input and outputs a recognized speaker by inputting each of the outputs as the speaker feature.

2. The system for creating a speaker model according to claim 1 , wherein the one or more processors are further configured to:

receive speech and converts the speech into a feature;

input the feature to a neural network and calculate a score that represents likelihood indicating whether the feature corresponds to one or more predetermined words; and

detect the one or more words from the speech using the score.

3. The system for creating a speaker model according to claim 2 , wherein the neural network used for calculating the score is the same as the first neural network.

4. The system for creating a speaker model according to claim 2 , wherein the neural network used for calculating the score is different from the first neural network.

5. The system for creating a speaker model according to claim 1 , wherein the one or more processors create Gaussian distribution represented by a mean and variance of the outputs as the speaker model.

6. The system for creating a speaker model according to claim 1 , wherein the one or more processors create the speaker model by learning using speech of a speaker and the outputs.

7. The system for creating a speaker model according to claim 1 , wherein the one or more processors create the speaker model for each partial section included in the one or more words.

8. The system for creating a speaker model according to claim 1 , wherein the one or more processors change, out of the network parameters from the input layer to the intermediate layer, weight of a part of the network parameters.

9. The system for creating a speaker model according to claim 1 , wherein the one or more processors add, out of the network parameters from the input layer to the intermediate layer, a random value to bias of a part of the network parameters.

10. The system for creating a speaker model according to claim 1 , wherein the network parameters include bias term parameters with respect to an input value to each layer from the input layer to the intermediate layer, and

the one or more processors add a random value to a part of the bias term parameters.

11. A recognition system comprising:

one or more processors configured to:

receive speech and converts the speech into a feature;

input the feature to a first neural network and calculate a score that represents likelihood indicating whether the feature corresponds to one or more predetermined words;

detect the one or more words from the speech using the score;

generate second neural networks, each of the second neural networks being generated by changing a part of network parameters from an input layer to a predetermined intermediate layer of the first neural network based on one of a plurality of patterns, the first neural network being a neural network for detecting one or more words without recognizing a speaker,

input a piece of the speech into the each of the second neural networks so as to obtain a plurality of outputs from the intermediate layer;

create a speaker model that receives a speaker feature comprising vocal sounds as input and outputs a recognized speaker by inputting each of the outputs as the speaker feature; and

recognize the speaker using the speaker model.

12. The recognition system according to claim 11 , wherein the one or more processors input the outputs from the intermediate layer with respect to speech input for recognition to the speaker model so as to recognize the speaker.

13. A computer program product having a non-transitory computer readable medium including programmed instructions stored therein, wherein the instructions, when executed by a computer, cause the computer to perform:

generating second neural networks, each of the second neural networks being generated by changing a part of network parameters from an input layer to a predetermined intermediate layer of a first neural network based on one of a plurality of patterns, the first neural network being a neural network for detecting one or more words without recognizing a speaker,

inputting a piece of speech into the each of the second neural networks so as to obtain a plurality of outputs from the intermediate layer; and

creating a speaker model that receives a speaker feature as input and outputs a recognized speaker by inputting each of the outputs as the speaker feature.

14. A controller comprising:

one or more processors configured to:

generate second neural networks, each of the second neural networks being generated by changing a part of network parameters from an input layer to a predetermined intermediate layer of a first neural network based on one of a plurality of patterns, the first neural network being a neural network for detecting one or more words without recognizing a speaker,

input a piece of speech into the each of the second neural networks so as to obtain a plurality of outputs from the intermediate layer; and

create a speaker model that receives a speaker feature comprising vocal sounds as input and outputs a recognized speaker by inputting each of the outputs as the speaker feature;

acquire speech of a user and detect one or more predetermined words;

determine whether the user is a predetermined user using the speaker model; and

output, when the user is the predetermined user, a control instruction that is defined to the one or more words.

15. The controller according to claim 14 , wherein the one or more processors determine whether the user is a predetermined user using a speaker model obtained by speech of the user and learning of the speech.

16. The controller according to claim 15 , wherein the speaker model is created by learning that uses a plurality of outputs obtained from speech of the user and an extended model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2019
From: FUJIMURA, HIROSHI
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 048910/0205 →
Priority Claims (1)
JP JP2018-118090 · Jun 21, 2018 · national
Continuity (1)
Related Publication 20190392839A1 · Dec 26, 2019