IP Library Granted Patent US 11,094,329
Granted Patent B2
US 11,094,329 · App. 16/134,529 · Granted Aug 17, 2021

Neural network device for speaker recognition, and method of operation thereof

Inventors: Sangha Park (Seoul, KR); Namsoo Kim (Seoul, KR); Hyungyong Kim (Seoul, KR); Sungchan Kang (Hwaseong-si, KR); Cheheung Kim (Yongin-si, KR); Yongseop Yoon (Seoul, KR); Choongho Rhee (Anyang-si, KR); Hyeokki Hong (Suwon-si, KR)
Assignees: SAMSUNG ELECTRONICS CO., LTD.; SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
G10L17/18G06N3/0454G06N3/08G10L17/00G10L17/06G10L25/30G10L15/02G10L15/063G10L15/16G10L17/02G10L17/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,094,329
App. No.
16/134,529
Granted
Aug 17, 2021
Kind
B2
Abstract

Provided are a neural network device and a method of operation thereof. The neural network device for speaker recognition may include: a memory configured to store one or more instructions; and a processor configured to generate a trained second neural network by training a first neural network, for separating a mixed voice signal into individual voice signals by executing the one or more instructions, generate a second neural network by adding at least one layer to the trained first neural network, and generate a trained second neural network by training the second neural network, for separating the mixed voice signal into the individual voice signals and for recognizing a speaker of each of the individual voice signals.

Claims (65)

1. A neural network device for speaker recognition, the neural network device comprising:

a memory configured to store one or more instructions; and

a processor configured, by executing the one or more instructions, to:

generate a trained first neural network by training a first neural network, for separating a training mixed voice signal into individual voice signals,

generate a second neural network by adding at least one layer to the trained first neural network,

generate a trained second neural network by training the second neural network, for separating the training mixed voice signal into the individual voice signals and for recognizing a speaker of each of the individual voice signals,

wherein the processor is further configured to add the at least one layer to the trained first neural network after the trained first neural network is generated, and

wherein the processor is further configured to:

obtain information about the training mixed voice signal of a plurality of speakers as input information for the second neural network,

obtain speaker identification information for an individual voice signal of each of the plurality of speakers as output information for the second neural network,

train the second neural network via the input information and the output information for the second neural network,

obtain feature information for speaker recognition about at least one individual voice signal included in a candidate mixed voice signal by extracting an output vector of a last hidden layer of the trained second neural network, in which the information about the candidate mixed voice signal is input,

compare the feature information with pre-registered feature information for speaker recognition, and

recognize, based on a result of comparing the feature information with pre-registered feature information, at least one speaker of the candidate mixed voice signal using the trained second neural network.

2. The neural network device of claim 1 , wherein the processor is further configured to:

obtain first information about the training mixed voice signal of the plurality of speakers as input information for the first neural network,

obtain second information about an individual voice signal of each of the plurality of speakers as output information for the first neural network, and

train the first neural network via the input information and the output information for the first neural network.

3. The neural network device of claim 1 , wherein the processor is further configured to generate the second neural network by removing a first output layer of the trained first neural network and connecting at least one hidden layer and a second output layer to the trained first neural network.

4. The neural network device of claim 1 , further comprising:

an acoustic sensor configured to sense the candidate mixed voice signal.

5. The neural network device of claim 4 , wherein the acoustic sensor comprises at least one of a wideband microphone, a resonator microphone, and a narrow band resonator microphone array.

6. The neural network device of claim 1 , further comprising:

an acoustic sensor configured to sense a voice signal of the at least one speaker,

wherein the processor is further configured to store the obtained feature information for speaker recognition together with the speaker identification information in the memory to register the at least one speaker.

7. A method of operating a neural network device for speaker recognition, the method comprising:

generating a trained first neural network by training a first neural network, for separating a training mixed voice signal into individual voice signals;

generating a second neural network by adding at least one layer to the trained first neural network;

generating a trained second neural network by training the second neural network, for separating the training mixed voice signal into the individual voice signals and for recognizing a speaker of each of the individual voice signals,

wherein the at least one layer is added to the trained first neural network after the trained first neural network is generated, and

wherein the method further comprises:

obtaining information about the training mixed voice signal of a plurality of speakers as input information for the second neural network,

obtaining speaker identification information for an individual voice signal of each of the plurality of speakers as output information for the second neural network,

training the second neural network via the input information and the output information for the second neural network,

obtaining feature information for speaker recognition about at least one individual voice signal included in a candidate mixed voice signal by extracting an output vector of a last hidden layer of the trained second neural network, in which the information about the candidate mixed voice signal is input,

comparing the feature information with pre-registered feature information for speaker recognition, and

recognizing, based on a result of comparing the feature information with pre-registered feature information, at least one speaker of the candidate mixed voice signal using the trained second neural network.

8. The method of claim 7 , wherein the generating the trained first neural network comprises:

obtaining first information about the training mixed voice signal of the plurality of speakers as input information for the first neural network;

obtaining second information about an individual voice signal of each of the plurality of speakers as output information for the first neural network; and

training the first neural network via the input information and the output information for the first neural network.

9. The method of claim 7 , wherein the generating the second neural network comprises generating the second neural network by removing an output layer of the trained first neural network and connecting at least one hidden layer to the output layer.

10. The method of claim 7 , further comprising:

obtaining the candidate mixed voice signal.

11. The method of claim 10 , wherein the obtaining of the mixed voice signal comprises sensing the mixed voice signal using at least one of a wideband microphone, a resonator microphone, and a narrow band resonator microphone array.

12. The method of claim 7 , further comprising:

obtaining a voice signal of the at least one speaker;

storing the obtained feature information for speaker recognition together with the speaker identification information in the memory to register the at least one speaker.

13. A non-transitory computer-readable recording medium having recorded thereon a program which, when executed by a computer, performs operations comprising:

generating a trained first neural network by training a first neural network, for separating a training mixed voice signal into individual voice signals;

generating a second neural network by adding at least one layer to the trained first neural network;

generating a trained second neural network by training the second neural network, for separating the training mixed voice signal into the individual voice signals and for recognizing a speaker of each of the individual voice signals

wherein the at least one layer is added to the trained first neural network after the trained first neural network is generated, and

wherein the operations further comprise:

obtaining information about the training mixed voice signal of a plurality of speakers as input information for the second neural network,

obtaining speaker identification information for an individual voice signal of each of the plurality of speakers as output information for the second neural network,

training the second neural network via the input information and the output information for the second neural network,

obtaining feature information for speaker recognition about at least one individual voice signal included in a candidate mixed voice signal by extracting an output vector of a last hidden layer of the trained second neural network, in which the information about the candidate mixed voice signal is input,

comparing the feature information with pre-registered feature information for speaker recognition, and

recognizing, based on a result of comparing the feature information with pre-registered feature information, at least one speaker of the candidate mixed voice signal using the trained second neural network.

14. The neural network device of claim 1 , wherein the processor is further configured to generate the trained second neural network after generating the trained first neural network.

15. The neural network device of claim 1 , wherein the processor is further configured to:

generate the trained first neural network based on a first data set comprising first output information; and

generate the trained second neural network based on a second data set comprising second output information different from the first output information.

16. The neural network device of claim 1 , wherein the first output information is individual voice signal information and the second output is speaker identification information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2018
From: PARK, SANGHA; KIM, NAMSOO; KIM, HYUNGYONG; KANG, SUNGCHAN; KIM, CHEHEUNG; YOON, YONGSEOP; RHEE, CHOONGHO; HONG, HYEOKKI
To: SAMSUNG ELECTRONICS CO., LTD.; SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 046903/0250 →
Priority Claims (1)
KR 10-2017-0157507 · Nov 23, 2017 · national
Continuity (1)
Related Publication 20190156837A1 · May 23, 2019