IP Library Granted Patent US 11,222,641
Granted Patent B2
US 11,222,641 · App. 16/576,170 · Granted Jan 11, 2022

Speaker recognition device, speaker recognition method, and recording medium

Inventor: Kousuke Itakura (Osaka, JP)
Assignee: PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AMERICA
G10L17/20G06N3/04G06N3/08G10L17/00G10L17/02G10L17/04G10L17/06G10L17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,222,641
App. No.
16/576,170
Granted
Jan 11, 2022
Kind
B2
Abstract

A speaker recognition device includes: a feature calculator that calculates two or more acoustic features of a voice of an utterance obtained; a similarity calculator that calculates two or more similarities, each being a similarity between one of one or more speaker-specific features of a target speaker for recognition and one of the two or more acoustic features; a combination unit that combines the two or more similarities to obtain a combined value; and a determiner that determines whether a speaker of the utterance is the target speaker based on the combined value. Here, (i) at least two of the two or more acoustic features have different properties, (ii) at least two of the two or more similarities have different properties, or (iii) at least two of the two or more acoustic features have different properties and at least two of the two or more similarities have different properties.

Claims (47)

1. A speaker recognition device, comprising:

a feature calculator that calculates two or more acoustic features of a voice of an utterance obtained;

a similarity calculator that calculates two or more similarities, each being a similarity between one of one or more speaker-specific features of a target speaker for recognition and one of the two or more acoustic features calculated by the feature calculator;

a combination unit that combines the two or more similarities calculated by the similarity calculator to obtain a combined value; and

a determiner that determines whether a speaker of the utterance is the target speaker for recognition based on the combined value obtained by the combination unit,

wherein (i) at least two of the two or more acoustic features have different properties, (ii) at least two of the two or more similarities have different properties, or (iii) at least two of the two or more acoustic features have different properties and at least two of the two or more similarities have different properties,

the at least two of the two or more similarities are a first similarity and a second similarity having different properties,

the first similarity is calculated from a first acoustic feature by probabilistic linear discriminant analysis by use of a trained calculation model that has been trained with a feature of the target speaker including how the target speaker speaks and that is used to calculate a first speaker-specific feature that is one of the one or more speaker-specific features, the first acoustic feature being one of the two or more acoustic features calculated by the feature calculator, and

the second similarity is calculated as a cosine distance between a second speaker-specific feature that is one of the one or more speaker-specific features and a second acoustic feature that is one of the two or more acoustic features calculated by the feature calculator.

2. The speaker recognition device according to claim 1 ,

wherein the first acoustic feature and the second acoustic feature have different properties,

the first acoustic feature is calculated by the feature calculator by applying linear transformation on a physical quantity of the voice of the utterance by use of an i-Vector, and

the second acoustic feature is calculated by the feature calculator by applying non-linear transformation on the physical quantity of the voice by use of a deep neural network (DNN).

3. The speaker recognition device according to claim 1 ,

wherein the first acoustic feature and the second acoustic feature have different properties,

the first acoustic feature is calculated by the feature calculator by applying non-linear transformation by use of a first model of a DNN,

the second acoustic feature is calculated by the feature calculator by applying non-linear transformation by use of a second model of the DNN that is different in property from the first model,

the first model is a model trained with first training data that includes a voice of the target speaker for recognition in a noise environment at or higher than a threshold level, and

the second model is a model trained with second training data that includes a voice of the target speaker for recognition in a noise environment below the threshold level.

4. The speaker recognition device according to claim 1 ,

wherein the first acoustic feature and the second acoustic feature are identical.

5. The speaker recognition device according to claim 1 ,

wherein the combination unit combines the two or more similarities by adding scores representing the two or more similarities calculated by the similarity calculator.

6. The speaker recognition device according to claim 1 ,

wherein the combination unit combines the two or more similarities calculated by the similarity calculator by normalizing the two or more similarities to cause a mean value to be zero and a variance to be one, and adding the two or more similarities having been normalized.

7. The speaker recognition device according to claim 1 ,

wherein the combination unit combines the two or more similarities calculated by the similarity calculator by normalizing the two or more similarities to cause a mean value to be zero and a variance to be one, and calculating a weighted sum of the two or more similarities having been normalized.

8. The speaker recognition device according to claim 7 ,

wherein the combination unit calculates the weighted sum by multiplying by a greater coefficient as a temporal length of the utterance obtained is longer.

9. A speaker recognition method performed by a computer, the speaker recognition method comprising:

calculating two or more acoustic features of a voice of an utterance obtained;

calculating two or more similarities, each being a similarity between one of one or more speaker-specific features of a target speaker for recognition and one of the two or more acoustic features calculated in the calculating of the two or more acoustic features;

combining the two or more similarities calculated in the calculating of the two or more similarities to obtain a combined value; and

determining whether a speaker of the utterance is the target speaker for recognition based on the combined value obtained in the combining,

wherein (i) at least two of the two or more acoustic features have different properties, (ii) at least two of the two or more similarities have different properties, or (iii) at least two of the two or more acoustic features have different properties and at least two of the two or more similarities have different properties,

the at least two of the two or more similarities are a first similarity and a second similarity having different properties,

the first similarity is calculated from a first acoustic feature by probabilistic linear discriminant analysis by use of a trained calculation model that has been trained with a feature of the target speaker including how the target speaker speaks and that is used to calculate a first speaker-specific feature that is one of the one or more speaker-specific features, the first acoustic feature being one of the two or more acoustic features calculated in the calculating of the two or more acoustic features, and

the second similarity is calculated as a cosine distance between a second speaker-specific feature that is one of the one or more speaker-specific features and a second acoustic feature that is one of the two or more acoustic features calculated in the calculating of the two or more acoustic features.

10. A non-transitory computer-readable recording medium for use in a computer, the recording medium having a computer program recorded thereon for causing the computer to execute:

calculating two or more acoustic features of a voice of an utterance obtained;

calculating two or more similarities, each being a similarity between one of one or more speaker-specific features of a target speaker for recognition and one of the two or more acoustic features calculated in the calculating of the two or more acoustic features;

combining the two or more similarities calculated in the calculating of the two or more similarities to obtain a combined value; and

determining whether a speaker of the utterance is the target speaker for recognition based on the combined value obtained in the combining,

wherein (i) at least two of the two or more acoustic features have different properties, (ii) at least two of the two or more similarities have different properties, or (iii) at least two of the two or more acoustic features have different properties and at least two of the two or more similarities have different properties,

the at least two of the two or more similarities are a first similarity and a second similarity having different properties,

the first similarity is calculated from a first acoustic feature by probabilistic linear discriminant analysis by use of a trained calculation model that has been trained with a feature of the target speaker including how the target speaker speaks and that is used to calculate a first speaker-specific feature that is one of the one or more speaker-specific features, the first acoustic feature being one of the two or more acoustic features calculated in the calculating of the two or more acoustic features, and

the second similarity is calculated as a cosine distance between a second speaker-specific feature that is one of the one or more speaker-specific features and a second acoustic feature that is one of the two or more acoustic features calculated in the calculating of the two or more acoustic features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2020
From: ITAKURA, KOUSUKE
To: PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AMERICA
Reel/Frame 051622/0647 →
Priority Claims (1)
JP JP2019-107341 · Jun 7, 2019 · national
Continuity (2)
Provisional Application 62741712 · Oct 5, 2018
Related Publication 20200111496A1 · Apr 9, 2020