IP Library Granted Patent US 9,959,873
Granted Patent B2
US 9,959,873 · App. 15/418,230 · Granted May 1, 2018

Method for generating unspecified speaker voice dictionary that is used in generating personal voice dictionary for identifying speaker to be identified

Inventor: Misaki Tsujikawa (Osaka, JP)
Assignee: PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AMERICA
G10L17/04G10L17/02G10L17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,959,873
App. No.
15/418,230
Granted
May 1, 2018
Kind
B2
Abstract

A method for generating voice dictionary is disclosed which makes it possible to improve the accuracy of speaker identification. A method according to an aspect of the present disclosure includes: acquiring voices of a plurality of unspecified speakers; acquiring noise in a predetermined place; superimposing the noise onto the voices of the plurality of unspecified speakers; and generating, on the basis of the features of the voices of the plurality of unspecified speakers, unspecified speaker voice dictionary that is used for generating personal voice dictionary for identifying a target speaker.

Claims (39)

1. A method comprising:

acquiring, via a processor, voices of a plurality of unspecified speakers;

acquiring, via the processor, noise in a predetermined place;

adjusting, via the processor, a sound pressure of the noise based on sound pressures of the voices of the plurality of unspecified speakers;

superimposing, via the processor, the noise whose sound pressure has been adjusted onto the voices of the plurality of unspecified speakers; and

generating, via the processor, an unspecified speaker voice dictionary from features of the voices of the plurality of unspecified speakers onto which the noise has been superimposed,

wherein the unspecified speaker voice dictionary is used in generating a personal voice dictionary for identifying a target speaker.

2. The method according to claim 1 , further comprising adjusting, via the processor, the sound pressure of the noise so that a sound pressure difference between (i) an average sound pressure of the voices of the plurality of unspecified speakers and (ii) the sound pressure of the noise takes is a predetermined value.

3. The method according to claim 2 , further comprising:

acquiring, via the processor, a first voice of the target speaker in a process of learning the personal voice dictionary;

generating, via the processor, the personal voice dictionary using the acquired voice of the target speaker and the generated unspecified speaker voice dictionary;

acquiring, via the processor, a second voice of the target speaker in a process of identifying the target speaker;

identifying, via the processor, the target speaker using the generated personal voice dictionary and the acquired second voice of the target speaker; and

changing, via the processor, the predetermined value to be larger when the identifying the target speaker fails.

4. The method according to claim 1 , further comprising:

acquiring, via the processor, the voices of the plurality of unspecified speakers from a first memory storing the voices of the plurality of unspecified speakers in advance; and

acquiring, via the processor, the noise from a second memory storing the noise in advance.

5. The method according to claim 4 , further comprising:

collecting, via the processor, noise of an environment surrounding a place where the target speaker is identified; and

storing the collected noise in the second memory.

6. The method according to claim 1 , further comprising:

acquiring, via the processor, a plurality of noises having different frequency characteristics; and

superimposing, via the processor, the plurality of noises onto the voices of the plurality of unspecified speakers.

7. An apparatus comprising:

a processor; and

a memory storing therein a computer program, which when executed by the processor, causes the processor to perform operations including:

acquiring voices of a plurality of unspecified speakers;

acquiring noise in a predetermined place;

adjusting a sound pressure of the noise based on sound pressures of the voices of the plurality of unspecified speakers;

superimposing the noise whose sound pressure has been adjusted onto the voices of the plurality of unspecified speakers; and

generating an unspecified speaker voice dictionary from features of the voices of the plurality of unspecified speakers onto which the noise has been superimposed,

wherein the unspecified speaker voice dictionary is used in generating personal voice dictionary for identifying a target speaker.

8. A non-transitory recording medium storing thereon a computer program, which when executed by a processor, causes the processor to perform operations comprising:

acquiring voices of a plurality of unspecified speakers;

acquiring noise in a predetermined place;

adjusting a sound pressure of the noise based on sound pressures of the voices of the plurality of unspecified speakers;

superimposing the noise whose sound pressure has been adjusted onto the voices of the plurality of unspecified speakers; and

generating an unspecified speaker voice dictionary from features of the voices of the plurality of unspecified speakers onto which the noise has been superimposed,

wherein the unspecified speaker voice dictionary is used in generating personal voice dictionary for identifying a target speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2017
From: TSUJIKAWA, MISAKI
To: PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AMERICA
Reel/Frame 041670/0010 →
Priority Claims (1)
JP 2016-048243 · Mar 11, 2016 · national
Continuity (1)
Related Publication 20170263257A1 · Sep 14, 2017