IP Library Granted Patent US 8,751,227
Granted Patent B2
US 8,751,227 · App. 12/921,062 · Granted Jun 10, 2014

Acoustic model learning device and speech recognition device

Inventor: Takafumi Koshinaka (Tokyo, JP)
Assignee: NEC Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,751,227
App. No.
12/921,062
Granted
Jun 10, 2014
Kind
B2
Abstract

Parameters of a first variation model, a second variation model and an environment-independent acoustic model are estimated in such a way that an integrated degree of fitness obtained by integrating a degree of fitness of the first variation model to the sample speech data, a degree of fitness of the second variation model to the sample speech data, and a degree of fitness of the environment-independent acoustic model to the sample speech data becomes the maximum. Therefore, when constructing an acoustic model by using sample speech data affected by a plurality of acoustic environments; the effect on a speech which is caused by each of the acoustic environments can be extracted with high accuracy.

Claims (49)

1. An acoustic model learning device comprising:

a first variation model learning unit that estimates a parameter defining a first variation model indicating a variation in a speech for each type of a first environment factor by using a plurality of sample speech data acquired for each combination of one of a plurality of types of the first environment factor and one of plurality of types of a second environment factor, the first environment factor being one of a plurality of environment factors that change and thereby cause a variation in a speech, and the second environment factor being another of the plurality of environment factors;

a second variation model learning unit that, using the plurality of sample speech data, with respect to each type of the second environment factor, estimates a parameter defining a second variation model indicating a variation in a speech; and

an environment-independent acoustic model learning unit that, using the plurality of sample speech data, estimates a parameter defining an environment-independent acoustic model not specified as any type of the first environment factor and the second environment factor, wherein

each of the learning units estimates each parameter in such a way that an integrated degree of fitness obtained by integrating a degree of fitness of the first variation model to the sample speech data, a degree of fitness of the second variation model to the sample speech data, and a degree of fitness of the environment-independent acoustic model to the sample speech data becomes the maximum,

wherein the first variation model and the second variation model are each defined by a two-stage affine transformation.

2. The acoustic model learning device according to claim 1 , wherein each of the learning units uses a probability that the sample speech data is observed, represented by the parameters of the first variation model, the second variation model and the environment-independent acoustic model, as the integrated degree of fitness.

3. The acoustic model learning device according to claim 2 , wherein each of the learning units estimates a parameter by using an iterative method based on any one of a maximum likelihood estimation method, a maximum a posteriori estimation method, and a Bayes estimation method.

4. The acoustic model learning device according to claim 3 , wherein the environment-independent acoustic model is a Gaussian mixture model or a hidden Markov model.

5. The acoustic model learning device according to claim 1 , wherein the environment-independent acoustic model is a Gaussian mixture model or a hidden Markov model.

6. The acoustic model learning device according to claim 1 , wherein each of the learning units estimates a parameter by using an iterative method based on any one of a maximum likelihood estimation method, a maximum a posteriori estimation method, and a Bayes estimation method.

7. The acoustic model learning device according to claim 6 , wherein the environment-independent acoustic model is a Gaussian mixture model or a hidden Markov model.

8. The acoustic model learning device according to claim 1 , wherein the environment-independent acoustic model is a Gaussian mixture model or a hidden Markov model.

9. A speech recognition device comprising:

a speech transformation unit that performs, on speech data as a recognition target acquired through the first environment factor of a given type, inverse transform of the variation indicated by the first variation model corresponding to the given type among first variation models obtained by the acoustic model learning device according to one of claim 1 , wherein

speech recognition is performed on speech data obtained by the speech transformation unit.

10. A speech recognition device comprising:

a speech transformation unit that performs, on speech data as a recognition target acquired through the second environment factor of a given type, inverse transform of the variation indicated by the second variation model corresponding to the given type among second variation models obtained by the acoustic model learning device according to one of claim 1 , wherein

speech recognition is performed on speech data obtained by the speech transformation unit.

11. An acoustic environment recognition device comprising:

a second speech transformation unit that performs, on speech data as a recognition target acquired through the second environment factor of a given type, inverse transform of the variation indicated by the second variation model corresponding to the given type among second variation models obtained by the acoustic model learning device according to one of claim 1 ;

a first speech transformation unit that sequentially performs, on speech data obtained by the second speech transformation unit, inverse transform of the variation indicated by each of first variation models obtained by the acoustic model learning device according to one of claim 1 and obtains a plurality of speech data; and

an identification unit that identifies a type of the first environment factor through which the speech data as a recognition target has passed by using the plurality of speech data obtained by the first speech transformation unit and the environment-independent acoustic model obtained by the acoustic model learning device according to one of claim 1 .

12. The acoustic environment recognition device according to claim 11 , wherein the first environment factor is a speaker, and the second environment factor is a transmission channel.

13. An acoustic model learning method comprising:

a first acoustic model learning step that estimates a parameter defining a first variation model indicating a variation in a speech for each type of a first environment factor by using a plurality of sample speech data acquired for each combination of one of a plurality of types of the first environment factor and one of plurality of types of a second environment factor, the first environment factor being one of a plurality of environment factors that change and thereby cause a variation in a speech, and the second environment factor being another of the plurality of environment factors;

a second variation model learning step that, using the plurality of sample speech data, with respect to each type of the second environment factor, estimates a parameter defining a second variation model indicating a variation in a speech; and

an environment-independent acoustic model learning step that, using the plurality of sample speech data, estimates a parameter defining an environment-independent acoustic model not specified as any type of the first environment factor and the second environment factor, wherein

each of the acoustic model learning steps estimates each parameter in such a way that an integrated degree of fitness obtained by integrating a degree of fitness of the first variation model to the sample speech data, a degree of fitness of the second variation model to the sample speech data, and a degree of fitness of the environment-independent acoustic model to the sample speech data becomes the maximum,

wherein the first variation model and the second variation model are each defined by a two-stage affine transformation.

14. An acoustic model learning method according to claim 13 , wherein each of the acoustic model learning steps uses a probability that the sample speech data is observed, represented by the parameters of the first variation model, the second variation model and the environment-independent acoustic model, as the integrated degree of fitness.

15. A non-transitory computer readable medium that records a program causing a computer to execute a process comprising:

a first acoustic model learning step that estimates a parameter defining a first variation model indicating a variation in a speech for each type of a first environment factor by using a plurality of sample speech data acquired for each combination of one of a plurality of types of the first environment factor and one of plurality of types of a second environment factor, the first environment factor being one of a plurality of environment factors that change and thereby cause a variation in a speech, and the second environment factor being another of the plurality of environment factors;

a second variation model learning step that, using the plurality of sample speech data, with respect to each type of the second environment factor, estimates a parameter defining a second variation model indicating a variation in a speech; and

an environment-independent acoustic model learning step that, using the plurality of sample speech data, estimates a parameter defining an environment-independent acoustic model not specified as any type of the first environment factor and the second environment factor, wherein

each of the acoustic model learning steps estimates each parameter in such a way that an integrated degree of fitness obtained by integrating a degree of fitness of the first variation model to the sample speech data, a degree of fitness of the second variation model to the sample speech data, and a degree of fitness of the environment-independent acoustic model to the sample speech data becomes the maximum,

wherein the first variation model and the second variation model are each defined by a two-stage affine transformation.

16. The non-transitory computer readable medium according to claim 9 , wherein each of the acoustic model learning steps uses a probability that the sample speech data is observed, represented by the parameters of the first variation model, the second variation model and the environment-independent acoustic model, as the integrated degree of fitness.

17. A speech recognition device comprising:

a speech transformation unit that performs, on speech data as a recognition target acquired through the first environment factor of a given type, inverse transform of the variation indicated by the first variation model corresponding to the given type among first variation models obtained by the acoustic model learning device according to claim 2 , wherein

speech recognition is performed on speech data obtained by the speech transformation unit.

18. A speech recognition device comprising:

a speech transformation unit that performs, on speech data as a recognition target acquired through the second environment factor of a given type, inverse transform of the variation indicated by the second variation model corresponding to the given type among second variation models obtained by the acoustic model learning device according to claim 2 , wherein

speech recognition is performed on speech data obtained by the speech transformation unit.

19. An acoustic environment recognition device comprising:

a second speech transformation unit that performs, on speech data as a recognition target acquired through the second environment factor of a given type, inverse transform of the variation indicated by the second variation model corresponding to the given type among second variation models obtained by the acoustic model learning device according to claim 2 ;

a first speech transformation unit that sequentially performs, on speech data obtained by the second speech transformation unit, inverse transform of the variation indicated by each of first variation models obtained by the acoustic model learning device according to claim 2 and obtains a plurality of speech data; and

an identification unit that identifies a type of the first environment factor through which the speech data as a recognition target has passed by using the plurality of speech data obtained by the first speech transformation unit and the environment-independent acoustic model obtained by the acoustic model learning device according to claim 2 .

20. The acoustic environment recognition device according to claim 19 , wherein the first is a speaker, and the second is a transmission channel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2010
From: KOSHINAKA, TAKAFUMI
To: NEC CORPORATION
Reel/Frame 024939/0246 →
Priority Claims (1)
JP 2008-118662 · Apr 30, 2008 · national
Continuity (1)
Related Publication 20110046952A1 · Feb 24, 2011