IP Library Granted Patent US 9,514,738
Granted Patent B2
US 9,514,738 · App. 14/442,191 · Granted Dec 6, 2016

Method and device for recognizing speech

Inventor: Yoichi Ando (Kobe, JP)
Assignees: Yoichi Ando; Yoshimasa Electronic Inc.
G10L15/01G10L15/02G10L15/04G10L15/10G10L25/06G10L2015/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,514,738
App. No.
14/442,191
Granted
Dec 6, 2016
Kind
B2
Abstract

A speech is recognized using ACF factors extracted from running autocorrelation functions calculated from the speech. The extracted ACF factors are a W φ(0) (width of ACF amplitude around zero-delay origin), a W φ(0)max (maximum value of the W φ(0) ), a τ 1 (pitch period), a φ 1 (pitch strength), and a Δφ 1 /Δt (rate of the pitch strength change). Syllables in the speech are identified by comparing the ACF factors with templates stored in a database.

Claims (56)

1. A method for recognizing speech, comprising steps of:

recording speech signals by a recording unit;

attenuating, by a LPF, a high frequency component of the speech signals received from the recording unit;

converting, by an AD converter, the speech signals received from the LPF from analog signal to digital signal;

storing, by a storage unit, the speech signals received from the AD converter;

reading out, by a processor, the speech signals stored in the storage unit;

calculating, by the processor, running autocorrelation functions from the speech signals read out from the storage unit;

extracting, by the processor, following ACF factors from the running autocorrelation functions:

a W φ(0) which is a width of ACF amplitude around zero-delay origin,

a W φ(0)max which is a maximum value of the W φ(0) ;

a τ 1 which is a pitch period;

a φ 1 which is a pitch strength; and

a Δφ 1 /Δt which is a rate of the pitch strength change;

identifying, by the processor, syllables in the speech signals by comparing the ACF factors with templates stored in a database.

2. The method according to claim 1 , further comprising the step of segmenting, by the processor, the speech signals into the syllables based on the ACF factors.

3. The method according to claim 1 , further comprising the steps of:

extracting ACF factors from the running autocorrelation functions:

a LL which is a listening level calculated by the amplitude at an origin of non-normalized running ACF;

a LL max which is a maximum value of the LL; and

a (τ e ) min which is a minimum value of effective duration τ e ; and

segmenting, by the processor, the speech signals into the syllables based on the LL, the (τ e ) min , the Δφ 1 /Δt, the τ 1 and the W φ(0) .

4. The method according to claim 1 , wherein the identifying step is performed in each of the syllables at time points after a (τ e ) min .

5. The method according to claim 2 , wherein the identifying step is performed in each of the syllables at time points after a (τ e ) min .

6. The method according to claim 3 , wherein the identifying step is performed in each of the syllables at time points after the (τ e ) min .

7. The method according to claim 1 , wherein the identifying step identifies the syllables in the speech signals based on a total distance between the ACF factors and the templates.

8. The method according to claim 2 , wherein the identifying step identifies the syllables in the speech signals based on a total distance between the ACF factors and the templates.

9. The method according to claim 3 , wherein the identifying step identifies the syllables in the speech signals based on a total distance between the ACF factors and the templates.

10. A speech recognition device, comprising:

a recording unit configured to record speech signals;

a LPF configured to attenuate a high frequency component of the speech signals received from the recording unit;

an AD converter configured to convert the speech signals received from the LPF from analog signal to digital signal;

a storage unit configured to store the speech signals received from the AD converter; and

a processor configured to read out the speech signals stored in the storage unit wherein the processor is configured to:

calculate running autocorrelation functions from the speech signals;

extract following ACF factors from the running autocorrelation functions:

a LL which is a listening level calculated by an amplitude at an origin of non-normalized running ACF;

a LL max which is a maximum value of LL;

a W φ(0) which is a width of ACF amplitude around zero-delay origin,

a W φ(0)max which is a maximum value of the W φ(0) ;

a τ 1 which is a pitch period;

a φ 1 which is a pitch strength; and

a Δφ 1 /Δt which is a rate of the pitch strength change; and

identify syllables in the speech signals by comparing the ACF factors with templates stored in a database.

11. The speech recognition device according to claim 10 , wherein the processor is further configured to segment the speech signals into syllables based on the ACF factors.

12. The speech recognition device according to claim 10 ,

wherein the processor further extracts

a LL which is a listening level from the running autocorrelation functions,

a LL max , which is a maximum value of LL, and

a (τ e ) min which is a minimum value of effective duration τ e , and

further configured to segment the speech signals into syllables based on the LL, the (τ e ) min , the Δφ 1 /Δt, the τ 1 and the W φ(0) .

13. The speech recognition device according to claim 10 , wherein the processor identifies each of the syllables at a time point after a (τ e ) min .

14. The speech recognition device according to claim 11 , wherein the processor identifies each of the syllables at a time point after a (τ e ) min .

15. The speech recognition device according to claim 12 , wherein the processor identifies each of the syllables at a time point after the (τ e ) min .

16. The speech recognition device according claim 10 , wherein the processor identifies the syllables in the speech signals based on a total distance between the ACF factors and the templates.

17. The speech recognition device according claim 11 , wherein the processor identifies the syllables in the speech signals based on a total distance between the ACF factors and the templates.

18. The speech recognition device according claim 12 , wherein the processor identifies the syllables in the speech signals based on a total distance between the ACF factors and the templates.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2015
From: SAKURAI, MASATSUGU
To: ANDO, YOICHI; YOSHIMASA ELECTRONIC INC.
Reel/Frame 035814/0758 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2015
From: ANDO, YOICHI
To: ANDO, YOICHI; SAKURAI, MASATSUGU
Reel/Frame 035617/0304 →
Continuity (1)
Related Publication 20150348536A1 · Dec 3, 2015