IP Library Granted Patent US 11,056,097
Granted Patent B2
US 11,056,097 · App. 16/520,104 · Granted Jul 6, 2021

Method and system for generating advanced feature discrimination vectors for use in speech recognition

Inventors: Kevin M. Short (Durham, NH); Brian Hone (Ipswich, MA)
Assignee: XMOS INC.
G10L15/02G10L25/03G10L25/18G10L25/21G10L25/24G10L25/93G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,056,097
App. No.
16/520,104
Granted
Jul 6, 2021
Kind
B2
Abstract

A computer-implemented method of generating advanced feature discrimination vectors (AFDVs) representing sounds forming part of an audio signal input to a device is provided. The method includes taking a plurality of samples of the audio signal, and for each sample of the audio signal taken: performing a signal analysis on the sample to extract one or more high resolution oscillator peaks therefrom; renormalizing the extracted oscillator peaks to eliminate variations in the fundamental frequency and time duration for each sample occurring over the window; normalizing the power of the renormalized extracted oscillator peaks; and forming the renormalized and power normalized extracted oscillator peaks into a respective AFDV for the sample. The method further includes outputting the respective AFDV to a comparison function configured to identify a characteristic of the sample based on a comparison of the respective AFDV with a library of AFDVs associated with known sounds and/or known speakers.

Claims (53)

1. A computer-implemented method of generating advanced feature discrimination vectors (AFDVs) representing sounds forming at least part of an audio signal input to a local device, the local device being in the same environment as at least one source of the audio signal, wherein the method is performed by a controller of the local device and comprises:

taking a plurality of samples of the audio signal, the plurality of samples being a portion of the audio signal as it evolves over a window of predetermined time;

for each sample of the audio signal taken:

performing a signal analysis on the sample to extract one or more high resolution oscillator peaks therefrom, the extracted oscillator peaks forming a spectral representation of the sample;

renormalizing the extracted oscillator peaks to eliminate variations in the fundamental frequency and time duration for each sample occurring over the window;

normalizing the power of the renormalized extracted oscillator peaks; and

forming the renormalized and power normalized extracted oscillator peaks into a respective AFDV for the sample; and

outputting the respective AFDV to a comparison function configured to identify a characteristic of the sample based on a comparison of the respective AFDV with a library of AFDVs associated with known sounds and/or known speakers, wherein the spectral representation of the sample is associated with a single glottal pulse period, and wherein each AFDV in the library also comprises a spectral representation associated with a single glottal pulse period.

2. The method of claim 1 , wherein the library of AFDVs is stored in memory of the local device, and wherein the comparison function is a function implemented by the local device.

3. The method of claim 1 , wherein the comparison function is a function implemented by a remote processor, wherein outputting the respective AFDV comprises transmitting the respective AFDV to the remote processor, and wherein the method comprises receiving the identified characteristic from the remote processor.

4. The method of claim 3 , wherein the respective AFDV is transmitted to the remote processor in an encrypted form.

5. The method of claim 1 , wherein at least some of the AFDVs in the library is associated with a known speech sound, and wherein the method comprises outputting the identified characteristic to a speech recognition engine configured to perform at least one of: translate spoken words into text; control automated systems through voice translation; or convert spoken words into outputs other than voice through an automated process.

6. The method of claim 5 , wherein at least some of the AFDVs in the library is associated with a single word or short phrase, and wherein the speech recognition engine is configured to identify a keyword based on the identified characteristic being a single word or short phrase.

7. The method of claim 6 , comprising using the identified keyword to control at least one of: the local device, an application controlled by the local device, or a remote device or system connected to the local device.

8. The method of claim 1 , wherein at least some of the AFDVs in the library is associated with a known speaker, and wherein the method comprises outputting the identified characteristic to a voice identification engine configured to identify a speaker of the sample.

9. The method of claim 8 , comprising using the identified characteristic to authenticate or verify the identity of the speaker of the sample as part of a security process.

10. The method of claim 1 , wherein the local device is a mobile device.

11. The method of claim 10 , wherein the mobile device is one of: a cellular phone, a smartphone, a tablet or a personal computer.

12. The method of claim 1 , wherein the signal analysis employs complex spectral phase evolution (CSPE).

13. A computer-implemented method of generating advanced feature discrimination vectors (AFDVs) representing sounds forming at least part of an audio signal input to a local device, the local device being in the same environment as at least one source of the audio signal, wherein the method is performed by a controller of the local device and comprises:

taking a plurality of samples of the audio signal, the plurality of samples being a portion of the audio signal as it evolves over a window of predetermined time;

for each sample of the audio signal taken:

performing a signal analysis on the sample to extract one or more high resolution oscillator peaks therefrom, the extracted oscillator peaks forming a spectral representation of the sample;

renormalizing the extracted oscillator peaks to eliminate variations in the fundamental frequency and time duration for each sample occurring over the window;

normalizing the power of the renormalized extracted oscillator peaks; and

forming the renormalized and power normalized extracted oscillator peaks into a respective AFDV for the sample; and

outputting the respective AFDV to a comparison function configured to identify a characteristic of the sample based on a comparison of the respective AFDV with a library of AFDVs associated with known sounds and/or known speakers, wherein said renormalizing of the extracted oscillator peaks comprises representing the oscillator peaks in a common coordinate system, wherein each of the AFDVS in the library is also represented in the common coordinate system, and wherein the common coordinate system comprises a common frequency scale or a common time scale.

14. The method of claim 13 , wherein the representation of each AFDV in the common coordinate system comprises a predetermined number of excitations periods.

15. The method of claim 13 , wherein said representing of the oscillator peaks in the common coordinate system comprises distributing the renormalized extracted oscillator peaks into slots of a comparator stack.

16. The method of claim 13 , wherein the library of AFDVs is represented as an array of AFDVs, and wherein the comparison function is configured to compare the respective AFDV with the array of AFDVs by evaluating a matrix-vector product of the array of AFDVs and the respective AFDV.

17. The method of claim 13 , wherein the library of AFDVs is stored in memory of the local device, and wherein the comparison function is a function implemented by the local device.

18. The method of claim 13 , wherein the comparison function is a function implemented by a remote processor, wherein outputting the respective AFDV comprises transmitting the respective AFDV to the remote processor, and wherein the method comprises receiving the identified characteristic from the remote processor.

19. The method of claim 18 , wherein the respective AFDV is transmitted to the remote processor in an encrypted form.

20. The method of claim 13 , wherein at least some of the AFDVs in the library is associated with a known speech sound, and wherein the method comprises outputting the identified characteristic to a speech recognition engine configured to perform at least one of: translate spoken words into text; control automated systems through voice translation; or convert spoken words into outputs other than voice through an automated process.

21. The method of claim 20 , wherein at least some of the AFDVs in the library is associated with a single word or short phrase, and wherein the speech recognition engine is configured to identify a keyword based on the identified characteristic being a single word or short phrase.

22. The method of claim 21 , comprising using the identified keyword to control at least one of: the local device, an application controlled by the local device, or a remote device or system connected to the local device.

23. The method of claim 13 , wherein at least some of the AFDVs in the library is associated with a known speaker, and wherein the method comprises outputting the identified characteristic to a voice identification engine configured to identify a speaker of the sample.

24. A device configured to generate advanced feature discrimination vectors (AFDVs) representing sounds forming at least part of an audio signal input to the device, the device being in the same environment as at least one source of the audio signal, wherein the device comprises at least one processor configured to:

obtain a plurality of samples of the audio signal, the plurality of samples being a portion of the audio signal as it evolves over a window of predetermined time;

for each sample of the audio signal taken:

perform a signal analysis on the sample to extract one or more high resolution oscillator peaks therefrom, the extracted oscillator peaks forming a spectral representation of the sample;

renormalize the extracted oscillator peaks to eliminate variations in the fundamental frequency and time duration for each sample occurring over the window;

normalize the power of the renormalized extracted oscillator peaks; and

form the renormalized and power normalized extracted oscillator peaks into a respective AFDV for the sample; and

output the respective AFDV to a comparison function configured to identify a characteristic of the sample based on a comparison of the respective AFDV with a library of AFDVs associated with known sounds and/or known speakers, wherein the spectral representation of the sample is associated with a single glottal pulse period, and wherein each AFDV in the library also comprises a spectral representation associated with a single glottal pulse period.

25. A device configured to generate advanced feature discrimination vectors (AFDVs) representing sounds forming at least part of an audio signal input to the device, the device being in the same environment as at least one source of the audio signal, wherein the device comprises at least one processor configured to:

obtain a plurality of samples of the audio signal, the plurality of samples being a portion of the audio signal as it evolves over a window of predetermined time;

for each sample of the audio signal taken:

perform a signal analysis on the sample to extract one or more high resolution oscillator peaks therefrom, the extracted oscillator peaks forming a spectral representation of the sample;

renormalize the extracted oscillator peaks to eliminate variations in the fundamental frequency and time duration for each sample occurring over the window;

normalize the power of the renormalized extracted oscillator peaks; and

form the renormalized and power normalized extracted oscillator peaks into a respective AFDV for the sample; and

output the respective AFDV to a comparison function configured to identify a characteristic of the sample based on a comparison of the respective AFDV with a library of AFDVs associated with known sounds and/or known speakers, wherein said renormalizing of the extracted oscillator peaks comprises representing the oscillator peaks in a common coordinate system, wherein each of the AFDVS in the library is also represented in the common coordinate system, and wherein the common coordinate system comprises a common frequency scale or a common time scale.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2019
From: SHORT, KEVIN M.; HONE, BRIAN T.
To: SETEM TECHNOLOGIES, INC.
Reel/Frame 049838/0290 →
CHANGE OF NAME Recorded Jul 23, 2019
From: SETEM TECHNOLOGIES, INC.
To: XMOS INC.
Reel/Frame 049838/0308 →
Continuity (5)
Continuation 15638627 · Jun 30, 2017
Continuation 14217198 · Mar 17, 2014
Provisional Application 61914002 · Dec 10, 2013
Provisional Application 61786888 · Mar 15, 2013
Related Publication 20200160839A1 · May 21, 2020
Cited By (1)
US 12,400,639