IP Library › Granted Patent US 11,837,253
Granted Patent B2
US 11,837,253 · App. 17/449,213 · Granted Dec 5, 2023

Distinguishing user speech from background speech in speech-dense environments

Inventor: David D. Hardek (Allison Park, PA)
Assignee: VOCOLLECT, INC.
G10L25/84G10L15/063G10L15/07G10L15/16G10L25/51G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,837,253
App. No.
17/449,213
Granted
Dec 5, 2023
Kind
B2
Abstract

A device, system, and method whereby a speech-driven system can distinguish speech obtained from users of the system from other speech spoken by background persons, as well as from background speech from public address systems. In one aspect, the present system and method prepares, in advance of field-use, a voice-data file which is created in a training environment. The training environment exhibits both desired user speech and unwanted background speech, including unwanted speech from persons other than a user and also speech from a PA system. The speech recognition system is trained or otherwise programmed to identify wanted user speech which may be spoken concurrently with the background sounds. In an embodiment, during the pre-field-use phase the training or programming may be accomplished by having persons who are training listeners audit the pre-recorded sounds to identify the desired user speech. A processor-based learning system is trained to duplicate the assessments made by the human listeners.

Claims (50)

1. A method of speech recognition, the method comprising:

generating a normalized audio input based on an accessed audio input;

determining if the normalized audio input matches a standardized user voice, the standardized user voice based on a plurality of training samples, wherein a single training word is selected from the plurality of training samples and normalized to generate the standardized user voice;

categorizing the received audio input as a user speech originating from an operator of the SRD in an instance in which it is determined that the normalized audio input matches the standardized user voice;

comparing the received audio input against a template comprising a plurality of samples of user speech in an instance in which it is determined that the normalized audio input does not match the standardized user voice; and

categorizing the received audio input as a background sound in an instance in which it is determined that the received audio input does not match a sample of user speech.

2. The method of claim 1 , wherein the single training word comprises at least one of a vowel sound, a consonant sound, a whole word sound, a phrase sound, a sentence fragment sound, or a whole sentence sound.

3. The method of claim 1 , further comprising generating the normalized audio input by applying a vocal length tract normalization (VLTN).

4. The method of claim 1 , further comprising generating the normalized audio input by applying a maximum likelihood linear regression (MLLR).

5. The method of claim 1 , further comprising:

generating a prompted user speech by prompting a user of the microphone to speak one or more specific words into the microphone;

normalizing the prompted user speech; and

storing the prompted user speech as normalized audio input.

6. The method of claim 1 , further comprising:

matching the normalized audio input to one or more background sample sounds, wherein the background sample sounds comprise at least one of: a speech originating from a person other than the operator of the SRD or a background environment sound; and

categorizing, based on the match of the normalized audio input to one or more background sample sounds, the received audio input as the background sound.

7. The method of claim 1 , further comprising:

converting the received audio input to a frequency domain; and

converting the received audio input in the frequency domain to a state vector.

8. The method of claim 7 , wherein the state vector comprises one or more amplitudes associated with one or more peak frequencies.

9. The method of claim 1 , further comprising digitizing the single training word to comprise multiple digital representations associated with a plurality of voices.

10. The method of claim 9 , wherein the plurality of voices may comprise a digital representation of a standardized voice associated with each gender and the digital representation of the standardized voice associated with at least one gender may be compared against the normalized audio input to categorize the received audio input as being associated with a specific gender.

11. A speech recognition device (SRD) comprising:

a headset comprising:

a microphone for receiving an audio input;

a memory;

a hardware processor, communicatively coupled to the memory and the microphone, configured to:

generate a normalized audio input based on an accessed audio input;

categorize the received audio input as user speech originating from an operator of the SRD in an instance in which it is determined that the normalized audio input matches a standardized user voice, the standardized user voice based on a plurality of training samples, wherein a single training word is selected from the plurality of training samples and normalized to generate the standardized user voice;

compare the received audio input against a template comprising a plurality of samples of user speech in an instance in which it is determined that the normalized audio input does not match the standardized user voice; and

categorize the received audio input as a background sound in an instance in which it is determined that the received audio input does not to match a sample of user speech.

12. The speech recognition device according to claim 11 , wherein the single training word comprises at least one of a vowel sound, a consonant sound, a whole word sound, a phrase sound, a sentence fragment sound, or a whole sentence sound.

13. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to:

generate the normalized audio input by applying a vocal length tract normalization (VLTN).

14. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to:

generate the normalized audio input by applying a maximum likelihood linear regression (MLLR).

15. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to:

generate a prompted user speech by prompting a user of the microphone to speak one or more specific words into the microphone;

normalize the prompted user speech; and

store the prompted user speech as normalized audio input.

16. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to:

match the normalized audio input to one or more background sample sounds, wherein the background sample sounds comprise at least one of a speech originating from a person other than the operator of the SRD or a background environment sound; and

categorize, based on the match of the normalized audio input to one or more background sample sounds, the received audio input as the background sound.

17. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to:

convert the received audio input to a frequency domain; and

convert the received audio input in the frequency domain to a state vector.

18. The speech recognition device according to claim 17 , wherein the conversion of the received audio input in the frequency domain to the state vector comprises amplitudes associated with one or more peak frequencies.

19. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to:

digitize the single training word to comprise multiple digital representations associated with a plurality of voices.

20. The speech recognition device according to claim 19 , wherein the plurality of voices may comprise a digital representation of a standardized voice associated with each gender and the digital representation of the standard voice associated with at least one gender may be compared against the normalized audio input to categorize the received audio input as being associated with a specific gender.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2021
From: HARDEK, DAVID D.
To: VOCOLLECT, INC.
Reel/Frame 057633/0590 →
Continuity (3)
Continuation 16695555 · Nov 26, 2019
Continuation 15220584 · Jul 27, 2016
Related Publication 20220013137A1 · Jan 13, 2022