IP Library › Granted Patent US 10,714,121
Granted Patent B2
US 10,714,121 · App. 15/220,584 · Granted Jul 14, 2020

Distinguishing user speech from background speech in speech-dense environments

Inventor: David D. Hardek (Allison Park, PA)
Assignee: VOCOLLECT, INC.
G10L25/84G10L15/063G10L15/07G10L15/16G10L25/51G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,714,121
App. No.
15/220,584
Granted
Jul 14, 2020
Kind
B2
Abstract

A device, system, and method whereby a speech-driven system can distinguish speech obtained from users of the system from other speech spoken by background persons, as well as from background speech from public address systems. In one aspect, the present system and method prepares, in advance of field-use, a voice-data file which is created in a training environment. The training environment exhibits both desired user speech and unwanted background speech, including unwanted speech from persons other than a user and also speech from a PA system. The speech recognition system is trained or otherwise programmed to identify wanted user speech which may be spoken concurrently with the background sounds. In an embodiment, during the pre-field-use phase the training or programming may be accomplished by having persons who are training listeners audit the pre-recorded sounds to identify the desired user speech. A processor-based learning system is trained to duplicate the assessments made by the human listeners.

Claims (50)

1. A method of speech recognition, the method comprising:

receiving at a microphone of a speech recognition device (SRD) an audio input;

identifying, via a hardware processor of the SRD which is communicatively coupled with the microphone, a vocalization of a human language in the received audio input; and

categorizing, via the hardware processor, the received vocalization of a human language, based on a speech difference between the received vocalization of human language and an audio mix which is stored in a memory of the SRD, the memory communicatively coupled with the hardware processor, as either one of:

a first vocalization originating from a user of the SRD in response to determining that an absolute value of the speech difference is less than a speech rejection threshold stored in the memory; or

a second vocalization originating from a background, in response to determining that the absolute value of the speech difference is greater than the speech rejection threshold, wherein the background speech is a non-user background speech in the vicinity of the SRD or a background environmental sound.

2. The method of claim 1 , further comprising:

categorizing the received vocalization based on a comparison of the received vocalization with a stored audio characterization model which is stored in the memory of the SRD, the memory communicatively coupled with the hardware processor, wherein the stored audio characterization model comprises:

a plurality of user speech samples; and

a background speech sample.

3. The method of claim 1 , wherein the audio mix comprises a concurrent sound of:

a user speech sample of the plurality of user speech samples; and

the background speech sample.

4. The method of claim 1 , wherein the stored speech rejection threshold comprises a pre-determined threshold calculated based on at least one of:

training samples of user speech; and

training samples of background speech.

5. The method of claim 4 , further comprising:

dynamically updating the speech rejection threshold, by the processor during field-use of the SRD, based upon the first vocalization originating from the user of the SRD.

6. The method of claim 2 , further comprising:

normalizing, via the hardware processor, the received vocalization; and

comparing the normalized received vocalization against a stored normalized user speech.

7. The method of claim 1 , further comprising:

determining, via the hardware processor, if the received vocalization matches an expected verbalization from among a vocabulary of one or more expected verbalizations stored in the memory;

upon determining that the received vocalization matches the expected verbalization, categorizing, via the processor, the received vocalization as the first vocalization originating from the user of the SRD; and

upon determining that the received vocalization does not match any of the one or more expected verbalizations, categorizing, via the processor, the received vocalization based on a comparison of the received vocalization with a stored characterization of user speech.

8. A speech recognition device (SRD), comprising:

a microphone for receiving an audio input;

a memory;

a hardware processor, communicatively coupled to the memory and the microphone, configured to:

identify a vocalization of a human language in the received audio input; and

categorize the received vocalization of a human language based on a speech difference between the received vocalization of human language and an audio mix which is stored in a memory of the SRD, the memory communicatively coupled with the hardware processor, as either one of:

a first vocalization originating from a user of the SRD, in response to determining that an absolute value of the speech difference is less than a speech rejection threshold stored in the memory; or

a second vocalization originating from a background speech, in response to determining that the absolute value of the speech difference is greater than the speech rejection threshold, wherein the background speech is a non-user background speech in the vicinity of the SRD or a background environmental sound.

9. The speech recognition device according to claim 8 , wherein the hardware processor is further configured to categorize the received vocalization based on a comparison of the received vocalization with a stored audio characterization model which is stored in the memory of the SRD, wherein the stored audio characterization model comprises:

a plurality of user speech samples; and

a background speech sample.

10. The speech recognition device according to claim 8 , wherein the audio mix comprises a concurrent sound of:

a user speech sample of the plurality of user speech samples; and

the background speech sample.

11. The speech recognition device according to claim 8 , wherein the stored speech rejection threshold comprises a pre-determined threshold calculated based on at least one of:

training samples of user speech; and

training samples of background speech.

12. The speech recognition device according to claim 11 , wherein the hardware processor is further configured to dynamically update the speech rejection threshold during field-use of the SRD, based upon the first vocalization originating from the user of the SRD.

13. The speech recognition device according to claim 9 , wherein the hardware processor is further configured to:

normalize the received vocalization; and

compare the normalized received vocalization against a stored normalized user speech.

14. The speech recognition device according to claim 8 , wherein the hardware processor is further configured to:

determine if the received vocalization matches an expected verbalization from among a vocabulary of one or more expected verbalizations stored in the memory;

upon determining that the received vocalization matches the expected verbalization, categorize the received vocalization as the first vocalization originating from the user of the SRD; and

upon determining that the received vocalization does not match any of the one or more expected verbalizations, categorize the received vocalization based on a comparison of the received vocalization with a stored characterization of user speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2016
From: HARDEK, DAVID D.
To: VOCOLLECT, INC.
Reel/Frame 039268/0149 →
Continuity (1)
Related Publication 20180033454A1 · Feb 1, 2018
Cited By (1)
US 12,400,678