IP Library › Granted Patent US 12,057,139
Granted Patent B2
US 12,057,139 · App. 18/328,034 · Granted Aug 6, 2024

Distinguishing user speech from background speech in speech-dense environments

Inventor: David D. Hardek (Allison Park, PA)
Assignee: VOCOLLECT, INC.
G10L25/84G10L15/063G10L15/07G10L15/16G10L25/51G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,057,139
App. No.
18/328,034
Granted
Aug 6, 2024
Kind
B2
Abstract

A device, system, and method whereby a speech-driven system can distinguish speech obtained from users of the system from other speech spoken by background persons, as well as from background speech from public address systems. In one aspect, the present system and method prepares, in advance of field-use, a voice-data file which is created in a training environment. The training environment exhibits both desired user speech and unwanted background speech, including unwanted speech from persons other than a user and also speech from a PA system. The speech recognition system is trained or otherwise programmed to identify wanted user speech which may be spoken concurrently with the background sounds. In an embodiment, during the pre-field-use phase the training or programming may be accomplished by having persons who are training listeners audit the pre-recorded sounds to identify the desired user speech. A processor-based learning system is trained to duplicate the assessments made by the human listeners.

Claims (62)

1. A method of speech recognition for use in warehousing picking operations, the method comprising:

receiving, at a microphone of a speech recognition device (SRD), an audio input;

identifying, via a hardware processor of the SRD which is communicatively coupled with the microphone, a vocalization in the received audio input;

categorizing, via the hardware processor, the received vocalization, based on a model stored in a memory of the SRD, the memory communicatively coupled with the hardware processor, as at least one of a first vocalization originating from a user of the SRD and a second vocalization originating from a background, wherein the background is at least one of non-user background speech and a background noise;

filtering the vocalization in the received audio input via the SRD to remove at least a portion of the second vocalization in an instance in which the vocalization is categorized as comprising the at least one of the second vocalization originating from a background of the SRD to create a filtered vocalization; and

processing the first vocalization in the filtered vocalization to generate at least one of words and phrases.

2. The method of claim 1 , wherein the model comprises:

a plurality of user speech samples; and

a background noise sample.

3. The method of claim 1 , wherein the second vocalization originating from a background is categorized based on the model having a pre-determined threshold calculated based on at least one of training samples of user speech, natural language, and training samples of background noise.

4. The method of claim 3 , further comprising dynamically updating the pre-determined threshold based upon the first vocalization originating from the user.

5. The method of claim 4 , further comprising normalizing the received vocalization; and comparing the normalized received vocalization against a stored normalized user speech.

6. The method of claim 1 , further comprising:

determining if the received vocalization matches an expected verbalization from among a vocabulary of one or more expected verbalizations stored in the model;

upon determining that the received vocalization matches the expected verbalization, categorizing the received vocalization as the first vocalization originating from the user; and

upon determining that the received vocalization does not match any of the one or more expected verbalizations, categorizing the received vocalization based on an understanding of a human language.

7. The method of claim 1 , wherein the second vocalization originating from a background is categorized based on a signal processing technique related to a difference between the first vocalization and the second vocalization.

8. The method of claim 1 , wherein the model is a machine learning model that is at least one of a neural network system, a support vector machine, and an inductive logic system.

9. The method of claim 1 , further comprising receiving a confirmation phrase, wherein the confirmation phrase is at least one of a confirmation of a recognition of a prompt, a confirmation of a completion of task, and a confirmation of an identification of at least of a location and an object.

10. The method of claim 1 , wherein the microphone is mounted on a headset and is configured for use in a warehouse environment, and wherein the background noise comprises at least one of an operation of vehicles in a warehouse and a movement of pallets in the warehouse.

11. A method of speech recognition for use in warehousing picking operations, the method comprising:

receiving, at a microphone of a speech recognition device (SRD) an audio input, wherein the audio input is related to a confirmation phrase uttered in response to an instruction to pick an item;

identifying, via a hardware processor of the SRD which is communicatively coupled with the microphone, a vocalization in the received audio input;

categorizing, via the hardware processor, the received vocalization of a human language, based on a model stored in a memory of the SRD, the memory communicatively coupled with the hardware processor, as at least one of a first vocalization originating from a user of the SRD and a second vocalization originating from a background, wherein the background is at least one of non-user background speech and a background noise;

filtering the vocalization in the received audio input via the SRD to remove at least a portion of the second vocalization in an instance in which the vocalization is categorized as comprising the at least one of the second vocalization originating from the user of the SRD to create a filtered vocalization; and

processing the first vocalization in the filtered vocalization to generate at least one of words and phrases.

12. The method of claim 11 , wherein the model comprises:

a plurality of user speech samples; and

a background noise sample.

13. The method of claim 11 , wherein the second vocalization originating from a background is categorized based on the model having a pre-determined threshold calculated based on at least one of training samples of user speech, natural language, and training samples of background noise.

14. The method of claim 13 , further comprising dynamically updating the pre-determined threshold based on the first vocalization originating from the user.

15. The method of claim 14 , further comprising normalizing the received vocalization; and comparing the normalized received vocalization against a stored normalized user speech.

16. The method of claim 11 , further comprising:

determining if the received vocalization matches an expected verbalization from among a vocabulary of one or more expected verbalizations stored in the model;

upon determining that the received vocalization matches the expected verbalization, categorizing the received vocalization as the first vocalization originating from the user; and

upon determining that the received vocalization does not match any of the one or more expected verbalizations, categorizing the received vocalization based on an understanding of the human language.

17. The method of claim 11 , wherein the second vocalization originating from a background is categorized based on a signal processing technique related to a difference between the first vocalization and the second vocalization.

18. The method of claim 11 , wherein the model is at least one of a neural network system, a support vector machine, and an inductive logic system.

19. The method of claim 11 , further comprising receiving a confirmation phrase, wherein the confirmation phrase is at least one of a confirmation of a recognition of a prompt, a confirmation of a completion of task, and a confirmation of an identification of at least of a location and an object.

20. The method of claim 11 , wherein the microphone is mounted on a headset and is configured for use in a warehouse environment, and wherein the background noise comprises at least one of an operation of vehicles in a warehouse and a movement of pallets in the warehouse.

21. A method of speech recognition for use in warehousing picking operations, the method comprising:

receiving at a microphone of a speech recognition device (SRD) an audio input, wherein the audio input is related to a confirmation phrase uttered in response to an instruction to pick an item;

identifying, via a hardware processor of the SRD which is communicatively coupled with the microphone, a vocalization in the received audio input;

categorizing, via the hardware processor, the received vocalization of a human language, based on a machine learning model trained to understand input natural language and stored in a memory of the SRD, the memory communicatively coupled with the hardware processor, as at least one of a first vocalization originating from a user of the SRD and a background noise;

filtering the received vocalization to remove the background noise to create a filtered vocalization; and

processing the first vocalization in the filtered vocalization to generate at least one of words and phrases.

22. The method of claim 21 , wherein the machine learning model is at least one of a neural network system, a support vector machine, and an inductive logic system.

23. The method of claim 21 , further comprising receiving the confirmation phrase, wherein the confirmation phrase is at least one of a confirmation of a recognition of a prompt, a confirmation of a completion of task, and a confirmation of an identification of at least of a location and an object.

24. The method of claim 21 , wherein the microphone is mounted on a headset and is configured for use in a warehouse environment.

25. The method of claim 21 , wherein the background noise comprises at least one of an operation of vehicles in a warehouse and a movement of pallets in the warehouse.

26. The method of claim 21 , wherein the machine learning model is trained to distinguish user speech from the background noise.

27. The method of claim 21 , wherein the machine learning model is trained based on a pre-field use process.

28. A speech recognition device (SRD) for use in warehousing picking operations comprising:

a microphone;

at least one processor; and

at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the SRD to at least:

receive, at the microphone, an audio input, wherein the audio input is related to a confirmation phrase uttered in response to an instruction to pick an item identify a vocalization in the received audio input;

categorize the received vocalization of a human language, based on a machine learning model trained to understand input natural language and stored in the at least one memory, as at least one of a first vocalization originating from a user and a background noise;

filter the received vocalization to remove the background noise to create a filtered vocalization; and

process the first vocalization in the filtered vocalization to generate at least one of words and phrases.

29. The SRD of claim 28 , wherein the machine learning model is at least one of a neural network system, a support vector machine, and an inductive logic system and wherein the machine learning model is trained to distinguish user speech from the background noise.

30. The SRD of claim 28 , wherein the machine learning model is trained based on a pre-field use process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2023
From: HARDEK, DAVID D.
To: VOCOLLECT, INC.
Reel/Frame 063838/0484 →
Continuity (4)
Continuation 17449213 · Sep 28, 2021
Continuation 16695555 · Nov 26, 2019
Continuation 15220584 · Jul 27, 2016
Related Publication 20230317101A1 · Oct 5, 2023