IP Library Granted Patent US 11,676,600
Granted Patent B2
US 11,676,600 · App. 17/397,522 · Granted Jun 13, 2023

Methods and apparatus for detecting a voice command

Inventors: William F. Ganong, III (Brookline, MA); Paul A. Van Mulbregt (Wayland, MA); Vladimir Sejnoha (Lexington, MA); Glen Wilson (Maynard, MA)
Assignee: CERENCE OPERATING COMPANY
G10L15/22G10L15/02G10L15/30H04M1/724H04W40/005H04W52/0251H04W52/0254H04W52/0261G10L2015/223G10L2015/226H04M2250/74H04W88/02Y02D30/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,676,600
App. No.
17/397,522
Granted
Jun 13, 2023
Kind
B2
Abstract

According to some aspects, a method of monitoring an acoustic environment of a mobile device, at least one computer readable medium encoded with instructions that, when executed, perform such a method and/or a mobile device configured to perform such a method is provided. The method comprises receiving acoustic input from the environment of the mobile device while the mobile device is operating in the low power mode, detecting whether the acoustic input includes a voice command based on performing a plurality of processing stages on the acoustic input, wherein at least one of the plurality of processing stages is performed while the mobile device is operating in the low power mode, and using at least one contextual cue to assist in detecting whether the acoustic input includes a voice command.

Claims (43)

1. A mobile device comprising:

at least one input configured to receive acoustic input from an environment of the mobile device while the mobile device is operating in a low power mode; and

at least one processor configured to

perform, while in the low power mode, at least one processing stage on the acoustic input to evaluate whether the acoustic input includes a voice command, the at least one processing stage including

transmitting at least a portion of the acoustic input from the mobile device to an automatic speech recognition (ASR) server via a network for processing by the ASR server to convert the at least a portion of the acoustic input into a text, and

transmitting the text from the ASR server to a natural language processing (NLP) server for processing by the NLP server to determine whether the text includes a voice command; and

responsive to receiving the voice command from the NLP server, initiate responding to the voice command.

2. The mobile device of claim 1 , wherein the at least one processor is further configured to:

while in the low power mode, pre-process the at least a portion of the acoustic input before transmitting to the ASR server.

3. The mobile device of claim 2 , wherein the pre-process includes at least one of: removing information from the at least a portion of the acoustic input, formatting the at least a portion of the acoustic input; or

modifying the at least a portion of the acoustic input.

4. The mobile device of claim 1 , wherein the at least one processor includes:

a first processor configured to perform the at least one processing stage while the mobile device is in the low power mode; and

a second processor configured to perform an operation corresponding to the voice command after the mobile device exiting from the low power mode.

5. The mobile device of claim 1 , wherein the at least one processor is further configured to perform an operation corresponding to the voice command without exiting from the low power mode.

6. The mobile device of claim 1 , wherein the NLP server is further configured to

process the text using a statistical model to extract semantic entities from the text based on probabilistic patterns observed in prior training inputs to determine the voice command.

7. A system, comprising:

one or more servers remotely connected to a mobile device, the one or more servers are configured to:

responsive to receiving an acoustic input from the mobile device via a mobile network while the mobile device is in a low power mode, process the acoustic input to determine whether the acoustic input include a user utterance,

responsive to determining the acoustic input includes the user utterance, generate a representation data for the user utterance,

process the representation data to determine if the user utterance include a voice command, and

responsive to determining the user utterance includes the voice command, send the voice command to the mobile device via the mobile network to wake up the mobile device from the low power mode and perform an operation corresponding to the voice command.

8. The system of claim 7 , wherein the one or more servers comprises:

a first server configured to process the acoustic input to determine the acoustic input includes the user utterance; and

a second server separated from the first server and configured to process the representation data to determine if the user utterance include a voice command.

9. The system of claim 7 , wherein the representation data includes a textual representation of the user utterance.

10. The system of claim 7 , wherein the representation data includes a sequence of abstract concept without a textual representation of the user utterance.

11. The system of claim 7 , wherein the acoustic input is a result of a pre-processing of an original input by the mobile device before sent to the one or more servers.

12. The system of claim 7 , wherein the pre-processing includes at least one of: removing information from the at least a portion of the original input, formatting the at least a portion of the original input; or modifying the at least a portion of the original input.

13. The system of claim 7 , one or more servers are further configured to process the representation data using a statistical model to extract semantic entities based on probabilistic patterns observed in prior training inputs to determine the voice command.

14. A method for a mobile device, comprising:

receiving, via at least one input, an acoustic input from an environment while the mobile device is operating in a low power mode;

while in the low power mode, pre-processing the at least a portion of the acoustic input, wherein the pre-process includes at least one of: removing information from the at least a portion of the acoustic input, formatting the at least a portion of the acoustic input; or modifying the at least a portion of the acoustic input;

transmitting, via a transceiver, at least a portion of the acoustic input as pre-processed from the mobile device to an automatic speech recognition (ASR) server via a network for processing by the ASR server to convert the at least a portion of the acoustic input into a text;

receiving, via the transceiver, a voice command from a natural language processing (NLP) server separated from the ASR server, wherein the voice command is determined from the text transmitted from the ASR server to the NLP server; and

initiating responding to the voice command.

15. The method of claim 14 , further comprising:

performing, via a first processor, the pre-processing while the mobile device is in the low power mode; and

performing, via a second processor, an operation corresponding to the voice command after the mobile device exiting from the low power mode.

16. The method of claim 14 , further comprising:

performing an operation corresponding to the voice command without exiting from the low power mode.

17. The mobile device of claim 1 , wherein the voice command is determined by the NLP server by processing the text using a statistical model to extract semantic entities from the text based on probabilistic patterns observed in prior training inputs to determine the voice command.

Assignments (2)
RELEASE (REEL 067417 / FRAME 0303) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0422 →
SECURITY AGREEMENT Recorded Apr 15, 2024
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 067417/0303 →