IP Library Granted Patent US 11,790,937
Granted Patent B2
US 11,790,937 · App. 17/303,001 · Granted Oct 17, 2023

Voice detection optimization using sound metadata

Inventors: Connor Kristopher Smith (New Hudson, MI); Kurt Thomas Soto (Ventura, CA); Charles Conor Sleith (Waltham, MA)
Assignee: Sonos, Inc.
G10L25/84G10L21/0208G10L25/03H04R3/00G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,937
App. No.
17/303,001
Granted
Oct 17, 2023
Kind
B2
Abstract

Systems and methods for optimizing voice detection via a network microphone device are disclosed herein. In one example, individual microphones of a network microphone device detect sound. The sound data is captured in a first buffer and analyzed to detect a trigger event. Metadata associated with the sound data is captured in a second buffer and provided to at least one network device to determine at least one characteristic of the detected sound based on the metadata. The network device provides a response that includes an instruction, based on the determined characteristic, to modify at least one performance parameter of the NMD. The NMD then modifies the at least one performance parameter based on the instruction.

Claims (92)

1. A network microphone device (NMD) comprising:

one or more processors;

one or more microphones; and

data storage having instructions thereon that, when executed by the one or more processors, cause the NMD to perform operations comprising:

detecting sound via the one or more microphones;

capturing sound data based on the detected sound, wherein the sound data includes a voice input;

analyzing the sound data to detect a wake word;

capturing metadata associated with the sound data, wherein the voice input is not derivable from the metadata;

transmitting the sound data to one or more remote computing devices to determine an intent based on the voice input;

transmitting the metadata to at least one of the one or more remote computing devices to determine at least one characteristic of the detected sound based on the metadata;

after transmitting the metadata, receiving a response from the at least one of the one or more remote computing devices, wherein the response includes an instruction, based on the determined characteristic, to modify at least one performance parameter of the NMD;

modifying the at least one performance parameter based on the instruction, wherein the modifying comprises at least one of:

adjusting a fixed gain of the NMD;

adjusting a wake-word-detection sensitivity parameter of the NMD;

adjusting a noise-reduction parameter of the NMD;

adjusting an acoustic echo cancellation parameter of the NMD;

adjusting a spatial processing algorithm of the NMD;

adjusting a localization algorithm of the NMD; or

disregarding input from a defective microphone; and

performing a command based on the determined intent.

2. The NMD of claim 1 , wherein transmitting the metadata occurs separately from transmitting the sound data.

3. The NMD of claim 1 , wherein transmitting the sound data comprises transmitting the sound data to one or more computing devices associated with a voice assistant service (VAS).

4. The NMD of claim 3 , wherein transmitting the metadata comprises transmitting the metadata to a remote evaluator distinct from the VAS.

5. The NMD of claim 1 , wherein the metadata comprises one or more of:

microphone frequency response data,

microphone spectral data,

acoustic echo cancellation (AEC) data,

echo return loss enhancement (ERLE) data,

arbitration data,

signal level data, or

direction detection data.

6. The NMD of claim 1 , wherein modifying the at least one performance parameter comprises adjusting one or more parameters of a multi-channel Wiener filter spatial processing algorithm of the NMD.

7. The NMD of claim 1 , wherein determining the at least one characteristic comprises comparing features of the metadata with averaged features of a sample population of sound metadata.

8. A tangible, non-transitory, computer-readable medium storing instructions that, when executed by one or more processors of a network microphone device (NMD), cause the NMD to perform operations comprising:

detecting sound via one or more microphones of the NMD;

capturing sound data based on the detected sound, wherein the sound data includes a voice input;

analyzing the sound data to detect a wake word;

capturing metadata associated with the sound data, wherein the voice input is not derivable from the metadata;

transmitting the sound data to one or more remote computing devices to determine an intent based on the voice input;

transmitting the metadata to at least one of the one or more remote computing devices to determine at least one characteristic of the detected sound based on the metadata;

after transmitting the metadata, receiving a response from the at least one of the one or more remote computing devices, wherein the response includes an instruction, based on the determined characteristic, to modify at least one performance parameter of the NMD;

modifying the at least one performance parameter based on the instruction, wherein the modifying comprises at least one of:

adjusting a fixed gain of the NMD;

adjusting a wake-word-detection sensitivity parameter of the NMD;

adjusting a noise-reduction parameter of the NMD;

adjusting an acoustic echo cancellation parameter of the NMD;

adjusting a spatial processing algorithm of the NMD;

adjusting a localization algorithm of the NMD; or

disregarding input from a defective microphone; and

performing a command based on the determined intent.

9. The computer-readable medium of claim 8 , wherein transmitting the metadata occurs separately from transmitting the sound data.

10. The computer-readable medium of claim 8 , wherein transmitting the sound data comprises transmitting the sound data to one or more computing devices associated with a voice assistant service (VAS).

11. The computer-readable medium of claim 10 , wherein transmitting the metadata comprises transmitting the metadata to a remote evaluator distinct from the VAS.

12. The computer-readable medium of claim 8 , wherein the metadata comprises one or more of:

microphone frequency response data,

microphone spectral data,

acoustic echo cancellation (AEC) data,

echo return loss enhancement (ERLE) data,

arbitration data,

signal level data, or

direction detection data.

13. The computer-readable medium of claim 8 , wherein modifying the at least one performance parameter comprises adjusting one or more parameters of a multi-channel Wiener filter spatial processing algorithm of the NMD.

14. The computer-readable medium of claim 8 , wherein determining the at least one characteristic comprises comparing features of the metadata with averaged features of a sample population of sound metadata.

15. A method performed by a network microphone device (NMD), the method comprising:

detecting sound via one or more microphones of the NMD;

capturing sound data based on the detected sound, wherein the sound data includes a voice input;

analyzing the sound data to detect a wake word;

capturing metadata associated with the sound data, wherein the voice input is not derivable from the metadata;

transmitting the sound data to one or more remote computing devices to determine an intent based on the voice input;

transmitting the metadata to at least one of the one or more remote computing devices to determine at least one characteristic of the detected sound based on the metadata;

after transmitting the metadata, receiving a response from the at least one of the one or more remote computing devices, wherein the response includes an instruction, based on the determined characteristic, to modify at least one performance parameter of the NMD;

modifying the at least one performance parameter based on the instruction, wherein the modifying comprises at least one of:

adjusting a fixed gain of the NMD;

adjusting a wake-word-detection sensitivity parameter of the NMD;

adjusting a noise-reduction parameter of the NMD;

adjusting an acoustic echo cancellation parameter of the NMD;

adjusting a spatial processing algorithm of the NMD;

adjusting a localization algorithm of the NMD; or

disregarding input from a defective microphone; and

performing a command based on the determined intent.

16. The method of claim 15 , wherein transmitting the metadata occurs separately from transmitting the sound data.

17. The method of claim 15 , wherein transmitting the sound data comprises transmitting the sound data to one or more computing devices associated with a voice assistant service (VAS).

18. The method of claim 17 , wherein transmitting the metadata comprises transmitting the metadata to a remote evaluator distinct from the VAS.

19. The method of claim 15 , wherein the metadata comprises one or more of:

microphone frequency response data,

microphone spectral data,

acoustic echo cancellation (AEC) data,

echo return loss enhancement (ERLE) data,

arbitration data,

signal level data, or

direction detection data.

20. The method of claim 15 , wherein modifying the at least one performance parameter comprises adjusting one or more parameters of a multi-channel Wiener filter spatial processing algorithm of the NMD.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: SMITH, CONNOR KRISTOPHER; SOTO, KURT THOMAS; SLEITH, CHARLES CONOR
To: SONOS, INC.
Reel/Frame 056286/0834 →
Continuity (2)
Continuation 16138111 · Sep 21, 2018
Related Publication 20210272586A1 · Sep 2, 2021
Cited By (30)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,231,859 US 12,236,932 US 12,288,558 US 12,314,633 US 12,322,390 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,424,220 US 12,438,977 US 12,462,802 US 12,498,899 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,562,167 US 12,578,779 US 12,579,978 US 12,626,717 US 12,699,543 US 12,711,962