IP Library Patent Application 19030415
Patent Application
App. No. 19/030,415

VOICE DETECTION OPTIMIZATION USING SOUND METADATA

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/030,415
Abstract

Systems and methods for optimizing voice detection via a network microphone device are disclosed herein. In one example, individual microphones of a network microphone device detect sound. The sound data is captured in a first buffer and analyzed to detect a trigger event. Metadata associated with the sound data is captured in a second buffer and provided to at least one network device to determine at least one characteristic of the detected sound based on the metadata. The network device provides a response that includes an instruction, based on the determined characteristic, to modify at least one performance parameter of the NMD. The NMD then modifies the at least one performance parameter based on the instruction.

Claims (64)

1 . A network microphone device (NMD) comprising:

one or more processors;

one or more microphones; and

data storage having instructions thereon that, when executed by the one or more processors, cause the NMD to perform operations comprising:

at a first time, capturing first sound data via the one or more microphones, wherein the first sound data includes a voice input;

analyzing the first sound data to detect a wake word;

obtaining an intent based on the voice input;

performing a command based on the intent;

at a second time, capturing second sound data via the one or more microphones;

capturing metadata associated with the second sound data, wherein the second sound data is not derivable from the metadata; and

based on the metadata, modifying at least one performance parameter of the NMD.

2 . The NMD of claim 1 , wherein the second sound data includes a voice utterance.

3 . The NMD of claim 1 , wherein the second sound data includes a wake word portion of a voice input.

4 . The NMD of claim 1 , wherein obtaining the intent comprises:

transmitting the first sound data to one or more remote computing devices associated with a voice assistant service (VAS) to determine the intent based on the voice input; and

receiving a response from at least one of the one or more remote computing devices associated with the VAS, wherein the response is indicative of the determined intent.

5 . The NMD of claim 1 , wherein modifying the at least one performance parameter comprises modifying audio playback characteristics of the NMD.

6 . The NMD of claim 1 , wherein modifying the at least one performance parameter comprises at least one of:

adjusting a fixed gain of the NMD;

adjusting a wake-word detection sensitivity parameter of the NMD;

adjusting a noise-reduction parameter of the NMD;

adjusting an acoustic echo cancellation parameter of the NMD;

adjusting a spatial processing algorithm of the NMD;

adjusting a localization algorithm of the NMD; or

disregarding input from a defective microphone.

7 . The NMD of claim 1 , wherein the operations further comprise determining at least one characteristic of the second sound data based on the metadata.

8 . A method to be performed by a network microphone device (NMD), the method comprising:

at a first time, capturing first sound data via one or more microphones of the NMD, wherein the first sound data includes a voice input;

analyzing the first sound data to detect a wake word;

obtaining an intent based on the voice input;

performing a command based on the intent;

at a second time, capturing second sound data via the one or more microphones;

capturing metadata associated with the second sound data, wherein the second sound data is not derivable from the metadata; and

based on the metadata, modifying at least one performance parameter of the NMD.

9 . The method of claim 8 , wherein the second sound data includes a voice utterance.

10 . The method of claim 8 , wherein the second sound data includes a wake word portion of a voice input.

11 . The method of claim 8 , wherein obtaining the intent comprises:

transmitting the first sound data to one or more remote computing devices associated with a voice assistant service (VAS) to determine the intent based on the voice input; and

receiving a response from at least one of the one or more remote computing devices associated with the VAS, wherein the response is indicative of the determined intent.

12 . The method of claim 8 , wherein modifying the at least one performance parameter comprises modifying audio playback characteristics of the NMD.

13 . The method of claim 8 , wherein modifying the at least one performance parameter comprises at least one of:

adjusting a fixed gain of the NMD;

adjusting a wake-word detection sensitivity parameter of the NMD;

adjusting a noise-reduction parameter of the NMD;

adjusting an acoustic echo cancellation parameter of the NMD;

adjusting a spatial processing algorithm of the NMD;

adjusting a localization algorithm of the NMD; or

disregarding input from a defective microphone.

14 . The method of claim 8 , further comprising determining at least one characteristic of the second sound data based on the metadata.

15 . A tangible, non-transitory, computer-readable medium having stored therein instructions executable by one or more processors to cause a network microphone device (NMD) to perform operations comprising:

at a first time, capturing first sound data via one or more microphones of the NMD, wherein the first sound data includes a voice input;

analyzing the first sound data to detect a wake word;

obtaining an intent based on the voice input;

performing a command based on the intent;

at a second time, capturing second sound data via the one or more microphones;

capturing metadata associated with the second sound data, wherein the second sound data is not derivable from the metadata; and

based on the metadata, modifying at least one performance parameter of the NMD.

16 . The computer-readable medium of claim 15 , wherein the second sound data includes a voice utterance.

17 . The computer-readable medium of claim 15 , wherein the second sound data includes a wake word portion of a voice input.

18 . The computer-readable medium of claim 15 , wherein obtaining the intent comprises:

transmitting the first sound data to one or more remote computing devices associated with a voice assistant service (VAS) to determine the intent based on the voice input; and

receiving a response from at least one of the one or more remote computing devices associated with the VAS, wherein the response is indicative of the determined intent.

19 . The computer-readable medium of claim 15 , wherein modifying the at least one performance parameter comprises modifying audio playback characteristics of the NMD.

20 . The computer-readable medium of claim 15 , wherein the operations further comprise determining at least one characteristic of the second sound data based on the metadata.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2025
From: SMITH, CONNOR KRISTOPHER; SOTO, KURT THOMAS; SLEITH, CHARLES CONOR
To: SONOS, INC.
Reel/Frame 069978/0735 →