IP Library Granted Patent US 10,811,015
Granted Patent B2
US 10,811,015 · App. 16/141,875 · Granted Oct 20, 2020

Voice detection optimization based on selected voice assistant service

Inventors: Connor Kristopher Smith (New Hudson, MI); Kurt Thomas Soto (Ventura, CA); Charles Conor Sleith (Waltham, MA)
Assignee: Sonos, Inc.
G10L15/30G10L15/08G10L15/22H04R1/406H04R3/005G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,811,015
App. No.
16/141,875
Filed
Sep 25, 2018
Granted
Oct 20, 2020
Kind
B2
Art Unit
2656
USPC
704/251
Abstract

Systems and methods for optimizing voice detection via a network microphone device (NMD) based on a selected voice-assistant service (VAS) are disclosed herein. In one example, the NMD detects sound via individual microphones and selects a first VAS to communicate with the NMD. The NMD produces a first sound-data stream based on the detected sound using a spatial processor in a first configuration. Once the NMD determines that a second VAS is to be selected over the first VAS, the spatial processor assumes a second configuration for producing a second sound-data stream based on the detected sound. The second sound-data stream is then transmitted to one or more remote computing devices associated with the second VAS.

Claims (39)

1. A method, comprising:

detecting sound via individual microphones of a network microphone device (NMD);

selecting a first voice assistant service (VAS) to communicate with the NMD;

producing a first sound-data stream based on the detected sound using a spatial processor, wherein the spatial processor is in a first configuration when the first VAS is selected;

determining that a second VAS is to be selected over the first VAS based on the detected sound;

selecting the second VAS to communicate with the NMD and selection of the first VAS;

producing a second sound-data stream based on the detected sound using the spatial processor, wherein the spatial processor is in a second configuration when the second VAS is selected; and

transmitting the second sound-data stream to a remote computing device associated with the second VAS.

2. The method of claim 1 , wherein the spatial processor is configured to process first and second numbers of channels of the detected sound in the first and second configurations, respectively.

3. The method of claim 1 , wherein:

the spatial processor comprises an algorithm having filter coefficients, and

the filter coefficients are different in the first configuration than in the second configuration.

4. The method of claim 1 , wherein the spatial processor utilizes a beamforming algorithm and a multi-channel wiener filter algorithm in the first and second configurations, respectively.

5. The method of claim 1 , further comprising refraining from transmitting the first or second sound-data stream until the determination is made.

6. The method of claim 1 , wherein determining that the second VAS is to be selected over the first VAS comprises identifying a wake-word associated with the second VAS in the first sound-data stream.

7. The method of claim 1 , further comprising:

causing, by a pipeline selector, the spatial processor to use a particular configuration corresponding to the selected VAS.

8. The method of claim 7 , wherein the pipeline selector is further configured to cause the spatial processor to modify at least one of:

a fixed gain of one or more microphones;

a wake-word sensitivity parameter; or

noise-reduction parameters.

9. The method of any claim 1 , wherein the first configuration of the spatial processor is selected by a manual user configuration.

10. The method of claim 1 , wherein the first configuration of the spatial processor is associated with the first VAS.

11. The method of claim 1 , further comprising:

transmitting the second sound stream to the second VAS;

receiving, from the second VAS, a command to be executed by the NMD.

12. The method of claim 1 , further comprising;

transmitting the second sound stream to the second VAS;

receiving, from the second VAS, a response to be output by the NMD.

13. The method of claim 1 , further comprising:

capturing metadata associated with the sound data in a lookback buffer;

providing the metadata to at least one network device to determine at least one characteristic of the detected sound based on the metadata;

after providing the metadata, receiving, from the at least one network device, a response including an instruction, based on the determined characteristic, to modify at least one performance parameter of the network microphone device; and

modifying the at least one performance parameter based on the instruction.

14. A non-transitory computer-readable medium comprising instructions for adjusting performance of a network microphone device (NMD), the instructions, when executed by a processor, causing the processor to perform the method of claim 1 .

15. A network microphone device (NMD) comprising:

one or more processors;

a microphone array comprising a plurality of individual microphones; and

a computer-readable medium according to claim 14 .

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2018
From: SMITH, CONNOR KRISTOPHER; SOTO, KURT THOMAS; SLEITH, CHARLES CONOR
To: SONOS, INC.
Reel/Frame 047344/0875 →
Continuity (1)
Related Publication 20200098372A1 · Mar 26, 2020
Cited By (11)
US 12,375,052 US 12,450,025 US 12,464,302 US 12,495,258 US 12,498,899 US 12,501,229 US 12,505,832 US 12,574,697 US 12,652,508 US 12,659,682 US 12,666,217