IP Library Granted Patent US 10,573,321
Granted Patent B1
US 10,573,321 · App. 16/434,426 · Granted Feb 25, 2020

Voice detection optimization based on selected voice assistant service

Inventors: Connor Kristopher Smith (New Hudson, MI); Kurt Thomas Soto (Ventura, CA); Charles Conor Sleith (Waltham, MA)
Assignee: Sonos, Inc.
G10L15/30G10L15/08G10L15/22H04R1/406H04R3/005G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,573,321
App. No.
16/434,426
Granted
Feb 25, 2020
Kind
B1
Abstract

Systems and methods for optimizing voice detection via a network microphone device (NMD) based on a selected voice-assistant service (VAS) are disclosed herein. In one example, the NMD detects sound via individual microphones and selects a first VAS to communicate with the NMD. The NMD produces a first sound-data stream based on the detected sound using a spatial processor in a first configuration. Once the NMD determines that a second VAS is to be selected over the first VAS, the spatial processor assumes a second configuration for producing a second sound-data stream based on the detected sound. The second sound-data stream is then transmitted to one or more remote computing devices associated with the second VAS.

Claims (48)

1. A playback device comprising:

a plurality of microphones;

a network interface;

one or more processors; and

tangible, non-transitory computer-readable media having stored therein instructions executable by the one or more processors to cause the playback device to perform a method comprising:

capturing audio via a first set of microphones selected from the plurality of microphones; analyzing the audio captured via the first set of microphones using a first wake-word engine on the playback device to detect a first wake word;

selecting a second wake-word engine on the playback device, wherein the second wake-word engine is different from the first wake-word engine;

after selecting the second wake-word engine, capturing audio via a second set of microphones selected from the plurality of microphones, wherein the second set of microphones is different from the first set of microphones;

analyzing the audio captured via the second set of microphones using the second wake-word engine to detect a second wake word;

detecting a wake word via one of the first wake-word engine or the second wake-word engine, wherein the detected wake word comprises one of the first wake word or the second wake word; and

transmitting, via the network interface, at least a voice utterance following the detected wake word to one or more remote servers corresponding to a particular voice assistant service associated with the detected wake word.

2. The playback device of claim 1 , wherein the first wake-word engine is associated with a first voice assistant service and the second wake-word engine is associated with a second voice assistant service, wherein capturing audio via the first set of microphones comprises capturing audio detected by the first set of microphones using a first signal processing scheme configured for the first voice assistant service, and wherein capturing audio via the second set of microphones comprises capturing audio detected by the second set of microphones using a second signal processing scheme configured for the second voice assistant service.

3. The playback device of claim 2 , wherein capturing audio via the second set of microphones further comprises capturing the audio detected by the second set of microphones using the second signal processing scheme while concurrently capturing the audio detected by the first set of microphones using the first signal processing scheme.

4. The playback device of claim 2 , wherein the method further comprises:

in response to selecting the second wake-word engine,

electing the second voice assistant service to process voice input over the first voice assistant service.

5. The playback device of claim 1 , wherein capturing audio via the second set of microphones comprises capturing audio via the second set of microphones while concurrently capturing the audio via the first set of microphones.

6. The playback device of claim 1 , wherein the second set of microphones comprises fewer microphones than the first set of microphones.

7. The playback device of claim 1 , wherein each microphone of the plurality of microphones is in one of the first set of microphones or the second set of microphones.

8. A tangible, non-transitory computer-readable medium having stored therein instructions executable by one or more processors to cause a playback device to perform a method comprising:

capturing audio via a first set of microphones selected from a plurality of microphones of the playback device;

analyzing the audio captured via the first set of microphones using a first wake-word engine on the playback device to detect a first wake word;

selecting a second wake-word engine on the playback device, wherein the second wake-word engine is different from the first wake-word engine;

after selecting the second wake-word engine, capturing audio via a second set of microphones selected from the plurality of microphones, wherein the second set of microphones is different from the first set of microphones;

analyzing the audio captured via the second set of microphones using the second wake-word engine to detect a second wake word;

detecting a wake word via one of the first wake-word engine or the second wake-word engine, wherein the detected wake word comprises one of the first wake word or the second wake word; and

transmitting, via a network interface of the playback device, at least a voice utterance following the detected wake word to one or more remote servers corresponding to a particular voice assistant service associated with the detected wake word.

9. The tangible, non-transitory computer-readable medium of claim 8 , wherein the first wake-word engine is associated with a first voice assistant service and the second wake-word engine is associated with a second voice assistant service, wherein capturing audio via the first set of microphones comprises capturing audio detected by the first set of microphones using a first signal processing scheme configured for the first voice assistant service, and wherein capturing audio via the second set of microphones comprises capturing audio detected by the second set of microphones using a second signal processing scheme configured for the second voice assistant service.

10. The tangible, non-transitory computer-readable medium of claim 9 , wherein capturing audio via the second set of microphones further comprises capturing the audio detected by the second set of microphones using the second signal processing scheme while concurrently capturing the audio detected by the first set of microphones using the first signal processing scheme.

11. The tangible, non-transitory computer-readable medium of claim 9 , wherein the method further comprises:

in response to enabling the second wake-word engine, electing the second voice assistant service to process voice input over the first voice assistant service.

12. The tangible, non-transitory computer-readable medium of claim 8 , wherein capturing audio via the second set of microphones comprises capturing audio via the second set of microphones while concurrently capturing the audio via the first set of microphones.

13. The tangible, non-transitory computer-readable medium of claim 8 , wherein the second set of microphones comprises fewer microphones than the first set of microphones.

14. The tangible, non-transitory computer-readable medium of claim 8 , wherein each microphone of the plurality of microphones is in one of the first set of microphones or the second set of microphones.

15. A method comprising:

capturing audio via a first set of microphones selected from a plurality of microphones of a playback device;

analyzing the audio captured via the first set of microphones using a first wake-word engine on the playback device to detect a first wake word;

selecting a second wake-word engine on the playback device, wherein the second wake-word engine is different from the first wake-word engine;

after selecting the second wake-word engine, capturing audio via a second set of microphones selected from the plurality of microphones, wherein the second set of microphones is different from the first set of microphones;

analyzing the audio captured via the second set of microphones using the second wake-word engine to detect a second wake word;

detecting a wake word via one of the first wake-word engine or the second wake-word engine, wherein the detected wake word comprises one of the first wake word or the second wake word; and

transmitting, via a network interface of the playback device, at least a voice utterance following the detected wake word to one or more remote servers corresponding to a particular voice assistant service associated with the detected wake word.

16. The method of claim 15 , wherein the first wake-word engine is associated with a first voice assistant service and the second wake-word engine is associated with a second voice assistant service, wherein capturing audio via the first set of microphones comprises capturing audio detected by the first set of microphones using a first signal processing scheme configured for the first voice assistant service, and wherein capturing audio via the second set of microphones comprises capturing audio detected by the second set of microphones using a second signal processing scheme configured for the second voice assistant service.

17. The method of claim 16 , wherein capturing audio via the second set of microphones further comprises capturing the audio detected by the second set of microphones using the second signal processing scheme while concurrently capturing the audio detected by the first set of microphones using the first signal processing scheme.

18. The method of claim 16 , wherein the method further comprises:

in response to selecting the second wake-word engine, electing the second voice assistant service to process voice input over the first voice assistant service.

19. The method of claim 15 , wherein capturing audio via the second set of microphones comprises capturing audio via the second set of microphones while concurrently capturing the audio via the first set of microphones.

20. The method of claim 15 , wherein the second set of microphones comprises fewer microphones than the first set of microphones.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2019
From: SMITH, CONNOR KRISTOPHER; SOTO, KURT THOMAS; SLEITH, CHARLES CONOR
To: SONOS, INC.
Reel/Frame 049422/0382 →
Cited By (27)
US 12,198,691 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,236,932 US 12,283,269 US 12,288,558 US 12,314,633 US 12,315,514 US 12,322,390 US 12,327,549 US 12,327,556 US 12,360,734 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,462,802 US 12,498,899 US 12,505,832 US 12,513,479 US 12,518,755 US 12,518,756 US 12,579,978 US 12,699,543 US 12,711,962