IP Library Granted Patent US 11,646,023
Granted Patent B2
US 11,646,023 · App. 17/247,507 · Granted May 9, 2023

Devices, systems, and methods for distributed voice processing

Inventors: Connor Kristopher Smith (New Hudson, MI); John Tolomei (Renton, WA); Betty Lee (Boston, MA)
Assignee: Sonos, Inc.
G10L15/22G10L15/08G10L15/30H04R1/406H04R3/005G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,646,023
App. No.
17/247,507
Granted
May 9, 2023
Kind
B2
Abstract

Systems and methods for distributed voice processing are disclosed herein. In one example, the method includes detecting sound via a microphone array of a first playback device and analyzing, via a first wake-word engine of the first playback device, the detected sound. The first playback device may transmit data associated with the detected sound to a second playback device over a local area network. A second wake-word engine of the second playback device may analyze the transmitted data associated with the detected sound. The method may further include identifying that the detected sound contains either a first wake word or a second wake word based on the analysis via the first and second wake-word engines, respectively. Based on the identification, sound data corresponding to the detected sound may be transmitted over a wide area network to a remote computing device associated with a particular voice assistant service.

Claims (49)

1. A system, comprising:

a network microphone device (NMD) and a playback device, the NMD comprising:

one or more processors;

one or more microphones; and

a first computer-readable medium storing instructions that, when executed by the one or more processors, cause the NMD to perform first operations, the first operations comprising:

detecting sound via the one or more microphones;

transmitting data associated with the detected sound to the playback device over a local area network;

the playback device comprising:

one or more processors; and

a second computer-readable medium storing instructions that, when executed by the one or more processors, cause the playback device to perform second operations, the second operations comprising:

identifying, via a wake word engine of the playback device, a wake word based on the transmitted data associated with the detected sound from the NMD;

based on the identification, transmitting sound data corresponding to the detected sound to one or more remote computing devices over a wide area network;

after the transmitting, receiving a response from the one or more remote computing devices; and

after receiving the response, transmitting a message to the NMD device over the local area network, wherein the message includes instructions to perform an action,

wherein the first computer-readable medium of the NMD causes the NMD to perform the action.

2. The system of claim 1 , wherein the NMD comprises one or more audio transducers, and wherein the action comprises playing back audio via the one or more audio transducers.

3. The system of claim 1 , wherein the playback device is a first playback device, and wherein a second playback device comprises the NMD.

4. The system of claim 1 , wherein the action is a first action and the second operations further comprise performing a second action via the playback device, wherein the second action is based on the response from the remote computing device.

5. The system of claim 1 , wherein the second operations further comprise disabling a wake word engine of the NMD in response to the identification of the wake word via the wake word engine of the playback device.

6. The system of claim 4 , wherein the second operations further comprise enabling the wake word engine of the NMD after the playback device receives the response from the one or more remote computing devices.

7. The system of claim 1 , wherein the one or more remote computing devices are associated with a particular voice assistant service.

8. A method comprising:

detecting sound via one or more microphones of a network microphone device (NMD);

transmitting data associated with the detected sound from the NMD to a playback device over a local area network;

identifying, via a wake word engine of the playback device, a wake word based on the transmitted data associated with the detected sound;

based on the identification, transmitting sound data corresponding to the detected sound from the playback device to one or more remote computing devices over a wide area network;

after the transmitting, receiving, via the playback device, a response from the one or more remote computing devices;

after receiving the response, transmitting a message from the playback device to the NMD over the local area network, wherein the message includes instructions to perform an action; and

performing the action via the NMD.

9. The method of claim 8 , wherein the NMD comprises one or more audio transducers, and wherein the action comprises playing back audio via the one or more audio transducers.

10. The method of claim 8 , wherein the playback device is a first playback device, and wherein a second playback device comprises the NMD.

11. The method of claim 8 , wherein the action is a first action and the method further comprises performing a second action via the playback device, wherein the second action is based on the response from the one or more remote computing devices.

12. The method of claim 8 , further comprising disabling a wake word engine of the NMD in response to the identification of the wake word via the wake word engine of the playback device.

13. The method of claim 10 , further comprising enabling a wake word engine of the NMD after the playback device receives the response from the one or more remote computing devices.

14. The method of claim 10 , wherein the wake word is a second wake word, and wherein the wake word engine of the NMD is configured to detect a first wake word that is different than the second wake word.

15. The method of claim 8 , wherein the one or more remote computing devices are associated with a particular voice assistant service.

16. A playback device comprising:

one or more processors; and

a computer-readable medium storing instructions that, when executed by the one or more processors, cause the playback device to perform operations comprising:

receiving, from a network microphone device (NMD) over a local area network, data associated with sound detected via one or more microphones of the NMD;

identifying, via a wake word engine of the playback device, a wake word based on the data associated with the detected sound;

after the identification, transmitting sound data corresponding to the detected sound to one or more remote computing devices over a wide area network;

after the transmitting, receiving a response from the one or more remote computing devices; and

after receiving the response, transmitting a message to the NMD over the local area network, wherein the message includes instructions for the playback device to perform an action.

17. The playback device of claim 16 , wherein the action is a first action and the operations further comprise performing a second action via the playback device, wherein the second action is based on the response from the one or more remote computing devices.

18. The playback device of claim 16 , wherein the operations further comprise disabling a wake word engine of the NMD in response to the identification of the wake word via the wake word engine of the playback device.

19. The playback device of claim 18 , wherein the operations further comprise enabling the wake word engine of the NMD after the playback device receives the response from the one or more remote computing devices.

20. The playback device of claim 18 , wherein the wake word is a first wake word, and wherein the wake word engine of the NMD is configured to detect a second wake word that is different than the first wake word.

21. The playback device of claim 16 , wherein the one or more remote computing devices are associated with a particular voice assistant service.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2020
From: SMITH, CONNOR KRISTOPHER; TOLOMEI, JOHN; LEE, BETTY
To: SONOS, INC.
Reel/Frame 054642/0797 →
Continuity (2)
Continuation 16271560 · Feb 8, 2019
Related Publication 20210210095A1 · Jul 8, 2021
Cited By (30)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,231,859 US 12,236,932 US 12,288,558 US 12,314,633 US 12,322,390 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,424,220 US 12,438,977 US 12,462,802 US 12,498,899 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,562,167 US 12,578,779 US 12,579,978 US 12,626,717 US 12,699,543 US 12,711,962