IP Library › Granted Patent US 10,152,969
Granted Patent B2
US 10,152,969 · App. 15/211,748 · Granted Dec 11, 2018

Voice detection by multiple devices

Inventors: Jonathon Reilly (Cambridge, MA); Gregory Burlingame (Woburn, MA); Christopher Butts (Evanston, IL); Romi Kadri (Cambridge, MA); Jonathan P. Lang (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L15/22G10L15/02G10L15/20G10L15/34G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,152,969
App. No.
15/211,748
Granted
Dec 11, 2018
Kind
B2
Abstract

Disclosed herein are example techniques for voice detection by multiple NMDs. An example implementation may involve receiving a set of voice recordings from a set of NMDs, and identifying a subset of voice recordings from which to determine a given voice command. The example implementation may further involve causing the identified subset of voice recordings to be analyzed to determine the given voice command.

Claims (57)

1. A first networked microphone device (NMD) comprising:

one or more amplifiers configured to drive one or more speakers;

a microphone array;

a network interface;

one or more processors;

tangible, non-transitory computer-readable media having stored therein instructions executable by the one or more processors to cause the first NMD to perform a method comprising:

continuously recording, via the microphone array, audio into a buffer;

detecting, in the recorded audio, a wake-word;

in response to detecting the wake-word, (i) listening, via the microphone array, for a voice command following the wake-word in the recorded audio and (ii) sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via a local area network, the instructions causing the one or more second NMDs to stop recording audio via respective microphone arrays of the one or more second NMDs for a pre-defined time period;

querying, via the network interface, one or more servers of a particular voice assistant service with the voice command following the detected wake-word within the recorded audio;

receiving, from one or more servers of the particular voice assistant service via the network interface in response to the query, a playback command corresponding to the voice command; and

playing back audio content according to the playback command via the one or more amplifiers configured to drive one or more speakers.

2. The first NMD of claim 1 , wherein the voice command includes an indication of a period of time for the first NMD to listen for the voice command, and wherein the method further comprises sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to stop recording audio via respective microphone arrays of the one or more second NMDs for the period of time indicated by the voice command.

3. The first NMD of claim 2 , wherein the method further comprises:

determining, based on the recorded audio in the buffer, that the first NMD is no longer receiving the voice command, and

based on determining that the first NMD is no longer receiving the voice command, sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to start recording audio via respective microphone arrays of the one or more second NMDs before the period of time indicated by the voice command has fully elapsed.

4. The first NMD of claim 1 , wherein the method further comprises:

determining, based on the recorded audio in the buffer, that the first NMD is no longer receiving the voice command, and

based on determining that the first NMD is no longer receiving the voice command, sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to start recording audio via respective microphone arrays of the one or more second NMDs before the pre-defined time period has fully elapsed.

5. The first NMD of claim 1 , wherein the playback command comprises a command to play back particular audio content in a first zone that includes the first NMD and a second zone that includes a second NMD, and wherein the method further comprises:

instructing, via the network interface, the second NMD of the second zone to play back the audio content according to the playback command in synchrony with playback of the audio content by the first NMD of the first zone.

6. The first NMD of claim 1 , wherein a first zone of a media playback system includes the first NMD, and wherein the first zone is configured into a zone group with a second zone that includes one or more playback devices, and wherein playing back the audio content according to the playback command comprises playing back the audio content in synchrony with one or more playback devices of the second zone.

7. The first NMD of claim 1 , wherein a first zone of a media playback system includes the first NMD and a second NMD in a bonded zone configuration in which the first NMD and the second NMD play respective channels of the audio content, and wherein playing back the audio content according to the playback command comprises playing back a first channel of the audio content in synchrony the second NMD playing back a second channel of the audio content.

8. Tangible, non-transitory, computer-readable media having instructions encoded therein, wherein the instructions, when executed by one or more processors, cause a first networked microphone device (NMD) to perform a method comprising:

continuously recording, via a microphone array of the first NMD, audio into a buffer;

detecting, in the recorded audio, a wake-word;

in response to detecting the wake-word, (i) listening, via a microphone of the first NMD, for a voice command following the wake-word in the recorded audio and (ii) sending, via a network interface, instructions to one or more second NMDs connected via to the first NMD via a local area network, the instructions causing the one or more second NMDs to stop recording audio via respective microphone arrays of the one or more second NMDs for a pre-defined time period;

querying, via a network interface of the first NMD, one or more servers of a particular voice assistant service with the voice command following the detected wake-word within the recorded audio;

receiving, from one or more servers of the particular voice assistant service via the network interface in response to the query, a playback command corresponding to the voice command; and

playing back audio content according to the playback command via one or more amplifiers configured to drive one or more speakers.

9. The tangible, computer readable media of claim 8 , wherein the voice command includes an indication of a period of time for the first NMD to listen for the voice command, and wherein the method further comprises sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to stop recording audio via respective microphone arrays of the one or more second NMDs for the period of time indicated by the voice command.

10. The tangible, computer readable media of claim 9 , wherein the method further comprises:

determining, based on the recorded audio in the buffer, that the first NMD is no longer receiving the voice command, and

based on determining that the first NMD is no longer receiving the voice command, sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to start recording audio via respective microphone arrays of the one or more second NMDs before the period of time indicated by the voice command has fully elapsed.

11. The tangible, computer readable media of claim 8 , wherein the method further comprises:

determining, based on the recorded audio in the buffer, that the first NMD is no longer receiving the voice command, and

based on determining that the first NMD is no longer receiving the voice command, sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to start recording audio via respective microphone arrays of the one or more second NMDs before the pre-defined time period has fully elapsed.

12. The tangible, computer readable media of claim 8 , wherein the playback command comprises a command to play back particular audio content in a first zone that includes the first NMD and a second zone that includes a second NMD, and wherein the method further comprises:

instructing, via the network interface, the second NMD of the second zone to play back the audio content according to the playback command in synchrony with playback of the audio content by the first NMD of the first zone.

13. The tangible, computer readable media of claim 8 , wherein a first zone of a media playback system includes the first NMD, and wherein the first zone is configured into a zone group with a second zone that includes one or more playback devices, and wherein playing back the audio content according to the playback command comprises playing back the audio content in synchrony with one or more playback devices of the second zone.

14. The tangible, computer readable media of claim 8 , wherein a first zone of a media playback system includes the first NMD and a second NMD in a bonded zone configuration in which the first NMD and the second NMD play respective channels of the audio content, and wherein playing back the audio content according to the playback command comprises playing back a first channel of the audio content in synchrony the second NMD playing back a second channel of the audio content.

15. A method comprising:

a first networked microphone device (NMD) continuously recording, via a microphone array of the first NMD, audio into a buffer;

the first NMD detecting, in the recorded audio, a wake-word;

in response to detecting the wake-word, the first NMD (i) listening, via a microphone of the first NMD, for a voice command following the wake-word in the recorded audio and (ii) sending, via a network interface, instructions to one or more second NMDs connected via to the first NMD via a local area network, the instructions causing the one or more second NMDs to stop recording audio via respective microphone arrays of the one or more second NMDs for a pre-defined time period;

the first NMD querying, via a network interface of the first NMD, one or more servers of a particular voice assistant service with the voice command following the detected wake-word within the recorded audio;

the first NMD receiving, from one or more servers of the particular voice assistant service via the network interface in response to the query, a playback command corresponding to the voice command; and

the first NMD playing back audio content according to the playback command via one or more amplifiers configured to drive one or more speakers.

16. The method of claim 15 , wherein the voice command includes an indication of a period of time for the first NMD to listen for the voice command, and wherein the method further comprises sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to stop recording audio via respective microphone arrays of the one or more second NMDs for the period of time indicated by the voice command.

17. The method of claim 15 , wherein the method further comprises:

determining, based on the recorded audio in the buffer, that the first NMD is no longer receiving the voice command, and

based on determining that the first NMD is no longer receiving the voice command, sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to start recording audio via respective microphone arrays of the one or more second NMDs before the pre-defined time period indicated by the voice command has fully elapsed.

18. The method of claim 17 , wherein the method further comprises:

determining, based on the recorded audio in the buffer, that the first NMD is no longer receiving the voice command, and

based on determining that the first NMD is no longer receiving the voice command, sending, via the network interface, instructions to one or more second NMDs connected via to the first NMD via the local area network, the instructions causing the one or more second NMDs to start recording audio via respective microphone arrays of the one or more second NMDs before the pre-defined time period has fully elapsed.

19. The method of claim 15 , wherein a first zone of a media playback system includes the first NMD, and wherein the first zone is configured into a zone group with a second zone that includes one or more playback devices, and wherein playing back the audio content according to the playback command comprises playing back the audio content in synchrony with one or more playback devices of the second zone.

20. The method of claim 15 , wherein a first zone of a media playback system includes the first NMD and a second NMD in a bonded zone configuration in which the first NMD and the second NMD play respective channels of the audio content, and wherein playing back the audio content according to the playback command comprises playing back a first channel of the audio content in synchrony the second NMD playing back a second channel of the audio content.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Oct 18, 2021
From: JPMORGAN CHASE BANK, N.A.
To: SONOS, INC.
Reel/Frame 058213/0597 →
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
SECURITY INTEREST Recorded Aug 30, 2018
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 046991/0433 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2016
From: REILLY, JONATHON; BURLINGAME, GREGORY; BUTTS, CHRISTOPHER; KADRI, ROMI; LANG, JONATHAN P
To: SONOS, INC.
Reel/Frame 040844/0935 →
Continuity (1)
Related Publication 20180018964A1 · Jan 18, 2018
Cited By (46)
US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,236,932 US 12,250,536 US 12,283,269 US 12,288,558 US 12,314,633 US 12,322,390 US 12,327,549 US 12,327,556 US 12,340,802 US 12,360,734 US 12,374,333 US 12,374,334 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,450,025 US 12,462,802 US 12,464,302 US 12,495,258 US 12,498,899 US 12,501,229 US 12,505,832 US 12,513,479 US 12,518,755 US 12,518,756 US 12,562,167 US 12,574,697 US 12,579,978 US 12,640,148 US 12,652,508 US 12,659,682 US 12,666,217 US 12,699,543 US 12,711,962 US 12,732,547 US 12,737,152 US 12,739,581 US 12,744,035 US 12,748,566 US 12,750,630