IP Library › Granted Patent US 11,664,023
Granted Patent B2
US 11,664,023 · App. 16/915,234 · Granted May 30, 2023

Voice detection by multiple devices

Inventors: Jonathon Reilly (Cambridge, MA); Gregory Burlingame (Woburn, MA); Christopher Butts (Evanston, IL); Romi Kadri (Cambridge, MA); Jonathan P. Lang (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L15/22G10L15/02G10L15/20G10L15/34G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,664,023
App. No.
16/915,234
Granted
May 30, 2023
Kind
B2
Abstract

Disclosed herein are example techniques for voice detection by multiple NMDs. An example implementation may involve one or more servers receiving, via a network interface, data representing multiple audio recordings of a voice input spoken by a given user, each audio recording recorded by a respective NMD of the multiple NMDs, wherein the voice input comprises a detected wake-word. Based on respective sound pressure levels of the multiple audio recordings of the voice input, the servers (i) select a particular NMD of the multiple NMDs and (ii) forego selection of other NMDs of the multiple NMDs. The servers send, via the network interface to the particular NMD, data representing a playback command that corresponds to a voice command in the voice input represented in the multiple audio recordings, wherein the data representing the playback command causes the particular NMD to play back audio content according to the playback command.

Claims (53)

1. A system comprising a first network microphone device (NMD) and a second NMD, the system configured to perform functions comprising:

detecting, via one or more microphones of the first NMD, first voice data representing a first portion of a voice input;

detecting, via one or more microphones of the second NMD, second voice data representing the first portion of the voice input;

based on (i) one or more characteristics of the first voice data and (ii) one or more characteristics of the second voice data, selecting the first voice data from among (a) the first voice data and (b) the second voice data;

processing, via one or more processors of the first NMD, the selected first voice data to determine a voice command;

detecting, via one or more microphones of the first NMD, third voice data representing a second portion of the voice input;

detecting, via one or more microphones of the second NMD, fourth voice data representing the second portion of the voice input;

based on (i) one or more characteristics of the third voice data and (ii) one or more characteristics of the fourth voice data, selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data;

processing, via one or more processors of the second NMD, the selected fourth voice data to determine the voice command; and

causing one or more devices to carry out the determined voice command.

2. The system of claim 1 , wherein the one or more characteristics of the first voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the first NMD, wherein the one or more characteristics of the second voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD, and wherein selecting the first voice data from among (a) the first voice data and (b) the second voice data comprises determining that the sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the first NMD are greater than then the sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD.

3. The system of claim 1 , wherein the one or more characteristics of the third voice data comprise sound pressure levels of the second portion of the voice input as detected by the one or more microphones of the first NMD, wherein the one or more characteristics of the fourth voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD, and wherein selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data comprises determining that the sound pressure levels of the second portion of the voice input as detected by the one or more microphones of the first NMD are less than a threshold level.

4. The system of claim 1 , wherein the functions further comprise:

sending, via a network interface of the first NMD, instructions to cause the second NMD to start recording the voice input via the one or more microphones of the second NMD.

5. The system of claim 1 , wherein selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data comprises:

sending, via a network interface of the first NMD, instructions to cause the second NMD to process the second portion of the voice input.

6. The system of claim 1 , wherein detecting, via one or more microphones of the first NMD, the first voice data representing the first portion of the voice input comprises:

detecting a wake word in the first voice data.

7. The system of claim 1 , wherein the first NMD is associated with a first room in a household, and wherein the second NMD is associated with a second room in the household.

8. The system of claim 1 , wherein a first playback device comprises the first NMD, wherein a second playback device comprises the second NMD, and wherein the first playback device and the second playback device are configured in a bonded zone of a media playback system that comprises the first playback device and the second playback device.

9. The system of claim 1 , wherein the first portion of the voice input and the second portion of the voice input at least partially overlap.

10. A method comprising:

detecting, via one or more microphones of a first network microphone device (NMD), first voice data representing a first portion of a voice input;

detecting, via one or more microphones of a second NMD, second voice data representing the first portion of the voice input;

based on (i) one or more characteristics of the first voice data and (ii) one or more characteristics of the second voice data, selecting the first voice data from among (a) the first voice data and (b) the second voice data;

processing, via one or more processors of the first NMD, the selected first voice data to determine a voice command;

detecting, via one or more microphones of the first NMD, third voice data representing a second portion of the voice input;

detecting, via one or more microphones of the second NMD, fourth voice data representing the second portion of the voice input;

based on (i) one or more characteristics of the third voice data and (ii) one or more characteristics of the fourth voice data, selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data;

processing, via one or more processors of the second NMD, the selected fourth voice data to determine the voice command; and

causing one or more devices to carry out the determined voice command.

11. The method of claim 10 , wherein the one or more characteristics of the first voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the first NMD, wherein the one or more characteristics of the second voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD, and wherein selecting the first voice data from among (a) the first voice data and (b) the second voice data comprises determining that the sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the first NMD are greater than then the sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD.

12. The method of claim 10 , wherein the one or more characteristics of the third voice data comprise sound pressure levels of the second portion of the voice input as detected by the one or more microphones of the first NMD, wherein the one or more characteristics of the fourth voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD, and wherein selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data comprises determining that the sound pressure levels of the second portion of the voice input as detected by the one or more microphones of the first NMD are less than a threshold level.

13. The method of claim 10 , further comprising:

sending, via a network interface of the first NMD, instructions to cause the second NMD to start recording the voice input via the one or more microphones of the second NMD.

14. The method of claim 10 , wherein selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data comprises:

sending, via a network interface of the first NMD, instructions to cause the second NMD to process the second portion of the voice input.

15. The method of claim 10 , wherein detecting, via one or more microphones of the first NMD, the first voice data representing the first portion of the voice input comprises:

detecting a wake word in the first voice data.

16. The method of claim 10 , wherein a first playback device comprises the first NMD, wherein a second playback device comprises the second NMD, and wherein the first playback device and the second playback device are configured in a bonded zone of a media playback system that comprises the first playback device and the second playback device.

17. The method of claim 10 , wherein the first portion of the voice input and the second portion of the voice input at least partially overlap.

18. A tangible, non-transitory, computer-readable media having instructions encoded therein, wherein the instructions, when executed by one or more processors, cause a system to perform functions comprising:

detecting, via one or more microphones of a first network microphone device (NMD), first voice data representing a first portion of a voice input;

detecting, via one or more microphones of a second NMD, second voice data representing the first portion of the voice input;

based on (i) one or more characteristics of the first voice data and (ii) one or more characteristics of the second voice data, selecting the first voice data from among (a) the first voice data and (b) the second voice data;

processing, via one or more processors of the first NMD, the selected first voice data to determine a voice command;

detecting, via one or more microphones of the first NMD, third voice data representing a second portion of the voice input;

detecting, via one or more microphones of the second NMD, fourth voice data representing the second portion of the voice input;

based on (i) one or more characteristics of the third voice data and (ii) one or more characteristics of the fourth voice data, selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data;

processing, via one or more processors of the second NMD, the selected fourth voice data to determine the voice command; and

causing one or more devices to carry out the determined voice command.

19. The tangible, non-transitory, computer-readable media of claim 18 , wherein the one or more characteristics of the first voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the first NMD, wherein the one or more characteristics of the second voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD, and wherein selecting the first voice data from among (a) the first voice data and (b) the second voice data comprises determining that the sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the first NMD are greater than then the sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD.

20. Tangible, non-transitory, computer-readable media of claim 18 , wherein the one or more characteristics of the third voice data comprise sound pressure levels of the second portion of the voice input as detected by the one or more microphones of the first NMD, wherein the one or more characteristics of the fourth voice data comprise sound pressure levels of the first portion of the voice input as detected by the one or more microphones of the second NMD, and wherein selecting the fourth voice data from among (a) the third voice data and (b) the fourth voice data comprises determining that the sound pressure levels of the second portion of the voice input as detected by the one or more microphones of the first NMD are less than a threshold level.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2020
From: REILLY, JONATHON; BURLINGAME, GREGORY; BUTTS, CHRISTOPHER; KADRI, ROMI; LANG, JONATHAN P.
To: SONOS, INC.
Reel/Frame 053075/0753 →
Continuity (4)
Continuation 16416752 · May 20, 2019
Continuation 16214666 · Dec 10, 2018
Continuation 15211748 · Jul 15, 2016
Related Publication 20200395015A1 · Dec 17, 2020
Cited By (33)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,231,859 US 12,236,932 US 12,288,558 US 12,314,633 US 12,322,390 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,424,220 US 12,438,977 US 12,462,802 US 12,498,899 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,562,167 US 12,578,779 US 12,579,978 US 12,626,717 US 12,699,543 US 12,711,962 US 12,732,547 US 12,744,035 US 12,749,486