IP Library Granted Patent US 11,354,092
Granted Patent B2
US 11,354,092 · App. 17/247,736 · Granted Jun 7, 2022

Noise classification for event detection

Inventors: Nick D'Amato (Santa Barbara, CA); Kurt Thomas Soto (Ventura, CA); Connor Kristopher Smith (New Hudson, MI)
Assignee: Sonos, Inc.
G06F3/167G06F3/162G06F3/165G10L15/22H04L12/2809H04R3/12H04R2227/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,354,092
App. No.
17/247,736
Granted
Jun 7, 2022
Kind
B2
Abstract

In one aspect, a network microphone device includes a plurality of microphones and is configured to detect sound via the one or more microphones. The network microphone device may capture sound data based on the detected sound in a first buffer, and capture metadata associated with the detected sound in a second buffer. The network microphone device may classify one or more noises in the detected sound and cause the network microphone device to perform an action based on the classification of the respective one or more noises.

Claims (54)

1. A playback device, comprising:

one or more processors;

one or more microphones;

a first buffer;

a second buffer;

a tangible, non-transitory, computer-readable medium storing instructions executable by the one or more processors to cause the playback device to perform operations comprising:

detecting sound via one or more microphones of the playback device, wherein the detected sound incudes a voice utterance;

capturing sound data one or more buffers of the playback device based on the detected sound;

analyzing, via the playback device, the sound data to detect a wake word;

based on the analyzed first sound data, detecting the wake word;

after detecting the wake word, transmitting at least a first portion of the sound data associated with the voice utterance to one or more remote computing devices associated with a voice assistant service;

capturing metadata associated with a second portion of the sound data via the playback device, wherein the wake word is not detected based on the second portion of the sound data;

processing the metadata to classify one or more noises in the sound data; and

performing an action based on the classification of the respective one or more noises.

2. The playback device of claim 1 , wherein the second portion of the sound data does not include the voice utterance.

3. The playback device of claim 1 , wherein processing the metadata comprises locally analyzing the metadata and classifying the one or more noises via the NMD.

4. The playback device of claim 1 , wherein classifying the one or more noises comprises comparing the metadata to reference metadata associated with known noise events.

5. The playback device of claim 1 , wherein performing an action comprises at least one of: playing back a sound via the playback device, sending a notification to a user's mobile computing device, or flashing a light.

6. The playback device of claim 1 , wherein causing the playback device to perform an action is based on at least one of a sound pressure level or a directionality of the sound.

7. The playback device of claim 1 , wherein processing the metadata comprises transmitting the metadata to one or more other remote servers for analyzing the metadata.

8. The playback device of claim 1 , wherein: the first portion of the sound data transmitted to the one or more servers comprises recorded audio; and the metadata comprises spectral information that is temporally disassociated from the recorded audio.

9. A method comprising:

detecting sound via one or more microphones of a playback device, wherein the detected sound incudes a voice utterance;

capturing sound data in one or more buffers of the playback device based on the detected sound;

analyzing, via the playback device, the sound data to detect a wake word;

based on the analyzed sound data, detecting the wake word;

after detecting the wake word, transmitting a first portion of the sound data associated with the voice utterance to one or more remote computing devices associated with a voice assistant service;

capturing metadata associated with a second portion of the sound data via the playback device, wherein the wake word is not detected based on the second portion of the sound data;

processing the metadata to classify one or more noises in the sound data; and

causing the playback device to perform an action based on the classification of the respective one or more noises.

10. The method of claim 9 , wherein the second portion of the sound data does not include the voice utterance.

11. The method of claim 9 , wherein:

the first portion of the sound data transmitted to the one or more servers comprises recorded audio; and

the metadata comprises spectral information that is temporally disassociated from the recorded audio.

12. The method of claim 9 , wherein processing the metadata comprises transmitting the metadata to one or more other remote servers for analyzing the metadata.

13. The method of claim 9 , wherein processing the metadata comprises locally analyzing the metadata and classifying the one or more noises via the playback device.

14. The method of claim 9 , wherein classifying the one or more noises comprises comparing the metadata to reference metadata associated with known noise events.

15. The method of claim 9 , wherein causing the playback device to perform an action comprises at least one of: playing back a sound via the playback device, sending a notification to a user's mobile computing device, or flashing a light.

16. The method of claim 9 , wherein causing the playback device to perform an action is based on at least one of a sound pressure level or a directionality of the sound.

17. Tangible, non-transitory, computer-readable medium storing instructions executable by one or more processors to cause a playback device to perform operations comprising:

detecting sound via one or more microphones of the playback device, wherein the detected sound incudes a voice utterance;

capturing sound data one or more buffers of the playback device based on the detected sound;

analyzing the first sound data to detect a wake word;

based on the analyzed first sound data, detecting the wake word;

after detecting the wake word, transmitting at least a first portion of the sound data associated with the voice utterance to one or more remote computing devices associated with a voice assistant service;

capturing metadata associated with a second portion of the sound data via the playback device, wherein the wake word is not detected based on the second portion of the sound data;

processing the metadata to classify one or more noises in the sound data; and

performing an action based on the classification of the respective one or more noises.

18. The tangible, non-transitory, computer-readable medium of claim 17 , wherein the second portion of the sound data does not include the voice utterance.

19. The tangible, non-transitory, computer-readable medium of claim 17 , wherein:

the first portion of the sound data transmitted to the one or more servers comprises recorded audio; and

the metadata comprises spectral information that is temporally disassociated from the recorded audio.

20. The tangible, non-transitory, computer-readable medium of claim 17 , wherein processing the metadata comprises transmitting the metadata to one or more other remote servers for analyzing the metadata.

21. The tangible, non-transitory, computer-readable medium of claim 17 , wherein processing the metadata comprises locally analyzing the metadata and classifying the one or more noises via the playback device.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2020
From: D'AMATO, NICK; SOTO, KURT THOMAS; SMITH, CONNOR KRISTOPHER
To: SONOS, INC.
Reel/Frame 054715/0791 →
Continuity (2)
Continuation 16528016 · Jul 31, 2019
Related Publication 20210216278A1 · Jul 15, 2021
Cited By (38)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,212,945 US 12,217,748 US 12,217,765 US 12,230,291 US 12,231,859 US 12,236,932 US 12,277,368 US 12,279,096 US 12,283,269 US 12,288,558 US 12,314,633 US 12,322,390 US 12,327,549 US 12,327,556 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,462,802 US 12,498,899 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,562,167 US 12,578,779 US 12,579,978 US 12,626,717 US 12,640,148 US 12,699,543 US 12,711,962