IP Library › Granted Patent US 11,308,962
Granted Patent B2
US 11,308,962 · App. 16/879,553 · Granted Apr 19, 2022

Input detection windowing

Inventors: Connor Kristopher Smith (New Hudson, MI); Matthew David Anderson (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L15/22G06F3/165G06F3/167G10L15/1815G10L15/30G10L25/51G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,308,962
App. No.
16/879,553
Granted
Apr 19, 2022
Kind
B2
Abstract

A device, such as Network Microphone Device or a playback device, detecting an event associated with the device or a system comprising the device. In response, an input detection window is opened for a given time period. During the given time period the device is arranged to receive an input sound data stream representing sound detected by a microphone. The input sound data stream is analyzed for a plurality of keywords and/or a wake-word for a Voice Assistant Service (VAS) and, based on the analysis, it is determined that the input sound data stream includes voice input data comprising a keyword or a wake-word for a VAS. In response, the device takes appropriate action such as causing the media playback system to perform a command corresponding to the keyword or sending at least part of the input sound data stream to the VAS.

Claims (64)

1. A playback device comprising:

at least one microphone configured to detect sound; at least one processor; and

data storage having instructions stored thereon that are executed by the at least one processor to cause the playback device to perform functions comprising:

detecting a first input sound data stream comprising a wake word and a voice input; performing an action based on the voice input, wherein the action comprises initiating

playback of a playback queue;

while playing back the playback queue, detecting an event, the event being associated with the playback device or a system comprising the playback device, wherein the event comprises an indication of a track change associated with the playback queue, wherein the indication corresponds to a user input to change a trackvia a controller; and

responsive to the event comprising the indication of the track change, opening an input detection window for a given time period;

wherein during the given time period the at least one processor is arranged to perform functions comprising:

receiving, by the at least one processor, a second input sound data stream from the at least one microphone, wherein the second input sound data stream represents sound detected by the at least one microphone;

analyzing the second input sound data stream by the at least one processor for a plurality of keywords supported by the playback device;

determining, based on the analysis, that the second input sound data stream includes voice input data comprising a keyword;

wherein the keyword is one of the plurality of keywords supported by the playback device; and

responsive to the determining that the second input sound data stream includes voice input data comprising a keyword, causing a command corresponding to the keyword to be performed.

2. The playback device according to claim 1 , further comprising:

storing historical data in the data storage, the historical data including determined events and any associated commands; and

analyzing the historical data to adjust the given time period associated with the input detection window.

3. The playback device according to claim 2 , wherein analyzing the historical data comprises:

determining a frequency of the associated commands and a time period, relative to a time of the event, during which the associated commands were received by the playback device or a system comprising the playback device; and

setting the given time period based on the frequency and the time period.

4. The playback device according to claim 2 , wherein the given time period is associated with a user account of the playback device or a system comprising the playback device.

5. The playback device according to claim 1 , wherein outside of the input detection window, the second input sound data stream is not analyzed.

6. The playback device according to claim 1 , further comprising:

a network interface, and

the instructions stored on the data storage that are executed by the at least one processor further cause the playback device to perform functions comprising:

transmitting, via the network interface and during the input detection window, at least part of the second input sound data stream to a remote server for analysis;

receiving data including a remote command from the remote server, the remote command being supported by the playback device; and causing the media playback system to perform the remote command.

7. The playback device according to claim 1 , wherein the analyzing the second input data stream comprises determining a keyword by at least one of:

natural language processing to determine an intent, based on an analysis of the keyword in the second input sound data stream; and pattern-matching based on a predefined library of keywords and associated keywords.

8. A method to be performed by a playback device comprising at least one microphone configured to detect sound, the method comprising:

detecting a first input sound data stream comprising a wake word and a voice input; performing an action based on the voice input wherein the action comprises initiating

playback of a playback queue;

while playing back the playback queue, detecting an event associated with the playback device or a system comprising the playback device, wherein the event comprises an indication of a track change associated with the playback queue wherein the indication corresponds to a user input to change a track via a controller device;

opening, upon the detection of the event comprising the indication of the track change, an input detection window for a given time period; and

during the given time period:

receiving a second input sound data stream representing the sound detected by the at least one microphone;

analyzing the second input sound data stream for a plurality of keywords supported by the playback device;

determining based on the analysis, that the second input sound data stream includes voice input data comprising a keyword; wherein the keyword is one of the plurality of keywords supported by the playback device; and

responsive to the determining that the second input sound data stream includes voice input data comprising a keyword, causing the playback device or a

system comprising the playback device to perform a command corresponding to the keyword.

9. The method according to claim 8 , further comprising:

storing historical data in the data storage, the historical data including detected events and any associated commands; and

analyzing the historical data to adjust the given time period associated with the input detection window.

10. The method according to claim 9 , wherein analyzing the historical data comprises:

determining a frequency of the associated commands and a time period, relative to a time of the detected event, during which the associated commands were received by the media playback system; and

setting the given time period based on the frequency and the time period.

11. The method according to claim 9 , wherein the given time period is associated with a user account of the playback device or a system comprising the playback device.

12. The method according to claim 8 , wherein outside of the input detection window, the second input sound data stream is not analyzed.

13. The method according to claim 8 , wherein the playback device comprises a network interface, and the method further comprises:

transmitting, via the network interface and during the input detection window, at least part of the second input sound data stream to a remote server for analysis;

receiving data including a remote command from the remote server, the remote command being supported by the playback device; and

causing the playback device or a system comprising the playback device to perform the remote command.

14. The method according to claim 8 , wherein the analyzing the second input data stream comprises determining a keyword by at least one of:

natural language processing to determine an intent, based on an analysis of the keyword in the input sound data stream; and

pattern-matching based on a predefined library of keywords and associated keywords.

15. A non-transitory computer-readable medium having instructions stored thereon that are executable by one or more processors to cause a playback device to perform functions, the playback device comprising at least one microphone configured to detect sound, the functions comprising:

detecting a first input sound data stream comprising a wake word and a voice input; performing an action based on the voice input, wherein the action comprises initiating

playback of a playback queue;

while playing back the playback queue, detecting an event associated with the playback device or a system comprising the playback device, wherein the event comprises an indication of a track change associated with the playback queue, wherein the indication corresponds to a user input to change a track via a controller device

opening, responsive to the detecting the event comprising the indication of the track change, an input detection window for a given time period; and

during the given time period:

receiving a second input sound data stream representing the sound detected by the at least one microphone;

analyzing the second input sound data stream for a plurality of keywords supported by the playback device;

determining based on the analysis, that the second input sound data stream includes voice input data comprising a keyword; wherein the keyword is one of the plurality of keywords supported by the playback device; and

responsive to the determining that the second input sound data stream includes voice input data comprising a keyword, causing the playback device or a system comprising the playback device to perform a command corresponding to the keyword.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2020
From: SMITH, CONNOR KRISTOPHER; ANDERSON, MATTHEW DAVID
To: SONOS, INC.
Reel/Frame 052735/0615 →
Continuity (1)
Related Publication 20210366477A1 · Nov 25, 2021
Cited By (2)
US 12,437,755 US 12,585,821