IP Library › Granted Patent US 11,361,756
Granted Patent B2
US 11,361,756 · App. 16/439,046 · Granted Jun 14, 2022

Conditional wake word eventing based on environment

Inventors: Connor Smith (New Hudson, MI); John Tolomei (Renton, WA); Kurt Soto (Ventura, CA)
Assignee: Sonos, Inc.
G10L15/083G10L15/20G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,361,756
App. No.
16/439,046
Granted
Jun 14, 2022
Kind
B2
Abstract

In one aspect, a playback device includes at least one microphone configured to detect sound. The playback detects sound via the one or more microphones and determines whether (i) the detected sound includes a voice input, (ii) the detected sound excludes background speech, and (iii) the voice input includes a command keyword. In response to the determining, the playback device performs a playback function corresponding to the command keyword.

Claims (69)

1. A playback device comprising:

one or more processors;

at least one microphone configured to detect sound;

at least one speaker; and

data storage having instructions stored thereon that are executable by the one or more processors to cause the playback device to perform functions comprising:

detecting a first sound via the at least one microphone;

determining that (i) the detected first sound includes a first voice input, (ii) the detected first sound excludes background speech, and (iii) the first voice input includes a particular command keyword from among a plurality of command keywords corresponding to respective functions, wherein voice inputs including the command keywords of the plurality are processable locally on the playback device, wherein a wake word corresponding to a voice assistant service is not detected in the first voice input, and wherein the plurality of command keywords exclude the wake word;

in response to determining (i) the detected first sound includes the first voice input, (ii) the detected first sound excludes background speech, and (iii) the first voice input includes the particular command keyword, performing a playback function corresponding to the particular command keyword without sending data corresponding to the first voice input to one or more servers of the voice assistant service for processing;

after performing the playback function corresponding to the particular command keyword, detecting a second sound via the at least one microphone;

determining that a second voice input in the detected second sound includes the wake word corresponding to the voice assistance service; and

in response to determining that the second voice input in the detected second sound includes the wake word corresponding to the voice assistance service, streaming, via a network interface of the playback device, the second voice input in the detected second sound to one or more remote servers of the voice assistant service for processing.

2. The playback device of claim 1 , wherein an additional command keyword within the plurality of command keywords corresponds to an TOT function, and wherein the functions further comprise:

detecting a third sound via the at least one microphone;

determining that (i) the detected third sound includes a third voice input, (ii) the detected third sound excludes background speech, and (iii) the third voice input includes a specific command keyword from among the plurality of command keywords corresponding to respective functions, wherein the wake word corresponding to the voice assistant service is not detected in the third voice input; and

in response to determining (i) the detected third sound includes the third voice input, (ii) the detected third sound excludes background speech, and (iii) the third voice input includes the particular command keyword, causing an TOT device to perform the TOT function corresponding to the specific command keyword without sending data corresponding to the third voice input to one or more servers of the voice assistant service for processing.

3. The playback device of claim 1 , wherein determining that the detected sound excludes background speech comprises:

determining sound metadata corresponding to the detected sound; and

analyzing the sound metadata to classify the detected sound according to one or more particular signatures selected from a plurality of signatures, wherein each signature of the plurality of signatures is associated with a noise source, and wherein at least one of the signatures of the plurality of signatures is a background speech signature indicative of background speech.

4. The playback device of claim 3 , wherein analyzing the sound metadata comprises:

classifying frames associated with the detected sound as having a particular speech signature other than the background speech signature; and

comparing a number of frames, if any, classified with the background speech signature to a number of frames classified with a signature other than the background speech signature.

5. The playback device of claim 3 , wherein determining that there is a voice input in the detected sound comprises detecting voice activity in the detected sound.

6. The playback device of claim 5 , wherein detecting voice activity in the detected sound comprises:

determining a number of first frames associated with the detected sound as containing speech; and

comparing the number of first frames to a number of second frames that are (a) associated with the detected sound and (b) not indicative of speech.

7. The playback device of claim 6 , wherein the first frames comprise one or more frames generated in response to near-field voice activity and one or more frames generated in response to far-field voice activity.

8. A method to be performed by a playback device comprising a network interface and at least one microphone configured to detect sound, the method comprising:

detecting a first sound via the at least one microphone;

determining that (i) the detected first sound includes a first voice input, (ii) the detected first sound excludes background speech, and (iii) the first voice input includes a particular command keyword from among a plurality of command keywords corresponding to respective functions, wherein voice inputs including the command keywords of the plurality are processable locally on the playback device, wherein a wake word corresponding to a voice assistant service is not detected in the first voice input, and wherein the plurality of command keywords exclude the wake word;

in response to determining (i) the detected first sound includes the first voice input, (ii) the detected first sound excludes background speech, and (iii) the first voice input includes the particular command keyword, performing a playback function corresponding to the particular command keyword without sending data corresponding to the first voice input to one or more servers of the voice assistant service for processing;

after performing the playback function corresponding to the particular command keyword, detecting a second sound via the at least one microphone;

determining that a second voice input in the detected second sound includes the wake word corresponding to the voice assistance service; and

in response to determining that the second voice input in the detected second sound includes the wake word corresponding to the voice assistance service, streaming, via the network interface of the playback device, the second voice input in the detected second sound to one or more remote servers of the voice assistant service for processing.

9. The method of claim 8 , further comprising:

detecting a third sound via the at least one microphone;

determining that (i) the detected third sound includes a third voice input, (ii) the detected third sound excludes background speech, and (iii) the third voice input includes a specific command keyword from among the plurality of command keywords corresponding to respective functions, wherein the wake word corresponding to the voice assistant service is not detected in the third voice input; and

in response to determining (i) the detected third sound includes the third voice input, (ii) the detected third sound excludes background speech, and (iii) the third voice input includes the particular command keyword, causing an TOT device to perform the TOT function corresponding to the specific command keyword without sending data corresponding to the third voice input to one or more servers of the voice assistant service for processing.

10. The method of claim 8 , wherein determining that there is an absence of background speech in the detected sound comprises:

determining sound metadata corresponding to the detected sound; and

analyzing the sound metadata to classify the detected sound according to one or more particular signatures selected from a plurality of signatures, wherein each signature of the plurality of signatures is associated with a noise source, and wherein at least one of the signatures of the plurality of signatures is a background speech signature indicative of background speech.

11. The method of claim 10 , wherein the analyzing comprises:

classifying frames associated with the detected sound as having a particular speech signature other than the background speech signature; and

comparing a number of frames, if any, classified with the background speech signature to a number of frames classified with a signature other than the background speech signature.

12. The method of claim 10 , wherein determining that there is a voice input in the detected sound comprises detecting voice activity in the detected sound.

13. The method of claim 12 , wherein detecting voice activity in the detected sound comprises:

determining a number of first frames associated with the detected sound as containing speech; and

comparing the number of first frames to a number of second frames that are (a) associated with the detected sound and (b) not indicative of speech.

14. The method of claim 13 , wherein the first frames comprise one or more frames generated in response to near-field voice activity and one or more frames generated in response to far-field voice activity.

15. A non-transitory computer-readable medium having instructions stored thereon that are executable by one or more processors to cause a playback device to perform functions, the playback device comprising a network interface and at least one microphone configured to detect sound, the functions comprising:

detecting a first sound via the at least one microphone;

determining that (i) the detected first sound includes a first voice input, (ii) the detected first sound excludes background speech, and (iii) the first voice input includes a particular command keyword from among a plurality of command keywords corresponding to respective functions, wherein voice inputs including the command keywords of the plurality are processable locally on the playback device, wherein a wake word corresponding to a voice assistant service is not detected in the first voice input, and wherein the plurality of command keywords exclude the wake word;

in response to determining (i) the detected first sound includes the first voice input, (ii) the detected first sound excludes background speech, and (iii) the first voice input includes the particular command keyword, performing a playback function corresponding to the particular command keyword without sending data corresponding to the first voice input to one or more servers of the voice assistant service for processing;

after performing the playback function corresponding to the particular command keyword, detecting a second sound via the at least one microphone;

determining that a second voice input in the detected second sound includes the wake word corresponding to the voice assistance service; and

in response to determining that the second voice input in the detected second sound includes the wake word corresponding to the voice assistance service, streaming, via a network interface of the playback device, the second voice input in the detected second sound to one or more remote servers of the voice assistant service for processing.

16. The non-transitory computer-readable medium of claim 15 , wherein the functions further comprise:

detecting a third sound via the at least one microphone;

determining that (i) the detected third sound includes a third voice input, (ii) the detected third sound excludes background speech, and (iii) the third voice input includes a specific command keyword from among the plurality of command keywords corresponding to respective functions, wherein the wake word corresponding to the voice assistant service is not detected in the third voice input; and

in response to determining (i) the detected third sound includes the third voice input, (ii) the detected third sound excludes background speech, and (iii) the third voice input includes the particular command keyword, causing an IOT device to perform the TOT function corresponding to the specific command keyword without sending data corresponding to the third voice input to one or more servers of the voice assistant service for processing.

17. The non-transitory computer-readable medium of claim 15 , wherein determining that there is an absence of background speech in the detected sound comprises:

determining sound metadata corresponding to the detected sound; and

analyzing the sound metadata to classify the detected sound according to one or more particular signatures selected from a plurality of signatures, wherein each signature of the plurality of signatures is associated with a noise source, and wherein at least one of the signatures of the plurality of signatures is a background speech signature indicative of background speech.

18. The non-transitory computer-readable medium of claim 17 , wherein the analyzing comprises:

classifying frames associated with the detected sound as having a particular speech signature other than the background speech signature; and

comparing a number of frames, if any, classified with the background speech signature to a number of frames classified with a signature other than the background speech signature.

19. The non-transitory computer-readable medium of claim 15 , wherein determining that there is a voice input in the detected sound comprises detecting voice activity in the detected sound.

20. The non-transitory computer-readable medium of claim 19 , wherein detecting voice activity in the detected sound comprises:

determining a number of first frames associated with the detected sound as containing speech; and

comparing the number of first frames to a number of second frames that are (a) associated with the detected sound and (b) not indicative of speech.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2019
From: SMITH, CONNOR; TOLOMEI, JOHN; SOTO, KURT
To: SONOS, INC.
Reel/Frame 049447/0753 →
Continuity (1)
Related Publication 20200395006A1 · Dec 17, 2020
Cited By (1)
US 12,749,486