IP Library Granted Patent US 11,445,301
Granted Patent B2
US 11,445,301 · App. 17/174,753 · Granted Sep 13, 2022

Portable playback devices with network operation modes

Inventors: Sangah Park (Somerville, MA); Ryan Myers (Santa Barbara, CA); John Tolomei (Renton, WA)
Assignee: Sonos, Inc.
H04R5/04G10L15/083G10L15/22H04R1/025H04R3/005G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,445,301
App. No.
17/174,753
Granted
Sep 13, 2022
Kind
B2
Abstract

Examples described herein relate to portable playback devices, such as smart headphones and earbuds, and ultra-portable devices having built-in voice assistants. Some example techniques relate to user interaction with voice assistants. Further example techniques relate to voice guidance played back by the headphones to guide the user under certain conditions.

Claims (86)

1. A system comprising a portable playback device, the portable playback device comprising:

at least one microphone;

at least one audio transducer;

at least one battery;

a first network interface comprising at least one first antenna;

a second network interface comprising at least one second antenna;

at least one processor;

a housing carrying the at least one microphone, the at least one audio transducer, the at least one battery, the first network interface, the second network interface, and the at least one processor;

data storage; and

program instructions stored on the data storage that, when executed by the at least one processor, cause the portable playback device to perform functions comprising:

while in a first connection mode of a set of connection modes, detecting a first input to invoke voice assistance, wherein the set of connection modes comprises: (i) the first connection mode wherein the portable playback device connects to a wireless local area network (WLAN) via the first network interface, (ii) a second connection mode wherein the portable playback device connects to a personal area network (PAN) via the second network interface, and (iii) a concurrent connection mode wherein the portable playback device concurrently connects to the WLAN via the first network interface and the PAN via the second network interface;

in response to detecting the first input, invoking a first voice assistant of a set of voice assistants, the first voice assistant corresponding to the first connection mode, wherein the first voice assistant is native to the portable playback device;

after invoking the first voice assistant, capturing a first voice input via the at least one microphone;

processing the first voice input via the first voice assistant;

switching to the second connection mode;

while in the second connection mode, detecting a second input to invoke voice assistance;

in response to detecting the second input, invoking a second voice assistant of the set of voice assistants, the second voice assistant corresponding to the second connection mode, wherein the second voice assistant is native to a mobile device connected to the portable playback device over the PAN;

after invoking the second voice assistant, capturing a second voice input via the at least one microphone; and

sending, over the PAN via the second network interface to the mobile device, data representing the second voice input for processing of the second voice input.

2. The system of claim 1 , wherein the first voice assistant comprises cloud-based natural language understanding (NLU), and wherein processing the first voice input via the first voice assistant comprises:

sending, via the first network interface, data representing the first voice input to a computing system for processing of the first voice input via the cloud-based NLU.

3. The system of claim 2 , wherein the first voice assistant further comprises local NLU configured to recognize a set of local keywords, and wherein processing the first voice input further comprises:

determining, via the local NLU, that portions of the captured first voice input correspond to keywords that are not in the set of local keywords, wherein the portable playback device sends the data representing the first voice input to the computing system in response to the determining.

4. The system of claim 3 , wherein the functions further comprise:

after invoking the first voice assistant, capturing a third voice input via the at least one microphone;

determining, via the local NLU, that portions of the captured third voice input correspond to keywords that in the set of local keywords; and

in response to the determining, processing the third voice input via the local NLU.

5. The system of claim 4 , wherein the local NLU comprises service keywords corresponding to control of playback from a particular streaming media service, and wherein processing the third voice input via the local NLU comprises:

determining that portions of the third voice input match one or more service keywords corresponding to one or more particular media playback commands; and

playing back audio content from one or more servers of the particular streaming media service according to the one or more particular media playback commands.

6. The system of claim 3 , wherein the set of local keywords correspond to media playback transport commands, and wherein the cloud-based NLU is configured to determine intent of voice inputs representing other commands.

7. The system of claim 3 , wherein the second voice assistant comprises (i) the local NLU and (ii) an additional cloud-based NLU, and wherein the functions further comprise:

determining, via the local NLU, that portions of the captured second voice input correspond to keywords that are not in the set of local keywords, wherein the portable playback device in response to the determining sends the data representing the second voice input to an additional computing system for processing via the additional cloud-based NLU.

8. The system of claim 1 , wherein the portable playback device comprises a touch control interface carried on an exterior surface of the housing, wherein detecting the first input to invoke voice assistance comprises detecting a particular set of one or more touch inputs to the touch control interface, and wherein detecting the second input to invoke voice assistance comprises detecting the same particular set of one or more touch inputs to the touch control interface.

9. The system of claim 1 , wherein detecting the first input to invoke voice assistance comprises receiving data representing a particular input to a control interface on the mobile device, and wherein detecting the second input to invoke voice assistance comprises detecting the same particular input to the control interface on the mobile device.

10. The system of claim 1 , wherein the functions further comprise:

switching to the concurrent connection mode;

while in the concurrent connection mode, detecting an additional input to invoke voice assistance;

in response to detecting the additional input, invoking the first voice assistant of the set of voice assistants, the first voice assistant corresponding to the concurrent connection mode;

after invoking the first voice assistant, capturing an additional voice input via the at least one microphone; and

processing the additional voice input via the first voice assistant.

11. A portable playback device comprising:

at least one microphone;

at least one audio transducer;

at least one battery;

a first network interface comprising at least one first antenna;

a second network interface comprising at least one second antenna;

at least one processor;

a housing carrying the at least one microphone, the at least one audio transducer, the at least one battery, the first network interface, the second network interface, and the at least one processor;

data storage; and

program instructions stored on the data storage that, when executed by the at least one processor, cause the portable playback device to perform functions comprising:

while in a first connection mode of a set of connection modes, detecting a first input to invoke voice assistance, wherein the set of connection modes comprises: (i) the first connection mode wherein the portable playback device connects to a wireless local area network (WLAN) via the first network interface, (ii) a second connection mode wherein the portable playback device connects to a personal area network (PAN) via the second network interface, and (iii) a concurrent connection mode wherein the portable playback device concurrently connects to the WLAN via the first network interface and the PAN via the second network interface;

in response to detecting the first input, invoking a first voice assistant of a set of voice assistants, the first voice assistant corresponding to the first connection mode, wherein the first voice assistant is native to the portable playback device;

after invoking the first voice assistant, capturing a first voice input via the at least one microphone;

processing the first voice input via the first voice assistant;

switching to the second connection mode;

while in the second connection mode, detecting a second input to invoke voice assistance;

in response to detecting the second input, invoking a second voice assistant of the set of voice assistants, the second voice assistant corresponding to the second connection mode, wherein the second voice assistant is native to a mobile device connected to the portable playback device over the PAN;

after invoking the second voice assistant, capturing a second voice input via the at least one microphone; and

sending, over the PAN via the second network interface to the mobile device, data representing the second voice input for processing of the second voice input.

12. The portable playback device of claim 11 , wherein the first voice assistant comprises cloud-based natural language understanding (NLU), and wherein processing the first voice input via the first voice assistant comprises:

sending, via the first network interface, data representing the first voice input to a computing system for processing of the first voice input via the cloud-based NLU.

13. The portable playback device of claim 12 , wherein the first voice assistant further comprises local NLU configured to recognize a set of local keywords, and wherein processing the first voice input further comprises:

determining, via the local NLU, that portions of the captured first voice input correspond to keywords that are not in the set of local keywords, wherein the portable playback device sends the data representing the first voice input to the computing system in response to the determining.

14. The portable playback device of claim 13 , wherein the functions further comprise:

after invoking the first voice assistant, capturing a third voice input via the at least one microphone;

determining, via the local NLU, that portions of the captured third voice input correspond to keywords that in the set of local keywords; and

in response to the determining, processing the third voice input via the local NLU.

15. The portable playback device of claim 14 , wherein the local NLU comprises service keywords corresponding to control of playback from a particular streaming media service, and wherein processing the third voice input via the local NLU comprises:

determining that portions of the third voice input match one or more service keywords corresponding to one or more particular media playback commands; and

playing back audio content from one or more servers of the particular streaming media service according to the one or more particular media playback commands.

16. The portable playback device of claim 13 , wherein the set of local keywords correspond to media playback transport commands, and wherein the cloud-based NLU is configured to determine intent of voice inputs representing other commands.

17. The portable playback device of claim 13 , wherein the second voice assistant comprises (i) the local NLU and (ii) an additional cloud-based NLU, and wherein the functions further comprise:

determining, via the local NLU, that portions of the captured second voice input correspond to keywords that are not in the set of local keywords, wherein the portable playback device in response to the determining sends the data representing the second voice input to an additional computing system for processing via the additional cloud-based NLU.

18. The portable playback device of claim 11 , wherein the portable playback device comprises a touch control interface carried on an exterior surface of the housing, wherein detecting the first input to invoke voice assistance comprises detecting a particular set of one or more touch inputs to the touch control interface, and wherein detecting the second input to invoke voice assistance comprises detecting the same particular set of one or more touch inputs to the touch control interface.

19. The portable playback device of claim 11 , wherein detecting the first input to invoke voice assistance comprises receiving data representing a particular input to a control interface on the mobile device, and wherein detecting the second input to invoke voice assistance comprises detecting the same particular input to the control interface on the mobile device.

20. A method to be performed by a portable playback device, the method comprising:

while in a first connection mode of a set of connection modes, detecting a first input to invoke voice assistance, wherein the set of connection modes comprises: (i) the first connection mode wherein the portable playback device connects to a wireless local area network (WLAN) via a first network interface, (ii) a second connection mode wherein the portable playback device connects to a personal area network (PAN) via a second network interface, and (iii) a concurrent connection mode wherein the portable playback device concurrently connects to the WLAN via the first network interface and the PAN via the second network interface;

in response to detecting the first input, invoking a first voice assistant of a set of voice assistants, the first voice assistant corresponding to the first connection mode, wherein the first voice assistant is native to the portable playback device;

after invoking the first voice assistant, capturing a first voice input via at least one microphone;

processing the first voice input via the first voice assistant;

switching to the second connection mode;

while in the second connection mode, detecting a second input to invoke voice assistance;

in response to detecting the second input, invoking a second voice assistant of the set of voice assistants, the second voice assistant corresponding to the second connection mode, wherein the second voice assistant is native to a mobile device connected to the portable playback device over the PAN;

after invoking the second voice assistant, capturing a second voice input via the at least one microphone; and

sending, over the PAN via the second network interface to the mobile device, data representing the second voice input for processing of the second voice input.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2021
From: PARK, SANGAH; MYERS, RYAN; TOLOMEI, JOHN
To: SONOS, INC.
Reel/Frame 056960/0374 →
Continuity (2)
Provisional Application 62975472 · Feb 12, 2020
Related Publication 20210281950A1 · Sep 9, 2021
Cited By (29)
US 12,192,713 US 12,211,490 US 12,212,945 US 12,217,748 US 12,230,291 US 12,236,932 US 12,277,368 US 12,279,096 US 12,283,269 US 12,288,558 US 12,322,390 US 12,327,549 US 12,327,556 US 12,360,734 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,579,978 US 12,603,093 US 12,626,717 US 12,640,148 US 12,699,543 US 12,711,962