IP Library › Granted Patent US 11,797,263
Granted Patent B2
US 11,797,263 · App. 17/453,632 · Granted Oct 24, 2023

Systems and methods for voice-assisted media content selection

Inventors: Sherwin Liu (Boston, MA); Paul Bates (Seattle, WA)
Assignee: Sonos, Inc.
G06F3/165G06F3/167G10L15/22G10L15/30G10L2015/221G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,797,263
App. No.
17/453,632
Granted
Oct 24, 2023
Kind
B2
Abstract

Systems and methods for media playback via a media playback system include (i) capturing a voice input comprising a request for media content, (ii) receiving information derived at least from the request for media content, (iii) requesting and receiving information from at least one remote computing device associated with a first media content service and at least one remote computing device associated with a second media content service, wherein (a) the information identifies first media content available via the first media content service for playback and identifies second media content available via the second media content service for playback, and (b) the first and second media content are related to the requested media content, and (iv) after receiving at least one of the first information and the second information, (a) selecting the first media content instead of the second media content, and (b) playing back the first media content.

Claims (48)

1. A method, comprising:

capturing voice input via a network microphone device (NMD) of a media playback system, the media playback system comprising one or more local network devices, including the network microphone device, within a physical environment and one or more first remote computing devices, wherein the voice input comprises a request for media content;

transmitting the voice input from the NMD to one or more second remote computing devices associated with a voice assistant service for deriving intent information regarding the request for media content based at least on the voice input;

receiving, at the media playback system, a response from the one or more second remote computing devices associated with the voice assistant service, wherein the response comprises the derived intent information and an identified media content service;

based at least in part on the derived intent information, requesting, via the media playback system and independent of the voice assistant service, media content information from one or more third remote computing devices associated with the identified media content service;

receiving, at the media playback system and independent of the voice assistant service, information from the one or more third remote computing devices, wherein the information identifies media content available via the media content service for playback; and

after receiving the information, (i) transmitting a uniform resource identifier (URI) or uniform resource locator (URL) associated with the media content from the one or more first remote computing devices of the media playback system to the NMD, and (ii) requesting, via the NMD, the media content, via the URI or URL, from the one or more third remote computing devices of the media content service for playback, and (iii) playing back the media content via the NMD.

2. The method of claim 1 , wherein the requesting, via the media playback system and independent of the voice assistant service, media content information from one or more third remote computing devices associated with the identified media content service comprises transmitting a request from the one or more first remote computing devices of the media playback system to the one or more third remote computing devices associated with the identified media content service.

3. The method of claim 1 , wherein the identified media content service is based at least in part on the derived intent information.

4. The method of claim 1 , further comprising:

transmitting, via the media playback system, a request for a voice response to the one or more second computing devices of the voice assistant service; and

receiving and playing back, via the media playback system, the voice response.

5. The method of claim 4 , wherein the voice response is at least one of (a) a request for additional information regarding the request for media content, and (b) an acknowledgement of receipt of the request for media content.

6. The method of claim 1 , further comprising, (i) after receiving the selection initiating the playback of the media content, and (ii) after initiating the playback of the media content, transmitting a request for a voice response to the one or more second remote computing devices of the voice assistant service.

7. The method of claim 1 , wherein the derived intent information comprises a predefined data structure including one or more media content attributes, and wherein requesting media content information from the media content service comprises querying the media content service for media corresponding to the media content attributes.

8. A media playback system, comprising:

one or more processors;

at least one network microphone device (NMD) comprising at least one microphone;

one or more first remote computing devices; and

tangible, non-transitory, computer-readable media storing instructions executable by one or more processors to cause the media playback system to perform operations comprising:

capturing voice input via the NMD, wherein the voice input comprises a request for media content;

transmitting the voice input to one or more second remote computing devices associated with a voice assistant service for deriving intent information regarding the request for media content based at least on the voice input;

receiving a response from the one or more second remote computing devices, wherein the response comprises the derived intent information and an identified media content service;

based at least in part on the derived intent information, requesting, independent of the voice assistant service, media content information from one or more third remote computing devices associated with the identified media content service;

receiving, independent of the voice assistant service, information from the one or more third remote computing devices, wherein the information identifies media content available via the media content service for playback; and

after receiving the information, (i) transmitting a uniform resource identifier (URI) or uniform resource locator (URL) associated with the media content from the one or more first remote computing devices of the media playback system to the NMD, and (ii) requesting, via the NMD, the media content, via the URI or URL, from the one or more third remote computing devices of the media content service for playback, and (iii) playing back the media content via the NMD.

9. The media playback system of claim 8 , wherein the requesting, via the media playback system and independent of the voice assistant service, media content information from one or more third remote computing devices associated with the identified media content service comprises transmitting a request from the one or more first remote computing devices of the media playback system to the one or more third remote computing devices associated with the identified media content service.

10. The media playback system of claim 8 , wherein the identified media content service is based at least in part on the derived intent information.

11. The media playback system of claim 8 , wherein the operations further comprise:

transmitting, via the media playback system, a request for a voice response to the one or more second computing devices of the voice assistant service; and

receiving and playing back, via the media playback system, the voice response.

12. The media playback system of claim 11 , wherein the voice response is at least one of (a) a request for additional information regarding the request for media content, and (b) an acknowledgement of receipt of the request for media content.

13. The media playback system of claim 8 , wherein the operations further comprise, (i) after receiving the selection initiating the playback of the media content, and (ii) after initiating the playback of the media content, transmitting a request for a voice response to the one or more second remote computing devices of the voice assistant service.

14. The media playback system of claim 8 , wherein the derived intent information comprises a predefined data structure including one or more media content attributes, and wherein requesting media content information from the media content service comprises querying the media content service for media corresponding to the media content attributes.

15. One or more tangible, non-transitory, computer-readable media storing instructions executable by one or more processors to cause a media playback system to perform operations comprising:

capturing voice input via a network microphone device (NMD) of a media playback system, the media playback system comprising one or more local network devices, including the network microphone device, within a physical environment and one or more first remote computing devices, wherein the voice input comprises a request for media content;

transmitting the voice input from the media playback system to one or more second remote computing devices associated with a voice assistant service for deriving intent information regarding the request for media content based at least on the voice input;

receiving, at the media playback system, a response from the one or more second remote computing devices associated with the voice assistant service, wherein the response comprises the derived intent information and an identified media content service;

based at least in part on the derived intent information, requesting, independent of the voice assistant service, media content information from one or more third remote computing devices associated with the identified media content service;

receiving, at the media playback system and independent of the voice assistant service, information from the one or more third remote computing devices, wherein the information identifies media content available via the media content service for playback; and

after receiving the information, (i) transmitting a uniform resource identifier (URI) or uniform resource locator (URL) associated with the media content from the one or more first remote computing devices of the media playback system to the NMD, and (ii) requesting, via the NMD, the media content, via the URI or URL, from the one or more third remote computing devices of the media content service for playback, and (iii) playing back the media content via the NMD.

16. The computer-readable media of claim 15 , wherein the requesting, via the media playback system and independent of the voice assistant service, media content information from one or more third remote computing devices associated with the identified media content service comprises transmitting a request from the one or more first remote computing devices of the media playback system to the one or more third remote computing devices associated with the identified media content service.

17. The computer-readable media of claim 15 , wherein the identified media content service is based at least in part on the derived intent information.

18. The computer-readable media of claim 15 , further comprising:

transmitting, via the media playback system, a request for a voice response to the one or more second computing devices of the voice assistant service; and

receiving and playing back, via the media playback system, the voice response.

19. The computer-readable media of claim 18 , wherein the voice response is at least one of (a) a request for additional information regarding the request for media content, and (b) an acknowledgement of receipt of the request for media content.

20. The computer-readable media of claim 15 , further comprising, (i) after receiving the selection initiating the playback of the media content, and (ii) after initiating the playback of the media content, transmitting a request for a voice response to the one or more second remote computing devices of the voice assistant service.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2021
From: LIU, SHERWIN; BATES, PAUL
To: SONOS, INC.
Reel/Frame 058025/0444 →
Continuity (3)
Continuation 16109375 · Aug 22, 2018
Provisional Application 62669385 · May 10, 2018
Related Publication 20220121418A1 · Apr 21, 2022
Cited By (2)
US 12,360,734 US 12,749,486