INTELLIGENT AUTOMATED ASSISTANT FOR MEDIA EXPLORATION
Systems and processes for operating an intelligent automated assistant are provided. In accordance with one example, a method includes, at an electronic device with one or more processors and memory, receiving a first natural-language speech input indicative of a request for media, where the first natural-language speech input comprises a first search parameter; providing, by a digital assistant, a first media item identified based on the first search parameter. The method further includes, while providing the first media item, receiving a second natural-language speech input and determining whether the second input corresponds to a user intent of refining the request for media. The method further includes, in accordance with a determination that the second speech input corresponds to a user intent of refining the request for media: identifying, based on the first parameter and the second speech input, a second media item and providing the second media item.
1 - 42 . (canceled)
43 . A method for operating a digital assistant, comprising:
at an electronic device with one or more processors and memory:
receiving a first natural-language speech input indicative of a request for media, wherein the first natural-language speech input comprises a first search parameter;
providing, by the digital assistant, playback of a first media item, wherein the first media item is identified based on the first search parameter;
while providing playback of the first media item, receiving a second natural-language speech input;
determining whether the second natural-language speech input corresponds to a user intent of refining the request for media;
in accordance with a determination that the second natural-language speech input corresponds to a user intent of refining the request for media:
identifying, based on the first search parameter and the second natural-language speech input, a second media item different from the first media item; and
providing, by the digital assistant, the second media item.
44 . The method of claim 43 , further comprising:
obtaining, based on the first natural-language speech input, a text string;
determining, based on the text string, a representation of user intent of obtaining recommendations for media items; and
determining, based on the representation of user intent, a task and one or more parameters for performing the task, wherein the one or more parameters include the first search parameter.
45 . The method of claim 43 , wherein providing playback of the first media item comprises:
providing, by the digital assistant, a speech output indicative of a verbal response associated with the first media item; and
while providing the speech output indicative of the verbal response, providing, by the digital assistant, playback of a portion of the first media item.
46 . The method of claim 43 , wherein providing playback of the first media item comprises:
providing, by the digital assistant, a plurality of media items, wherein the plurality of media items include the first media item.
47 . The method of claim 43 , further comprising:
in response to receiving the second natural-language speech input, adjusting the manner in which playback of the first media item is provided.
48 . The method of claim 43 , wherein determining whether the second natural-language speech input corresponds to a user intent of refining the request for media comprises: deriving a representation of user intent of refining the request for media based on one or more predefined phrases and natural-language equivalents of the one or more phrases.
49 . The method of claim 43 , further comprising:
obtaining, based on the second natural-language speech input, one or more parameters for refining the request for media.
50 . The method of claim 49 , wherein determining whether the second natural-language speech input corresponds to a user intent of refining the request for media comprises: deriving a representation of user intent of refining the request for media based on context information.
51 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to lyrical content of a media item.
52 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to an occasion or a time period.
53 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to an activity.
54 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to a location.
55 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to a mood.
56 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to a release date within a predetermined time frame.
57 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to an intended audience.
58 . The method of claim 49 , wherein a parameter of the one or more parameters corresponds to a collection of media items.
59 . The method of claim 49 ,
wherein the second natural-language speech input is associated with a first user; and
wherein a parameter of the one or more parameters corresponds to a second user different from the first user.
60 . The method of claim 49 , wherein obtaining the one or more parameters for refining the request for media comprises: determining the one or more parameters based on context information.
61 . The method of claim 60 , wherein the context information comprises information related to the first media item.
62 . The method of claim 60 , further comprising:
detecting physical presence of one or more users,
wherein the context information comprises information related to the one or more users.
63 . The method of claim 60 , wherein the context information comprises a setting associated with one or more users of the electronic device.
64 . The method of claim 49 , further comprising:
obtaining, based on the first natural-language speech input, a first set of media items;
selecting the first media item from the first set of media items;
obtaining, based on the second natural-language speech input, a second set of media items, wherein the second set of media items is a subset of the first set of media items; and
selecting the second media item from the second set of media items.
65 . The method of claim 64 , wherein obtaining the second set of media items comprises:
selecting, from the first set of media items, one or more media items based on the one or more parameters for refining the request for media.
66 . The method of claim 49 , wherein identifying the second media item comprises:
determining whether content associated with the second media item matches at least one of the one or more parameters.
67 . The method of claim 49 , wherein identifying the second media item comprises:
determining whether metadata associated with the second media item matches at least one of the one or more parameters.
68 . The method of claim 43 , further comprising:
obtaining the second media item from a user-specific corpus of media items, the user-specific corpus of media items generated based on data associated with a user.
69 . The method of claim 68 , further comprising:
identifying the user-specific corpus of media items based on acoustic information associated with the second natural-language speech input.
70 . The method of claim 68 , wherein a media item in the user-specific corpus of media items includes metadata indicative of: an activity; a mood; an occasion; a location; a time; a curator; a playlist; one or more previous user inputs; or any combination thereof.
71 . The method of claim 70 , wherein at least a portion of the metadata is based on information from a second user different from the first user.
72 . The method of claim 43 , further comprising:
receiving a third natural-language speech input;
determining, based on the third natural-language speech input, a representation of user intent of associating the second media item with a collection of media items;
associating the second media item with the collection of media items; and
providing, by the digital assistant, an audio output indicative of the association.
73 . The method of claim 43 , further comprising:
while providing the second media item, receiving a fourth natural-language speech input;
determining, based on the fourth natural-language speech input, a representation of user intent of obtaining information related to a particular media item;
providing, by the digital assistant, the information related to the particular media item.
74 . The method of claim 73 , further comprising: selecting the particular media item based on context information.
75 . The method of claim 43 , further comprising:
while providing the second media item, providing, by the digital assistant, a speech output indicative of a third media item;
after providing the second media item, providing playback of the third media item.
76 . The method of claim 43 , wherein providing the second media item comprises:
providing, by the digital assistant, a speech output indicative of a verbal response associated with the second media item; and
while providing the speech output indicative of the verbal response, providing, by the digital assistant, playback of a portion of the second media item.
77 . The method of claim 43 , wherein providing the second media item comprises:
providing, by the digital assistant, playback of the second media item.
78 . The method of claim 43 , wherein providing the second media item comprises:
providing, by the digital assistant, a speech output indicative of a prompt for a user selection among a plurality of media items, wherein the plurality of media items includes the second media item.
79 . The method of claim 43 , wherein the first media item is a song, an audio book, a podcast, a station, a playlist, or any combination thereof.
80 . The method of claim 43 , wherein the second media item is a song, an audio book, a podcast, a station, a playlist, or any combination thereof.
81 . The method of claim 43 , wherein the electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or a combination thereof.