Method and device for presenting an audio and synthesized reality experience
In various implementations, methods of presenting an audio/SR experience are disclosed. In one embodiment, while playing an audio file in an environment, in response to determining that the respective temporal criterion and the respective environmental criterion of an SR content event is met, the SR content event is displayed in association with the environment. In one embodiment, SR content is obtained and displayed in association with an environment based on an audio file and a 3D point cloud of the environment. In one embodiment, SR content is obtained and displayed in association with an environment based on spoken words of a real sound of the environment.
1 . A method comprising:
at a device including a processor, non-transitory memory, a microphone, and a display:
recording, via the microphone, a sound produced in an environment while displaying, on the display, a volumetric environment based on the environment;
performing object detection to identify in the volumetric environment representations of physical objects in the environment;
detecting, using the one or more processors, one or more spoken words in the sound;
obtaining, based on the one or more spoken words, mixed reality (MR) content; and
modifying the volumetric environment dynamically while the one or more spoken words are playing, including adding the MR content corresponding to the one or more spoken words and concurrently presenting supplemental content selected based on at least one of a tempo, a volume dynamic, or a frequency dynamic of audio data played in the volumetric environment, wherein the MR content is displayed on a portion of the display that is selected based on a location of one of the detected representations as indicated by the one or more spoken words.
2 . The method of claim 1 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a trigger word and obtaining the MR content based on the trigger word.
3 . The method of claim 2 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a modifier word associated with the trigger word and obtaining the MR content based on the modifier word.
4 . The method of claim 1 , wherein obtaining, based on the one or more spoken words, the MR content includes selecting the MR content from a library of labeled MR content elements based on at least one of the one or more spoken words.
5 . The method of claim 1 , further comprising: playing, via a speaker, an audio file associated with the MR content.
6 . The method of claim 1 , wherein obtaining the MR content is further based on one or more spatial characteristics of the environment.
7 . The method of claim 1 , wherein obtaining the MR content is based on an environmental class of the environment.
8 . The method of claim 1 , wherein obtaining the MR content is based on an object of a particular shape detected in the environment.
9 . The method of claim 1 , wherein obtaining the MR content is based on an object of a particular type detected in the environment.
10 . A device comprising:
one or more processors;
a non-transitory memory;
a microphone;
a display; and
one or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the device to:
record, via the microphone, a sound produced in an environment while displaying, on the display, a volumetric environment based on the environment;
perform object detection to identify in the volumetric environment representations of physical objects in the environment;
detect, using the one or more processors, one or more spoken words in the sound;
obtain, based on the one or more spoken words, mixed reality (MR) content; and
modify the volumetric environment dynamically while the one or more spoken words are playing, including adding the MR content corresponding to the one or more spoken words and concurrently presenting supplemental content selected based on at least one of a tempo, a volume dynamic, or a frequency dynamic of audio data played in the volumetric environment, wherein the MR content is displayed on a portion of the display that is selected based on a location of one of the detected representations as indicated by the one or more spoken words.
11 . The device of claim 10 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a trigger word and obtaining the MR content based on the trigger word.
12 . The device of claim 11 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a modifier word associated with the trigger word and obtaining the MR content based on the modifier word.
13 . The device of claim 10 , wherein obtaining, based on the one or more spoken words, the MR content includes selecting the MR content from a library of labeled MR content elements based on at least one of the one or more spoken words.
14 . The device of claim 10 , wherein the one or more programs, which, when executed by the one or more processors, further cause the device to play, via a speaker, an audio file associated with the MR content.
15 . The device of claim 10 , wherein obtaining the MR content is further based on one or more spatial characteristics of the environment.
16 . The device of claim 10 , wherein obtaining the MR content is based on an environmental class of the environment.
17 . The device of claim 10 , wherein obtaining the MR content is based on an object of a particular shape detected in the environment.
18 . The device of claim 10 , wherein obtaining the MR content is based on an object of a particular type detected in the environment.
19 . A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device with a microphone and a display, cause the device to:
record, via the microphone, a sound produced in an environment while displaying, on the display, a volumetric environment based on the environment;
perform object detection to identify in the volumetric environment representations of physical objects in the environment;
detect, using the one or more processors, one or more spoken words in the sound;
obtain, based on the one or more spoken words, mixed reality (MR) content; and
modify the volumetric environment dynamically while the one or more spoken words are playing, including adding the MR content corresponding to the one or more spoken words and concurrently presenting supplemental content selected based on at least one of a tempo, a volume dynamic, or a frequency dynamic of audio data played in the volumetric environment, wherein the MR content is displayed on a portion of the display that is selected based on a location of one of the detected representations as indicated by the one or more spoken words.
20 . The non-transitory memory of claim 19 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a trigger word and obtaining the MR content based on the trigger word.