IP Library Granted Patent US 12,620,182
Granted Patent B2
US 12,620,182 · App. 18/237,666 · Granted May 5, 2026

Method and device for presenting an audio and synthesized reality experience

Inventor: Ian M. Richter (Los Angeles, CA)
Assignee: APPLE INC.
G06T19/006G06F3/165G06T7/60G06V20/10H04N21/43074
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,182
App. No.
18/237,666
Granted
May 5, 2026
Kind
B2
Abstract

In various implementations, methods of presenting an audio/SR experience are disclosed. In one embodiment, while playing an audio file in an environment, in response to determining that the respective temporal criterion and the respective environmental criterion of an SR content event is met, the SR content event is displayed in association with the environment. In one embodiment, SR content is obtained and displayed in association with an environment based on an audio file and a 3D point cloud of the environment. In one embodiment, SR content is obtained and displayed in association with an environment based on spoken words of a real sound of the environment.

Claims (41)

1 . A method comprising:

at a device including a processor, non-transitory memory, a microphone, and a display:

recording, via the microphone, a sound produced in an environment while displaying, on the display, a volumetric environment based on the environment;

performing object detection to identify in the volumetric environment representations of physical objects in the environment;

detecting, using the one or more processors, one or more spoken words in the sound;

obtaining, based on the one or more spoken words, mixed reality (MR) content; and

modifying the volumetric environment dynamically while the one or more spoken words are playing, including adding the MR content corresponding to the one or more spoken words and concurrently presenting supplemental content selected based on at least one of a tempo, a volume dynamic, or a frequency dynamic of audio data played in the volumetric environment, wherein the MR content is displayed on a portion of the display that is selected based on a location of one of the detected representations as indicated by the one or more spoken words.

2 . The method of claim 1 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a trigger word and obtaining the MR content based on the trigger word.

3 . The method of claim 2 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a modifier word associated with the trigger word and obtaining the MR content based on the modifier word.

4 . The method of claim 1 , wherein obtaining, based on the one or more spoken words, the MR content includes selecting the MR content from a library of labeled MR content elements based on at least one of the one or more spoken words.

5 . The method of claim 1 , further comprising: playing, via a speaker, an audio file associated with the MR content.

6 . The method of claim 1 , wherein obtaining the MR content is further based on one or more spatial characteristics of the environment.

7 . The method of claim 1 , wherein obtaining the MR content is based on an environmental class of the environment.

8 . The method of claim 1 , wherein obtaining the MR content is based on an object of a particular shape detected in the environment.

9 . The method of claim 1 , wherein obtaining the MR content is based on an object of a particular type detected in the environment.

10 . A device comprising:

one or more processors;

a non-transitory memory;

a microphone;

a display; and

one or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the device to:

record, via the microphone, a sound produced in an environment while displaying, on the display, a volumetric environment based on the environment;

perform object detection to identify in the volumetric environment representations of physical objects in the environment;

detect, using the one or more processors, one or more spoken words in the sound;

obtain, based on the one or more spoken words, mixed reality (MR) content; and

modify the volumetric environment dynamically while the one or more spoken words are playing, including adding the MR content corresponding to the one or more spoken words and concurrently presenting supplemental content selected based on at least one of a tempo, a volume dynamic, or a frequency dynamic of audio data played in the volumetric environment, wherein the MR content is displayed on a portion of the display that is selected based on a location of one of the detected representations as indicated by the one or more spoken words.

11 . The device of claim 10 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a trigger word and obtaining the MR content based on the trigger word.

12 . The device of claim 11 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a modifier word associated with the trigger word and obtaining the MR content based on the modifier word.

13 . The device of claim 10 , wherein obtaining, based on the one or more spoken words, the MR content includes selecting the MR content from a library of labeled MR content elements based on at least one of the one or more spoken words.

14 . The device of claim 10 , wherein the one or more programs, which, when executed by the one or more processors, further cause the device to play, via a speaker, an audio file associated with the MR content.

15 . The device of claim 10 , wherein obtaining the MR content is further based on one or more spatial characteristics of the environment.

16 . The device of claim 10 , wherein obtaining the MR content is based on an environmental class of the environment.

17 . The device of claim 10 , wherein obtaining the MR content is based on an object of a particular shape detected in the environment.

18 . The device of claim 10 , wherein obtaining the MR content is based on an object of a particular type detected in the environment.

19 . A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device with a microphone and a display, cause the device to:

record, via the microphone, a sound produced in an environment while displaying, on the display, a volumetric environment based on the environment;

perform object detection to identify in the volumetric environment representations of physical objects in the environment;

detect, using the one or more processors, one or more spoken words in the sound;

obtain, based on the one or more spoken words, mixed reality (MR) content; and

modify the volumetric environment dynamically while the one or more spoken words are playing, including adding the MR content corresponding to the one or more spoken words and concurrently presenting supplemental content selected based on at least one of a tempo, a volume dynamic, or a frequency dynamic of audio data played in the volumetric environment, wherein the MR content is displayed on a portion of the display that is selected based on a location of one of the detected representations as indicated by the one or more spoken words.

20 . The non-transitory memory of claim 19 , wherein obtaining, based on the one or more spoken words, the MR content includes detecting, in the one or more spoken words, a trigger word and obtaining the MR content based on the trigger word.

Continuity (3)
Continuation 17053676
Provisional Application 62677904 · May 30, 2018
Related Publication 20240054734A1 · Feb 15, 2024
References Cited (26)
US 8984405B1 · Geller et al. · 2015 [cited by applicant]
US 9143742B1 · Amira et al. · 2015 [cited by applicant]
US 9704298B2 · Espeset et al. · 2017 [cited by applicant]
US 10284809B1 · Noel · 2019 [cited by applicant]
US 10503964B1 · Valgardsson et al. · 2019 [cited by applicant]
US 10555023B1 · Mccarthy et al. · 2020 [cited by applicant]
US 10979676B1 · Kelly et al. · 2021 [cited by applicant]
US 20130042296A1 · Hastings · 2013 [cited by examiner]
US 20130044128A1 · Liu · 2013 [cited by examiner]
US 20140075317A1 · Dugan et al. · 2014 [cited by applicant]
US 20140176604A1 · Venkitarman et al. · 2014 [cited by applicant]
US 20150382079A1 · Lister et al. · 2015 [cited by applicant]
US 20160012609A1 · Laska et al. · 2016 [cited by applicant]
US 20170078825A1 · Mangiat et al. · 2017 [cited by applicant]
US 20180035137A1 · Chen et al. · 2018 [cited by applicant]
US 20180046256A1 · Holz · 2018 [cited by applicant]
US 20180082117A1 · Sharma et al. · 2018 [cited by applicant]
US 20180089935A1 · Froy, Jr. · 2018 [cited by applicant]
US 20180095542A1 · Mallinson · 2018 [cited by applicant]
US 20180284955A1 · Canavor · 2018 [cited by examiner]
US 20190176027A1 · Smith · 2019 [cited by examiner]
CN 106448687A · 2017 [cited by applicant]
CN 108062796A · 2018 [cited by applicant]
Angel X. Chang, Mihail Eric, Manolis Savva, and Christopher D. Manning. 2017. SceneSeer: 3D Scene Design with Natural Language. CoRR abs/1703.00050 (2017). http://arxiv.org/abs/1703.00050. [cited by examiner]
International Search Report and Written Opinion dated Nov. 11, 2019, International Application No. PCT/US2019/034324, pp. 1-14. [cited by applicant]
First Chinese Office Action dated Apr. 2, 2024, Chinese Application No. 2019800323475, pp. 1-8 (Includes English Translation of Reporting Letter). [cited by applicant]