IP Library › Granted Patent US 12,284,508
Granted Patent B2
US 12,284,508 · App. 18/079,669 · Granted Apr 22, 2025

Visual content presentation with viewer position-based audio

Inventors: Shai Messingher Lang (Santa Clara, CA); Alexandre Da Veiga (San Francisco, CA); Spencer H. Ray (San Jose, CA); Symeon Delikaris Manias (Playa Vista, CA)
Assignee: Apple Inc.
H04S7/303G06F3/011H04S3/008H04S2400/01H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,284,508
App. No.
18/079,669
Granted
Apr 22, 2025
Kind
B2
Abstract

Various implementations disclosed herein include devices, systems, and methods that display visual content as part of a 3D environment and add audio corresponding to the visual content. The audio may be spatialized to be from one or more audio source locations within the 3D environment. For example, a video may be presented on a virtual surface within an extended reality (XR) environment while audio associated with the video is spatialized to sound as if it is produced from an audio source location corresponding to that virtual surface. How the audio is provided may be determined based on the position of the viewer (e.g., the user or his/her device) relative to the presented visual content.

Claims (50)

1. A method comprising:

at a device having a processor:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises a distance of the viewer from the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a single point source audio mode and wherein the single point source audio mode is selected based on the distance exceeding a threshold; and

presenting audio content with the visual content according to the audio mode.

2. The method of claim 1 , wherein a position of the point source is determined based on the position of the visual content.

3. A method comprising:

at a device having a processor:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises the viewer being located outside of a shape associated with the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a single point source audio mode and wherein the single point source audio mode is selected based on the viewer being located outside of the shape; and

presenting audio content with the visual content according to the audio mode.

4. A method comprising:

at a device having a processor:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises a distance of the viewer from the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a multi-channel audio mode, wherein the multi-channel audio mode is selected based on the distance being less than a threshold; and

presenting audio content with the visual content according to the audio mode.

5. A method comprising:

at a device having a processor:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises the viewer being located outside of a shape associated with the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a multi-channel audio mode, wherein the multi-channel audio mode is selected based on the viewer being located outside of the shape; and

presenting audio content with the visual content according to the audio mode.

6. A method comprising:

at a device having a processor:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises the viewer being located within a shape associated with the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a spatialized audio mode, wherein the spatialized audio mode is selected based on the viewer being located within the shape; and

presenting audio content with the visual content according to the audio mode.

7. A method comprising:

at a device having a processor:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment;

determining an audio mode based on the positional relationship;

presenting audio content with the visual content according to the audio mode; and

increasing audio spatialization based on detecting the viewer approaching the visual content.

8. A system comprising:

a non-transitory computer-readable storage medium; and

one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises a distance of the viewer from the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a single point source audio mode and wherein the single point source audio mode is selected based on the distance exceeding a threshold; and

presenting audio content with the visual content according to the audio mode.

9. A non-transitory computer-readable storage medium storing program instructions executable on a device to perform operations comprising:

determining a position in a three-dimensional (3D) environment to display visual content;

determining a positional relationship of a viewer relative to the visual content in the 3D environment, wherein the positional relationship comprises a distance of the viewer from the visual content;

determining an audio mode based on the positional relationship, wherein the audio mode is a single point source audio mode and wherein the single point source audio mode is selected based on the distance exceeding a threshold; and

presenting audio content with the visual content according to the audio mode.

Continuity (3)
Continuation PCTUS2021035573 · Jun 3, 2021
Provisional Application 63038961 · Jun 15, 2020
Related Publication 20230262406A1 · Aug 17, 2023
References Cited (3)
US 20180332420A1 · Salume et al. · 2018 [cited by applicant]
US 20220092862A1 · Faulkner · 2022 [cited by examiner]
International Search Report and Written Opinion of the International Searching Authority, International Application No. PCT/US2021/035573, 10 pages, Sep. 23, 2021. [cited by applicant]