IP Library › Granted Patent US 11,184,731
Granted Patent B2
US 11,184,731 · App. 16/822,556 · Granted Nov 23, 2021

Rendering metadata to control user movement based audio rendering

Inventors: Nils Günther Peters (San Diego, CA); Moo Young Kim (San Diego, CA); S M Akramus Salehin (San Diego, CA); Siddhartha Goutham Swaminathan (San Diego, CA); Isaac Garcia Munoz (San Diego, CA); Dipanjan Sen (Dublin, CA)
H04S7/304H04R5/033H04R5/04H04S3/008H04S2400/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,184,731
App. No.
16/822,556
Filed
Mar 18, 2020
Granted
Nov 23, 2021
Kind
B2
Art Unit
2688
USPC
381/303
Abstract

In general, techniques are described for rendering metadata to control user movement based audio rendering. A device comprising a memory and one or more processors may be configured to perform the techniques. The memory may be configured to store audio data representative of a soundfield. The one or more processors may be coupled to the memory, and configured to obtain rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device, of a renderer used to render audio data representative of a soundfield, specify, in a bitstream representative of the audio data, the rendering metadata, and output the bitstream.

Claims (88)

1. A device comprising:

a memory configured to store audio data representative of a soundfield; and

one or more processors coupled to the memory, and configured to:

obtain, from a bitstream representative of the audio data, rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device

obtain the indication of the movement of the user;

obtain delay metadata indicative of controls for enabling or disabling delay adaptations of the renderer that adjust a speed with which movement of the user results in adaptations to the render;

obtain a renderer, by which to render the audio data into one or more speaker feeds, based on the rendering metadata, the delay metadata, and the indication; and

apply the renderer to the audio data to generate the speaker feeds.

2. The device of claim 1 , wherein the one or more processors are configured to obtain translational rendering metadata indicative of controls for enabling or disabling translational adaptations, based on translational movement of the user, of the renderer.

3. The device of claim 1 , wherein the one or more processors are configured to obtain rotational rendering metadata indicative of controls for enabling or disabling rotational adaptations, based on rotational movement of the user, of the renderer.

4. The device of claim 1 , wherein the one or more processors are configured to obtain one or more of:

six degrees of freedom rendering metadata indicative of controls for enabling or disabling six degrees of freedom adaptations, based on translational movement and rotational movement of the user, of the renderer;

three degrees of freedom plus rendering metadata indicative of controls for enabling or disabling three degrees of freedom adaptations, based on translational movement of a head of the user and rotational movement of the user, of the renderer; or

three degrees of freedom rendering metadata indicative of controls for enabling or disabling three degrees of freedom adaptations, based on rotational movement of the user, of the renderer.

5. The device of claim 1 , wherein the one or more processors are configured to obtain one or more of:

x-axis rendering metadata indicative of controls for enabling or disabling x-axis adaptations, based on x-axis movement of the user of the device, of the renderer;

y-axis rendering metadata indicative of controls for enabling or disabling y-axis adaptations, based on y-axis movement of the user of the device, of the renderer; and

z-axis rendering metadata indicative of controls for enabling or disabling z-axis adaptations, based on z-axis movement of the user of the device, of the renderer;

yaw rendering metadata indicative of controls for enabling or disabling yaw adaptations, based on yaw movement of the user of the device, of the renderer;

pitch rendering metadata indicative of controls for enabling or disabling pitch adaptations, based on pitch movement of the user of the device, of the renderer; and

roll rendering metadata indicative of controls for enabling or disabling roll adaptations, based on roll movement of the user of the device, of the renderer.

6. The device of claim 1 , wherein the one or more processors are configured to obtain distance rendering metadata indicative of controls for enabling or disabling distance adaptations, based on a distance between a sound source and a location of the user of the device in the soundfield as modified by the indication of the movement of the user of the device, of the renderer.

7. The device of claim 1 , wherein the one or more processors are configured to obtain translational threshold metadata indicative of controls for enabling or disabling application of a translational threshold when performing translation adaptions, based on translation movement of the user, with respect to the renderer.

8. The device of claim 1 , wherein the device includes a virtual reality headset coupled to one or more speakers configured to reproduce, based on the speaker feeds, the soundfield.

9. The device of claim 1 , wherein the device includes an augmented reality headset coupled to one or more speakers configured to reproduce, based on the speaker feeds, the soundfield.

10. The device of claim 1 , wherein the device further includes one or more speakers configured to reproduce, based on the speaker feeds, the soundfield.

11. A device comprising:

a memory configured to store audio data representative of a soundfield; and

one or more processors coupled to the memory, and configured to:

obtain rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device, of a renderer used to render audio data representative of a soundfield;

obtain delay metadata, of the renderer, indicative of controls for enabling or disabling doppler adaptations, based on a speed of the user in a virtual environment presented to the user;

specify, in a bitstream representative of the audio data, the rendering metadata and the delay metadata; and

output the bitstream.

12. The device of claim 11 , wherein the one or more processors are configured to obtain translational rendering metadata indicative of controls for enable or disabling translational adaptations, based on translational movement of the user, of the renderer.

13. The device of claim 11 , wherein the one or more processors are configured to obtain rotational rendering metadata indicative of controls for enabling or disabling rotational adaptations, based on rotational movement of the user, of the renderer.

14. The device of claim 11 , wherein the one or more processors are configured to obtain one or more of:

six degrees of freedom rendering metadata indicative of controls for enabling or disabling six degrees of freedom adaptations, based on translational movement and rotational movement of the user, of the renderer;

three degrees of freedom plus rendering metadata indicative of controls for enabling or disabling three degrees of freedom adaptations, based on translation movement of a head of the user and rotational movement of the user, of the renderer; or

three degrees of freedom rendering metadata indicative of controls for enabling or disabling three degrees of freedom adaptations, based on rotational movement of the user, of the renderer.

15. The device of claim 11 , wherein the one or more processors are configured to obtain one or more of:

x-axis rendering metadata indicative of controls for enabling or disabling x-axis adaptations, based on x-axis movement of the user of the device, of the renderer;

y-axis rendering metadata indicative of controls for enabling or disabling y-axis adaptations, based on y-axis movement of the user of the device, of the renderer; and

z-axis rendering metadata indicative of controls for enabling or disabling z-axis adaptations, based on z-axis movement of the user of the device, of the renderer;

yaw rendering metadata indicative of controls for enabling or disabling yaw adaptations, based on yaw movement of the user of the device, of the renderer;

pitch rendering metadata indicative of controls for enabling or disabling pitch adaptations, based on pitch movement of the user of the device, of the renderer; and

roll rendering metadata indicative of controls for enabling or disabling roll adaptations, based on roll movement of the user of the device, of the renderer.

16. The device of claim 11 , wherein the one or more processors are configured to obtain distance rendering metadata indicative of controls for enabling or disabling distance adaptations, based on a distance between a sound source and the location of the device within the soundfield, of the renderer.

17. The device of claim 11 , wherein the one or more processors are configured to obtain translational threshold metadata indicative of controls for enabling or disabling application of a translational threshold when performing translation adaptions, based on translation movement of the user, with respect to the renderer.

18. A device comprising:

a memory configured to store audio data representative of a soundfield; and

one or more processors coupled to the memory, and configured to:

obtain, from a bitstream representative of the audio data, rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device;

obtain the indication of the movement of the user;

obtain doppler effect metadata indicative of controls for enabling or disabling doppler adaptations, based on a speed of the user in a virtual environment presented to the user;

obtain a renderer, by which to render the audio data into one or more speaker feeds, based on the rendering metadata, the doppler effect metadata, and the indication; and

apply the renderer to the audio data to generate the speaker feeds.

19. A device comprising:

a memory configured to store audio data representative of a soundfield; and

one or more processors coupled to the memory, and configured to:

obtain, from a bitstream representative of the audio data, rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device;

obtain the indication of the movement of the user;

obtain six degrees of freedom rendering metadata indicative of controls for enabling or disabling six degrees of freedom adaptations, based on translational movement and rotational movement of the user, of the renderer;

obtain, when the six degrees of freedom rendering metadata indicates that the six degrees of freedom adaptation of the renderer is disabled, rotational rendering metadata indicative of controls for enabling or disabling rotational adaptation, based on rotational movement of the user, of the renderer; and

obtain, when the rotational rendering metadata indicates that the rotational adaptations of the renderer is disabled, one or more of:

yaw rendering metadata indicative of controls for enabling or disabling yaw adaptations, based on yaw movement of the user of the device, of the renderer;

pitch rendering metadata indicative of controls for enabling or disabling pitch adaptations, based on pitch movement of the user of the device, of the renderer; and

roll rendering metadata indicative of controls for enabling or disabling roll adaptations, based on roll movement of the user of the device, of the renderer;

obtain a renderer, by which to render the audio data into one or more speaker feeds, based on the rendering metadata, the six degrees of freedom rendering metadata, and the indication; and

apply the renderer to the audio data to generate the speaker feeds.

20. A device comprising:

a memory configured to store audio data representative of a soundfield; and

one or more processors coupled to the memory, and configured to:

obtain rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device, of a renderer used to render audio data representative of a soundfield;

obtain doppler effect rendering metadata, of the renderer, indicative of controls for enabling or disabling doppler adaptations, based on a speed of the user in a virtual environment presented to the user, of the renderer;

specify, in a bitstream representative of the audio data, the rendering metadata and the doppler effect rendering metadata; and

output the bitstream.

21. A device comprising:

a memory configured to store audio data representative of a soundfield; and

one or more processors coupled to the memory, and configured to:

obtain rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device, of a renderer used to render audio data representative of a soundfield;

obtain six degrees of freedom rendering metadata, of the renderer, indicative of controls for enabling or disabling six degrees of freedom adaptations, based on translational movement and rotational movement of the user, of the renderer;

obtain, when the six degrees of freedom rendering metadata indicates that the six degrees of freedom adaptation of the renderer is disabled, rotational rendering metadata indicative of controls for enable or disabling rotational adaptations, based on rotational movement of the user, of the renderer; and

obtain, when the rotational rendering metadata indicates that the rotational adaptation of the renderer is disabled, one or more of:

yaw rendering metadata indicative of controls for enabling or disabling yaw adaptations, based on yaw movement of the user of the device, of the renderer;

pitch rendering metadata indicative of controls for enabling or disabling pitch adaptations, based on pitch movement of the user of the device, of the renderer; and

roll rendering metadata indicative of controls for enabling or disabling roll adaptations, based on roll movement of the user of the device, of the renderer;

specify, in a bitstream representative of the audio data, the rendering data and the six degrees of freedom rendering metadata; and

output the bitstream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: PETERS, NILS GUNTHER; KIM, MOO YOUNG; SALEHIN, S M AKRAMUS; SWAMINATHAN, SIDDHARTHA GOUTHAM; MUNOZ, ISAAC GARCIA; SEN, DIPANJAN
To: QUALCOMM INCORPORATED
Reel/Frame 056067/0274 →
Continuity (2)
Provisional Application 62821190 · Mar 20, 2019
Related Publication 20200304935A1 · Sep 24, 2020