IP Library Granted Patent US 12,268,959
Granted Patent B2
US 12,268,959 · App. 17/208,991 · Granted Apr 8, 2025

Methods, apparatus and systems for optimizing communication between sender(s) and receiver(s) in computer-mediated reality applications

Inventors: Christof Fersch (Neumarkt, DE); Nicolas R. Tsingos (San Francisco, CA)
Assignees: DOLBY INTERNATIONAL AB; DOLBY LABORATORIES LICENSING CORPORATION
A63F13/428A63F13/213G06F3/012H04L67/131H04S3/008H04S7/303H04S2400/01H04S2400/11H04S2420/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,268,959
App. No.
17/208,991
Granted
Apr 8, 2025
Kind
B2
Abstract

The present invention is directed to systems, methods and apparatus for processing media content for reproduction by a first apparatus. The method includes obtaining pose information indicative of a position and/or orientation of a user. The pose information is transmitted to a second apparatus that provides the media content. The media content is rendered based on the pose information to obtain rendered media content. The rendered media content is transmitted to the first apparatus for reproduction. The present invention may include a first apparatus for reproducing media content and a second apparatus storing the media content. The first apparatus is configured to obtain pose information indicative and transmit the pose information to the second apparatus; and the second apparatus is adapted to: render the media content based on the pose information to obtain rendered media content; and transmit the rendered media content to the first apparatus for reproduction.

Claims (68)

1. A system comprising a first apparatus for reproducing media content and a second apparatus storing the media content,

wherein the first apparatus is adapted to:

obtain pose information indicative of a position or orientation of a user;

transmit the pose information to the second apparatus;

receive rendered media content and predicted pose information from the second apparatus;

compare the predicted pose information to actual pose information, wherein the actual pose information is pose information obtained at a timing at which the rendered media content is actually processed by the first apparatus for reproduction;

update the rendered media content based on a result of the comparison;

the second apparatus is adapted to:

determine the predicted pose information based on the pose information received from the first apparatus and previous pose information received from the first apparatus at a previous timing;

render the media content based on the predicted pose information to obtain the rendered media content; and

transmit the rendered media content and the predicted pose information to the first apparatus for reproduction.

2. The system according to claim 1 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content; and/or

the media content comprises video content and the rendered media content comprises rendered video content.

3. The system according to claim 1 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content; and

the first apparatus is further adapted to generate an audible representation of the rendered audio content.

4. The system according to claim 2 , wherein the audio content is one of First Order Ambisonics, FOA-based, Higher Order Ambisonics, HOA-based, object-based, or channel based audio content, or a combination of two or more of FOA-based, HOA-based, object-based, or channel based audio content.

5. The system according to claim 2 , wherein the rendered audio content is one of binaural audio content, FOA audio content, HOA audio content, or channel-based audio content, or a combination of two or more of binaural audio content, FOA audio content, HOA audio content, or channel-based audio content.

6. The system according to claim 1 , wherein the predicted pose information is predicted for an estimate of a timing at which the rendered media content is expected to be processed by the first apparatus for reproduction.

7. The system according to claim 1 , wherein the rendered media content is transmitted to the first apparatus in uncompressed form.

8. The system according to claim 1 , wherein the second apparatus is further adapted to encode the rendered media content before transmission to the first apparatus; and

the first apparatus is further adapted to decode the encoded rendered media content after reception at the first apparatus.

9. The system according to claim 6 , wherein the estimate of the timing at which the rendered media content is expected to be processed by the first apparatus for reproduction includes an estimation of a time that is necessary for encoding and decoding the rendered media content and/or an estimate of a time that is necessary for transmitting the rendered media content to the first apparatus.

10. The system according to claim 1 , wherein the predicted pose information is obtained further based on an estimate of a time that is necessary for encoding and decoding the rendered media content and/or an estimate of a time that is necessary for transmitting the rendered media content to the first apparatus.

11. The system according to claim 1 , wherein the second apparatus is further adapted to:

determine gradient information indicative of how the rendered media content changes in response to changes of the pose information; and

transmit the gradient information to the first apparatus together with the rendered media content; and

the first apparatus is further adapted to:

update the rendered media content based on the gradient information and the result of the comparison.

12. The system according to claim 1 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content;

the first apparatus is further adapted to transmit environmental information indicative of acoustic characteristics of an environment in which the first apparatus is located to the second apparatus; and

the rendering the media content is further based on the environmental information.

13. The system according to claim 1 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content;

the first apparatus is further adapted to transmit morphologic information indicative of a morphology of the user or part of the user to the second apparatus; and

the rendering the media content is further based on the morphologic information.

14. A second apparatus for providing media content for reproduction by a first apparatus, the second apparatus adapted to:

receive pose information indicative of a position or orientation of a user of the first apparatus,

wherein the second apparatus comprises:

a position predictor configured to determine predicted pose information based on the pose information received from the first apparatus and previous pose information received from the first apparatus at a previous timing;

a renderer configured to render the media content based on the predicted pose information to obtain rendered media content; and

a transmission unit configured to transmit the rendered media content and the predicted pose information to the first apparatus for updating the rendered media content for reproduction based on a comparison of the predicted pose information with actual pose information determined at the first apparatus, wherein the actual pose information is pose information obtained at a timing at which the rendered media content is actually processed by the first apparatus for reproduction.

15. The second apparatus according to claim 14 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content; and/or

the media content comprises video content and the rendered media content comprises rendered video content.

16. The second apparatus according to claim 15 , wherein the audio content is one of First Order Ambisonics, FOA-based, Higher Order Ambisonics, HOA-based, object-based, or channel based audio content, or a combination of two or more of FOA-based, HOA-based, object-based, or channel based audio content.

17. The second apparatus according to claim 15 , wherein the rendered audio content is one of binaural audio content, FOA audio content, HOA audio content, or channel-based audio content, or a combination of two or more of binaural audio content, FOA audio content, HOA audio content, or channel-based audio content.

18. The second apparatus according to claim 14 , wherein the predicted pose information is predicted for an estimate of a timing at which the rendered media content is expected to be processed by the first apparatus for reproduction.

19. The second apparatus according to claim 14 , wherein the rendered media content is transmitted to the first apparatus in uncompressed form.

20. The second apparatus according to claim 14 , further adapted to encode the rendered media content before transmission to the first apparatus.

21. The second apparatus according to claim 18 , wherein the estimate of the timing at which the rendered media content is expected to be processed by the first apparatus for reproduction includes an estimation of a time that is necessary for encoding and decoding the rendered media content and/or an estimate of a time that is necessary for transmitting the rendered media content to the first apparatus.

22. The second apparatus according to claim 14 , wherein the predicted pose information is obtained further based on an estimate of a time that is necessary for encoding and decoding the rendered media content and/or an estimate of a time that is necessary for transmitting the rendered media content to the first apparatus.

23. The second apparatus according to claim 14 , further adapted to:

determine gradient information indicative of how the rendered media content changes in response to changes of the pose information; and

transmit the gradient information to the first apparatus together with the rendered media content.

24. The second apparatus according to claim 14 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content;

the second apparatus is further adapted to receive environmental information indicative of acoustic characteristics of an environment in which the first apparatus is located from the first apparatus; and

the rendering the media content is further based on the environmental information.

25. The second apparatus according to claim 14 , wherein the media content comprises audio content and the rendered media content comprises rendered audio content;

the second apparatus is further adapted to receive morphologic information indicative of a morphology of the user or part of the user from the first apparatus; and

the rendering the media content is further based on the morphologic information.

26. A first apparatus for reproducing media content provided by a second apparatus, the first apparatus adapted to:

obtain pose information indicative of a position and/or orientation of a user of the first apparatus;

transmit the pose information to the second apparatus;

receive rendered media content and predicted pose information from the second apparatus, wherein the rendered media content has been obtained by rendering the media content based on the predicted pose information;

compare the predicted pose information to an actual pose information, wherein the actual pose information is pose information obtained at a timing at which the rendered media content is actually processed by the first apparatus for reproduction;

update the rendered media content based on a result of the comparison; and

reproduce the updated rendered media content.

27. The system according to claim 1 , wherein update the rendered media content comprises extrapolating the rendered media content.

28. The system according to claim 27 , wherein extrapolating the rendered media content comprises applying blind upmixing to the rendered media content.

29. The system according to claim 27 , wherein extrapolating the rendered media content comprises applying a local rotation to the rendered media content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2021
From: FERSCH, CHRISTOF; TSINGOS, NICOLAS R.
To: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
Reel/Frame 055701/0723 →
Priority Claims (1)
EP 17176248 · Jun 15, 2017 · regional
Continuity (4)
Division 16485928
Provisional Application 62519952 · Jun 15, 2017
Provisional Application 62680678 · Jun 5, 2018
Related Publication 20210275915A1 · Sep 9, 2021
References Cited (35)
US 6259795B1 · McGrath · 2001 [cited by applicant]
US 6331851B1 · Suzuki · 2001 [cited by examiner]
US 9396588B1 · Li · 2016 [cited by examiner]
US 20040098462A1 · Horvitz · 2004 [cited by applicant]
US 20080144794A1 · Gardner · 2008 [cited by applicant]
US 20110066262A1 · Kelly · 2011 [cited by examiner]
US 20110210982A1 · Sylvan · 2011 [cited by examiner]
US 20140222439A1 · Jung · 2014 [cited by applicant]
US 20140355825A1 · Kim · 2014 [cited by examiner]
US 20150012466A1 · Sapiro · 2015 [cited by examiner]
US 20150029218A1 · Williams · 2015 [cited by applicant]
US 20150109415A1 · Son · 2015 [cited by applicant]
US 20150304634A1 · Karvounis · 2015 [cited by examiner]
US 20150350846A1 · Chen · 2015 [cited by applicant]
US 20160361658A1 · Osman · 2016 [cited by applicant]
US 20170018121A1 · Lawson · 2017 [cited by applicant]
US 20170115488A1 · Ambrus · 2017 [cited by examiner]
US 20170295446A1 · Thagadur Shivappa · 2017 [cited by examiner]
US 20220021996A1 · Brimijoin, II · 2022 [cited by examiner]
JP 2012518313A · 2012 [cited by applicant]
JP 2014513367A · 2014 [cited by applicant]
JP 2015233252A · 2015 [cited by applicant]
JP 2017079457A · 2017 [cited by applicant]
RU 2560340C2 · 2015 [cited by applicant]
RU 2605370C2 · 2016 [cited by applicant]
WO 200155833W · 2001 [cited by applicant]
WO 2013064914A1 · 2013 [cited by applicant]
WO 2023220024A1 · 2023 [cited by applicant]
Chan, K. et al “Distributed Sound Rendering for Interactive Virtual Environments” IEEE International Conference on Multimedia and Expo ,Jun. 27-30, 2004, pp. 1-4. [cited by applicant]
Mariette, N. et al “SoundDelta” A Study of Audio Augmented Reality using WIFI-distributed Ambisonic Cell Rendering AES, presented at the 128th Convention, May 22-25, 2010, London, UK, pp. 1-15. [cited by applicant]
Warusfel, O. et al “LISTEN Augmenting Everyday Environments Through Interactive Soundscapes” Cordis Eu Research, start date Jan. 1, 2001. [cited by applicant]
Breebaart Jet Al: “Multi-channel goes mobile: MPEG surround binaural rendering”, AES International Conference. Audio for Mobile and Handhelddevices, Sep. 2, 2006 (Sep. 2, 2006), pp. 1-13. 13 pages. [cited by applicant]
J Breebaart et al.: “Binaural Cues for Multiple Sound Sources” In: “Spatial Audio Processing: MPEG Surround and Other Applications”, Jan. 1, 2007 (Jan. 1, 2007), John Wiley & Sons. 16 pages. [cited by applicant]
Minnaar Pauli et al.: “The importance of head movements for binaural room synthesis—a pilot experiment”, Jan. 1, 2000 (Jan. 1, 2000). 6 pages. [cited by applicant]
Natural listening over headphones in augmented reality using adaptive filtering techniques. Rishabh Ranjan, Woon- Seng Gan. IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 23, Issue 11, Nov. 2015 ht… [cited by applicant]