IP Library Granted Patent US 12,561,921
Granted Patent B2
US 12,561,921 · App. 18/638,144 · Granted Feb 24, 2026

Scene encoding for multi-user mixed-reality telepresence

Inventors: Eugene Chai (Murray Hill, NJ); Kittipat Apicharttrisorn (Murray Hill, NJ); Sarit Mukherjee (Murray Hill, NJ); Limin Wang (Plainsboro, NJ); Hyunseok Chang (Holmdel, NJ)
Assignee: Nokia Solutions and Networks Oy
G06T19/006G06V10/25G06V20/20H04N19/132G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,921
App. No.
18/638,144
Granted
Feb 24, 2026
Kind
B2
Abstract

A scene encoding subsystem for multi-user mixed-reality telepresence has at least one encoder that processes signals from a plurality of cameras. A scene manager determines relative priorities of the cameras and allocates CPU resources and GPU resources in the at least one encoder based on the relative priorities of the cameras, where (i) greater CPU resources are allocated to higher priority cameras than to lower priority cameras and (ii) greater GPU resources are allocated to lower priority cameras than to higher priority cameras. The scene manager instructs a CPU to process a skipped frame using a bounding-box expansion algorithm based on object motion determined using GPU object detection and segmentation processing of previous non-skipped frames. The scene manager multiplicatively decreases and additively increases CPU and GPU resources for a set of multiple cameras based on performance of any one camera in the set.

Claims (34)

1 . A scene manager for a scene encoding subsystem for multi-user mixed-reality telepresence, the scene encoding subsystem comprising a plurality of cameras and at least one encoder adapted to process signals from the plurality of cameras, the at least one encoder comprising at least one central processing unit (CPU) and at least one graphics processing unit (GPU), the scene manager comprising:

at least one processor; and

at least one memory storing instructions that, upon being executed by the at least one processor, cause the scene manager at least to:

determine relative priorities of the cameras; and

allocate CPU resources and GPU resources in the at least one encoder based on the relative priorities of the cameras, wherein:

greater CPU resources are allocated to higher priority cameras than to lower priority cameras; and

greater GPU resources are allocated to lower priority cameras than to higher priority cameras.

2 . The scene manager of claim 1 , wherein the scene manager is adapted to modify the allocations of the CPU and GPU resources in the at least one encoder over time as the determined relative priorities of the cameras change over time.

3 . The scene manager of claim 1 , wherein the scene manager is adapted to determine the relative priorities of the cameras based on user position and field of view of each camera wherein cameras having fields of view that capture more of the user have higher priority than cameras having fields of view that capture less of the user.

4 . The scene manager of claim 3 , wherein the scene manager is adapted to instruct the at least one encoder to ignore signals from at least one camera whose field of view captures none of the user.

5 . The scene manager of claim 1 , wherein the scene manager is adapted to instruct a GPU to skip certain frames from at least one higher priority camera.

6 . The scene manager of claim 5 , wherein the scene manager is adapted to instruct a CPU to process a skipped frame using a bounding-box expansion algorithm based on object motion determined using GPU object detection and segmentation processing of previous non-skipped frames.

7 . The scene manager of claim 1 , wherein the scene manager is adapted to allocate higher point-cloud capture densities to higher-priority cameras.

8 . The scene manager of claim 1 , wherein the scene manager is adapted to perform an additive increase, multiplicative decrease (AIMD) technique to allocate the CPU and GPU resources based on the determined relative priorities of the cameras.

9 . The scene manager of claim 8 , wherein the scene manager is adapted to multiplicatively decrease and additively increase CPU and GPU resources for a set of multiple cameras based on performance of any one camera in the set.

10 . The scene manager of claim 1 , wherein the scene manager is adapted to:

receive feedback from the at least one encoder; and

control the processing of the at least one encoder based on the feedback.

11 . The scene manager of claim 10 , wherein the feedback comprises CPU and GPU processing latencies for different cameras.

12 . A method for a scene manager of a scene encoding subsystem for multi-user mixed-reality telepresence, the scene encoding subsystem comprising a plurality of cameras and at least one encoder adapted to process signals from the plurality of cameras, the at least one encoder comprising at least one CPU and at least one GPU, the method comprising the scene manager:

determining relative priorities of the cameras; and

allocating CPU resources and GPU resources in the at least one encoder based on the relative priorities of the cameras, wherein:

greater CPU resources are allocated to higher priority cameras than to lower priority cameras; and

greater GPU resources are allocated to lower priority cameras than to higher priority cameras.

13 . The method of claim 12 , wherein the scene manager modifies the allocations of the CPU and GPU resources in the at least one encoder over time as the determined relative priorities of the cameras change over time.

14 . The method of claim 12 , wherein the scene manager determines the relative priorities of the cameras based on user position and field of view of each camera wherein cameras having fields of view that capture more of the user have higher priority than cameras having fields of view that capture less of the user.

15 . The method of claim 12 , wherein the scene manager instructs a GPU to skip certain frames from at least one higher priority camera.

16 . The method of claim 15 , wherein the scene manager instructs a CPU to process a skipped frame using a bounding-box expansion algorithm based on object motion determined using GPU object detection and segmentation processing of previous non-skipped frames.

17 . The method of claim 12 , wherein the scene manager allocates higher point-cloud capture densities to higher-priority cameras.

18 . The method of claim 12 , wherein the scene manager performs an AIMD technique to allocate the CPU and GPU resources based on the determined relative priorities of the cameras.

19 . The method of claim 18 , wherein the scene manager multiplicatively decreases and additively increases CPU and GPU resources for a set of multiple cameras based on performance of any one camera in the set.

20 . The method of claim 12 , wherein the scene manager:

receives feedback from the at least one encoder; and

controls the processing of the at least one encoder based on the feedback.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2024
From: CHAI, EUGENE; APICHARTTRISORN, KITTIPAT; MUKHERJEE, SARIT; WANG, LIMIN; CHANG, HYUNSEOK
To: NOKIA OF AMERICA CORPORATION
Reel/Frame 067680/0444 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2024
From: NOKIA OF AMERICA CORPORATION
To: NOKIA SOLUTIONS AND NETWORKS OY
Reel/Frame 067680/0476 →
Continuity (1)
Related Publication 20250329117A1 · Oct 23, 2025
References Cited (36)
US 10757423B2 · Abbas · 2020 [cited by applicant]
US 10972768B2 · Zou et al. · 2021 [cited by applicant]
US 11363240B2 · Valli · 2022 [cited by applicant]
US 20230229282A1 · Reynolds et al. · 2023 [cited by applicant]
CN 115426488A · 2022 [cited by applicant]
EP 3547276A1 · 2019 [cited by applicant]
Xiang et al., “Pipelined Data-Parallel CPU/GPU Scheduling for Multi-DNN Real-Time Inference”, IEEE, 2019. (Year: 2019). [cited by examiner]
Choi, Seungbeom, et al. “Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing.” 2022 USENIX Annual Technical Conference. Carlsbad, CA, USA. (2022): 199-216. [cited by applicant]
Dhakal, Aditya, et al. “GSLICE: Controlled Spatial Sharing of GPUs for a Scalable Inference Platform.” Proceedings of the 11th ACM Symposium on Cloud Computing. Virtual Event, USA. (2020): 492-506. [cited by applicant]
Ghoshal, Moinak, et al. “Performance of Cellular Networks on the Wheels.” Proceedings of the 2023 ACM on Internet Measurement Conference. Montreal, QC, Canada (2023): 678-695. [cited by applicant]
Ginzburg, Samuel, et al. “Serverless Isn't Server-Less: Measuring and Exploiting Resource Variability on Cloud FaaS Platforms.” Proceedings of the 2020 Sixth International Workshop on Serverless Computing. Delft, Nether… [cited by applicant]
Guan, Yongjie, et al. “MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real Time.” Proceedings of the 29th Annual International Conference on Mobile Computing and Networking. Madrid, Sp… [cited by applicant]
Gül, Serhan, et al. “Cloud rendering-based volumetric video streaming system for mixed reality services.” Proceedings of the 11th ACM Multimedia Systems Conference. Istanbul, Turkey (2020): 357-360. [cited by applicant]
Gül, Serhan, et al. “Low-latency cloud-based volumetric video streaming using head motion prediction.” Proceedings of the 30th ACM Workshop on Network and Operating Systems Support for Digital Audio and Video. Istanbul,… [cited by applicant]
Han, Bo, et al. “ViVo: Visibility-aware mobile volumetric video streaming.” Proceedings of the 26th Annual International Conference on Mobile Computing and Networking. London, United Kingdom. (2020): 1-13. [cited by applicant]
Illahi, Gazi Karam, et al. “Foveated Streaming of Real-Time Graphics.” Proceedings of the 12th ACM Multimedia Systems Conference. Istanbul, Turkey (2021): 214-226. [cited by applicant]
Krajancich, Brooke, et al. “Towards attention-aware foveated rendering.” ACM Transactions on Graphics (TOG) 42.4 Article 77 (2023): 1-10. [cited by applicant]
Lawrence, Jason, et al. “Project Starline: a high-fidelity telepresence system.” ACM Transactions on Graphics (TOG) 40.6 Article 242 (2021): 1-16. [cited by applicant]
Liu, Yu, et al. “Vues: Practical Mobile Volumetric Video Streaming Through Multiview Transcoding.” Proceedings of the 28th Annual International Conference on Mobile Computing and Networking. Sydney, NSW, Australia. (202… [cited by applicant]
Natarajan, Prabhu, et al. “Multi-Camera Coordination and Control in Surveillance Systems: a Survey.” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 11.4 Article 57 (2015): 1-30. [cited by applicant]
Pasandi, Hannaneh Barahouei, et al. “CONVINCE: Collaborative Cross-Camera Video Analytics at the Edge.” 2020 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops). Austin, … [cited by applicant]
Tursun, Okan Tarhan, et al. “Luminance-Contrast-Aware Foveated Rendering.” ACM Transactions on Graphics (TOG) 38.4 Article 98 (2019): 1-14. [cited by applicant]
Xiang, Donglai, et al. “Dressing Avatars: Deep Photorealistic Appearance for Physically Simulated Clothing .” ACM Transactions on Graphics (TOG) 41.6 Article 222 (2022): 1-15. [cited by applicant]
Yu, Fuxun, et al. “Automated Runtime-Aware Scheduling for Multi-Tenant DNN Inference on GPU .” 2021 IEEE/ACM International Conference on Computer Aided Design (ICCAD), Munich, Germany (2021): 1-9. [cited by applicant]
Zhang, Anlan, et al. “YuZu: Neural-Enhanced Volumetric Video Streaming.” 19th USENIX Symposium on Networked Systems Design and Implementation Renton, WA, USA. (2022): 137-154. [cited by applicant]
Zhang, Xiaojie “Resource Management in Mobile Edge Computing for Compute-intensive Application.” Dissertation, The City University of New York (2023): 1-185. [cited by applicant]
Zhang, Letian, et al. “E3 Pose: Energy-Efficient Edge-assisted Multi-camera System for Multi-human 3D Pose Estimation.” arXiv preprint arXiv:2301.09015 (2023): 1-13. [cited by applicant]
Apple Vision Pro, www.apple.com, 2023 [retrieved on Jul. 17, 2024] Retrieved from the Internet: <URL: https://www.apple.com/apple-vision-pro/> (34 pages). [cited by applicant]
Metaverse Market Forecast, www.gminsights.com, 2022 [retrieved on Jul. 17, 2024] Retrieved from the Internet: <URL: https://www.gminsights.com/industry-analysis/metaverse-market> (8 pages). [cited by applicant]
Draco 3D Data Compression, www.github.io, 2023 [retrieved on Dec. 5, 2023] Retrieved from the Internet: <URL: https://google.github.io/draco/> (3 pages). [cited by applicant]
Project Starline: Feel like you're there, together, www.blog.google, 2023 [retrieved on Jul. 17, 2024] Retrieved from the Internet: <URL: https://blog.google/technology/research/project-starline/ > (3 pages). [cited by applicant]
Magic Leap, www.magicleap.com, 2023 [retrieved on Jul. 8, 2023] Retrieved from the Internet: <URL: https://www.magicleap.com/en-us/> (9 pages). [cited by applicant]
Microsoft Hololens 2, www.microsoft.com, 2023 [retrieved on Jul. 10, 2024] Retrieved from the Internet: <URL: https://www.microsoft.com/en-us/hololens> (8 pages). [cited by applicant]
Nvidia PeopleNet Model, www.nvidia.com, 2023 [retrieved on Dec. 11, 2023] Retrieved from the Internet: <URL: https://catalog.ngc.nvidia.com/orgs/nvidia/ teams/tao/models/peoplenet> (2 pages). [cited by applicant]
StereoLabs, www.stereolabscom, 2023 [retrieved on Nov. 15, 2023] Retrieved from the Internet: <URL: https://stereolabs.com> (10 pages). [cited by applicant]
Telefónica showcases its holographic telepresence with 3D capture at MWC, , www.telefonica.com, 2023 [retrieved on Sep. 26, 2023] Retrieved from the Internet: <URL: https://www.telefonica.com/en/communication-room/press… [cited by applicant]