IP Library › Granted Patent US 12,573,184
Granted Patent B2
US 12,573,184 · App. 18/453,331 · Granted Mar 10, 2026

System and method for presenting three-dimensional content and three-dimensional content calculation apparatus

Inventors: Kai-Hsiang Lin (New Taipei City, TW); Hung-Chun Chou (New Taipei City, TW); Tung-Chan Tsai (New Taipei City, TW); Chieh-Sheng Wang (New Taipei City, TW); Shih-Hao Lin (New Taipei City, TW); Wen-Cheng Hsu (New Taipei City, TW)
Assignee: Acer Incorporated
G06V10/776G06T7/10G06T7/50G06V10/82G06V20/70G06T2207/20021G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,184
App. No.
18/453,331
Granted
Mar 10, 2026
Kind
B2
Abstract

A system and a method for presenting three-dimensional content and a three-dimensional content calculation apparatus are provided. In the method, the calculation apparatus receives a request for presentation content including one or more images from a client device, receives the presentation content from a content delivery network according to the request, processes the images using a first machine-learning model to generate a first predicted result, processes the images using multiple machine-learning models to generate at least a second predicted result and a third predicted result, selects a second machine-learning model from the machine-learning models based on a comparison of the first predicted result with the at least the second predicted result and the third predicted result, processes the images using the second machine-learning model and sends a processing result to the client device. Accordingly, the client device generates a three-dimensional presentation of the presentation content.

Claims (66)

1 . A system for presenting three-dimensional content, comprising:

a content delivery network comprising a server, configured to provide a presentation content consisting of one or more images;

a client device comprising a first processor, configured to send a request for the presentation content; and

a computing device comprising a second processor, connected to the content delivery device and the client device, is configured to:

receiving the request from the client device;

receiving the presentation content from the content delivery network according to the request;

processing one or more images in the presentation content using a first machine-learning model to generate a first predicted result;

processing the one or more images using multiple machine-learning models to generate at least a second predicted result and a third predicted result, wherein the first machine-learning model is larger than the machine-learning models;

selecting a second machine-learning model from the machine-learning models based on a comparison of the first predicted result with the at least the second predicted result and the third predicted result, wherein

for said comparison, the first predicted result generated from the first machine-learning model for the one or more images is utilized as a ground truth to evaluate at least the second predicted result and the third predicted result generated for the same one or more images; and

processing the images using the second machine-learning model and sends a processing result to the client device,

wherein the client device uses the processing result to generate a three-dimensional presentation of the presentation content.

2 . The system for presenting three-dimensional content according to claim 1 , wherein the computing device comprises sending the presentation content and a depth map obtained by using the second machine-learning model for processing the images to the client device, and the client device comprises generating a three-dimensional presentation of the presentation content using the presentation content and the depth map.

3 . The system for presenting three-dimensional content according to claim 1 , wherein the computing device comprises sending a depth map obtained by using the second machine-learning model for processing the images to the client device, and the client device comprises receiving the presentation content from the content delivery network, and generating a three-dimensional presentation of the presentation content using the presentation content and the depth map.

4 . The system for presenting three-dimensional content according to claim 1 , wherein the computing device comprises obtaining a depth map by using the second machine-learning model for processing the images, generating a three-dimensional presentation of the presentation content using the presentation content and the depth map, and transmitting or streaming the three-dimensional presentation of the presentation content to the client device for display.

5 . The system for presenting three-dimensional content according to claim 1 , wherein the computing device comprises:

generating a first accuracy value of the first predicted result, a second accuracy value of the second predicted result, and a third accuracy value of the third predicted result; and

comparing the first accuracy value, the second accuracy value, and the third accuracy value, wherein the second machine-learning model is selected based on the fact that the second accuracy value is higher than the first accuracy value and the third accuracy value.

6 . The system for presenting three-dimensional content according to claim 1 , wherein the computing device comprises:

using a loss function to generate a first value of the first predicted result, a second value of the second predicted result, and a third value of the third predicted result; and

comparing the first value, the second value, and the third value, wherein the second machine-learning model is selected based on the fact that the second value is higher than the first value and the third value.

7 . The system for presenting three-dimensional content according to claim 1 , wherein the first machine-learning model comprises more layers than each of the machine-learning models.

8 . A method for presenting three-dimensional content suitable for a system for presenting three-dimensional content including a content delivery network comprising a server, a client device comprising a first processor, and a computing device comprising a second processor, the method comprising the following steps:

the computing device receives a request for a presentation content from the client device, the presentation content comprises one or more images;

the computing device receives the presentation content from the content delivery network according to the request;

the computing device processes the one or more images using a first machine-learning model to generate a first predicted result;

the computing device processes the images using multiple machine-learning models to generate at least a second predicted result and a third predicted result, wherein the first machine-learning model is larger than the machine-learning models;

the computing device selects a second machine-learning model from the machine-learning models based on a comparison of the first predicted result with the at least the second predicted result and the third predicted result, wherein

for said comparison, the first predicted result generated from the first machine-learning model for the one or more images is utilized as a ground truth to evaluate at least the second predicted result and the third predicted result generated for the same one or more images;

the computing device processes the images using the second machine-learning model and sends a processing result to the client device; and

the client device uses the processing result to generate a three-dimensional presentation of the presentation content.

9 . The method according to claim 8 , further comprises:

the computing device sends the presentation content and a depth map obtained by using the second machine-learning model for processing the images to the client device; and

the client device generates a three-dimensional presentation of the presentation content using the presentation content and the depth map.

10 . The method according to claim 8 , further comprises:

the computing device sends a depth map obtained by using the second machine-learning model for processing the images to the client device; and

the client device receives the presentation content from the content delivery network, and generates a three-dimensional presentation of the presentation content using the presentation content and the depth map.

11 . The method according to claim 8 , further comprises:

the computing device obtains a depth map by using the second machine-learning model for processing the images, generates a three-dimensional presentation of the presentation content using the presentation content and the depth map, and transmits or streams the three-dimensional presentation of the presentation content to the client device for display.

12 . The method according to claim 8 , wherein the step of selecting the second machine-learning model from the machine-learning models comprises:

generating a first accuracy value of the first predicted result, a second accuracy value of the second predicted result, and a third accuracy value of the third predicted result; and

comparing the first accuracy value, the second accuracy value, and the third accuracy value, wherein the second machine-learning model is selected based on the fact that the second accuracy value is higher than the first accuracy value and the third accuracy value.

13 . The method according to claim 8 , wherein the step of selecting the second machine-learning model comprises:

using a loss function to generate a first value of the first predicted result, a second value of the second predicted result, and a third value of the third predicted result; and

comparing the first value, the second value, and the third value, wherein the second machine-learning model is selected based on the fact that the second value is higher than the first value and the third value.

14 . The method according to claim 8 , wherein the first machine-learning model comprises more layers than each of the machine-learning models.

15 . A three-dimensional content calculation apparatus, comprising:

a communications interface applying one or more communication protocols, configured to communicate with a client device comprising a first processor and a content delivery network comprising a server;

a non-transitory storage medium, configured to store instructions; and

a second processor, coupled to the communications interface and the non-transitory storage medium, and configured to access and execute instructions stored by the non-transitory storage medium to:

receiving a request for a presentation content from the client device through the communications interface, and the presentation content comprises one or more images;

receiving the presentation content from the content delivery network through the communications interface according to the request;

processing one or more images in the presentation content using a first machine-learning model to generate a first predicted result;

processing the one or more images using multiple machine-learning models to generate at least a second predicted result and a third predicted result, wherein the first machine-learning model is larger than the machine-learning models;

selecting a second machine-learning model from the machine-learning models based on a comparison of the first predicted result with the at least the second predicted result and the third predicted result, wherein

for said comparison, the first predicted result generated from the first machine-learning model for the one or more images is utilized as a ground truth to evaluate at least the second predicted result and the third predicted result generated for the same one or more images; and

processing the one or more images using the second machine-learning model and sends a processing result to the client device, so that the client device uses the processing result to generate a three-dimensional presentation of the presentation content.

16 . The three-dimensional content calculation apparatus according to claim 15 , wherein the second processor comprises sending the presentation content and a depth map obtained by using the second machine-learning model for processing the images to the client device through the communications interface, and the client device comprises generating a three-dimensional presentation of the presentation content using the presentation content and the depth map.

17 . The three-dimensional content calculation apparatus according to claim 15 , wherein the second processor comprises sending a depth map obtained by using the second machine-learning model for processing the images to the client device through the communications interface, and the client device comprises receiving the presentation content from the content delivery network, and generating a three-dimensional presentation of the presentation content using the presentation content and the depth map.

18 . The three-dimensional content calculation apparatus according to claim 15 , wherein the second processor comprises obtaining a depth map by using the second machine-learning model for processing the images, generating a three-dimensional presentation of the presentation content using the presentation content and the depth map, and transmitting or streaming the three-dimensional presentation of the presentation content to the client device for display.

19 . The three-dimensional content calculation apparatus according to claim 15 , wherein the second processor comprises:

generating a first accuracy value of the first predicted result, a second accuracy value of the second predicted result, and a third accuracy value of the third predicted result; and

comparing the first accuracy value, the second accuracy value, and the third accuracy value, wherein the second machine-learning model is selected based on the fact that the second accuracy value is higher than the first accuracy value and the third accuracy value.

20 . The three-dimensional content calculation apparatus according to claim 15 , wherein the second processor comprises:

using a loss function to generate a first value of the first predicted result, a second value of the second predicted result, and a third value of the third predicted result; and

comparing the first value, the second value, and the third value, wherein the second machine-learning model is selected based on the fact that the second value is higher than the first value and the third value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2023
From: LIN, KAI-HSIANG; CHOU, HUNG-CHUN; TSAI, TUNG-CHAN; WANG, CHIEH-SHENG; LIN, SHIH-HAO; HSU, WEN-CHENG
To: ACER INCORPORATED
Reel/Frame 064686/0983 →
Continuity (2)
Continuation 18325976 · May 30, 2023
Related Publication 20240404259A1 · Dec 5, 2024
References Cited (44)
US 10440088B2 · Jayachandran · 2019 [cited by examiner]
US 10911468B2 · Muddu et al. · 2021 [cited by applicant]
US 11308365B2 · Wen · 2022 [cited by examiner]
US 11816801B1 · Bhushan · 2023 [cited by examiner]
US 12254566B2 · Ramirez Solorzano · 2025 [cited by examiner]
US 20150145889A1 · Hanai · 2015 [cited by examiner]
US 20180300610A1 · Pye · 2018 [cited by examiner]
US 20190034784A1 · Li · 2019 [cited by examiner]
US 20190273902A1 · Varekamp · 2019 [cited by examiner]
US 20210166117A1 · Chu et al. · 2021 [cited by applicant]
US 20220101098A1 · Li · 2022 [cited by examiner]
US 20220108131A1 · Kuen · 2022 [cited by examiner]
US 20220230421A1 · Yang · 2022 [cited by examiner]
US 20220366526A1 · Das et al. · 2022 [cited by applicant]
US 20230075836A1 · Tang · 2023 [cited by examiner]
US 20230110925A1 · Akbari · 2023 [cited by examiner]
US 20230198855A1 · Ganesan et al. · 2023 [cited by applicant]
US 20230222817A1 · Sun · 2023 [cited by examiner]
US 20230232080A1 · Kim · 2023 [cited by examiner]
US 20230244910A1 · Kim · 2023 [cited by examiner]
US 20230252294A1 · Gu · 2023 [cited by examiner]
US 20230324553A1 · Timmer · 2023 [cited by examiner]
US 20230401831A1 · Krishnan · 2023 [cited by examiner]
US 20230409876A1 · Agrawal · 2023 [cited by examiner]
US 20240073452A1 · Li · 2024 [cited by examiner]
US 20240086709A1 · Donderici · 2024 [cited by examiner]
US 20240105296A1 · Fu · 2024 [cited by examiner]
US 20240112445A1 · Im · 2024 [cited by examiner]
US 20240249182A1 · Kirshenboim · 2024 [cited by examiner]
US 20240296335A1 · Jandial · 2024 [cited by examiner]
US 20240296338A1 · Yeom · 2024 [cited by examiner]
US 20240346326A1 · Tyou · 2024 [cited by examiner]
US 20240394509A1 · Cheon · 2024 [cited by examiner]
US 20240412075A1 · Zhao · 2024 [cited by examiner]
US 20250165782A1 · Sun · 2025 [cited by examiner]
US 20250259068A1 · Zisserman · 2025 [cited by examiner]
CN 114787833 · 2022 [cited by applicant]
CN 114998716 · 2022 [cited by applicant]
CN 115699204 · 2023 [cited by applicant]
TW 202145084 · 2021 [cited by applicant]
TW 202247650 · 2022 [cited by applicant]
Yinhui Ren et al, “3D Reconstruction From Monocular Images Based on Deep Convolutional Networks”, 2020 13th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), Oct.… [cited by applicant]
Daniel Kang et al., “NoScope: Optimizing Neural Network Queries over Video at Scale”, Proceedings of the VLDB Endowment, Aug. 1, 2017, pp. 1-12, vol. 10, Issue 11. [cited by applicant]
Carlos Riquelme et al., “Scaling Vision with Sparse Mixture of Experts”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 10, 2021, pp. 1-43, XP081987921. [cited by applicant]