IP Library Granted Patent US 12,423,966
Granted Patent B2
US 12,423,966 · App. 18/203,695 · Granted Sep 23, 2025

Deep neural network-based real-time inference method, and cloud device and edge device performing deep neural network-based real-time inference method

Inventors: Joo Chan Lee (Suwon-si, KR); Jong Hwan Ko (Suwon-si, KR)
Assignee: Research & Business Foundation SUNGKYUNKWAN UNIVERSITY
G06V10/82G06V10/28G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,966
App. No.
18/203,695
Granted
Sep 23, 2025
Kind
B2
Abstract

The present disclosure relates to a deep neural network-based real-time inference apparatus, system, and method, and more particularly, to a deep neural network-based real-time inference apparatus, system, and method capable of accelerating image inference.

Claims (33)

1. A deep neural network-based real-time inference apparatus including a cloud server configured to infer an acquired image along with an edge device in a split manner, the apparatus comprising:

a memory configured to store information of a second artificial intelligence model identical to a first artificial intelligence model of the edge device; and

a processor executing one or more instructions stored in the memory, wherein the instructions, when executed by the processor, cause the processor to receive a quantized feature of an output of a first layer corresponding to a predetermined split point among a plurality of layers included in the first artificial intelligence model, and determine a processing result for the image based on the second artificial intelligence model by inputting the quantized feature to a second layer of the second artificial intelligence model corresponding a layer immediately after the first layer.

2. The deep neural network-based real-time inference apparatus of claim 1 , wherein the cloud server determines an object in the image by inputting the quantized feature to the artificial intelligence model.

3. The deep neural network-based real-time inference apparatus of claim 1 , wherein the first artificial intelligence model and the second artificial intelligence model include a deep neural network trained both with quantized and non-quantized features for each predetermined split layer.

4. The deep neural network-based real-time inference apparatus of claim 1 , wherein the processor is configured to analyze at least one of a network resource between the edge device and the cloud server, a computing resource of the edge device, and a computing resource of the cloud server, and determine a location of the split point with respect to the first layer based on an analysis result.

5. The deep neural network-based real-time inference apparatus of claim 1 , wherein the first artificial intelligence model and the second artificial intelligence model further include a quantization switch for switching between a path for quantizing a feature of each layer and a path for passing the feature of each layer without being quantized, and

wherein the quantization switch is provided for each layer of the artificial intelligence models.

6. The deep neural network-based real-time inference apparatus of claim 1 , wherein the first artificial intelligence model and the second artificial intelligence model dynamically apply a feature distribution matching unit that normalizes a distribution of features output from the first layer immediately after the split point through mean and variance.

7. The deep neural network-based real-time inference apparatus of claim 6 , wherein the feature distribution matching unit applies a convolution layer for restoring feature precision to the quantized feature.

8. A deep neural network-based real-time inference apparatus including an edge device configured to transmit a quantized feature obtained by processing an acquired image to a cloud server, the apparatus comprising:

a memory configured to store information of a first artificial intelligence model identical to a second artificial intelligence model of the cloud server; and

a processor executing one or more instructions stored in the memory, wherein the instructions, when executed by the processor, cause the processor to analyze at least one of resources, select a first layer corresponding to a predetermined split point from among a plurality of layers included in the first artificial intelligence model according to the at least one of resources, quantize only a feature of the first layer, and transmit the quantized feature to the cloud server.

9. The deep neural network-based real-time inference apparatus of claim 8 , wherein the cloud server determines a processing result for the image based on the second artificial intelligence model by inputting the quantized feature to a second layer of the second artificial intelligence model corresponding to a layer immediately after the first layer.

10. The deep neural network-based real-time inference apparatus of claim 8 , wherein the first artificial intelligence model and the second artificial intelligence model includes a deep neural network trained both with quantized and non-quantized features for each predetermined split layer.

11. The deep neural network-based real-time inference apparatus of claim 8 , wherein the at least one of resources includes a network resource between the edge device and the cloud server, a computing resource of the edge device, and a computing resource of the cloud server.

12. The deep neural network-based real-time inference apparatus of claim 8 , wherein the first artificial intelligence model and the second artificial intelligence model further include a quantization switch for switching between a path for quantizing a feature of each layer and a path for passing the feature of each layer without being quantized, and

wherein the quantization switch is provided for each layer of the artificial intelligence models.

13. The deep neural network-based real-time inference apparatus of claim 8 , wherein the first artificial intelligence model and the second artificial intelligence model dynamically apply a feature distribution matching unit that normalizes a distribution of features output from the first layer immediately after the split point through mean and variance.

14. A deep neural network-based real-time inference method, the method comprising:

acquiring, by a processor included in a cloud server, an image from an edge device;

analyzing, by a processor included in the edge device, at least one of resources related to the edge device and a cloud server;

selecting, by the processor included in the edge device, a first layer corresponding to a split point from among a plurality of layers included in a pre-trained first artificial intelligence model of the edge device according to the at least one of resources;

quantizing, by the processor included in the edge device, only a feature of the first layer corresponding to the split point;

transmitting, by the processor included in the edge device, the quantized feature to the cloud server;

inputting, by the processor included in the cloud server, the quantized feature to a second layer corresponding to a layer immediately after the first layer among a plurality of layers included in a second artificial intelligence model identical to the first artificial intelligence model; and

determining, by the processor included in the cloud server, a processing result for the image based on the second artificial intelligence model.

15. The deep neural network-based real-time inference method of claim 14 , wherein the at least one of resources includes a network resource between the edge device and the cloud server, a computing resource of the edge device, and a computing of the cloud server.

16. The deep neural network-based real-time inference method of claim 15 , wherein the first artificial intelligence model and the second artificial intelligence model further include a quantization switch for switching between a path for quantizing a feature of each layer and a path for passing the feature of each layer without being quantized, and

wherein the quantization switch is provided for each split layer of the artificial intelligence models.

17. The deep neural network-based real-time inference method of claim 15 , wherein the first artificial intelligence model and the second artificial intelligence model dynamically apply a feature distribution matching unit that normalizes a distribution of features output from the first layer immediately after the split point through mean and variance.

18. The deep neural network-based real-time inference method of claim 17 , wherein the feature distribution matching unit applies a convolution layer for restoring feature precision to the quantized feature.

19. The deep neural network-based real-time inference method of claim 14 , wherein the first artificial intelligence model and the second artificial intelligence model include a deep neural network trained both with quantized and non-quantized features for each predetermined split layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2023
From: LEE, JOO CHAN; KO, JONG HWAN
To: RESEARCH & BUSINESS FOUNDATION SUNGKYUNKWAN UNIVERSITY
Reel/Frame 063805/0174 →
Priority Claims (1)
KR 10-2022-0066704 · May 31, 2022 · national
Continuity (1)
Related Publication 20230386192A1 · Nov 30, 2023
References Cited (6)
US 20210035330A1 · Xie · 2021 [cited by examiner]
KR 1020190087264A · 2019 [cited by applicant]
KR 1020190093204A · 2019 [cited by applicant]
KR 1020200113744A · 2020 [cited by applicant]
KR 1020210062346A · 2021 [cited by applicant]
Request for the Submission of an Opinion for Korean Patent Application No. 10-2022-0066704, dated Jun. 29, 2025. [cited by applicant]