IP Library › Granted Patent US 12,505,663
Granted Patent B2
US 12,505,663 · App. 18/266,734 · Granted Dec 23, 2025

Method for compressing an AI-based object detection model for deployment on resource-limited devices

Inventors: Devesh Walawalkar (Pittsburgh, PA); Marios Savvides (Pittsburgh, PA)
Assignee: CARNEGIE MELLON UNIVERSITY
G06V10/87G06N3/082G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,663
App. No.
18/266,734
Granted
Dec 23, 2025
Kind
B2
Abstract

Disclosed herein is a method for efficiently reducing the computational footprint of any AI-based object detection model, so as to enable its real-time deployment on computing resource-limited (i.e., low-power, embedded) devices. The disclosed method provides a step-by-step framework using an optimized combination of compression techniques to effectively compress any given AI-based object detection model.

Claims (19)

1 . A method for compressing an AI-based object detector comprising:

creating a training dataset of images with labelled images from one or more object classes of interest, the images having an initial size;

choosing an object detection model having a backbone component and a detection branch;

replacing the backbone component of the chosen object detection model with a smaller model;

reducing the size of the images in the training dataset;

training the object detection model on the reduced-size training dataset;

pruning the object detection model using a pruning algorithm;

quantizing the object detection model; and

deploying the object detection model;

wherein the object detection model with the smaller model backbone component is first trained on a diverse dataset of images containing many object classes, including both object classes of interest and object classes not of interest, to establish a base set of weights and subsequently trained on the training dataset of reduced size images containing labelled objects from the one or more object classes of interest.

2 . The method of claim 1 wherein the smaller model used to replace the backbone of the chosen object detection model is MobileNetv2 or EfficientNet.

3 . The method of claim 1 wherein the initial size of the images in the training dataset is 240×240 and further wherein the images are reduced to a 180×180 resolution.

4 . The method of claim 1 wherein the size of the images in the training dataset is reduced by at least 30%.

5 . The method of claim 1 wherein the pruning algorithm used to prune the object detection model is Network Slimming.

6 . The method of claim 1 wherein quantizing the object detection model comprises replacing float32 precision weights with float16 weights, int8 weights or a combination of float16 and int8 weights.

7 . A system comprising:

a processor;

memory, storing software that, when executed by the processor, performs the method of claim 1 .

8 . The method of claim 1 wherein objects in the object classes of interest are labelled in the reduced-size training dataset by placement of the objects in a bounding box.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2023
From: WALAWALKAR, DEVESH; SAVVIDES, MARIOS
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 065378/0358 →
Continuity (3)
Provisional Application 63150777 · Feb 18, 2021
Provisional Application 63146780 · Feb 8, 2021
Related Publication 20240096085A1 · Mar 21, 2024
References Cited (13)
US 20200349711A1 · Duke · 2020 [cited by examiner]
US 20210201661A1 · Al Jazaery · 2021 [cited by examiner]
CN 111553387 · 2020 [cited by examiner]
CN 111553387A · 2020 [cited by applicant]
CN 112001477 · 2020 [cited by examiner]
CN 112001477A · 2020 [cited by applicant]
CN 112200187 · 2021 [cited by examiner]
CN 112200187A · 2021 [cited by applicant]
Wenzheng et al, (“A Lightweight Convolutional Neural Network Flame Detection Algorithm”, Faculty of Information Technology Beijing University of Technology Beijing, China, IEEE 2021, pp. 83-86), (hereafter Wenzheng) (Ye… [cited by examiner]
International Search Report and Written Opinion for the International Application No. PCT/US22/14189, mailed Sep. 22, 2022, 14 pages. [cited by applicant]
Cai et al., “YOLOv4-5D: An Effective and Efficient Object Detector for Autonomous Driving,” IEEE Transactions on Instrumentation and Measurement, vol. 70, No. 4503613: 13 pages (2021). [cited by applicant]
Hong et al., “A New Semantic Segmentation Network with FPFN and Dense ASPP,” In IEEE International Conference on Control, Automation and Information Sciences (ICCAIS): pp. 799-804 (2021). [cited by applicant]
Written Opinion of the International Preliminary Examing Authority for Application No. PCT/Us2022/014189, mailed Apr. 16, 2024, 6 pages. [cited by applicant]