IP Library › Granted Patent US 12,505,649
Granted Patent B2
US 12,505,649 · App. 17/727,649 · Granted Dec 23, 2025

Object detection with a deep learning accelerator of artificial neural networks

Inventor: Sheik Dawood Beer Mohideen (Seattle, WA)
Assignee: Micron Technology, Inc.
G06V10/764G06F17/16G06V10/25G06V10/40G06V10/82G06V10/955
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,649
App. No.
17/727,649
Granted
Dec 23, 2025
Kind
B2
Abstract

Systems, devices, and methods related to an object detector and a deep learning accelerator are described. For example, a computing apparatus has an integrated circuit device with the deep learning accelerator configured to execute instructions generated by a compiler from a description of an artificial neural network of the object detector. The artificial neural network includes a first cross stage partial network to extract features from an image and a second cross stage partial network to combine the features to identify a region of interest in the image showing an object. The artificial neural network uses a technique of minimum cost assignment in assigning a classification to the object and thus avoids post processing of non-maximum suppression.

Claims (45)

1 . A method, comprising:

receiving, in a computing apparatus, data representative of an image;

extracting, by the computing apparatus from the data using a first cross stage partial network, a plurality of features;

combining, by the computing apparatus, the plurality of features to identify a single region of interest in the image via a second cross stage partial network; and

determining, by the computing apparatus, a classification of an object shown in the single region of interest in the image using a technique of minimum cost assignment.

2 . The method of claim 1 , further comprising:

receiving data representative of a description of an artificial neural network having the first cross stage partial network and the second cross stage partial network; and

generating, from the data representative of the description of the artificial neural network, a compiler output configured to be executed on the computing apparatus to perform the extracting, the combining, and the determining.

3 . The method of claim 2 , wherein the computing apparatus includes an integrated circuit die of a field-programmable gate array or application specific integrated circuit implementing a deep learning accelerator, the deep learning accelerator comprising at least one processing unit configured to perform matrix operations and a control unit configured to load instructions from random access memory for execution.

4 . The method of claim 3 , wherein the compiler output includes the instructions executable by the deep learning accelerator to implement operations of the artificial neural network and matrices used by the instructions during execution of the instructions to implement the operations of the artificial neural network.

5 . The method of claim 4 , wherein the at least one processing unit includes a matrix-matrix unit configured to operate on two matrix operands of an instruction.

6 . The method of claim 5 , wherein:

the matrix-matrix unit includes a plurality of matrix-vector units configured to operate in parallel;

each of the plurality of matrix-vector units includes a plurality of vector-vector units configured to operate in parallel; and

each of the plurality of vector-vector units includes a plurality of multiply-accumulate units configured to operate in parallel.

7 . A computing apparatus, comprising:

memory; and

a plurality of processing units configured to:

extract a plurality of features from an image using a first cross stage partial network;

combine the plurality of features to identify a single region of interest in the image via a second cross stage partial network; and

determine a classification of an object shown in the single region of interest in the image using a technique of minimum cost assignment.

8 . The computing apparatus of claim 7 , wherein the plurality of processing units are configured via a compiler output generated by a compiler from data representative of a description of an artificial neural network having the first cross stage partial network and the second cross stage partial network.

9 . The computing apparatus of claim 8 , wherein the compiler output includes instructions executable by the plurality of processing units to implement operations of the artificial neural network and matrices used by the instructions during execution of the instructions to implement the operations of the artificial neural network.

10 . The computing apparatus of claim 9 , further comprising:

an integrated circuit package configured to enclose the computing apparatus.

11 . The computing apparatus of claim 10 , further comprising:

an integrated circuit die of a field-programmable gate array or application specific integrated circuit implementing a deep learning accelerator having the plurality of processing units, including at least one processing unit configured to perform matrix operations and a control unit configured to load instructions from the memory for execution.

12 . The computing apparatus of claim 11 , wherein the at least one processing unit includes a matrix-matrix unit configured to operate on two matrix operands of an instruction.

13 . The computing apparatus of claim 12 , wherein:

the matrix-matrix unit includes a plurality of matrix-vector units configured to operate in parallel;

each of the plurality of matrix-vector units includes a plurality of vector-vector units configured to operate in parallel; and

each of the plurality of vector-vector units includes a plurality of multiply-accumulate units configured to operate in parallel.

14 . A non-transitory computer storage medium storing instructions which when executed by a computing apparatus cause the computing apparatus to perform a method, the method comprising:

extracting, by the computing apparatus from an image using a first cross stage partial network, a plurality of features;

combining, by the computing apparatus, the plurality of features to identify a single region of interest in the image via a second cross stage partial network; and

determining, by the computing apparatus, a classification of an object shown in the single region of interest in the image using a technique of minimum cost assignment.

15 . The non-transitory computer storage medium of claim 14 , wherein the instructions are generated by a compiler from a description of an artificial neural network having the first cross stage partial network and the second cross stage partial network.

16 . The non-transitory computer storage medium of claim 15 , wherein the compiler is configured for an integrated circuit die of a field-programmable gate array or application specific integrated circuit implementing a deep learning accelerator, the deep learning accelerator comprising at least one processing unit configured to perform matrix operations and a control unit configured to load instructions from random access memory for execution.

17 . The non-transitory computer storage medium of claim 16 , wherein the at least one processing unit includes a matrix-matrix unit configured to operate on two matrix operands of an instruction.

18 . The non-transitory computer storage medium of claim 17 , wherein:

the matrix-matrix unit includes a plurality of matrix-vector units configured to operate in parallel;

each of the plurality of matrix-vector units includes a plurality of vector-vector units configured to operate in parallel; and

each of the plurality of vector-vector units includes a plurality of multiply-accumulate units configured to operate in parallel.

19 . The non-transitory computer storage medium of claim 18 , wherein the compiler is further configured to generate, from the description of the artificial neural network, matrices used by the instructions during execution of the instructions to implement operations of the artificial neural network.

20 . The non-transitory computer storage medium of claim 19 , further storing the matrices generated by the compiler from the description of the artificial neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2022
From: BEER MOHIDEEN, SHEIK DAWOOD
To: MICRON TECHNOLOGY, INC.
Reel/Frame 059848/0208 →
Continuity (2)
Provisional Application 63185280 · May 6, 2021
Related Publication 20220358748A1 · Nov 10, 2022
References Cited (15)
US 10234990B2 · Hoch · 2019 [cited by examiner]
US 20200034645A1 · Fan et al. · 2020 [cited by applicant]
US 20200394458A1 · Yu · 2020 [cited by examiner]
CN 112257727 · 2021 [cited by applicant]
Guo, Chaoxu, et al. “Augfpn: Improving multi-scale feature learning for object detection.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020. (Year: 2020). [cited by examiner]
Wang, Chien-Yao, et al. “CSPNet: A new backbone that can enhance learning capability of CNN.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops. 2020. (Year: 2020). [cited by examiner]
Yufei Ma, Naveen Suda, Yu Cao, Sarma Vrudhula, Jae-sun Seo, ALAMO: FPGA acceleration of deep learning algorithms with a modularized RTL compiler, Integration, vol. 62, 2018, pp. 14-23, ISSN 0167-9260, https://doi.org/10… [cited by examiner]
A. Shawahna, S. M. Sait and A. El-Maleh, “FPGA-Based Accelerators of Deep Learning Networks for Learning and Classification: A Review,” in IEEE Access, vol. 7, pp. 7823-7859, 2019, doi: 10.1109/ACCESS.2018.2890150 (Year… [cited by examiner]
Guo, Chaoxu, et al., “AugFPN: Improving Multi-Scale Feature Learning for Object Detection.” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Aug. 5, 2020. [cited by applicant]
International Search Report and Written Opinion, PCT/US2022/026746, mailed on Aug. 12, 2022. [cited by applicant]
Ma, Yufei, et al., “ALAMO: FPGA acceleration of deep learning algorithms with a modularized RTL compiler.” Integration, the VLSI Journal vol. 62, Jan. 5, 2018. [cited by applicant]
Wang, Chien-Yao, et al., “CSPNet: A New Backbone that can Enhance Learning Capability of CNN.” IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jul. 28, 2020. [cited by applicant]
Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao, “Scaled-YOLOv4: Scaling Cross Stage Partial Network”, arXiv:2011.08036v2 [cs.CV], Feb. 22, 2021. [cited by applicant]
Chien-Yao Wang, Hong-Yuan Mark Liao, I-Hau Yeh, Yueh-Hua Wu, Ping-Yang Chen, Jun-Wei Hsieh, “CSPNet: A New Backbone that can Enhance Learning Capability of CNN”, arXiv:1911.11929v1 [cs.CV] Nov. 27, 2019. [cited by applicant]
Peize Sun, Yi Jiang, Enze Xie, Zehuan Yuan, Changhu Wang, Ping Luo, “OneNet: Towards End-to-End One-Stage Object Detection”, arXiv:2012.05780v1 [cs.CV], Dec. 10, 2020. [cited by applicant]