IP Library › Granted Patent US 12,651,452
Granted Patent B2
US 12,651,452 · App. 18/593,037 · Granted Jun 9, 2026

Accelerator circuitry for acceleration of non-maximum suppression for object detection

Inventors: Chunyun Chen (Singapore, SG); Mohamed Mostafa Sabry Aly (Singapore, SG)
Assignee: Nanyang Technological University
G06V10/955G06V10/22G06V10/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,452
App. No.
18/593,037
Granted
Jun 9, 2026
Kind
B2
Abstract

This document describes an accelerator circuitry for facilitating acceleration of non-maximum suppression (NMS) for detection of objections within an image.

Claims (108)

1 . An accelerator circuitry for facilitating acceleration of object detection operations of an image, the circuitry comprising:

a first set of processing elements performing a plurality of first computations in parallel, each first computation comprising:

receiving a unique bounding box associated with detected features within the image, and projecting the received bounding box to a score map cell in a first-stage three-dimensional confidence score map based on dimensions, a confidence score and a spatial location of the bounding box, and a down sampling ratio of the first-stage three-dimensional confidence score map;

a second set of processing elements communicatively coupled to the first set of processing elements, and to first and second data buffers, the second set of processing elements performing a plurality of second computations comprising:

partitioning each channel of the first-stage three-dimensional confidence score map into a plurality of regions, whereby dimensions of regions in each channel are different from dimensions of regions in other channels of the first-stage three-dimensional confidence score map,

mapping each region in each of the channels of the first-stage three-dimensional confidence score map to a corresponding kernel index in the first data buffer, and for each region, storing, at the corresponding kernel index mapped to the region, a score map cell that has a highest score in the region,

a third set of processing elements communicatively coupled to the first and second data buffers, the third set of processing elements performing a plurality of third computations comprising:

retrieving the score map cells stored in the first data buffer,

forming a padded post first-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post first-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the padded post first-stage three-dimensional confidence score map to a corresponding kernel index in the second data buffer, and for each region, storing, at the corresponding kernel index mapped to the region, a score map cell that has a highest score in the region,

wherein the accelerator circuitry accelerates detection of objects in the image based at least in part on the score map cells stored in the second data buffer.

2 . The accelerator circuitry according to claim 1 , wherein before the accelerator circuitry accelerates detection of objects in the image based at least in part on the score map cells stored in the second data buffer, the plurality of second computations performed by the second set of processing elements further comprises:

retrieving the score map cells stored in the second data buffer,

generating a second-stage three-dimensional confidence score map based on the retrieved score map cells,

concatenating channels of the second-stage three-dimensional confidence score map that each have a similar scale to form a plurality of scale-concatenated channels, wherein each scale-concatenated channel is associated with a scale of a channel of the second-stage three-dimensional confidence score map;

partitioning each of the plurality of scale-concatenated channels into regions,

mapping each region in each of the plurality of scale-concatenated channels to its corresponding kernel index in the first data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region, and

wherein the plurality of third computations performed by the third set of processing elements further comprises:

retrieving the score map cells stored in the first data buffer,

forming a padded second-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post second-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the padded post second-stage three-dimensional confidence score map to its corresponding kernel index in the second data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region.

3 . The accelerator circuitry according to claim 2 , whereby the plurality of second computations performed by the second set of processing elements further comprises:

retrieving the score map cells stored in the second data buffer,

generating a third-stage three-dimensional confidence score map based on the retrieved score map cells,

concatenating channels of the third-stage three-dimensional confidence score map that each have a similar ratio to form a plurality of ratio-concatenated channels, wherein each ratio-concatenated channel is associated with a ratio of a channel of the third-stage three-dimensional confidence score map;

partitioning each of the plurality of ratio-concatenated channels into regions,

mapping each region in each of the plurality of ratio-concatenated channels to its corresponding kernel index in the first data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region, and

wherein the plurality of third computations performed by the third set of processing elements further comprises:

retrieving the score map cells stored in the first data buffer,

forming a padded post third-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post third-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the padded post third-stage three-dimensional confidence score map to its corresponding kernel index in the second data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region.

4 . The accelerator circuitry according to claim 3 , whereby the plurality of second computations performed by the second set of processing elements further comprises:

retrieving the score map cells stored in the second data buffer,

generating a fourth-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each of the channels in the fourth-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels to its corresponding kernel index in the first data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region, and

wherein the plurality of third computations performed by the third set of processing elements further comprises:

retrieving the score map cells stored in the first data buffer,

forming a padded post fourth-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post fourth-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the post fourth-stage three-dimensional confidence score map to its corresponding kernel index in the second data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region.

5 . The accelerator circuitry according to claim 4 whereby each of the channels of the padded post first-stage three-dimensional confidence score map, the padded post second-stage three-dimensional confidence score map, the padded post third-stage three-dimensional confidence score map and the padded post fourth-stage three-dimensional confidence score map comprise border score map cells padded with zeroes.

6 . The accelerator circuitry according to claim 1 , whereby each first computation performed by the first set of processing elements to project the received bounding box to the score map cell in the first-stage three-dimensional confidence score map comprises each first computation:

performing channel recovery on the received bounding box based on the dimensions of the bounding box to determine a channel of the first-stage three-dimensional confidence score map for the received bounding box; and

performing spatial recovery on the received bounding box based on the spatial location of the receiving bounding box, and the down sampling ratio of the first-stage three-dimensional confidence score map to determine a spatial location of a score map cell on the channel of the first-stage three-dimensional confidence score map associated with the received bounding box.

7 . The accelerator circuitry according to claim 6 , whereby each first computation performed by the first set of processing elements to perform channel recovery on the received bounding box based on the dimensions of the bounding box comprises each first computation:

identifying a channel of the first-stage three-dimensional confidence score map to be used as the channel for the received bounding box based on Euclidean-distances of the channels of the first-stage three-dimensional confidence score map to the dimensions of the bounding box.

8 . The accelerator circuitry according to claim 1 , whereby kernel indices in the first and second data buffers are grouped into read groups in each data buffer, whereby kernel indices in each read group are read sequentially when it is determined that the read group contains a valid score map cell.

9 . The accelerator circuitry according to claim 1 , whereby kernel indices in the first and second data buffers are grouped into read groups in each data buffer, whereby kernel indices in each read group are skipped when it is determined that the read group does not contain a valid score map cell.

10 . The accelerator circuitry according to claim 1 further comprising a fourth set of processing elements communicatively provided between the second and third sets of processing elements, and the first and second data buffers, the fourth set of processing elements performing a plurality of fourth computations comprising:

arbitrating kernel index conflicts at the first and second data buffers.

11 . A method to facilitate acceleration of object detection operations of an image, the method comprising:

performing a plurality of first computations in parallel using a first set of processing elements, each first computation comprising the steps of:

receiving a unique bounding box associated with detected features within the image, and projecting the received bounding box to a score map cell in a first-stage three-dimensional confidence score map based on dimensions, a confidence score and a spatial location of the bounding box, and a down sampling ratio of the first-stage three-dimensional confidence score map;

performing a plurality of second computations using a second set of processing elements communicatively coupled to the first set of processing elements, and to first and second data buffers, the second set of second computations comprising the steps of:

partitioning each channel of the first-stage three-dimensional confidence score map into a plurality of regions, whereby dimensions of regions in each channel are different from dimensions of regions in other channels of the first-stage three-dimensional confidence score map,

mapping each region in each of the channels of the first-stage three-dimensional confidence score map to a corresponding kernel index in the first data buffer, and for each region, storing, at the corresponding kernel index mapped to the region, a score map cell that has a highest score in the region,

performing a plurality of third computations using a third set of processing elements communicatively coupled to the first and second data buffers, the third set of computations comprising the steps of:

retrieving the score map cells stored in the first data buffer,

forming a padded post first-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post first-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the padded post first-stage three-dimensional confidence score map to a corresponding kernel index in the second data buffer, and for each region, storing, at the corresponding kernel index mapped to the region, a score map cell that has a highest score in the region,

wherein detection of objects in the image are accelerated based at least in part on the score map cells stored in the second data buffer.

12 . The method according to claim 11 , wherein before the method of detecting objects in the image are accelerated based at least in part on the score map cells stored in the second data buffer, the plurality of second computations further comprises the steps of:

retrieving the score map cells stored in the second data buffer,

generating a second-stage three-dimensional confidence score map based on the retrieved score map cells,

concatenating channels of the second-stage three-dimensional confidence score map that each have a similar scale to form a plurality of scale-concatenated channels, wherein each scale-concatenated channel is associated with a scale of a channel of the second-stage three-dimensional confidence score map;

partitioning each of the plurality of scale-concatenated channels into regions,

mapping each region in each of the plurality of scale-concatenated channels to its corresponding kernel index in the first data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region, and

wherein the plurality of third computations further comprises the steps of:

retrieving the score map cells stored in the first data buffer,

forming a padded second-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post second-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the padded post second-stage three-dimensional confidence score map to its corresponding kernel index in the second data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region.

13 . The method according to claim 12 , whereby the plurality of second computations further comprises the steps of:

retrieving the score map cells stored in the second data buffer,

generating a third-stage three-dimensional confidence score map based on the retrieved score map cells,

concatenating channels of the third-stage three-dimensional confidence score map that each have a similar ratio to form a plurality of ratio-concatenated channels, wherein each ratio-concatenated channel is associated with a ratio of a channel of the third-stage three-dimensional confidence score map;

partitioning each of the plurality of ratio-concatenated channels into regions, mapping each region in each of the plurality of ratio-concatenated channels to its corresponding kernel index in the first data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region, and

wherein the plurality of third computations further comprises the steps of:

retrieving the score map cells stored in the first data buffer,

forming a padded post third-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post third-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the padded post third-stage three-dimensional confidence score map to its corresponding kernel index in the second data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region.

14 . The method according to claim 13 , the plurality of second computations further comprises the steps of:

retrieving the score map cells stored in the second data buffer,

generating a fourth-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each of the channels in the fourth-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels to its corresponding kernel index in the first data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region, and

wherein the plurality of third computations further comprises the steps of:

retrieving the score map cells stored in the first data buffer,

forming a padded post fourth-stage three-dimensional confidence score map based on the retrieved score map cells,

partitioning each channel of the padded post fourth-stage three-dimensional confidence score map into regions,

mapping each region in each of the channels of the post fourth-stage three-dimensional confidence score map to its corresponding kernel index in the second data buffer, and for each region, storing, at its corresponding kernel index mapped to the region, a score map cell that has a highest score in the region.

15 . The method according to claim 14 whereby each of the channels of the padded post first-stage three-dimensional confidence score map, the padded post second-stage three-dimensional confidence score map, the padded post third-stage three-dimensional confidence score map and the padded post fourth-stage three-dimensional confidence score map comprise border score map cells padded with zeroes.

16 . The method according to claim 11 , whereby the step of projecting the received bounding box to the score map cell in the first-stage three-dimensional confidence score map by each of the first computations further comprises each first computation:

performing channel recovery on the received bounding box based on the dimensions of the bounding box to determine a channel of the first-stage three-dimensional confidence score map for the received bounding box; and

performing spatial recovery on the received bounding box based on the spatial location of the receiving bounding box, and the down sampling ratio of the first-stage three-dimensional confidence score map to determine a spatial location of a score map cell on the channel of the first-stage three-dimensional confidence score map associated with the received bounding box.

17 . The method according to claim 16 , whereby the step of performing channel recovery on the received bounding box based on the scale and ratio of the bounding box by each of the first computation further comprises each first computation:

identifying a channel of the first-stage three-dimensional confidence score map to be used as the channel for the received bounding box based on Euclidean-distances of the channels of the first-stage three-dimensional confidence score map to the dimensions of the bounding box.

18 . The method according to claim 11 , whereby kernel indices in the first and second data buffers are grouped into read groups in each data buffer, whereby kernel indices in each read group are read sequentially when it is determined that the read group contains a valid score map cell.

19 . The method according to claim 11 , whereby kernel indices in the first and second data buffers are grouped into read groups in each data buffer, whereby kernel indices in each read group are skipped when it is determined that the read group does not contain a valid score map cell.

20 . The method according to claim 11 further comprising the steps of:

performing a plurality of fourth computations using a fourth set of processing elements communicatively provided between the second and third sets of processing elements, and the first and second data buffers, the plurality of fourth computations comprising the steps of:

arbitrating kernel index conflicts at the first and second data buffers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2024
From: CHEN, CHUNYUN; ALY, MOHAMED MOSTAFA SABRY
To: NANYANG TECHNOLOGICAL UNIVERSITY
Reel/Frame 066843/0882 →
Priority Claims (1)
SG 10202300546V · Mar 1, 2023 · national
Continuity (1)
Related Publication 20240296670A1 · Sep 5, 2024
References Cited (50)
US 7551785B2 · Qian · 2009 [cited by examiner]
US 9542751B2 · Mannino · 2017 [cited by examiner]
US 10769493B2 · Yu · 2020 [cited by examiner]
US 10861217B2 · Moloney · 2020 [cited by examiner]
US 11481862B2 · Mao · 2022 [cited by examiner]
US 20160328856A1 · Mannino · 2016 [cited by examiner]
US 20170061229A1 · Rastgar · 2017 [cited by examiner]
US 20190057507A1 · El-Khamy · 2019 [cited by examiner]
US 20190130189A1 · Zhou · 2019 [cited by examiner]
US 20190258878A1 · Koivisto · 2019 [cited by examiner]
US 20220222477A1 · Shen · 2022 [cited by examiner]
EP 3876206A2 · 2021 [cited by examiner]
Guo et al., 2019, “Creating 3D Bounding Box Hypothesis from Deep Network Score-Maps” (pp. 904-908) (Year: 2019). [cited by examiner]
Bachrach et al., “Chisel: Constructing Hardware in a Scala Embedded Language,” in DAC Design automation conference 2012. IEEE, pp. 1212-1221, 2012. [cited by applicant]
Bodla et al., “Soft-NMS—Improving Object Detection With One Line of Code,” in ICCV, pp. 5561-5569, 2017. [cited by applicant]
Cai et al., “MaxpoolNMS: Getting Rid of NMS Bottlenecks in Two-Stage Object Detectors,” in CVPR, 2019, pp. 9356-9364. [cited by applicant]
Dalal & Triggs, “Histograms of Oriented Gradients for Human Detection,” in 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05), vol. 1. Ieee, pp. 1-8, 2005. [cited by applicant]
Deng et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition,” in CVPR, pp. 4690-4699, 2019. [cited by applicant]
Deng et al., “RetinaFace: Single-stage Dense Face Localisation in the Wild,” arXiv preprint arXiv:1905.00641, 2019. [cited by applicant]
Desai et al., “Discriminative Models for Multi-Class Object Layout,” Int J Comput Vis, pp. 1-12, 2011. [cited by applicant]
Everingham et al., “The PASCAL Visual Object Classes (VOC) Challenge,” Int J Comput Vis, pp. 1-36, 2009. [cited by applicant]
Felzenszwalb et al., “A Discriminatively Trained, Multiscale, Deformable Part Model,” in 2008 IEEE conference on computer vision and pattern recognition. Ieee, pp. 1-8, 2008. [cited by applicant]
Girshick et al., “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR, pp. 580-587, 2014. [cited by applicant]
Girshick, “Fast R-CNN,” in CVPR, pp. 1440-1448, 2015. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” in CVPR, 2016, pp. 770-778. [cited by applicant]
Henderson & Ferrari, “End-to-end training of object class detectors for mean average precision,” in Asian Conference on Computer Vision. Springer, 1-15, 2016. [cited by applicant]
Hosang et al., “A convnet for non-maximum suppression,” in GCPR, pp. 1-14, 2016. [cited by applicant]
Hosang et al., “Learning non-maximum suppression,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4507-4515, 2017. [cited by applicant]
Jiang et al., “Acquisition of Localization Confidence for Accurate Object Detection,” in Proceedings of the European conference on computer vision (ECCV), pp. 1-16, 2018. [cited by applicant]
Lecun et al., “Deep learning,” Nature, vol. 521, No. 7553, pp. 436-444, 2015. [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2117-2125, 2017. [cited by applicant]
Lin et al., “Focal Loss for Dense Object Detection,” in Proceedings of the IEEE international conference on computer vision, pp. 2980-2988, 2017. [cited by applicant]
Liu et al., “Adaptive NMS: Refining Pedestrian Detection in a Crowd,” in CVPR, pp. 6459-6468, 2019. [cited by applicant]
Liu et al., “SSD: Single Shot MultiBox Detector,” in ECCV, pp. 21-37, 2016. [cited by applicant]
Mo et al., “A Multi-Task Hardwired Accelerator for Face Detection and Alignment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 11, pp. 4284-4298, 2020. [cited by applicant]
Muralimanohar et al., “CACTI 6.0: A Tool to Understand Large Caches,” University of Utah and Hewlett Packard Laboratories, Tech. Rep, vol. 147, pp. 1-20, 2009. [cited by applicant]
Oro et al., “Work-efficient parallel non-maximum suppression for embedded gpu architectures,” in ICASSP, pp. 1-5, 2016. [cited by applicant]
Oro et al., “Work-Efficient Parallel Non-Maximum Suppression Kernels,” The Computer Journal, pp. 1-5, 2020. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection,” in CVPR, pp. 779-788, 2016. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Adv Neural Inf Process Syst, vol. 28, pp. 91-99, 2015. [cited by applicant]
Rothe et al., “Non-Maximum Suppression for Object Detection by Passing Messages between Windows,” in ACCV, pp. 1-16, 2015. [cited by applicant]
Salscheider, “FeatureNMS: Non-Maximum Suppression by Learning Feature Embeddings,” in ICPR, pp. 7848-7854, 2021. [cited by applicant]
Sermanet et al., “OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks,” arXiv preprint arXiv:1312.6229, 2013. [cited by applicant]
Shi et al., “A Fast and Power-Efficient Hardware Architecture for Non-Maximum Suppression,” IEEE Transactions on Circuits and Systems—II: Express Briefs, vol. 66, No. 11, pp. 1870-1874, 2019. [cited by applicant]
Snyder, “Verilator and systemperl,” in North American SystemC Users' Group, Design Automation Conference, 2004. [cited by applicant]
Viola & Jones, “Rapid Object Detection Using a Boosted Cascade of Simple Features,” IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 1-13, 2001. [cited by applicant]
Viola & Jones, “Robust Real-Time Face Detection,” International Journal of Computer Vision, pp. 1-18, 2004. [cited by applicant]
Wan et al., “End-to-End Integration of a Convolutional Network, Deformable Parts Model and Non-Maximum Suppression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 851-859, 2015. [cited by applicant]
Zhang et al., “PSRR-MaxpooINMS: Pyramid Shifted MaxpoolNMS with Relationship Recovery,” in CVPR, pp. 15840-15848, 2021. [cited by applicant]
Zou et al., “Object Detection in 20 Years: A Survey,” arXiv preprint arXiv:1905.05055, 2019. [cited by applicant]